AaronTekle commited on
Commit
ad20200
Β·
verified Β·
1 Parent(s): d59eafd

Create Technical_README.md

Browse files
Files changed (1) hide show
  1. Technical_README.md +1726 -0
Technical_README.md ADDED
@@ -0,0 +1,1726 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # **Agentic Machine Learning Engineer (Traditional ML Engineering) - Qwen3-Coder-30B-A3B-Instruct**
2
+
3
+ #### Note 1: Do not upload confidential, regulated, proprietary, customer, or personally identifiable production data into a public Hugging Face Space
4
+
5
+ #### Note 2: This project is an automated traditional machine learning engineering experiment focused on supervised classification and regression workflows
6
+
7
+ #### Note 3: This project uses a bounded single-agent function-calling / tool-use loop rather than an unlimited autonomous agent
8
+ ---
9
+ ### (this) Bounded Tool-Calling Agent vs Open-Ended Autonomous Agent
10
+
11
+ **Bounded Tool-Calling Agent (my agent)**
12
+
13
+ > β€œReason about the task, call the available ML tools when needed, and stop within a controlled number of steps.”
14
+
15
+ **Open-Ended Autonomous Agent**
16
+
17
+ > β€œContinue deciding and acting until the broader goal is achieved.”
18
+
19
+ | | Bounded Tool-Calling Agent | Open-Ended Autonomous Agent |
20
+ | ------------------------- | --------------------------------------------------- | ------------------------------------------- |
21
+ | **Consistency** | higher because tools and loop limits are controlled | can vary more between runs |
22
+ | **Traceability** | tool calls and observations are exposed | can become harder to trace across long runs |
23
+ | **Resource control** | hard max agent-step limit | may require additional stopping logic |
24
+ | **Decision-making** | dynamic tool selection within a constrained toolset | can make broader autonomous decisions |
25
+ | **Flexibility** | controlled | higher |
26
+ | **Risk of runaway loops** | reduced through `MAX_AGENT_STEPS` param | higher without explicit limits |
27
+
28
+ bounded agent **decides which ML operation to perform inside a controlled tool loop**, while a fully open-ended autonomous agent may continue planning and acting without the same fixed execution boundary.
29
+
30
+ ---
31
+
32
+ ### Agent Goals:
33
+
34
+ * agentic machine learning engineering app for profiling datasets, automated supervised-learning modeling, preprocessing data, training baseline models, comparing algorithms, evaluating real model performance, and generating reusable scikit-learn pipelines
35
+
36
+ * agent loads a real-world dataset, creates the modeling context, gives a Qwen3-Coder agent access to machine-learning tools, and allows the model to decide which tools to call before producing a final ML assessment and reusable pipeline code
37
+
38
+ * reduce repetitive setup work when starting supervised machine learning experiments
39
+
40
+ * ground model-performance claims in actual scikit-learn execution
41
+
42
+ * expose tool calls and observations for auditability
43
+
44
+ * demonstrate how an LLM can orchestrate traditional machine learning workflows without performing the numerical fitting itself
45
+
46
+ ---
47
+
48
+ ### Machine Learning Agent - Tasks
49
+
50
+ Machine Learning Agent can analyze datasets and execute supervised-learning workflows by first understanding the dataset, target, feature types, and modeling objective.
51
+
52
+ #### Dataset inspection
53
+
54
+ `inspect_dataset` tool can analyze:
55
+
56
+ * **Shape**: number of rows and columns
57
+ * **Columns**: available feature / target candidates
58
+ * **Types**: pandas data types
59
+ * **Numeric features**: identified numerical columns
60
+ * **Categorical features**: identified categorical columns
61
+ * **Nulls**: missing-value information
62
+ * **Cardinality**: number of unique values per column
63
+ * **Examples**: representative values from each column
64
+ * **Targets**: possible prediction-target columns
65
+
66
+ #### Modeling tasks
67
+
68
+ the ML workflow can:
69
+
70
+ * inspect the dataset
71
+ * identify a prediction target
72
+ * infer classification vs regression
73
+ * separate numeric and categorical features
74
+ * construct preprocessing transformations
75
+ * train baseline estimators
76
+ * calculate holdout metrics
77
+ * compare compatible algorithms
78
+ * inspect feature importance / coefficients where supported
79
+ * generate reusable scikit-learn pipeline code
80
+ * export fitted pipelines as Joblib artifacts
81
+
82
+ findings can then be converted into:
83
+
84
+ ```text
85
+ Dataset assessment
86
+ ↓
87
+ Recommended modeling setup
88
+ ↓
89
+ Preprocessing pipeline
90
+ ↓
91
+ Model training
92
+ ↓
93
+ Holdout evaluation
94
+ ↓
95
+ Algorithm comparison
96
+ ↓
97
+ Reusable scikit-learn pipeline
98
+ ```
99
+
100
+ ## **What, Why, and General Objectives**
101
+
102
+ ### What
103
+
104
+ Tool-using Machine Learning Engineering AI agent that profiles datasets, determines supervised-learning setup, trains real scikit-learn models, compares algorithms, evaluates model performance, and generates reusable ML pipelines.
105
+
106
+ ### Why
107
+
108
+ Machine learning engineers repeatedly perform the same early-stage workflow when starting supervised learning related tasks/projects:
109
+
110
+ ```text
111
+ load dataset
112
+ ↓
113
+ inspect features + target
114
+ ↓
115
+ determine problem type
116
+ ↓
117
+ build preprocessing
118
+ ↓
119
+ select algorithm
120
+ ↓
121
+ train model
122
+ ↓
123
+ evaluate performance
124
+ ↓
125
+ compare alternatives
126
+ ↓
127
+ generate reusable pipeline
128
+ ```
129
+
130
+ automating this workflow reduces repetitive machine learning setup, baseline experimentation, preprocessing, evaluation, and pipeline-generation work
131
+
132
+ ### General Objectives
133
+
134
+ * automate traditional ML modeling
135
+
136
+ * make model evaluation traceable through tool outputs
137
+
138
+ * provide a consistent preprocessing and model-training workflow
139
+
140
+ * compare baseline algorithms quickly
141
+
142
+ * calculate metrics from real fitted models rather than LLM estimates
143
+
144
+ * generate reusable scikit-learn pipeline code
145
+
146
+ * expose tool calls and observations for auditability
147
+
148
+ * demonstrate how an LLM agent can orchestrate traditional machine learning tools
149
+
150
+ * keep local ML execution lightweight enough for Hugging Face Spaces
151
+
152
+ ## **LLM / Agent Stack**
153
+
154
+ * **Agent LLM:** [**Qwen/Qwen3-Coder-30B-A3B-Instruct**](https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct)
155
+
156
+ * **Inference:** `huggingface_hub.InferenceClient`
157
+
158
+ * **Provider routing:** Hugging Face Inference Providers using `provider="auto"` by default
159
+
160
+ * **Agent pattern:** bounded single-agent function-calling / tool-use loop
161
+
162
+ * **Tool routing:** `tool_choice="auto"`
163
+
164
+ * **Maximum agent steps:** `4` by default
165
+
166
+ * **Prevent Infinite Loops:** keeps the agent from repeatedly calling the same tool or getting stuck in an unsuccessful reasoning loop
167
+
168
+ * **Control Costs:** limits unnecessary inference-provider requests and token usage
169
+
170
+ * **Manage Context Windows:** prevents tool observations from growing indefinitely inside the conversation context
171
+
172
+ * **Reduce Latency:** prevents unnecessarily long tool-calling runs
173
+
174
+ * **Improve Debuggability:** provides a predictable execution boundary when analyzing the tool trace
175
+
176
+ * **Local ML execution:** scikit-learn
177
+
178
+ * **Dataset processing:** pandas
179
+
180
+ * **Model artifact export:** Joblib
181
+
182
+ * **Parquet support:** PyArrow
183
+
184
+ * **Excel support:** OpenPyXL
185
+
186
+ * **HF sample dataset:** `scikit-learn/adult-census-income`
187
+
188
+
189
+ ## **Goal:**
190
+
191
+ machine learning engineering is not only model generation
192
+
193
+ a usable supervised ML workflow requires:
194
+
195
+ * understanding the dataset
196
+
197
+ * selecting a prediction target
198
+
199
+ * deciding whether the task is classification or regression
200
+
201
+ * identifying numerical and categorical features
202
+
203
+ * preprocessing missing and categorical data
204
+
205
+ * selecting an estimator
206
+
207
+ * training models on real data
208
+
209
+ * calculating ML evaluation metrics
210
+
211
+ * comparing alternatives options (ML models)
212
+
213
+ * packaging preprocessing + model logic into reusable code
214
+
215
+ doing this manually for every new dataset creates repetitive experimentation and setup work.
216
+
217
+ this Machine Learning Engineering Agent is designed to:
218
+
219
+ * reduce time spent manually inspecting new datasets
220
+
221
+ * establish a consistent supervised-learning baseline workflow
222
+
223
+ * identify numerical and categorical feature requirements
224
+
225
+ * infer or use a selected classification / regression problem type
226
+
227
+ * allow the LLM to call machine-learning execution functions
228
+
229
+ * train real models through scikit-learn
230
+
231
+ * calculate real holdout evaluation metrics
232
+
233
+ * compare compatible baseline algorithms
234
+
235
+ * generate production-oriented scikit-learn pipelines
236
+
237
+ * export fitted pipelines as Joblib artifacts
238
+
239
+ * expose the agent's tool calls through a trace for auditability
240
+
241
+ goal of the app is to combine:
242
+
243
+ ```text
244
+ LLM reasoning
245
+ +
246
+ ML execution
247
+ +
248
+ tool calling
249
+ +
250
+ reusable pipeline generation
251
+ ```
252
+
253
+ ## **Agentic Workflow:**
254
+
255
+ ```text
256
+ Dataset
257
+ ↓
258
+ app.py
259
+ ↓
260
+ Dataset loading + session state
261
+ ↓
262
+ Dataset profile
263
+ ↓
264
+ Target selection
265
+ ↓
266
+ Problem-type inference / selection
267
+ ↓
268
+ Algorithm selection
269
+ ↓
270
+ MachineLearningAgent
271
+ ↓
272
+ Qwen3-Coder-30B-A3B-Instruct
273
+ ↓
274
+ HF InferenceClient
275
+ ↓
276
+ tool_choice="auto"
277
+ ↓
278
+ Agent decides which ML tool(s) to call
279
+ ↓
280
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
281
+ β”‚ inspect_dataset β”‚
282
+ β”‚ recommend_modeling_setup β”‚
283
+ β”‚ train_candidate_model β”‚
284
+ β”‚ compare_algorithms β”‚
285
+ β”‚ generate_pipeline_code β”‚
286
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
287
+ ↓
288
+ Tool observations returned to Qwen
289
+ ↓
290
+ Agent evaluates whether another tool is needed
291
+ ↓
292
+ Bounded tool-use loop
293
+ ↓
294
+ ML assessment
295
+ ↓
296
+ Modeling risks
297
+ ↓
298
+ Recommended modeling setup
299
+ ↓
300
+ Evaluation findings
301
+ ↓
302
+ Model comparison
303
+ ↓
304
+ Generated scikit-learn pipeline
305
+ ↓
306
+ Tool Trace + Diagnostics
307
+ ```
308
+
309
+ ## **Agent Loop Type**
310
+
311
+ ### **Bounded Single-Agent Function-Calling / Tool-Use Loop**
312
+
313
+ * agentic pattern where one AI agent can autonomously choose from external ML tools, process their outputs, and repeat the process inside a strictly defined step limit
314
+
315
+ * unlike open-ended agents that can continue indefinitely, this bounded loop places a hard-fixed limit on execution
316
+
317
+ * the project uses a **single agent**, not a multi-agent architecture
318
+
319
+ ### **Architecture:**
320
+
321
+ **observe task β†’ choose tool β†’ execute ML operation β†’ return observation β†’ repeat if needed β†’ synthesize final answer**
322
+
323
+ the LLM does not directly calculate every machine learning result itself
324
+
325
+ `MachineLearningAgent` exposes Python / scikit-learn tools to Qwen through Hugging Face function calling
326
+
327
+ the model receives:
328
+
329
+ * system instructions
330
+
331
+ * user machine learning task
332
+
333
+ * dataset profile
334
+
335
+ * selected target
336
+
337
+ * inferred or selected problem type
338
+
339
+ * selected algorithm
340
+
341
+ * available tool definitions
342
+
343
+ the model can then request one or more tool calls.
344
+
345
+ ### Example Input Workflow
346
+
347
+ ```text
348
+ Qwen
349
+ ↓
350
+ inspect_dataset()
351
+ ↓
352
+ dataset profile returned to Qwen
353
+ ↓
354
+ recommend_modeling_setup()
355
+ ↓
356
+ recommended setup returned to Qwen
357
+ ↓
358
+ train_candidate_model()
359
+ ↓
360
+ real training metrics returned to Qwen
361
+ ↓
362
+ compare_algorithms()
363
+ ↓
364
+ algorithm leaderboard returned to Qwen
365
+ ↓
366
+ generate_pipeline_code()
367
+ ↓
368
+ scikit-learn pipeline code returned to Qwen
369
+ ↓
370
+ Final ML response
371
+ ```
372
+
373
+ loop is bounded by:
374
+
375
+ ```text
376
+ MAX_AGENT_STEPS = 4
377
+ ```
378
+
379
+ if the model does not finish within the configured tool loop, the application moves toward final response synthesis rather than allowing unlimited additional tool calls.
380
+
381
+ this makes the architecture a controlled **bounded tool-calling machine learning agent**.
382
+
383
+ ## **Agent Architecture:**
384
+
385
+ ```text
386
+ USER
387
+ β”‚
388
+ β–Ό
389
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
390
+ β”‚ app.py β”‚
391
+ β”‚ Gradio UI β”‚
392
+ β”‚ Session state β”‚
393
+ β”‚ Dataset loading β”‚
394
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
395
+ β”‚
396
+ β”‚ task + dataset + target
397
+ β–Ό
398
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
399
+ β”‚ MachineLearningAgent β”‚
400
+ β”‚ agent.py β”‚
401
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
402
+ β”‚
403
+ β–Ό
404
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
405
+ β”‚ Qwen3-Coder 30B A3B β”‚
406
+ β”‚ HF InferenceClient β”‚
407
+ β”‚ tool_choice="auto" β”‚
408
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
409
+ β”‚
410
+ decides which tool
411
+ to execute
412
+ β”‚
413
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
414
+ β”‚ β”‚ β”‚
415
+ β–Ό β–Ό β–Ό
416
+ inspect_dataset recommend_modeling_setup train_candidate_model
417
+ β”‚ β”‚ β”‚
418
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
419
+ β”‚
420
+ β–Ό
421
+ compare_algorithms
422
+ β”‚
423
+ β–Ό
424
+ generate_pipeline_code
425
+ β”‚
426
+ β–Ό
427
+ Tool observation
428
+ β”‚
429
+ β–Ό
430
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
431
+ β”‚ Observation added back β”‚
432
+ β”‚ to LLM context β”‚
433
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
434
+ β”‚
435
+ β–Ό
436
+ More tools needed?
437
+ / \
438
+ yes no
439
+ β”‚ β”‚
440
+ └──────┐ β”‚
441
+ β”‚ β”‚
442
+ β–Ό β–Ό
443
+ Tool loop Final synthesis
444
+ β”‚
445
+ β–Ό
446
+ ML assessment
447
+ Modeling risks
448
+ Recommended setup
449
+ Evaluation findings
450
+ Model comparison
451
+ Pipeline code
452
+ Tool trace
453
+ ```
454
+
455
+
456
+ ### **Core Architectural Pattern:**
457
+
458
+ ```text
459
+ app.py
460
+ ↓
461
+ agent.py
462
+ ↓
463
+ MachineLearningAgent
464
+ ↓
465
+ Qwen3-Coder + HF tool calling
466
+ ↓
467
+ ml_engine.py
468
+ ↓
469
+ ML operations
470
+ ↓
471
+ tool observations
472
+ ↓
473
+ agent.py
474
+ ↓
475
+ final ML assessment + pipeline code
476
+ ```
477
+
478
+
479
+ ## **Agent Tools**
480
+
481
+ ### **1. `inspect_dataset`**
482
+
483
+ analyzes the loaded dataset and returns structural information used by the agent before modeling
484
+
485
+ typical outputs include:
486
+
487
+ * source
488
+
489
+ * row count
490
+
491
+ * column count
492
+
493
+ * column names
494
+
495
+ * feature types
496
+
497
+ * pandas dtypes
498
+
499
+ * null information
500
+
501
+ * unique-value information
502
+
503
+ * possible target columns
504
+
505
+ * example values
506
+
507
+ used when the agent needs to understand the structure of the dataset before recommending a modeling approach
508
+
509
+ ---
510
+
511
+ ### **2. `recommend_modeling_setup`**
512
+
513
+ evaluates the dataset and selected target to recommend a supervised-learning configuration
514
+
515
+ recommendation can include:
516
+
517
+ * target column
518
+
519
+ * inferred problem type
520
+
521
+ * candidate algorithms
522
+
523
+ * dataset characteristics relevant to modeling
524
+
525
+ * preprocessing expectations
526
+
527
+ * numerical feature information
528
+
529
+ * categorical feature information
530
+
531
+ goal: setup a reasonable baseline modeling configuration before training
532
+
533
+ ---
534
+
535
+ ### **3. `train_candidate_model`**
536
+
537
+ builds and trains the selected scikit-learn pipeline using the actual loaded dataset
538
+
539
+ training workflow:
540
+
541
+ ```text
542
+ dataset
543
+ ↓
544
+ feature / target split
545
+ ↓
546
+ train / holdout split
547
+ ↓
548
+ numeric preprocessing
549
+ ↓
550
+ categorical preprocessing
551
+ ↓
552
+ selected estimator
553
+ ↓
554
+ fit
555
+ ↓
556
+ holdout prediction
557
+ ↓
558
+ holdout evaluation
559
+ ```
560
+
561
+ the agent receives metrics produced by the trained model
562
+
563
+ the agent is instructed not to invent performance values
564
+
565
+ ---
566
+
567
+ ### **4. `compare_algorithms`**
568
+
569
+ trains compatible baseline algorithms and returns a comparison leaderboard based on real holdout results.
570
+
571
+ classification, candidate algorithms include:
572
+
573
+ * Logistic Regression
574
+
575
+ * Random Forest Classifier
576
+
577
+ * SGD Classifier
578
+
579
+ * Linear SVM Classifier
580
+
581
+ for regression, candidate algorithms include:
582
+
583
+ * Ridge Regression
584
+
585
+ * Random Forest Regressor
586
+
587
+ * SGD Regressor
588
+
589
+ this allows the agent to compare alternatives using actual evaluation results instead of guessing which model performs best.
590
+
591
+ ---
592
+
593
+ ### **5. `generate_pipeline_code`**
594
+
595
+ generates reusable scikit-learn pipeline code grounded in:
596
+
597
+ * loaded dataset
598
+
599
+ * target column
600
+
601
+ * problem type
602
+
603
+ * selected algorithm
604
+
605
+ * numeric features
606
+
607
+ * categorical features
608
+
609
+ generated code includes preprocessing and estimator construction so the workflow can be moved beyond the interactive application.
610
+
611
+ ## **App Fallback**
612
+
613
+ app also includes a non-agent fallback path.
614
+
615
+ ```text
616
+ HF_TOKEN configured?
617
+ β”‚
618
+ β”Œβ”€β”€β”€β”΄β”€β”€β”€β”
619
+ β”‚ β”‚
620
+ yes no
621
+ β”‚ β”‚
622
+ β–Ό β–Ό
623
+ Qwen agent fallback
624
+ tool loop β”‚
625
+ β”‚ β”œβ”€β”€ inspect dataset
626
+ β”‚ β”œβ”€β”€ infer modeling setup
627
+ β”‚ β”œβ”€β”€ train baseline model
628
+ β”‚ β”œβ”€β”€ calculate real metrics
629
+ β”‚ └── generate pipeline code
630
+ β”‚
631
+ └──────────┬──────────
632
+ β–Ό
633
+ UI output
634
+ ```
635
+
636
+ fallback can also be used if the Hugging Face model request fails.
637
+
638
+ this means the application can still:
639
+
640
+ * profile the dataset
641
+
642
+ * infer a modeling setup
643
+
644
+ * train the selected baseline
645
+
646
+ * calculate real holdout metrics
647
+
648
+ * generate scikit-learn pipeline code
649
+
650
+ without requiring the Qwen agent to successfully complete the tool-calling workflow.
651
+
652
+ ## **Dataset Support**
653
+
654
+ app can preload a public Hugging Face dataset or accept a user-uploaded dataset.
655
+
656
+ uploaded dataset files can include:
657
+
658
+ 1. CSV
659
+ 2. Parquet
660
+ 3. JSON
661
+ 4. JSONL
662
+ 5. XLSX
663
+ 6. XLS
664
+
665
+ the app profiles the dataset before exposing it to the ML workflow.
666
+
667
+ ### **Hugging Face Example Dataset**
668
+
669
+ default public example dataset:
670
+
671
+ ```text
672
+ scikit-learn/adult-census-income
673
+ ```
674
+
675
+ default prediction target:
676
+
677
+ ```text
678
+ income
679
+ ```
680
+
681
+ the sample dataset is intended for binary classification.
682
+
683
+ users can replace the preloaded example with their own supported dataset through the upload interface.
684
+
685
+
686
+ ## **Preprocessing Pipeline**
687
+
688
+ the scikit-learn workflow separates numerical and categorical processing.
689
+
690
+ ```text
691
+ numeric features
692
+ ↓
693
+ median imputation
694
+ ↓
695
+ scaling when appropriate
696
+
697
+
698
+ categorical features
699
+ ↓
700
+ most-frequent imputation
701
+ ↓
702
+ one-hot encoding
703
+
704
+
705
+ numeric + categorical features
706
+ ↓
707
+ ColumnTransformer
708
+ ↓
709
+ selected estimator
710
+ ↓
711
+ trained Pipeline
712
+ ```
713
+
714
+ this keeps preprocessing and model inference inside the same reusable scikit-learn pipeline.
715
+
716
+
717
+ ## **Model Lab**
718
+
719
+ ### **Classification Algorithms**
720
+
721
+ * Logistic Regression
722
+
723
+ * Random Forest Classifier
724
+
725
+ * SGD Classifier
726
+
727
+ * Linear SVM Classifier
728
+
729
+ ### **Regression Algorithms**
730
+
731
+ * Ridge Regression
732
+
733
+ * Random Forest Regressor
734
+
735
+ * SGD Regressor
736
+
737
+ the selected algorithm is trained locally through automated scikit-learn pipelines.
738
+
739
+
740
+ ## **Evaluation**
741
+
742
+ ### **Classification Metrics**
743
+
744
+ classification evaluation can include:
745
+
746
+ * [accuracy](https://www.geeksforgeeks.org/machine-learning/metrics-for-machine-learning-model/)
747
+
748
+ * [weighted precision](https://www.geeksforgeeks.org/machine-learning/metrics-for-machine-learning-model/)
749
+
750
+ * [weighted recall](https://www.geeksforgeeks.org/machine-learning/metrics-for-machine-learning-model/)
751
+
752
+ * [weighted F1](https://www.geeksforgeeks.org/machine-learning/metrics-for-machine-learning-model/)
753
+
754
+ * [ROC-AUC](https://www.geeksforgeeks.org/machine-learning/metrics-for-machine-learning-model/)
755
+
756
+ ### **Regression Metrics**
757
+
758
+ regression evaluation includes:
759
+
760
+ * [**MAE**](https://www.datacamp.com/tutorial/mean-absolute-error)
761
+ * ml regression evaluation metric used to measure the average magnitude of prediction errors in machine learning regression models
762
+
763
+ * [**RMSE**](https://c3.ai/resources/glossary/data-science/root-mean-square-error-rmse)
764
+ * ml regression evaluation metric used to measure the average magnitude of error in regression machine learning models
765
+
766
+ * [**RΒ²**](https://www.geeksforgeeks.org/machine-learning/ml-r-squared-in-regression-analysis/)
767
+ * ml regression evaluation metric that measures the goodness of fit for a regression model by showing the proportion of variance in the dependent target variable that is explained by the independent features
768
+
769
+ **note:** evaluation values are produced by the local scikit-learn execution layer rather than estimated by the LLM.
770
+
771
+ ## **Agent Outputs**
772
+
773
+ final agent response is designed to contain:
774
+
775
+ 1. ML assessment
776
+ 2. Modeling risks
777
+ 3. Recommended modeling setup
778
+ 4. Evaluation findings
779
+ 5. Model comparison
780
+ 6. Pipeline recommendation
781
+ 7. Generated scikit-learn code
782
+ 8. Tool trace
783
+
784
+ app separately exposes structured modeling outputs such as:
785
+
786
+ * metrics
787
+
788
+ * leaderboard
789
+
790
+ * feature importance / coefficients
791
+
792
+ * pipeline code
793
+
794
+ * model artifact
795
+
796
+ * tool trace
797
+
798
+ * runtime diagnostics
799
+
800
+
801
+ ## **Agent Metrics / UI**
802
+
803
+ ### **1. Dataset Workspace**
804
+
805
+ dataset workspace shows:
806
+
807
+ **Dataset source**
808
+
809
+ * Hugging Face example-dataset loading
810
+
811
+ * local file upload
812
+
813
+ **Data preview**
814
+
815
+ * preview of the currently loaded dataset
816
+
817
+ **Inferred schema / feature information**
818
+
819
+ * column
820
+
821
+ * pandas dtype
822
+
823
+ * numeric / categorical feature information
824
+
825
+ * null information
826
+
827
+ * unique-value information
828
+
829
+ * example values
830
+
831
+ * possible target columns
832
+
833
+ ---
834
+
835
+ ### **2. Model Lab**
836
+
837
+ model workspace allows the user to configure:
838
+
839
+ **Target column**
840
+
841
+ * selected prediction target
842
+
843
+ **Problem type**
844
+
845
+ * classification
846
+
847
+ * regression
848
+
849
+ **Algorithmic Selection:**
850
+
851
+ classification:
852
+
853
+ * [**Logistic Regression**](https://www.geeksforgeeks.org/machine-learning/understanding-logistic-regression/)
854
+ * classification based supervised ML algorithm used to predict the probability of a categorical target variable
855
+
856
+ * [**Random Forest Classifier**](https://www.geeksforgeeks.org/random-forest-classifier-using-scikit-learn/)
857
+ * ensemble ML algorithm that builds multiple decision trees and combines their predictions to improve accuracy and reduce overfitting
858
+
859
+ * [**SGD Classifier**](https://scikit-learn.org/stable/modules/sgd.html)
860
+ * efficient ML estimator that optimizes regularized linear models using a first-order optimization routine. Instead of a distinct machine learning algorithm itself, it represents an optimization methodology used to train traditional linear models like Support Vector Machines (SVM) or Logistic Regression
861
+
862
+ * [**Linear SVM Classifier**](https://www.geeksforgeeks.org/machine-learning/support-vector-machine-algorithm/)
863
+ * supervised ML algorithm used to sort data into discrete categories by drawing a straight decision boundary
864
+
865
+ regression:
866
+
867
+ * [**Ridge Regression**](https://www.geeksforgeeks.org/machine-learning/what-is-ridge-regression/)
868
+ * L2 regularization, is a modified version of linear regression designed to improve a model's stability and prevent overfitting.
869
+ * standard linear regression only focuses on minimizing the difference between predicted and actual values, it often becomes unstable when dealing with highly correlated variables (multicollinearity) or a large number of features. **Ridge regression fixes this by penalizing the model for having excessively large coefficients, forcing them to shrink toward zero.**
870
+
871
+ * [**Random Forest Regressor**](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)
872
+ * ensemble machine learning algorithm used to predict continuous numerical values (e.g. housing prices, temperatures, or stock trends). algorithm works by building multiple independent decision trees at training time and outputting the average prediction of all the individual trees. this method, known as Bootstrap Aggregating (Bagging), reduces overfitting and improves model accuracy
873
+
874
+ * [**SGD Regressor**](https://www.geeksforgeeks.org/python/stochastic-gradient-descent-regressor/):
875
+ * linear regression model trained using Stochastic Gradient Descent (SGD), an iterative optimization technique that updates model parameters one sample at a time rather than using the entire dataset at once
876
+ * how SGD Regressor Works
877
+ * Sample-by-Sample Updates: estimates the gradient of the loss function using a single randomly chosen data point (or a small mini-batch) per step, then adjusts the weights immediately.
878
+ * Loss Functions: supports ordinary least squares (squared_error), robust regression (huber), and linear support vector regression (epsilon_insensitive).
879
+ * Regularization: includes penalties like L2 (ridge), L1 (lasso), or a mix of both (elasticnet) to shrink weights and avoid overfitting.
880
+ * Key Advantages:
881
+ * Speed with Big Data: trains much faster than traditional batch gradient descent when dealing with massive datasets containing over 10,000 samples.
882
+ * Online Learning: supports partial_fit, meaning it can continuously learn from real-time data streams as new information arrives.
883
+ * Escaping Local Minima: inherent noise in single-sample updates helps the algorithm jump out of shallow local minimums in complex cost surfaces.
884
+
885
+ app can then expose:
886
+
887
+ * fitted model metrics
888
+
889
+ * feature importance or coefficients
890
+
891
+ * generated pipeline artifact
892
+
893
+ * generated pipeline code
894
+
895
+ ---
896
+
897
+ ### **3. Model Comparison**
898
+
899
+ comparison workflow trains compatible baseline estimators and returns a leaderboard using real holdout results.
900
+
901
+ this provides a comparison layer that the LLM can reference during final synthesis.
902
+
903
+ the comparison values are derived from actual model execution instead of LLM-generated estimates.
904
+
905
+ ---
906
+
907
+ ### **4. Machine Learning Agent**
908
+
909
+ agent workspace accepts a natural-language ML engineering task.
910
+
911
+ Qwen agent can then:
912
+
913
+ * inspect the dataset
914
+
915
+ * recommend a modeling setup
916
+
917
+ * train a candidate model
918
+
919
+ * compare algorithms
920
+
921
+ * generate pipeline code
922
+
923
+ * synthesize findings from tool observations
924
+
925
+ ---
926
+
927
+ ### **5. Tool Trace**
928
+
929
+ every agent tool invocation is appended to the tool trace.
930
+
931
+ Example trace:
932
+
933
+ ```json
934
+ [
935
+ {
936
+ "tool": "inspect_dataset",
937
+ "arguments": {},
938
+ "result": {
939
+ "rows": 32561,
940
+ "columns": 15
941
+ }
942
+ },
943
+ {
944
+ "tool": "train_candidate_model",
945
+ "arguments": {
946
+ "target": "income",
947
+ "algorithm": "Logistic Regression"
948
+ },
949
+ "result": {
950
+ "metrics": "holdout metrics"
951
+ }
952
+ }
953
+ ]
954
+ ```
955
+
956
+ trace allows the user to analyze / inspect:
957
+
958
+ * which tools were called
959
+
960
+ * arguments passed to each tool
961
+
962
+ * tool observations
963
+
964
+ * number of tool calls
965
+
966
+ * modeling operations performed during the agent loop
967
+
968
+ ---
969
+
970
+ ### **6. Diagnostics**
971
+
972
+ runtime diagnostics can include:
973
+
974
+ * model
975
+
976
+ * provider
977
+
978
+ * HF token configured
979
+
980
+ * tool-call count
981
+
982
+ * tools used
983
+
984
+ * selected algorithm
985
+
986
+ * problem type
987
+
988
+ * target
989
+
990
+ Example:
991
+
992
+ ```json
993
+ {
994
+ "model": "Qwen/Qwen3-Coder-30B-A3B-Instruct",
995
+ "provider": "HF automatic routing",
996
+ "hf_token_configured": true,
997
+ "tool_calls": 4,
998
+ "tools_used": [
999
+ "inspect_dataset",
1000
+ "recommend_modeling_setup",
1001
+ "train_candidate_model",
1002
+ "compare_algorithms"
1003
+ ],
1004
+ "target": "income",
1005
+ "problem_type": "classification",
1006
+ "algorithm": "Logistic Regression"
1007
+ }
1008
+ ```
1009
+
1010
+ the provider can appear as:
1011
+
1012
+ ```text
1013
+ HF automatic routing
1014
+ ```
1015
+
1016
+ when a specific inference provider is not explicitly configured.
1017
+
1018
+
1019
+ ## **Local ML Execution Model**
1020
+
1021
+ model training is performed directly inside the application with scikit-learn.
1022
+
1023
+ ```text
1024
+ Loaded dataset
1025
+ ↓
1026
+ pandas DataFrame
1027
+ ↓
1028
+ feature / target split
1029
+ ↓
1030
+ scikit-learn preprocessing
1031
+ ↓
1032
+ scikit-learn estimator
1033
+ ↓
1034
+ local CPU training
1035
+ ↓
1036
+ holdout predictions
1037
+ ↓
1038
+ evaluation metrics
1039
+ ↓
1040
+ Joblib artifact
1041
+ ```
1042
+
1043
+ Qwen3-Coder does **not** perform the numerical model fitting itself.
1044
+
1045
+ ```text
1046
+ Qwen3-Coder
1047
+ ↓
1048
+ reasoning + tool selection + synthesis
1049
+
1050
+ scikit-learn
1051
+ ↓
1052
+ preprocessing + training + prediction + evaluation
1053
+ ```
1054
+
1055
+ this keeps model-performance claims grounded in real execution.
1056
+
1057
+ ## **Machine Learning Execution Separation**
1058
+
1059
+ ### Qwen3-Coder
1060
+
1061
+ responsible for:
1062
+
1063
+ * understanding the user's ML task
1064
+
1065
+ * selecting appropriate tools
1066
+
1067
+ * interpreting dataset observations
1068
+
1069
+ * reasoning about modeling risks
1070
+
1071
+ * deciding whether additional tools are required
1072
+
1073
+ * synthesizing the final response
1074
+
1075
+ ### scikit-learn
1076
+
1077
+ responsible for:
1078
+
1079
+ * preprocessing
1080
+
1081
+ * train / holdout splitting
1082
+
1083
+ * fitting estimators
1084
+
1085
+ * making predictions
1086
+
1087
+ * calculating evaluation metrics
1088
+
1089
+ * model comparison
1090
+
1091
+ * fitted-pipeline generation
1092
+
1093
+ conceptually:
1094
+
1095
+ ```text
1096
+ User task
1097
+ ↓
1098
+ Qwen3-Coder
1099
+ ↓
1100
+ Select ML operation
1101
+ ↓
1102
+ scikit-learn tool
1103
+ ↓
1104
+ Real execution result
1105
+ ↓
1106
+ Qwen3-Coder
1107
+ ↓
1108
+ Interpret result / select next tool
1109
+ ↓
1110
+ Final ML assessment
1111
+ ```
1112
+
1113
+ ## **File Descriptions**
1114
+
1115
+ ### **`agent.py`**
1116
+
1117
+ main agent implementation.
1118
+
1119
+ contains the logic for:
1120
+
1121
+ * system prompt
1122
+
1123
+ * tool schemas
1124
+
1125
+ * `MachineLearningAgent`
1126
+
1127
+ * tool dispatch
1128
+
1129
+ * fallback
1130
+
1131
+ * bounded agent loop
1132
+
1133
+ * `HF InferenceClient` requests
1134
+
1135
+ * `tool_choice="auto"`
1136
+
1137
+ * tool observations
1138
+
1139
+ * final response synthesis
1140
+
1141
+ architecturally:
1142
+
1143
+ ```text
1144
+ agent.py
1145
+ ↓
1146
+ decides WHAT ML action is needed
1147
+ ↓
1148
+ ml_engine.py
1149
+ ↓
1150
+ executes the ML operation
1151
+ ```
1152
+
1153
+ ---
1154
+
1155
+ ### **`ml_engine.py`**
1156
+
1157
+ machine learning execution layer.
1158
+
1159
+ contains:
1160
+
1161
+ * dataset loading
1162
+
1163
+ * dataset profiling
1164
+
1165
+ * target inference
1166
+
1167
+ * problem-type inference
1168
+
1169
+ * feature separation
1170
+
1171
+ * preprocessing construction
1172
+
1173
+ * algorithm construction
1174
+
1175
+ * train / holdout splitting
1176
+
1177
+ * model fitting
1178
+
1179
+ * prediction
1180
+
1181
+ * evaluation
1182
+
1183
+ * model comparison
1184
+
1185
+ * feature importance / coefficients
1186
+
1187
+ * pipeline-code generation
1188
+
1189
+ * Joblib artifact generation
1190
+
1191
+ architecturally:
1192
+
1193
+ ```text
1194
+ agent.py
1195
+ ↓
1196
+ decides WHAT ML action is needed
1197
+ ↓
1198
+ ml_engine.py
1199
+ ↓
1200
+ executes preprocessing / training / evaluation
1201
+ ↓
1202
+ returns structured observation
1203
+ ```
1204
+ ---
1205
+
1206
+ ### **`app.py`**
1207
+
1208
+ responsibilities include:
1209
+
1210
+ * session state
1211
+
1212
+ * Hugging Face sample loading
1213
+
1214
+ * local dataset upload
1215
+
1216
+ * dataset preview
1217
+
1218
+ * target controls
1219
+
1220
+ * problem-type controls
1221
+
1222
+ * algorithm controls
1223
+
1224
+ * training callbacks
1225
+
1226
+ * comparison callbacks
1227
+
1228
+ * generated-code display
1229
+
1230
+ * downloadable model artifact
1231
+
1232
+ * agent execution
1233
+
1234
+ * tool-trace display
1235
+
1236
+ * runtime diagnostics
1237
+
1238
+ ---
1239
+
1240
+ ### **`config.py`**
1241
+
1242
+ agent and application runtime configuration.
1243
+
1244
+ controls values such as:
1245
+
1246
+ ```text
1247
+ HF_MODEL_ID
1248
+
1249
+ HF_PROVIDER
1250
+
1251
+ HF_TOKEN
1252
+
1253
+ MAX_UPLOAD_MB
1254
+
1255
+ MAX_PROFILE_ROWS
1256
+
1257
+ MAX_TRAIN_ROWS
1258
+
1259
+ MAX_AGENT_STEPS
1260
+
1261
+ DEFAULT_MAX_TOKENS
1262
+
1263
+ RANDOM_STATE
1264
+ ```
1265
+
1266
+ ## **Modeling Guardrails**
1267
+
1268
+ application separates LLM reasoning from numerical machine learning execution.
1269
+
1270
+ Qwen agent can decide which operation should be performed, but:
1271
+
1272
+ ```text
1273
+ training
1274
+
1275
+ prediction
1276
+
1277
+ evaluation
1278
+
1279
+ metric calculation
1280
+
1281
+ model comparison
1282
+ ```
1283
+
1284
+ are handled by scikit-learn execution tools.
1285
+
1286
+ this architecture helps prevent fabricated model-performance claims.
1287
+
1288
+ additional controls include:
1289
+
1290
+ * bounded agent steps
1291
+
1292
+ * configurable maximum model tokens
1293
+
1294
+ * fixed random state by default
1295
+
1296
+ * controlled algorithm list
1297
+
1298
+ * maximum training-row limit
1299
+
1300
+ * deterministic / structured preprocessing behavior
1301
+
1302
+ * fallback execution path
1303
+
1304
+ * tool trace
1305
+
1306
+ * structured model observations
1307
+
1308
+ ## **Model Performance Grounding**
1309
+
1310
+ the system is designed so the LLM does not invent model metrics.
1311
+
1312
+ ```text
1313
+ Dataset
1314
+ ↓
1315
+ scikit-learn model training
1316
+ ↓
1317
+ Prediction
1318
+ ↓
1319
+ Metric calculation
1320
+ ↓
1321
+ Structured tool observation
1322
+ ↓
1323
+ Qwen3-Coder
1324
+ ↓
1325
+ Final interpretation
1326
+ ```
1327
+
1328
+ performance values should only be presented as real model results when the associated training / evaluation tool successfully executes.
1329
+
1330
+ if the tool does not run successfully, the final response should not claim that a model was trained or evaluated.
1331
+
1332
+
1333
+ ## **Model Artifact**
1334
+
1335
+ trained preprocessing + estimator pipelines can be serialized through Joblib.
1336
+
1337
+ ```text
1338
+ Numeric preprocessing
1339
+ +
1340
+ Categorical preprocessing
1341
+ +
1342
+ Estimator
1343
+ ↓
1344
+ scikit-learn Pipeline
1345
+ ↓
1346
+ fit
1347
+ ↓
1348
+ Joblib serialization
1349
+ ↓
1350
+ .joblib artifact
1351
+ ```
1352
+
1353
+ this allows the fitted pipeline to be reused later without manually rebuilding preprocessing and estimator logic.
1354
+
1355
+
1356
+ ## **Agent Pattern Workflow:**
1357
+
1358
+ ```text
1359
+ Single Agent
1360
+ +
1361
+ Function Calling
1362
+ +
1363
+ Machine Learning Tools
1364
+ +
1365
+ Bounded Tool Loop
1366
+ +
1367
+ Tool Observations
1368
+ +
1369
+ Final LLM Synthesis
1370
+ ```
1371
+
1372
+ full execution:
1373
+
1374
+ ```text
1375
+ USER
1376
+ ↓
1377
+ Gradio / app.py
1378
+ ↓
1379
+ MachineLearningAgent / agent.py
1380
+ ↓
1381
+ Qwen3-Coder
1382
+ ↓
1383
+ tool_choice="auto"
1384
+ ↓
1385
+ ML Tool
1386
+ ↓
1387
+ ml_engine.py
1388
+ ↓
1389
+ Observation
1390
+ ↓
1391
+ Qwen3-Coder
1392
+ ↓
1393
+ repeat up to MAX_AGENT_STEPS
1394
+ ↓
1395
+ Final ML Assessment
1396
+ ↓
1397
+ Recommended modeling setup
1398
+ ↓
1399
+ Evaluation findings
1400
+ ↓
1401
+ Model comparison
1402
+ ↓
1403
+ Pipeline code
1404
+ ↓
1405
+ Tool Trace
1406
+ ```
1407
+
1408
+
1409
+ # **Core Runtime Components**
1410
+
1411
+ ## [**scikit-learn**](https://scikit-learn.org/stable/)
1412
+
1413
+ ### **What it is:**
1414
+
1415
+ scikit-learn ml framework used to build preprocessing pipelines, estimators, training workflows, predictions, and evaluation metrics.
1416
+
1417
+ ### **Why it is used in this project:**
1418
+
1419
+ the application needs a real ML execution layer so model-performance values come from fitted models rather than LLM estimates.
1420
+
1421
+ ### **How it is used here:**
1422
+
1423
+ ```text
1424
+ Dataset
1425
+ ↓
1426
+ Preprocessing
1427
+ ↓
1428
+ Estimator
1429
+ ↓
1430
+ scikit-learn Pipeline
1431
+ ↓
1432
+ Training
1433
+ ↓
1434
+ Holdout prediction
1435
+ ↓
1436
+ Holdout evaluation
1437
+ ↓
1438
+ Metrics + fitted artifact
1439
+ ```
1440
+
1441
+ scikit-learn is used to:
1442
+
1443
+ * build numeric preprocessing
1444
+
1445
+ * build categorical preprocessing
1446
+
1447
+ * combine transformations through `ColumnTransformer`
1448
+
1449
+ * build estimator pipelines
1450
+
1451
+ * split datasets into train and holdout sets
1452
+
1453
+ * fit models
1454
+
1455
+ * generate predictions
1456
+
1457
+ * calculate classification metrics
1458
+
1459
+ * calculate regression metrics
1460
+
1461
+ * compare candidate estimators
1462
+
1463
+ * expose supported feature importance / coefficient values
1464
+
1465
+ * produce reusable fitted model pipelines
1466
+
1467
+
1468
+ ## [**pandas**](https://pandas.pydata.org/docs/)
1469
+
1470
+ ### **Why it is used in this project:**
1471
+
1472
+ the application needs a lightweight in-memory tabular representation before datasets can be profiled and passed into scikit-learn.
1473
+
1474
+ ### **How it is used here:**
1475
+
1476
+ * load supported dataset formats
1477
+
1478
+ * hold the active dataset in memory
1479
+
1480
+ * inspect columns and pandas data types
1481
+
1482
+ * determine numerical and categorical features
1483
+
1484
+ * analyze null values
1485
+
1486
+ * inspect unique values
1487
+
1488
+ * separate prediction target from model features
1489
+
1490
+ * provide training data to scikit-learn
1491
+
1492
+ workflow:
1493
+
1494
+ ```text
1495
+ Uploaded / HF dataset
1496
+ ↓
1497
+ pandas DataFrame
1498
+ ↓
1499
+ Dataset profiling
1500
+ ↓
1501
+ Feature + target identification
1502
+ ↓
1503
+ scikit-learn
1504
+ ```
1505
+
1506
+ ## **Joblib**
1507
+
1508
+ ### **What it is:**
1509
+
1510
+ Joblib is used to serialize fitted Python machine learning objects
1511
+
1512
+ ### **Why it is used in this project:**
1513
+
1514
+ a trained model is more useful when the complete preprocessing + estimator pipeline can be reused outside the interactive session.
1515
+
1516
+ ### **How it is used here:**
1517
+
1518
+ ```text
1519
+ Fitted preprocessing
1520
+ +
1521
+ Fitted estimator
1522
+ ↓
1523
+ scikit-learn Pipeline
1524
+ ↓
1525
+ Joblib
1526
+ ↓
1527
+ .joblib model artifact
1528
+ ```
1529
+
1530
+ after training, the fitted preprocessing + estimator pipeline can be exported as a `.joblib` artifact for later reuse.
1531
+
1532
+ ## **PyArrow**
1533
+
1534
+ ### **What it is:**
1535
+
1536
+ PyArrow provides support for Apache Arrow / Parquet-based data interchange.
1537
+
1538
+ ### **How it is used here:**
1539
+
1540
+ PyArrow supports loading Parquet datasets into the application's pandas-based dataset workflow.
1541
+
1542
+ ```text
1543
+ Parquet file
1544
+ ↓
1545
+ PyArrow support
1546
+ ↓
1547
+ pandas DataFrame
1548
+ ↓
1549
+ ML workflow
1550
+ ```
1551
+
1552
+
1553
+ ## **OpenPyXL**
1554
+
1555
+ ### **What it is:**
1556
+
1557
+ OpenPyXL provides Python support for reading Excel workbook formats used by the dataset upload workflow.
1558
+
1559
+ ### **How it is used here:**
1560
+
1561
+ ```text
1562
+ XLSX dataset
1563
+ ↓
1564
+ OpenPyXL
1565
+ ↓
1566
+ pandas
1567
+ ↓
1568
+ Dataset profile
1569
+ ↓
1570
+ ML workflow
1571
+ ```
1572
+
1573
+ ## **HF Automatic Routing**
1574
+
1575
+ ### **What it is:**
1576
+
1577
+ Hugging Face automatic routing allows `InferenceClient` to route the Qwen model request through an available Hugging Face Inference Provider when a specific provider is not explicitly configured.
1578
+
1579
+ ### **Why it is used:**
1580
+
1581
+ it avoids coupling the application to one inference backend and simplifies deployment of the agent.
1582
+
1583
+ ### **How it is used here:**
1584
+
1585
+ ```text
1586
+ MachineLearningAgent
1587
+ ↓
1588
+ huggingface_hub.InferenceClient
1589
+ ↓
1590
+ provider="auto"
1591
+ ↓
1592
+ Hugging Face Inference Provider
1593
+ ↓
1594
+ Qwen3-Coder
1595
+ ↓
1596
+ tool request or final synthesis
1597
+ ```
1598
+
1599
+ model configuration is controlled through:
1600
+
1601
+ ```text
1602
+ HF_MODEL_ID
1603
+
1604
+ HF_PROVIDER
1605
+
1606
+ HF_TOKEN
1607
+ ```
1608
+
1609
+ within this project:
1610
+
1611
+ * `HF_MODEL_ID` identifies the Qwen model
1612
+
1613
+ * `HF_PROVIDER` can configure provider behavior
1614
+
1615
+ * automatic routing can be used when a specific backend is not forced
1616
+
1617
+ * the routed model performs tool selection, reasoning, and final synthesis
1618
+
1619
+ * scikit-learn performs the actual numerical ML operations
1620
+
1621
+
1622
+ ## **Training / Evaluation Architecture**
1623
+
1624
+ ```text
1625
+ Dataset
1626
+ ↓
1627
+ pandas
1628
+ ↓
1629
+ Feature / target split
1630
+ ↓
1631
+ Train / holdout split
1632
+ ↓
1633
+ ColumnTransformer
1634
+ ↓
1635
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
1636
+ β”‚ Numeric preprocessing β”‚
1637
+ β”‚ β”‚
1638
+ β”‚ median imputation β”‚
1639
+ β”‚ optional scaling β”‚
1640
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
1641
+ +
1642
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
1643
+ β”‚ Categorical preprocessing β”‚
1644
+ β”‚ β”‚
1645
+ β”‚ most-frequent imputation β”‚
1646
+ β”‚ one-hot encoding β”‚
1647
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
1648
+ ↓
1649
+ Estimator
1650
+ ↓
1651
+ scikit-learn Pipeline
1652
+ ↓
1653
+ fit
1654
+ ↓
1655
+ holdout prediction
1656
+ ↓
1657
+ evaluation metrics
1658
+ ↓
1659
+ Agent observation
1660
+ ```
1661
+
1662
+
1663
+ ## **Classification Workflow**
1664
+
1665
+ ```text
1666
+ Classification Dataset
1667
+ ↓
1668
+ Target Selection
1669
+ ↓
1670
+ Preprocessing
1671
+ ↓
1672
+ Candidate classifier
1673
+ ↓
1674
+ Training
1675
+ ↓
1676
+ Holdout predictions
1677
+ ↓
1678
+ accuracy
1679
+ weighted precision
1680
+ weighted recall
1681
+ weighted F1
1682
+ ROC-AUC where supported
1683
+ ↓
1684
+ Agent interpretation
1685
+ ```
1686
+
1687
+ supported baseline classifiers include:
1688
+
1689
+ * Logistic Regression
1690
+
1691
+ * Random Forest Classifier
1692
+
1693
+ * SGD Classifier
1694
+
1695
+ * Linear SVM Classifier
1696
+
1697
+ ## **Regression Workflow**
1698
+
1699
+ ```text
1700
+ Regression dataset
1701
+ ↓
1702
+ Target selection
1703
+ ↓
1704
+ Preprocessing
1705
+ ↓
1706
+ Candidate regressor
1707
+ ↓
1708
+ Training
1709
+ ↓
1710
+ Holdout predictions
1711
+ ↓
1712
+ MAE
1713
+ RMSE
1714
+ RΒ²
1715
+ ↓
1716
+ Agent interpretation
1717
+ ```
1718
+
1719
+ supported baseline regressors include:
1720
+
1721
+ * Ridge Regression
1722
+
1723
+ * Random Forest Regressor
1724
+
1725
+ * SGD Regressor
1726
+