# Phase 0 Command Contract This document captures baseline command contracts without changing runtime behavior. ## Contract Scope Commands covered: 1. run-lang-task 2. run-lang-benchmark 3. eval-v2 4. view-task-v2 5. eval-comprehensive 6. tag-failures 7. view-exec-report ## run-lang-task - Command: `dabench run-lang-task --config [--display-mode technical|executive]` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `task_ids` required positional list. 2. `--config` required file path. 3. `--display-mode` default `technical`. - Backend call path: 1. `run_lang_task_command` -> `create_run_output_dir`. 2. Loop task IDs in provided order. 3. Per task call `run_langgraph_single_task`. 4. Optional executive rendering via `generate_executive_report` when display mode is `executive` and `trace.json` exists. - Artifacts produced: 1. Shared run directory under configured output dir. 2. Per task: `trace.json`, `prediction.csv` (or fallback minimal CSV), task status fields in trace. - Error and exit semantics: 1. Invalid/duplicate run id from output dir creation surfaces as Typer bad parameter. 2. Per-task failures are represented in artifacts; command still summarizes all selected tasks. - Compatibility constraints: 1. Tasks share a single run directory. 2. Task order is preserved. 3. Must continue to support both technical and executive modes. ## run-lang-benchmark - Command: `dabench run-lang-benchmark --config [--limit N]` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `--config` required file path. 2. `--limit` optional positive int. - Backend call path: 1. `run_lang_benchmark_command` -> dataset task count for progress setup. 2. Delegates to `run_langgraph_benchmark(config, limit, progress_callback)`. - Artifacts produced: 1. Run directory with per-task task folders. 2. Run-level `summary.json` from langgraph runner. - Error and exit semantics: 1. Invalid run id/max workers bubbles as command error. 2. Task failures do not abort benchmark loop by default. - Compatibility constraints: 1. Uses existing benchmark runner implementation. 2. Keeps progress and summary behavior. ## eval-v2 - Command: `dabench eval-v2 [--gold-root path] [--task-root path] [--lambda-penalty float] [--mode standard|verbose|research]` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `run_id` required. 2. `--gold-root` default `/data3/dataFAIR/kdd-dev/public/output`. 3. `--task-root` optional, auto-detected from input_full/input_med/input_hard when omitted. 4. `--lambda-penalty` default `0.1`. 5. `--mode` default `standard`. - Backend call path: 1. Resolve run path. 2. `evaluate_run_v2`. 3. `write_evaluation_v2` and `write_auditor_validation_report`. 4. Hardening: `validate_evaluation_run`, `generate_health_report`, `generate_replay_artifacts`. 5. Consistency checks: `validate_evaluation_consistency`, `write_validation_report`. 6. Render report via `render_evaluation_report`. - Artifacts produced at run root: 1. `task_metrics.csv` 2. `trajectory.csv` 3. `tool_calls.csv` 4. `comprehensive_evaluation.csv` 5. `auditor_validation_report.md` 6. `validation_report.md` 7. `artifact_reconciliation_report.txt` 8. `engineering_health_report.txt` 9. Per-task `task_replay.json` - Error and exit semantics: 1. Missing run path exits non-zero. 2. Validator warnings do not fail command. 3. Validator errors fail command with non-zero exit. - Compatibility constraints: 1. Keep run-root outputs for compatibility. 2. Keep hardening pipeline invocation order and report generation. ## view-task-v2 - Command: `dabench view-task-v2 ` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `task_id` required. 2. `run_id` required. - Backend call path: 1. Resolve run path and task directory. 2. Read `task_metrics.csv` (required). 3. Read optional `trajectory.csv` and `tool_calls.csv`. 4. Render via `render_verbose_task_detail`. - Artifacts read: 1. Run-root metrics CSVs. 2. Task directory existence check. - Error and exit semantics: 1. Missing task directory exits non-zero. 2. Missing `task_metrics.csv` exits non-zero with guidance to run eval-v2. 3. Missing row for task exits non-zero. - Compatibility constraints: 1. Run-root layout remains required. 2. Optional trajectory/tool calls remain best-effort. ## eval-comprehensive - Command: `dabench eval-comprehensive [--gold-root path] [--lambda-penalty float] [--output-csv path]` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `run_id` required. 2. `--gold-root` default `/data3/dataFAIR/kdd-dev/public/output`. 3. `--lambda-penalty` default `0.1`. 4. `--output-csv` optional; defaults to `/comprehensive_evaluation.csv`. - Backend call path: 1. Resolve run path. 2. Call `evaluate_run_comprehensive`. 3. Write CSV (drop `action_counts` if present). - Artifacts produced: 1. `comprehensive_evaluation.csv` or custom output path. - Error and exit semantics: 1. Missing run path exits non-zero. 2. Empty evaluated task set exits non-zero. - Compatibility constraints: 1. Preserve current CSV columns and summary rendering behavior. ## tag-failures - Command: `dabench tag-failures [--gold-root path] [--eval-csv path]` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `run_id` required. 2. `--gold-root` default `/data3/dataFAIR/kdd-dev/public/output`. 3. `--eval-csv` optional; defaults to `/evaluation.csv` if present. - Backend call path: 1. Resolve run path. 2. Call `tag_run`. 3. Print bucket and phase summaries. 4. Write `failure_tags.csv`. - Artifacts produced: 1. `failure_tags.csv` at run root. - Error and exit semantics: 1. Missing run path exits non-zero. 2. No tasks found exits non-zero. - Compatibility constraints: 1. Preserve existing bucket taxonomy output contract. ## view-exec-report - Command: `dabench view-exec-report ` - Definition: `src/data_agent_baseline/cli.py` - Arguments and defaults: 1. `task_id` required. 2. `run_id` required. - Backend call path: 1. Resolve run path. 2. Resolve task directory. 3. Require `trace.json`. 4. Call `generate_executive_report(trace_path, task_dir, console)`. - Artifacts read: 1. Per-task `trace.json`. - Error and exit semantics: 1. Missing run, task, or trace exits non-zero. - Compatibility constraints: 1. Preserve report generation from existing trace artifact. ## Eval-v2 Artifact Schema — Current Production Names This section documents the **current actual column names** as exported by `eval-v2` at Phase 0 baseline. The schema is frozen for Phase 0. Schema renaming or migration is **deferred and is not part of Phase 0**. ### Source constants Canonical column lists are defined in `src/data_agent_baseline/langgraph_agent/eval_v2.py` as: - `TASK_METRICS_COLUMNS` - `TRAJECTORY_COLUMNS` - `TOOL_CALLS_COLUMNS` ### task_metrics.csv / comprehensive_evaluation.csv `comprehensive_evaluation.csv` is written by `write_evaluation_v2` as an alias for `task_metrics.csv` (same data, same columns). Key columns at Phase 0: - Identity: `run_id`, `task_id`, `trace_id`, `difficulty`, `timestamp`, `task_type` - Outcome: **`execution_success`**, **`final_score`**, `failure_reason`, `failure_category`, `outcome_status` - Answer accuracy: `answer_precision`, `answer_recall`, `answer_f1`, `column_precision`, `column_recall`, `column_f1`, `row_precision`, `row_recall`, `row_f1`, `recall` - Verification: `verification_triggered`, `verification_score`, `aat_verification_triggered`, `aat_verification_score` - Agent behaviour: `replan_count`, `retry_count`, `total_replans`, `autonomy_score`, `tool_efficiency`, `data_understanding_score` - Cost: `total_tokens`, `llm_calls`, `execution_time` - Full column list: see `TASK_METRICS_COLUMNS` in `eval_v2.py` (238 columns at Phase 0 baseline). The names `execution_success` and `final_score` are the current production names. **No renaming is performed in Phase 0.** ### trajectory.csv Columns from `TRAJECTORY_COLUMNS` (29 columns at Phase 0 baseline): `run_id`, `task_id`, `trace_id`, `step_id`, `parent_step_id`, `stage`, `legacy_stage`, `phase`, `agent`, `trajectory_agent`, `agent_type`, `tool`, `action`, `coordinator_checkpoint`, `confidence`, `decision`, `review_type`, `verification_status`, `specialist_selected`, `success`, `duration_seconds`, `tokens`, `retries`, `timestamp`, `elapsed_seconds`, `thought`, `action_input`, `observation`, `raw_response`, `metadata_json` ### tool_calls.csv Columns from `TOOL_CALLS_COLUMNS` (17 columns at Phase 0 baseline): `run_id`, `task_id`, `trace_id`, `step_id`, `call_id`, `tool_name`, `tool_category`, `success`, `latency_seconds`, `tokens`, `retry_count`, `input_size`, `output_size`, `error_type`, `error_message`, `timestamp`, `metadata_json` ### Deferred work - Column renaming (e.g., `execution_success` → `task_success`, `final_score` → `score`) is explicitly deferred. - No schema migration, column aliasing, or DataFrame reshaping is performed in Phase 0. - Tests reference only the current names.