DAO_kdd26 / docs /implementation /PHASE0_COMMAND_CONTRACT.md
sipe5001's picture
Add Hugging Face Docker Space configuration
d3d0e0e
|
Raw
History Blame Contribute Delete
9.18 kB

Phase 0 Command Contract

This document captures baseline command contracts without changing runtime behavior.

Contract Scope

Commands covered:

  1. run-lang-task
  2. run-lang-benchmark
  3. eval-v2
  4. view-task-v2
  5. eval-comprehensive
  6. tag-failures
  7. view-exec-report

run-lang-task

  • Command: dabench run-lang-task <task_ids...> --config <path> [--display-mode technical|executive]
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. task_ids required positional list.
  2. --config required file path.
  3. --display-mode default technical.
  • Backend call path:
  1. run_lang_task_command -> create_run_output_dir.
  2. Loop task IDs in provided order.
  3. Per task call run_langgraph_single_task.
  4. Optional executive rendering via generate_executive_report when display mode is executive and trace.json exists.
  • Artifacts produced:
  1. Shared run directory under configured output dir.
  2. Per task: trace.json, prediction.csv (or fallback minimal CSV), task status fields in trace.
  • Error and exit semantics:
  1. Invalid/duplicate run id from output dir creation surfaces as Typer bad parameter.
  2. Per-task failures are represented in artifacts; command still summarizes all selected tasks.
  • Compatibility constraints:
  1. Tasks share a single run directory.
  2. Task order is preserved.
  3. Must continue to support both technical and executive modes.

run-lang-benchmark

  • Command: dabench run-lang-benchmark --config <path> [--limit N]
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. --config required file path.
  2. --limit optional positive int.
  • Backend call path:
  1. run_lang_benchmark_command -> dataset task count for progress setup.
  2. Delegates to run_langgraph_benchmark(config, limit, progress_callback).
  • Artifacts produced:
  1. Run directory with per-task task folders.
  2. Run-level summary.json from langgraph runner.
  • Error and exit semantics:
  1. Invalid run id/max workers bubbles as command error.
  2. Task failures do not abort benchmark loop by default.
  • Compatibility constraints:
  1. Uses existing benchmark runner implementation.
  2. Keeps progress and summary behavior.

eval-v2

  • Command: dabench eval-v2 <run_id_or_path> [--gold-root path] [--task-root path] [--lambda-penalty float] [--mode standard|verbose|research]
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. run_id required.
  2. --gold-root default /data3/dataFAIR/kdd-dev/public/output.
  3. --task-root optional, auto-detected from input_full/input_med/input_hard when omitted.
  4. --lambda-penalty default 0.1.
  5. --mode default standard.
  • Backend call path:
  1. Resolve run path.
  2. evaluate_run_v2.
  3. write_evaluation_v2 and write_auditor_validation_report.
  4. Hardening: validate_evaluation_run, generate_health_report, generate_replay_artifacts.
  5. Consistency checks: validate_evaluation_consistency, write_validation_report.
  6. Render report via render_evaluation_report.
  • Artifacts produced at run root:
  1. task_metrics.csv
  2. trajectory.csv
  3. tool_calls.csv
  4. comprehensive_evaluation.csv
  5. auditor_validation_report.md
  6. validation_report.md
  7. artifact_reconciliation_report.txt
  8. engineering_health_report.txt
  9. Per-task task_replay.json
  • Error and exit semantics:
  1. Missing run path exits non-zero.
  2. Validator warnings do not fail command.
  3. Validator errors fail command with non-zero exit.
  • Compatibility constraints:
  1. Keep run-root outputs for compatibility.
  2. Keep hardening pipeline invocation order and report generation.

view-task-v2

  • Command: dabench view-task-v2 <task_id> <run_id_or_path>
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. task_id required.
  2. run_id required.
  • Backend call path:
  1. Resolve run path and task directory.
  2. Read task_metrics.csv (required).
  3. Read optional trajectory.csv and tool_calls.csv.
  4. Render via render_verbose_task_detail.
  • Artifacts read:
  1. Run-root metrics CSVs.
  2. Task directory existence check.
  • Error and exit semantics:
  1. Missing task directory exits non-zero.
  2. Missing task_metrics.csv exits non-zero with guidance to run eval-v2.
  3. Missing row for task exits non-zero.
  • Compatibility constraints:
  1. Run-root layout remains required.
  2. Optional trajectory/tool calls remain best-effort.

eval-comprehensive

  • Command: dabench eval-comprehensive <run_id_or_path> [--gold-root path] [--lambda-penalty float] [--output-csv path]
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. run_id required.
  2. --gold-root default /data3/dataFAIR/kdd-dev/public/output.
  3. --lambda-penalty default 0.1.
  4. --output-csv optional; defaults to <run>/comprehensive_evaluation.csv.
  • Backend call path:
  1. Resolve run path.
  2. Call evaluate_run_comprehensive.
  3. Write CSV (drop action_counts if present).
  • Artifacts produced:
  1. comprehensive_evaluation.csv or custom output path.
  • Error and exit semantics:
  1. Missing run path exits non-zero.
  2. Empty evaluated task set exits non-zero.
  • Compatibility constraints:
  1. Preserve current CSV columns and summary rendering behavior.

tag-failures

  • Command: dabench tag-failures <run_id_or_path> [--gold-root path] [--eval-csv path]
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. run_id required.
  2. --gold-root default /data3/dataFAIR/kdd-dev/public/output.
  3. --eval-csv optional; defaults to <run>/evaluation.csv if present.
  • Backend call path:
  1. Resolve run path.
  2. Call tag_run.
  3. Print bucket and phase summaries.
  4. Write failure_tags.csv.
  • Artifacts produced:
  1. failure_tags.csv at run root.
  • Error and exit semantics:
  1. Missing run path exits non-zero.
  2. No tasks found exits non-zero.
  • Compatibility constraints:
  1. Preserve existing bucket taxonomy output contract.

view-exec-report

  • Command: dabench view-exec-report <task_id> <run_id_or_path>
  • Definition: src/data_agent_baseline/cli.py
  • Arguments and defaults:
  1. task_id required.
  2. run_id required.
  • Backend call path:
  1. Resolve run path.
  2. Resolve task directory.
  3. Require trace.json.
  4. Call generate_executive_report(trace_path, task_dir, console).
  • Artifacts read:
  1. Per-task trace.json.
  • Error and exit semantics:
  1. Missing run, task, or trace exits non-zero.
  • Compatibility constraints:
  1. Preserve report generation from existing trace artifact.

Eval-v2 Artifact Schema — Current Production Names

This section documents the current actual column names as exported by eval-v2 at Phase 0 baseline. The schema is frozen for Phase 0. Schema renaming or migration is deferred and is not part of Phase 0.

Source constants

Canonical column lists are defined in src/data_agent_baseline/langgraph_agent/eval_v2.py as:

  • TASK_METRICS_COLUMNS
  • TRAJECTORY_COLUMNS
  • TOOL_CALLS_COLUMNS

task_metrics.csv / comprehensive_evaluation.csv

comprehensive_evaluation.csv is written by write_evaluation_v2 as an alias for task_metrics.csv (same data, same columns). Key columns at Phase 0:

  • Identity: run_id, task_id, trace_id, difficulty, timestamp, task_type
  • Outcome: execution_success, final_score, failure_reason, failure_category, outcome_status
  • Answer accuracy: answer_precision, answer_recall, answer_f1, column_precision, column_recall, column_f1, row_precision, row_recall, row_f1, recall
  • Verification: verification_triggered, verification_score, aat_verification_triggered, aat_verification_score
  • Agent behaviour: replan_count, retry_count, total_replans, autonomy_score, tool_efficiency, data_understanding_score
  • Cost: total_tokens, llm_calls, execution_time
  • Full column list: see TASK_METRICS_COLUMNS in eval_v2.py (238 columns at Phase 0 baseline).

The names execution_success and final_score are the current production names. No renaming is performed in Phase 0.

trajectory.csv

Columns from TRAJECTORY_COLUMNS (29 columns at Phase 0 baseline):

run_id, task_id, trace_id, step_id, parent_step_id, stage, legacy_stage, phase, agent, trajectory_agent, agent_type, tool, action, coordinator_checkpoint, confidence, decision, review_type, verification_status, specialist_selected, success, duration_seconds, tokens, retries, timestamp, elapsed_seconds, thought, action_input, observation, raw_response, metadata_json

tool_calls.csv

Columns from TOOL_CALLS_COLUMNS (17 columns at Phase 0 baseline):

run_id, task_id, trace_id, step_id, call_id, tool_name, tool_category, success, latency_seconds, tokens, retry_count, input_size, output_size, error_type, error_message, timestamp, metadata_json

Deferred work

  • Column renaming (e.g., execution_successtask_success, final_scorescore) is explicitly deferred.
  • No schema migration, column aliasing, or DataFrame reshaping is performed in Phase 0.
  • Tests reference only the current names.