Spaces:
Sleeping
DABench Evaluation System
Complete guide to the evaluation harness with integrated hardening features.
⚠️ Recent Fixes (June 2026)
Validation and Reporting Improvements:
- Fixed
compute_verification_passed()bug: returns boolean (0/1), resolvingverification_passed=2errors. - Coordinator consistency reconciliation (
replan_count_consistent,retry_count_consistent,coordinator_metrics_consistent) is now diagnostic-only warning output, not a hard eval-v2 failure condition. - Meaningful disagreement reporting is hardened to reduce confidence-only inflation.
- Result: eval-v2 completes across modes with reconciliation diagnostics reported as warnings when present.
🚀 Quick Start
Run Evaluation (with automatic hardening)
cd /data3/dataFAIR/kdd-dev/public
# Standard mode (quick overview)
dabench eval-v2 <RUN_ID> --mode standard
# Verbose mode (detailed analysis)
dabench eval-v2 <RUN_ID> --mode verbose
# Research mode (all metrics for papers)
dabench eval-v2 <RUN_ID> --mode research
What runs automatically:
- ✅ Standard evaluation (task_metrics, trajectory, tool_calls CSVs)
- ✅ Artifact reconciliation validation
- ✅ Replay artifact generation (debug snapshots)
- ✅ Engineering health report
Report Display Features:
- Separated Health Assessments:
- Harness Health (9.7/10 ✓ HEALTHY): Infrastructure quality (reconciliation, validators, attribution coverage)
- Run Quality (⚠ DEGRADED 64% accuracy): Outcome metrics (answer accuracy, execution success)
- Per-Task Results Table: Shows all tasks with execution success, answer quality, timing, and trajectory
- Exec column: Execution success (✓ = code ran, ✗ = crash)
- Root Cause column: Displays failure diagnostics for failed tasks (e.g., filter_logic_error, schema_misunderstanding)
- Overall Summary: Distinguishes execution success (code ran) from answer accuracy (correct results)
- Difficulty Breakdown: Three-column view:
- Execution Success Rate: % tasks that ran without crashes
- Answer Accuracy: % tasks with final_score ≥ 0.8
- Mean Final Score: Average correctness
- Mean Runtime: Average execution time per difficulty
- Mean Tokens: Average token usage per difficulty
- MAS Effectiveness: Shows first attempt vs final accuracy and recovery gain
- Coordinator Intervention Effectiveness: Replan and retry success rates
- Specialist Agent Value Analysis: Automatic ablation showing agent impact
- MAS Failure Analysis: Debugging-focused breakdown showing:
- MAS Failure Categories: Structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, etc.)
- Failure Distribution by Stage: AAT phases (UNDERSTAND/PLAN/EXECUTE/VERIFY/AGGREGATE)
- Outcome Error Types: Evaluation buckets (low_recall, wrong_schema, etc.)
- AAT Metrics: Coordinator decisions, specialist activation, verification outcomes
- Analyst Team Summary (verbose/research):
- mean agreement score
- tasks with meaningful disagreement
- disagreement type distribution
- coordinator override count
- verifier disagreement count
- critical disagreement score distribution (0-3)
- auditor trigger totals + precision/recall (warning->failure, failure->warning)
- Verification Timeline: Separates execution approval from ground truth correctness
- Phase Timing: UNDERSTAND → PLAN → EXECUTE → VERIFY → SUMMARIZE with reconciliation view
📁 Generated Artifacts
Every eval-v2 run produces:
artifacts/runs/<RUN_ID>/
├── task_metrics.csv # Per-task metrics (100+ columns)
├── trajectory.csv # Step-by-step execution trace
├── tool_calls.csv # Per-tool-call analysis
├── comprehensive_evaluation.csv # Backward compatibility
│
├── artifact_reconciliation_report.txt # Validation results
├── engineering_health_report.txt # System health diagnostics (includes answer accuracy, attribution coverage)
├── auditor_validation_report.md # Auditor effectiveness and failure-prevention diagnostics
│
└── task_*/
├── trace.json # Raw execution trace
├── answer.csv # Generated answer
└── task_replay.json # Complete debug context with failure attribution
📊 Key Metrics (100+ Total)
The evaluation system provides comprehensive metrics across 11 dimensions suitable for academic publication.
1. Correctness Metrics (Multi-Level F1)
| Metric | Formula | Range | Description |
|---|---|---|---|
answer_precision |
matched_cells / pred_cells | [0, 1] | Cell-level precision |
answer_recall |
matched_cells / gold_cells | [0, 1] | Cell-level recall |
answer_f1 |
2·P·R/(P+R) | [0, 1] | Cell-level F1 score |
column_precision |
matched_cols / pred_cols | [0, 1] | Column-level precision |
column_recall |
matched_cols / gold_cols | [0, 1] | Column-level recall |
column_f1 |
2·P·R/(P+R) | [0, 1] | Column-level F1 score |
row_precision |
min(pred, gold) / pred | [0, 1] | Row-level precision |
row_recall |
min(pred, gold) / gold | [0, 1] | Row-level recall |
row_f1 |
2·P·R/(P+R) | [0, 1] | Row-level F1 score |
final_score |
From evaluator | [0, 1] | Legacy overall score |
2. Autonomy Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
first_try_success |
succeeded ∧ attempts=1 | {0, 1} | Success without retry |
replan_count |
max(0, plan_attempts - 1) | [0, ∞) | Number of replans |
autonomy_score |
1.0 - replans/max_replans | [0, 1] | Independence measure |
coordinator_interventions |
coordinator_calls - 3 | [0, ∞) | Beyond baseline |
user_intervention_count |
0 (autonomous) | 0 | Manual interventions |
3. Planning Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
planner_steps |
count(action="planner") | [0, ∞) | Planning invocations |
planner_revisions |
max(0, plan_attempts - 1) | [0, ∞) | Plan revisions |
planner_dead_ends |
execution_attempts - 1 | [0, ∞) | Failed plans |
plan_execution_alignment |
1.0 if first_try else decay | [0, 1] | Plan-execution match |
4. Tool Usage Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
unique_tools_used |
|{tools}| | [0, ∞) | Tool diversity count |
tool_diversity |
unique_tools / total_calls | [0, 1] | Diversity ratio |
useful_tool_calls |
explore + final_exec | [0, ∞) | Contributory calls |
wasted_tool_calls |
total - useful | [0, ∞) | Non-contributory |
tool_efficiency |
useful / total | [0, 1] | Efficiency ratio |
tool_selection_accuracy |
(useful - failures) / total | [0, 1] | Selection quality |
tool_calls |
From trace | [0, ∞) | Total invocations |
tool_failures |
From trace | [0, ∞) | Failed invocations |
5. Data Understanding Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
tables_discovered |
From explore phase | [0, ∞) | Tables found |
columns_discovered |
From explore phase | [0, ∞) | Columns found |
relevant_tables_found |
Heuristic: all discovered | [0, ∞) | Relevant tables |
relevant_columns_found |
matched_columns | [0, ∞) | Relevant columns |
schema_exploration_steps |
count(phase="explore") | [0, ∞) | Exploration steps |
data_understanding_score |
(table_score + col_score) / 2 | [0, 1] | Composite score |
Formula for data_understanding_score (weighted composite):
# Specialist activation component
required_specialists = 2 + (1 if documents_required else 0) # schema, domain, [document]
activated_specialists = schema_used + domain_used + (document_used if docs_required else 0)
specialist_activation_score = activated_specialists / required_specialists
# Discovery component (from exploration phase)
table_score = relevant_tables_found / tables_discovered
column_score = relevant_columns_found / columns_discovered
discovery_score = (table_score + column_score) / 2
# Weighted formula (can exceed pure specialist score)
data_understanding_score = 0.5 * specialist_activation_score + 0.5 * discovery_score
# Bounds enforced: [0, 1]
Note: Score may exceed simple specialist calculation (e.g., 0.833 with 2/3 specialists if discovery_score is high).
6. Verification Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
verification_triggered |
1 if critic steps > 0 | {0, 1} | Verification used |
verification_steps |
count(critic actions) | [0, ∞) | Verification count |
critic_verification_passed |
From critic trace | {0, 1} | Critic passed check |
critic_verification_score |
critic_passed / critic_checks | [0, 1] | Critic quality |
critic_failures_detected |
From critic | [0, ∞) | Critic issues detected |
aat_verification_triggered |
1 if AAT verifier ran | {0, 1} | AAT verifier used |
aat_verification_passed |
From coordinator | {0, 1} | AAT verifier result |
aat_verification_score |
AAT verifier confidence | [0, 1] | AAT verification quality |
verification_passed |
Legacy (= critic_passed) | {0, 1} | Backward compat (critic) |
verification_score |
Legacy (= critic_score) | [0, 1] | Backward compat (critic) |
Critical Distinction:
critic_verification_passed: Step-level critic checks (legacy ReAct agent)aat_verification_passed: Final AAT coordinator approval (multi-agent system)- They are independent: Critic may pass but coordinator may still request retry
Invariant: aat_verification_passed=1 ↔ coordinator_final_decision="APPROVE_FINAL"
Backward Compatibility: verification_passed and verification_score maintain critic semantics for legacy comparisons.
7. Recovery Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
failure_detected |
From trace | {0, 1} | Failure occurred |
failure_stage |
First failure phase | str | Stage of failure |
recovery_success |
recovered after failure | {0, 1} | Recovery outcome |
recovery_depth |
attempts until success | [0, ∞) | Recovery iterations |
recovery_success_rate |
successes / attempts | [0, 1] | Recovery rate |
8. Trajectory Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
trajectory_length |
count(steps) | [0, ∞) | Total steps |
branching_factor |
avg(children per node) | [1, ∞) | Execution branches |
max_execution_depth |
max(depth in tree) | [0, ∞) | Deepest path |
critic_loops |
count(critic iterations) | [0, ∞) | Critic cycles |
trajectory_efficiency |
useful_steps / total_steps | [0, 1] | Step efficiency |
trajectory_summary |
Pattern string | str | Execution pattern |
Common trajectory patterns:
UNDERSTAND→PLAN→EXECUTE→VERIFY→SUMMARIZEUNDERSTAND→PLAN→EXECUTE→VERIFY→RETRY→SUMMARIZEUNDERSTAND→PLAN→REPLAN→EXECUTE→VERIFY→SUMMARIZE
9. Failure Taxonomy
| Metric | Type | Description |
|---|---|---|
failure_category |
enum | PLANNING / DATA_UNDERSTANDING / TOOL_EXECUTION / etc. |
root_cause |
str | Specific error cause |
severity |
enum | CRITICAL / HIGH / MEDIUM / LOW |
recoverable_failure |
bool | Can be recovered |
Categories:
PLANNING_FAILURE- Bad plan generationDATA_UNDERSTANDING_FAILURE- Schema/data misunderstandingTOOL_SELECTION_FAILURE- Wrong tool chosenTOOL_EXECUTION_FAILURE- Tool crash/errorREASONING_FAILURE- Logic errorsVERIFICATION_FAILURE- Verifier malfunctionAGGREGATION_FAILURE- Data aggregation errorsANSWER_FORMAT_FAILURE- Wrong output formatSCHEMA_MISMATCH_FAILURE- Schema mismatchUNKNOWN_FAILURE- Needs manual inspection
10. Confidence & Calibration Metrics
| Metric | Formula | Range | Description |
|---|---|---|---|
confidence_score |
From agent output | [0, 1] | Agent confidence |
confidence_correct |
conf ≥ 0.8 ∧ correct | {0, 1} | High conf + correct |
confidence_error |
|conf - correctness| | [0, 1] | Calibration error |
calibration_bucket |
Binned by confidence | str | Calibration bin |
Calibration buckets: very_low (0-0.2), low (0.2-0.4), medium (0.4-0.6), high (0.6-0.8), very_high (0.8-1.0)
11. Composite Scores (For Ranking)
| Metric | Formula | Range | Description |
|---|---|---|---|
analyst_score |
w₁·autonomy + w₂·efficiency + w₃·verification | [0, 1] | Weighted composite |
Analyst Score Formula:
analyst_score = 0.4·autonomy_score + 0.3·tool_efficiency + 0.3·verification_score
This composite metric balances:
- 40% Autonomy: Independence and minimal human intervention
- 30% Efficiency: Effective tool usage and resource management
- 30% Verification: Quality assurance and self-checking
Execution Metrics
| Metric | Type | Description |
|---|---|---|
execution_success |
bool | Code ran without crashes |
execution_time |
float | Wall clock time (seconds) |
total_tokens |
int | Total LLM tokens used |
llm_calls |
int | Number of LLM invocations |
tool_calls |
int | Number of tool invocations |
tool_failures |
int | Failed tool calls |
trajectory_length |
int | Number of execution steps |
AAT Architecture Metrics
| Metric | Type | Description |
|---|---|---|
coordinator_calls |
int | Strategic coordinator invocations |
coordinator_final_decision |
str | APPROVE_FINAL / RETRY_EXECUTION / REPLAN |
aat_verification_passed |
bool | AAT verifier approval |
schema_agent_called |
bool | Schema specialist invoked |
domain_agent_called |
bool | Domain specialist invoked |
document_agent_called |
bool | Document specialist invoked |
specialist_participation |
float | Proportion of specialists used |
Research Metrics (KDD Paper)
| Metric | Formula | Range | Description |
|---|---|---|---|
cross_source_reasoning_success |
From task metadata | {0, 1} | Multi-source reasoning |
explanation_quality_score |
From output | [0, 1] | Explanation quality |
reproducibility_score |
Deterministic replay | [0, 1] | Result stability |
� Statistical Analysis for Papers
Recommended Metrics for Publication
Primary Metrics (Table 1 - Main Results):
final_score(correctness) - Mean ± Stdanalyst_score(composite) - Mean ± Stdanswer_f1(cell-level) - Mean ± Stdautonomy_score- Mean ± Stdtool_efficiency- Mean ± Std
Breakdown by Difficulty (Table 2):
Report now distinguishes execution success from answer quality:
- Execution Success Rate: Tasks that ran without crashes (execution_success=1)
- Answer Accuracy: Tasks with correct answers (final_score ≥ 0.8)
- Mean Final Score: Average correctness score
# Compute breakdown
SUCCESS_THRESHOLD = 0.8
for difficulty in ['Easy', 'Medium', 'Hard', 'Extreme']:
subset = df[df['difficulty'] == difficulty]
print(f"{difficulty}:")
print(f" Execution success: {(subset['execution_success']==1).mean():.1%}")
print(f" Answer accuracy: {(subset['final_score']>=SUCCESS_THRESHOLD).mean():.1%}")
print(f" Mean score: {subset['final_score'].mean():.3f}")
Breakdown by Task Type (Table 3):
df.groupby('task_type')[['final_score', 'analyst_score']].agg(['mean', 'std', 'count'])
Multi-Agent Performance (Table 4):
# Compare specialist participation
df.groupby('specialist_participation')[['final_score', 'autonomy_score']].mean()
Ablation Studies
Ablation 1: Impact of Verification
with_verification = df[df['verification_triggered'] == 1]
without_verification = df[df['verification_triggered'] == 0]
print("With verification:", with_verification['final_score'].mean())
print("Without verification:", without_verification['final_score'].mean())
# Statistical test
from scipy.stats import mannwhitneyu
stat, p_value = mannwhitneyu(with_verification['final_score'],
without_verification['final_score'])
Ablation 2: Impact of Replanning
first_try = df[df['first_try_success'] == 1]
with_replans = df[df['replan_count'] > 0]
print("First try success rate:", first_try['final_score'].mean())
print("After replanning:", with_replans['final_score'].mean())
Ablation 3: Impact of Tool Efficiency
# Quartile analysis
df['efficiency_quartile'] = pd.qcut(df['tool_efficiency'], q=4, labels=['Q1', 'Q2', 'Q3', 'Q4'])
df.groupby('efficiency_quartile')['final_score'].agg(['mean', 'std', 'count'])
Correlation Analysis
import seaborn as sns
import matplotlib.pyplot as plt
# Select key metrics for correlation
metrics = ['final_score', 'analyst_score', 'autonomy_score',
'tool_efficiency', 'verification_score', 'data_understanding_score']
# Compute correlation matrix
corr = df[metrics].corr()
# Visualize
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, cmap='coolwarm', center=0,
square=True, linewidths=1)
plt.title('Metric Correlation Matrix')
plt.tight_layout()
plt.savefig('correlation_matrix.png', dpi=300)
Statistical Significance Testing
from scipy.stats import wilcoxon, mannwhitneyu
# Compare two systems (e.g., baseline vs. proposed)
baseline_df = pd.read_csv('baseline_run/task_metrics.csv')
proposed_df = pd.read_csv('proposed_run/task_metrics.csv')
# Paired test (same tasks)
merged = baseline_df.merge(proposed_df, on='task_id', suffixes=('_baseline', '_proposed'))
stat, p_value = wilcoxon(merged['final_score_baseline'], merged['final_score_proposed'])
print(f"Wilcoxon signed-rank test: p={p_value:.4f}")
# Effect size (Cohen's d)
mean_diff = merged['final_score_proposed'].mean() - merged['final_score_baseline'].mean()
pooled_std = np.sqrt((merged['final_score_proposed'].std()**2 +
merged['final_score_baseline'].std()**2) / 2)
cohens_d = mean_diff / pooled_std
print(f"Effect size (Cohen's d): {cohens_d:.3f}")
Failure Analysis for Papers
# Failure distribution (Figure 2)
failure_dist = df[df['final_score'] < 0.8]['failure_category'].value_counts()
plt.figure(figsize=(10, 6))
failure_dist.plot(kind='bar')
plt.xlabel('Failure Category')
plt.ylabel('Count')
plt.title('Failure Distribution')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.savefig('failure_distribution.png', dpi=300)
# Root cause analysis (Table 5)
root_causes = df[df['final_score'] < 0.8]['root_cause'].value_counts().head(10)
print(root_causes)
Reporting Template
Results Section:
We evaluate our system on the DABench benchmark containing 50 tasks of varying
difficulty (Easy: 15, Medium: 20, Hard: 15). Our system achieves a mean final
score of X.XX ± Y.YY (mean ± std), significantly outperforming the baseline
(p < 0.001, Wilcoxon signed-rank test). The analyst score, a composite metric
combining autonomy (weight=0.4), tool efficiency (weight=0.3), and verification
quality (weight=0.3), reaches Z.ZZ ± W.WW.
Breakdown by difficulty reveals consistent performance across all levels:
- Easy: X1 ± Y1 (n=15)
- Medium: X2 ± Y2 (n=20)
- Hard: X3 ± Y3 (n=15)
Our multi-agent architecture demonstrates strong autonomy with AA% first-try
success rate and an average of B.B replanning operations per task. Tool
efficiency reaches C.C ± D.D, indicating effective tool selection. Verification
mechanisms trigger in VV% of executions and detect EE failures, contributing to
improved final scores.
Failure analysis (Figure 2) shows the primary failure categories are:
1. TOOL_EXECUTION_FAILURE (XX%)
2. DATA_UNDERSTANDING_FAILURE (YY%)
3. SCHEMA_MISMATCH_FAILURE (ZZ%)
🎓 Experimental Methodology
Dataset Preparation
- Task Selection: Use stratified sampling by difficulty
- Data Splits: Train/Val/Test or K-fold cross-validation
- Seed Control: Fix random seeds for reproducibility
# Stratified sampling
from sklearn.model_selection import train_test_split
df = pd.read_csv('all_tasks.csv')
train, test = train_test_split(df, test_size=0.3,
stratify=df['difficulty'],
random_state=42)
Baseline Comparisons
Recommended Baselines:
- Random tool selection
- Fixed planning strategy
- No verification
- Single-agent (no specialists)
- Prior work (if available)
Reproducibility
Report:
- Hardware (GPU type, RAM)
- Software versions (Python, LLM API version)
- Random seeds
- Hyperparameters
- Number of runs (recommend 3-5 for variance)
Provide:
- Code repository
- Trained model weights (if applicable)
- Full evaluation CSVs (task_metrics.csv, trajectory.csv)
- Configuration files
Ethical Considerations
- Data privacy: Ensure benchmark tasks don't contain PII
- Computational cost: Report total compute time and carbon footprint
- Failure modes: Document dangerous failure patterns
- Limitations: Clearly state what the system cannot do
�🔍 Quick Analysis Examples
Load and Analyze
import pandas as pd
# Load metrics
df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")
# Success rate
success_rate = (df['final_score'] >= 0.8).mean()
print(f"Success rate: {success_rate:.1%}")
# By difficulty
print("\nScores by difficulty:")
print(df.groupby("difficulty")[["final_score", "analyst_score"]].mean())
# Failed tasks
failed = df[df['final_score'] < 0.8]
print(f"\nFailed: {len(failed)} tasks")
print(failed[['task_id', 'final_score', 'failure_category', 'root_cause']])
Debug Failed Task
import json
# Load replay artifact
with open("artifacts/runs/<RUN_ID>/task_38/task_replay.json") as f:
replay = json.load(f)
# Check failure
if replay['failure_attribution']:
fa = replay['failure_attribution']
print(f"Category: {fa['failure_category']}")
print(f"Root cause: {fa['root_cause']}")
print(f"Stage: {fa['failure_stage']}")
print(f"Reason: {fa['failure_reason']}")
print(f"Suggested fix: {fa['suggested_fix']}")
# Review execution
print(f"\nFinal score: {replay['evaluation_result']['final_score']}")
print(f"Coordinator decision: {replay['coordinator_final_decision']}")
print(f"Verification: {replay['verification_passed']}")
View Trajectory
import pandas as pd
# Load trajectory for specific task
traj = pd.read_csv("artifacts/runs/<RUN_ID>/trajectory.csv")
task_traj = traj[traj['task_id'] == 'task_38']
# View execution flow
print(task_traj[['step_id', 'phase', 'agent', 'tool', 'success', 'tokens']])
# Analyze failures
failures = task_traj[~task_traj['success']]
print(f"\nFailures: {len(failures)}")
print(failures[['step_id', 'tool', 'observation']])
🛠️ Hardening Features (Integrated)
1. Artifact Reconciliation
Validates:
- ✅ Tool call counts match across artifacts
- ✅ Token counts reconcile (trajectory vs metrics)
- ✅ Verification semantics consistent (
verification_passed ↔ coordinator_final_decision) - ✅ Time accounting (wall clock ≥ component time)
- ✅ Trajectory completeness
Report: artifact_reconciliation_report.txt
2. Replay Artifacts
Complete debug snapshots per task:
- Question & context
- All execution attempts (plan, code, stdout, stderr)
- Agent executions (MAS observability)
- Tool calls with success/failure
- Coordinator decisions
- Verifier output
- Final answer
- Evaluation result
- Structured failure attribution
Location: task_*/task_replay.json
3. Engineering Health Report
System diagnostics:
- Reconciliation pass/fail status
- Invariant violations (verification, tool calls, tokens)
- Time accounting gaps (overhead analysis)
- Failure taxonomy distribution
- Top recurring root causes
- System health indicators
- Health score (0-10)
Report: engineering_health_report.txt
🏥 Health Score Interpretation
| Score | Status | Action |
|---|---|---|
| 9-10 | ✅ HEALTHY | Ready to use |
| 7-8 | ⚠️ GOOD | Review warnings |
| 5-6 | ⚠️ FAIR | Fix issues before publication |
| 3-4 | ❌ POOR | Investigation required |
| 0-2 | ❌ UNHEALTHY | Do not use |
🔧 Implementation Details
Evaluation Pipeline Architecture
┌─────────────────────────────────────────────────────────────┐
│ Evaluation Harness V2 │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Data Collection │
│ • Load trace.json (agent execution trace) │
│ • Load prediction.csv (agent output) │
│ • Load gold.csv (ground truth) │
│ • Load task.json (metadata) │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Metric Computation │
│ • Correctness (multi-level F1: answer/column/row) │
│ • Autonomy (first-try success, replans, autonomy score) │
│ • Planning (revisions, dead ends, alignment) │
│ • Tool Usage (diversity, efficiency, selection accuracy) │
│ • Data Understanding (tables/columns, exploration) │
│ • Verification (triggered, passed, failures detected) │
│ • Recovery (detected, stage, success, depth) │
│ • Trajectory (length, branches, efficiency, patterns) │
│ • Failure Taxonomy (category, root cause, severity) │
│ • Confidence & Calibration (score, error, bucket) │
│ • Composite Scores (analyst_score = weighted blend) │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Normalized Storage (3 CSV Files) │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ task_metrics.csv (Primary Evaluation Table) │ │
│ │ • One row per task execution │ │
│ │ • 100+ columns covering all metric dimensions │ │
│ │ • Granularity: Task-level │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ trajectory.csv (Trajectory Trace) │ │
│ │ • One row per trajectory step │ │
│ │ • Enables process mining and step-level debugging │ │
│ │ • Granularity: Step-level │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ tool_calls.csv (Tool Usage Analysis) │ │
│ │ • One row per tool invocation │ │
│ │ • Tracks latency, tokens, retries, errors │ │
│ │ • Granularity: Tool-call-level │ │
│ └───────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Evaluation Hardening Suite (Integrated) │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 1. Artifact Reconciliation │ │
│ │ • Validates CSV consistency │ │
│ │ • Enforces invariants │ │
│ │ • Generates validation report │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 2. Replay Artifact Generation │ │
│ │ • Complete debug snapshots per task │ │
│ │ • Structured failure attribution │ │
│ │ • MAS observability tracking │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 3. Engineering Health Report │ │
│ │ • System health diagnostics │ │
│ │ • Time accounting analysis │ │
│ │ • Health score (0-10) │ │
│ └───────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Terminal Visualization (3 Modes) │
│ • Standard: Core metrics + summary │
│ • Verbose: + Agent behavior analysis │
│ • Research: + All metrics for papers (mean, std) │
└─────────────────────────────────────────────────────────────┘
Data Model Schema
1. task_metrics.csv (Primary Evaluation Table)
Purpose: Comprehensive per-task evaluation metrics for academic publication.
Granularity: One row per task execution (50-500 tasks typical).
Column Count: 100+ columns organized into 14 categories.
Schema Categories:
- Identification (5 cols): run_id, task_id, trace_id, difficulty, timestamp
- Task Metadata (5 cols): task_type, source_count, source_types, requires_cross_source_reasoning, ground_truth_available
- Correctness (15 cols): final_score, answer_precision/recall/f1, column_precision/recall/f1, row_precision/recall/f1, matched_columns, pred_rows/cols, gold_rows/cols
- Execution (6 cols): execution_success, execution_time, total_tokens, llm_calls, tool_calls, tool_failures
- Autonomy (5 cols): first_try_success, replan_count, user_intervention_count, autonomy_score, coordinator_interventions
- Planning (5 cols): planner_steps, planner_revisions, planner_dead_ends, plan_execution_alignment, plan_attempts
- Tool Usage (10 cols): unique_tools_used, tool_diversity, useful_tool_calls, wasted_tool_calls, tool_efficiency, tool_retry_count, tool_selection_accuracy
- Data Understanding (7 cols): tables_discovered, columns_discovered, relevant_tables_found, relevant_columns_found, schema_exploration_steps, data_understanding_score
- Verification (6 cols): verification_triggered, verification_steps, verification_passed, verification_failures_detected, verification_score, aat_verification_passed
- Recovery (6 cols): failure_detected, failure_stage, recovery_success, recovery_depth, recovery_success_rate
- Trajectory (7 cols): trajectory_length, trajectory_summary, branching_factor, max_execution_depth, critic_loops, trajectory_efficiency
- Failure Taxonomy (6 cols): failure_category, root_cause, recoverable_failure, severity, failure_reason, failure_agent
- Confidence (5 cols): confidence_score, confidence_correct, confidence_error, calibration_bucket
- Composite (1 col): analyst_score
- AAT Architecture (10 cols): coordinator_calls, coordinator_final_decision, coordinator_checkpoints, schema_agent_called, domain_agent_called, document_agent_called, specialist_participation
- Per-Stage Metrics (25 cols): understanding_time/calls/tokens, planning_time/calls/tokens, execution_time/calls/tokens, verification_time/calls/tokens, summary_time/calls/tokens
- Per-Action Metrics (28 cols): list_context_calls/time, read_json_calls/time, read_knowledge_calls/time, execute_python_calls/time, etc.
Total: 121 columns
2. trajectory.csv (Step-by-Step Trace)
Purpose: Detailed execution trace for process mining and debugging.
Granularity: One row per trajectory step (10-50 steps per task typical).
Columns (24):
- Identification: run_id, task_id, trace_id, step_id, parent_step_id
- Execution Context: stage, legacy_stage, phase, agent, trajectory_agent, agent_type
- Action: tool, action, coordinator_checkpoint
- Decision: confidence, decision, review_type, verification_status, specialist_selected
- Outcome: success, duration_seconds, tokens, retries
- Timing: timestamp, elapsed_seconds
- Details: thought, action_input, observation, raw_response, metadata_json
Use Cases:
- Process mining (find common execution patterns)
- Step-level debugging (identify exact failure point)
- Agent behavior analysis (tool selection patterns)
- Performance profiling (step latencies)
3. tool_calls.csv (Tool-Level Analysis)
Purpose: Per-tool-invocation metrics for optimization.
Granularity: One row per tool call (matches trajectory tool steps).
Columns (15):
- Identification: run_id, task_id, trace_id, step_id, call_id
- Tool: tool_name, tool_category
- Performance: success, latency_seconds, tokens, retry_count
- Data: input_size, output_size
- Errors: error_type, error_message
- Metadata: timestamp, metadata_json
Use Cases:
- Tool efficiency analysis (which tools are slow?)
- Error rate tracking (which tools fail most?)
- Token cost analysis (which tools are expensive?)
- Tool selection optimization (which tools work best for what?)
4. task_replay.json (Debug Snapshot - Per Task)
Purpose: Complete execution context for offline debugging without re-running.
Granularity: One JSON file per task.
Structure:
{
"task_id": "task_38",
"trace_id": "...",
"run_id": "...",
"difficulty": "Medium",
"question": "...",
"available_sources": [...],
"execution_attempts": [
{
"attempt_number": 1,
"plan": {...},
"code_executed": "...",
"stdout": "...",
"stderr": "...",
"success": false
}
],
"agent_executions": [
{
"agent_name": "StrategicCoordinator",
"checkpoint": "UNDERSTANDING",
"decision": "PROCEED",
"confidence": 0.85,
"duration": 5.2,
"tokens": 1500
}
],
"coordinator_decisions": [...],
"verifier_output": {...},
"final_answer_csv": "...",
"evaluation_result": {
"final_score": 0.0,
"answer_f1": 0.0
},
"failure_attribution": {
"failure_category": "TOOL_EXECUTION_FAILURE",
"root_cause": "KeyError: 'column_name'",
"failure_stage": "execution",
"failure_agent": "Executor",
"evidence_trace_step_ids": [12, 14],
"suggested_fix": "Validate column existence"
}
}
Module Structure
src/data_agent_baseline/langgraph_agent/
├── eval_v2.py # Main orchestrator (650 lines)
│ ├── evaluate_task_v2() # Single task evaluation
│ ├── evaluate_run_v2() # Full run evaluation
│ ├── write_evaluation_v2() # CSV output
│ └── extract_trajectory/tools() # Trace extraction
│
├── eval_v2_metrics.py # Metric computations (1280 lines)
│ ├── compute_correctness() # Multi-level F1
│ ├── compute_autonomy() # Independence metrics
│ ├── compute_planning() # Planning quality
│ ├── compute_tool_usage() # Tool efficiency
│ ├── compute_data_understanding() # Schema understanding
│ ├── compute_verification() # Quality assurance
│ ├── compute_recovery() # Error recovery
│ ├── compute_trajectory() # Execution path analysis
│ ├── compute_failure_taxonomy() # Error classification
│ ├── compute_confidence() # Calibration metrics
│ └── compute_analyst_score() # Composite score
│
├── eval_v2_viz.py # Visualization (580 lines)
│ ├── render_evaluation_report() # Main report
│ ├── render_task_table() # Per-task table
│ ├── render_summary_sections() # Analysis sections
│ └── render_verbose_task_detail() # Task drill-down
│
├── eval_artifact_reconciliation.py # Validation (435 lines)
│ ├── ArtifactReconciliator # Cross-artifact validation
│ ├── validate_run() # Run-level validation
│ └── format_report() # Validation report
│
├── eval_replay_artifacts.py # Debug snapshots (630 lines)
│ ├── ReplayArtifactGenerator # Snapshot generation
│ ├── FailureAttributor # Root cause analysis
│ └── generate_all_replays() # Batch generation
│
└── eval_health_report.py # Health diagnostics (490 lines)
├── EngineeringHealthReport # Health metrics
├── TimeAccountingGap # Overhead analysis
└── generate_health_report() # Report generation
Key Files
Core Evaluation:
eval_v2.py- Main evaluation engine (scoring, metrics)eval_v2_metrics.py- Individual metric calculatorseval_v2_viz.py- Rendering & visualization
Hardening Suite:
eval_artifact_reconciliation.py- Cross-artifact validationeval_replay_artifacts.py- Debug snapshot generationeval_health_report.py- System health diagnostics
CLI Integration:
cli.py:eval_v2_command()- Orchestrates entire pipeline
Critical Invariants
# Verification consistency
aat_verification_passed = 1 ↔ coordinator_final_decision = "APPROVE_FINAL"
# Note: critic_verification_passed is independent of AAT coordinator decision
# Tool call reconciliation
len(tool_calls_df) == task_metrics['tool_calls']
tool_failures_count == task_metrics['tool_failures']
# Token conservation
abs(trajectory_tokens - metrics_tokens) < 100
# Time accounting (with overhead)
wall_clock_time = execution_time # End-to-end elapsed
component_compute_time = sum(agent/tool tracked durations)
unaccounted_overhead = wall_clock_time - component_compute_time
time_accounting_ratio = component_compute_time / wall_clock_time
# Expected: component_time < wall_clock_time (overhead exists)
# Framework overhead, I/O wait, concurrency gaps, logging typically 20-50%
# Warning only if time_accounting_ratio < 0.05 (95% unaccounted)
# Info if component_time > wall_clock_time (indicates concurrent execution)
Comparison with Previous Systems
| Feature | Traditional Eval | DABench V1 | DABench V2 (Ours) |
|---|---|---|---|
| Storage Model | Single CSV | Single CSV | 3 normalized CSVs |
| Metrics | 5-10 basic | ~30 metrics | 100+ comprehensive |
| Correctness | Overall F1 | Overall F1 | Multi-level F1 (answer/column/row) |
| Autonomy | Not measured | Retry count | Composite autonomy score |
| Planning | Not measured | Step count | Revisions, alignment, dead ends |
| Tool Analysis | Call count | Call count | Efficiency, diversity, selection accuracy |
| Trajectory | Not captured | Basic steps | Full trace with branching, patterns |
| Failure Analysis | Error message | Simple bucket | Structured taxonomy + root cause |
| Verification | Not measured | Basic check | Multi-stage verification score |
| Confidence | Not measured | Not measured | Calibration metrics |
| Composite Scores | None | None | Analyst score (weighted) |
| Debugging | Manual | Manual | Automated replay artifacts |
| Validation | None | Basic | Comprehensive reconciliation |
| Process Mining | Not supported | Not supported | Full trajectory CSV |
| MAS Observability | Not supported | Not supported | Per-agent tracking |
| Time Accounting | Wall clock only | Wall clock only | Component + overhead |
Key Innovations:
- Multi-level correctness: Separate precision/recall/F1 at answer/column/row levels
- Autonomy quantification: First-try success, replan count, composite autonomy score
- Composite analyst score: Weighted blend of autonomy, efficiency, verification (suitable for ranking)
- Structured failure taxonomy: 10 failure categories with root cause attribution
- Integrated validation: Automatic artifact reconciliation with invariant enforcement
- Debug-ready artifacts: Complete replay snapshots for offline debugging
- Process mining support: Full trajectory CSV for pattern discovery
- MAS observability: Per-agent execution tracking in multi-agent systems
🎯 Failure Attribution
Failure Categories
PLANNING_FAILURE # Bad plan generation
DATA_UNDERSTANDING_FAILURE # Misunderstood data/schema
TOOL_SELECTION_FAILURE # Wrong tool chosen
TOOL_EXECUTION_FAILURE # Tool crashed/errored
REASONING_FAILURE # Logic errors
VERIFICATION_FAILURE # Verifier malfunction
AGGREGATION_FAILURE # Data aggregation errors
ANSWER_FORMAT_FAILURE # Wrong output format
SCHEMA_MISMATCH_FAILURE # Schema mismatch
UNKNOWN_FAILURE # Needs manual inspection
Attribution Structure
{
"failure_category": "TOOL_EXECUTION_FAILURE",
"root_cause": "KeyError: 'column_name'",
"failure_stage": "execution",
"failure_agent": "Executor",
"failure_reason": "Attempted to access non-existent column",
"evidence_trace_step_ids": [12, 14],
"suggested_fix": "Validate column existence before access",
"evidence_summary": "Step 12: execute_python failed with KeyError",
"execution_success": false,
"tool_failure_count": 1
}
📖 Common Workflows
1. Evaluate New Run
# Run prediction first (if not already done)
dabench run input_full --agent aat
# Evaluate
dabench eval-v2 <RUN_ID> --mode research
# Review health
cat artifacts/runs/<RUN_ID>/engineering_health_report.txt
2. Debug Failed Tasks
# Find failed tasks
python3 -c "
import pandas as pd
df = pd.read_csv('artifacts/runs/<RUN_ID>/task_metrics.csv')
failed = df[df['final_score'] < 0.8]
print(failed[['task_id', 'failure_category', 'root_cause']])
"
# Debug specific task
cat artifacts/runs/<RUN_ID>/task_38/task_replay.json | jq '.failure_attribution'
3. Compare Runs
import pandas as pd
# Load two runs
run1 = pd.read_csv("artifacts/runs/RUN_A/task_metrics.csv")
run2 = pd.read_csv("artifacts/runs/RUN_B/task_metrics.csv")
# Merge on task_id
merged = run1.merge(run2, on='task_id', suffixes=('_A', '_B'))
# Compare
print(f"Run A mean: {merged['final_score_A'].mean():.3f}")
print(f"Run B mean: {merged['final_score_B'].mean():.3f}")
# Tasks improved in Run B
improved = merged[merged['final_score_B'] > merged['final_score_A']]
print(f"\nImproved: {len(improved)} tasks")
4. Generate Paper Figures
import pandas as pd
import matplotlib.pyplot as plt
df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")
# Score distribution
plt.figure(figsize=(10, 6))
plt.hist(df['final_score'], bins=20, edgecolor='black')
plt.xlabel('Final Score')
plt.ylabel('Count')
plt.title('Score Distribution')
plt.savefig('score_distribution.png')
# Autonomy vs Efficiency
plt.figure(figsize=(10, 6))
plt.scatter(df['autonomy_score'], df['tool_efficiency'],
c=df['final_score'], cmap='viridis')
plt.xlabel('Autonomy Score')
plt.ylabel('Tool Efficiency')
plt.colorbar(label='Final Score')
plt.savefig('autonomy_vs_efficiency.png')
🧪 Testing
Run Test Suite
cd /workspace/ainn-cm-poc-data-agent
# Test evaluation harness
pytest tests/test_eval_harness.py -v
# Test specific validation
pytest tests/test_eval_harness.py::TestVerificationConsistency -v
Manual Validation
# Re-validate existing run
python3 -c "
from pathlib import Path
from src.data_agent_baseline.langgraph_agent.eval_artifact_reconciliation import validate_evaluation_run
passed, report = validate_evaluation_run(Path('artifacts/runs/<RUN_ID>'))
print(report)
print(f'\nPassed: {passed}')
"
🐛 Troubleshooting
Issue: "Evaluation produced inconsistent metrics"
Cause: Old consistency validator found errors
Solution: Check artifact_reconciliation_report.txt for details:
cat artifacts/runs/<RUN_ID>/artifact_reconciliation_report.txt
Common issues:
verification_outcome_mismatch: Verification flag doesn't match coordinator decisiontool_call_count_mismatch: Tool calls CSV doesn't match metricsdata_understanding_inflation: Score exceeds theoretical maximum
Issue: Health score < 7
Cause: System detected quality issues
Solution: Review engineering_health_report.txt:
cat artifacts/runs/<RUN_ID>/engineering_health_report.txt
Look for:
- Time accounting gaps > 50%
- High tool failure rate > 10%
- Missing failure attribution
Issue: Missing trajectory duration
Symptom: Warnings about "missing time data"
Cause: Trajectory extraction didn't populate duration_seconds column
Impact: Time validation skipped (not critical)
📚 Additional Documentation
Architecture Details:
- See
src/data_agent_baseline/langgraph_agent/eval_v2.pyfor scoring logic - See
src/data_agent_baseline/langgraph_agent/eval_v2_metrics.pyfor metric definitions
Test Coverage:
- See
tests/test_eval_harness.pyfor validation tests
CLI Integration:
- See
src/data_agent_baseline/cli.py:eval_v2_command()for integration
🔄 Version History
V2.1 (Current) - June 14, 2026
Phase 2: MAS Debugging Enhancements
Focused improvements for multi-agent system debugging and actionable diagnostics:
Deterministic Failure Attribution: Maps evaluation buckets (low_recall, wrong_schema, etc.) to structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, AGGREGATION_FAILURE)
- Eliminates UNKNOWN_FAILURE when bucket exists
- Infers failure_stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
- Identifies failure_agent (Schema Agent, Planner, Executor, etc.)
Separated Health Assessments: Clear distinction between infrastructure health and run quality
- Harness Health: Infrastructure metrics (reconciliation pass rate, validator errors, attribution coverage)
- Run Quality: Outcome metrics (answer accuracy, execution success rate)
- Status thresholds: HEALTHY (9.0+, no errors), DEGRADED (7.0+), UNHEALTHY (else)
Enhanced Visualization:
- Per-task table: Renamed "Succ" → "Exec" to clarify execution vs. correctness
- Added "Root Cause" column showing failure diagnostics (filter_logic_error, schema_misunderstanding, etc.)
- New "MAS Failure Analysis" section with:
- MAS Failure Categories table (structured categories: REASONING_FAILURE, etc.)
- Failure Distribution by Stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
- Outcome Error Types (evaluation buckets)
Partial-Correct Task Handling: New
outcome_statusfield distinguishes:- "correct": final_score ≥ 0.8
- "partial": succeeded but 0 < final_score < 0.8
- "failed": final_score < 0.8 or execution failure
Improved Terminology: Renamed "time gaps" → "Unaccounted Overhead Time" with clear explanation (framework overhead, I/O wait, async queuing)
Complete Replay Artifacts: Every task_replay.json includes:
- Full failure diagnostics (failure_category, root_cause, failure_stage, failure_agent)
- Run context (harness_health_status, run_quality_status, outcome_status)
- Time accounting (unaccounted_overhead_time_seconds, time_accounting_ratio)
Design Philosophy: Optimized for debugging and MAS improvement, not paper metrics. All changes maintain backward compatibility.
Phase 2.1 Cleanup (Jan 2026):
- Separated harness health (infrastructure) from run quality (outcomes)
- Fixed MAS Failure Categories display to show structured categories instead of buckets
- Added outcome_status field for partial-correct handling
- Enhanced replay artifacts with health/quality context
- Improved time accounting terminology
Phase 0: MAS Debugging Enhancements - Final Round (Jan 2026)
Goal: Make evaluation harness maximally useful for all future phases (Baseline ReAct, MAS, DAG Visualization, Replay/Time Travel, Confidence & Verification, Research/Ablation Studies).
Key Improvements:
MAS Recovery Effectiveness Metrics (Task 2):
- New fields:
initial_answer_correct,final_answer_correct,recovered_after_replan,recovered_after_retry - Display: "MAS Effectiveness" section showing:
- First Attempt Accuracy: Initial correct rate
- Final Accuracy: Final correct rate
- Recovered Tasks: Count of tasks improved through MAS interventions
- MAS Recovery Gain: Percentage improvement (e.g., +14%)
- Impact: Directly answers "Did MAS actually improve answers?"
- New fields:
Replan/Retry Effectiveness Tracking (Task 3):
- New fields:
replan_requested,replan_successful,retry_requested,retry_successful - Display: "Coordinator Intervention Effectiveness" section showing:
- Replans: Requested count, successful count, success rate
- Retries: Requested count, successful count, success rate
- Impact: Shows which coordinator interventions actually help
- New fields:
Specialist Agent Value Analysis (Task 4):
- Existing fields:
schema_agent_used,domain_agent_used,document_agent_used - Display: "Specialist Agent Value Analysis" section showing for each agent:
- Tasks Used
- Accuracy With Agent
- Accuracy Without Agent
- Impact (delta %)
- Impact: Automatic ablation showing which specialists add value
- Existing fields:
Expanded Failure Stage Taxonomy (Task 5):
- Updated
FailureStageenum: UNDERSTAND, PLAN, EXECUTE, VERIFY, AGGREGATE - Replaced coarse stages (EXPLORATION, PLANNING, EXECUTION) with AAT-aligned taxonomy
- Impact: Finer-grained debugging for AAT phase-specific failures
- Updated
Cost by Difficulty (Task 6):
- Difficulty breakdown now includes Mean Runtime and Mean Tokens columns
- Impact: Required for Baseline vs MAS vs Future comparisons
Verification Timeline Clarity (Task 7):
- Separated "Execution Approval" (coordinator decision) from "Ground Truth Result" (evaluation correctness)
- Added explanatory note distinguishing verification from correctness
- Impact: Eliminates confusion between process approval and actual correctness
Comprehensive CSV Storage (Task 8):
- All new fields stored in task_metrics.csv:
outcome_error_type, MAS recovery fields, specialist usage - Impact: Future phases can run aggregations without parsing replay artifacts
- All new fields stored in task_metrics.csv:
Removed Duplicate Reporting (Task 1):
- Eliminated redundant "Failure Categories" section
- Kept: "MAS Failure Categories" (structured) and "Outcome Error Types" (buckets)
Design Philosophy: Every change focused on making the evaluation harness more actionable for debugging and improving the MAS, not for paper-writing. Provides automatic ablation studies and directly answers key questions about MAS effectiveness.
V2.0 - June 2026
- ✅ Integrated hardening suite (automatic reconciliation, replay, health)
- ✅ 100+ comprehensive metrics for KDD Creative Track
- ✅ AAT architecture observability
- ✅ Structured failure attribution
- ✅ MAS-aware trajectory extraction
- ✅ Time accounting with overhead tracking
V1 (Legacy)
- Basic metrics (precision, recall, F1)
- Manual validation required
- Limited debugging support
📝 Summary
One command does it all:
dabench eval-v2 <RUN_ID> --mode standard
Automatically provides:
- ✅ Comprehensive metrics (100+ columns)
- ✅ Complete validation & reconciliation
- ✅ Full debug snapshots (replay artifacts)
- ✅ System health diagnostics
- ✅ Failure attribution & root cause analysis
No manual steps required.