DAO_kdd26 / docs /implementation /EVALUATION.md
sipe5001's picture
Add Hugging Face Docker Space configuration
d3d0e0e
|
Raw
History Blame Contribute Delete
55.9 kB

DABench Evaluation System

Complete guide to the evaluation harness with integrated hardening features.


⚠️ Recent Fixes (June 2026)

Validation and Reporting Improvements:

  • Fixed compute_verification_passed() bug: returns boolean (0/1), resolving verification_passed=2 errors.
  • Coordinator consistency reconciliation (replan_count_consistent, retry_count_consistent, coordinator_metrics_consistent) is now diagnostic-only warning output, not a hard eval-v2 failure condition.
  • Meaningful disagreement reporting is hardened to reduce confidence-only inflation.
  • Result: eval-v2 completes across modes with reconciliation diagnostics reported as warnings when present.

🚀 Quick Start

Run Evaluation (with automatic hardening)

cd /data3/dataFAIR/kdd-dev/public

# Standard mode (quick overview)
dabench eval-v2 <RUN_ID> --mode standard

# Verbose mode (detailed analysis)
dabench eval-v2 <RUN_ID> --mode verbose

# Research mode (all metrics for papers)
dabench eval-v2 <RUN_ID> --mode research

What runs automatically:

  1. ✅ Standard evaluation (task_metrics, trajectory, tool_calls CSVs)
  2. ✅ Artifact reconciliation validation
  3. ✅ Replay artifact generation (debug snapshots)
  4. ✅ Engineering health report

Report Display Features:

  • Separated Health Assessments:
    • Harness Health (9.7/10 ✓ HEALTHY): Infrastructure quality (reconciliation, validators, attribution coverage)
    • Run Quality (⚠ DEGRADED 64% accuracy): Outcome metrics (answer accuracy, execution success)
  • Per-Task Results Table: Shows all tasks with execution success, answer quality, timing, and trajectory
    • Exec column: Execution success (✓ = code ran, ✗ = crash)
    • Root Cause column: Displays failure diagnostics for failed tasks (e.g., filter_logic_error, schema_misunderstanding)
  • Overall Summary: Distinguishes execution success (code ran) from answer accuracy (correct results)
  • Difficulty Breakdown: Three-column view:
    • Execution Success Rate: % tasks that ran without crashes
    • Answer Accuracy: % tasks with final_score ≥ 0.8
    • Mean Final Score: Average correctness
    • Mean Runtime: Average execution time per difficulty
    • Mean Tokens: Average token usage per difficulty
  • MAS Effectiveness: Shows first attempt vs final accuracy and recovery gain
  • Coordinator Intervention Effectiveness: Replan and retry success rates
  • Specialist Agent Value Analysis: Automatic ablation showing agent impact
  • MAS Failure Analysis: Debugging-focused breakdown showing:
    • MAS Failure Categories: Structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, etc.)
    • Failure Distribution by Stage: AAT phases (UNDERSTAND/PLAN/EXECUTE/VERIFY/AGGREGATE)
    • Outcome Error Types: Evaluation buckets (low_recall, wrong_schema, etc.)
  • AAT Metrics: Coordinator decisions, specialist activation, verification outcomes
  • Analyst Team Summary (verbose/research):
    • mean agreement score
    • tasks with meaningful disagreement
    • disagreement type distribution
    • coordinator override count
    • verifier disagreement count
    • critical disagreement score distribution (0-3)
    • auditor trigger totals + precision/recall (warning->failure, failure->warning)
  • Verification Timeline: Separates execution approval from ground truth correctness
  • Phase Timing: UNDERSTAND → PLAN → EXECUTE → VERIFY → SUMMARIZE with reconciliation view

📁 Generated Artifacts

Every eval-v2 run produces:

artifacts/runs/<RUN_ID>/
├── task_metrics.csv                          # Per-task metrics (100+ columns)
├── trajectory.csv                            # Step-by-step execution trace
├── tool_calls.csv                            # Per-tool-call analysis
├── comprehensive_evaluation.csv              # Backward compatibility
│
├── artifact_reconciliation_report.txt        # Validation results
├── engineering_health_report.txt             # System health diagnostics (includes answer accuracy, attribution coverage)
├── auditor_validation_report.md              # Auditor effectiveness and failure-prevention diagnostics
│
└── task_*/
    ├── trace.json                            # Raw execution trace
    ├── answer.csv                            # Generated answer
    └── task_replay.json                      # Complete debug context with failure attribution

📊 Key Metrics (100+ Total)

The evaluation system provides comprehensive metrics across 11 dimensions suitable for academic publication.

1. Correctness Metrics (Multi-Level F1)

Metric Formula Range Description
answer_precision matched_cells / pred_cells [0, 1] Cell-level precision
answer_recall matched_cells / gold_cells [0, 1] Cell-level recall
answer_f1 2·P·R/(P+R) [0, 1] Cell-level F1 score
column_precision matched_cols / pred_cols [0, 1] Column-level precision
column_recall matched_cols / gold_cols [0, 1] Column-level recall
column_f1 2·P·R/(P+R) [0, 1] Column-level F1 score
row_precision min(pred, gold) / pred [0, 1] Row-level precision
row_recall min(pred, gold) / gold [0, 1] Row-level recall
row_f1 2·P·R/(P+R) [0, 1] Row-level F1 score
final_score From evaluator [0, 1] Legacy overall score

2. Autonomy Metrics

Metric Formula Range Description
first_try_success succeeded ∧ attempts=1 {0, 1} Success without retry
replan_count max(0, plan_attempts - 1) [0, ∞) Number of replans
autonomy_score 1.0 - replans/max_replans [0, 1] Independence measure
coordinator_interventions coordinator_calls - 3 [0, ∞) Beyond baseline
user_intervention_count 0 (autonomous) 0 Manual interventions

3. Planning Metrics

Metric Formula Range Description
planner_steps count(action="planner") [0, ∞) Planning invocations
planner_revisions max(0, plan_attempts - 1) [0, ∞) Plan revisions
planner_dead_ends execution_attempts - 1 [0, ∞) Failed plans
plan_execution_alignment 1.0 if first_try else decay [0, 1] Plan-execution match

4. Tool Usage Metrics

Metric Formula Range Description
unique_tools_used |{tools}| [0, ∞) Tool diversity count
tool_diversity unique_tools / total_calls [0, 1] Diversity ratio
useful_tool_calls explore + final_exec [0, ∞) Contributory calls
wasted_tool_calls total - useful [0, ∞) Non-contributory
tool_efficiency useful / total [0, 1] Efficiency ratio
tool_selection_accuracy (useful - failures) / total [0, 1] Selection quality
tool_calls From trace [0, ∞) Total invocations
tool_failures From trace [0, ∞) Failed invocations

5. Data Understanding Metrics

Metric Formula Range Description
tables_discovered From explore phase [0, ∞) Tables found
columns_discovered From explore phase [0, ∞) Columns found
relevant_tables_found Heuristic: all discovered [0, ∞) Relevant tables
relevant_columns_found matched_columns [0, ∞) Relevant columns
schema_exploration_steps count(phase="explore") [0, ∞) Exploration steps
data_understanding_score (table_score + col_score) / 2 [0, 1] Composite score

Formula for data_understanding_score (weighted composite):

# Specialist activation component
required_specialists = 2 + (1 if documents_required else 0)  # schema, domain, [document]
activated_specialists = schema_used + domain_used + (document_used if docs_required else 0)
specialist_activation_score = activated_specialists / required_specialists

# Discovery component (from exploration phase)
table_score = relevant_tables_found / tables_discovered
column_score = relevant_columns_found / columns_discovered
discovery_score = (table_score + column_score) / 2

# Weighted formula (can exceed pure specialist score)
data_understanding_score = 0.5 * specialist_activation_score + 0.5 * discovery_score

# Bounds enforced: [0, 1]

Note: Score may exceed simple specialist calculation (e.g., 0.833 with 2/3 specialists if discovery_score is high).

6. Verification Metrics

Metric Formula Range Description
verification_triggered 1 if critic steps > 0 {0, 1} Verification used
verification_steps count(critic actions) [0, ∞) Verification count
critic_verification_passed From critic trace {0, 1} Critic passed check
critic_verification_score critic_passed / critic_checks [0, 1] Critic quality
critic_failures_detected From critic [0, ∞) Critic issues detected
aat_verification_triggered 1 if AAT verifier ran {0, 1} AAT verifier used
aat_verification_passed From coordinator {0, 1} AAT verifier result
aat_verification_score AAT verifier confidence [0, 1] AAT verification quality
verification_passed Legacy (= critic_passed) {0, 1} Backward compat (critic)
verification_score Legacy (= critic_score) [0, 1] Backward compat (critic)

Critical Distinction:

  • critic_verification_passed: Step-level critic checks (legacy ReAct agent)
  • aat_verification_passed: Final AAT coordinator approval (multi-agent system)
  • They are independent: Critic may pass but coordinator may still request retry

Invariant: aat_verification_passed=1 ↔ coordinator_final_decision="APPROVE_FINAL"

Backward Compatibility: verification_passed and verification_score maintain critic semantics for legacy comparisons.

7. Recovery Metrics

Metric Formula Range Description
failure_detected From trace {0, 1} Failure occurred
failure_stage First failure phase str Stage of failure
recovery_success recovered after failure {0, 1} Recovery outcome
recovery_depth attempts until success [0, ∞) Recovery iterations
recovery_success_rate successes / attempts [0, 1] Recovery rate

8. Trajectory Metrics

Metric Formula Range Description
trajectory_length count(steps) [0, ∞) Total steps
branching_factor avg(children per node) [1, ∞) Execution branches
max_execution_depth max(depth in tree) [0, ∞) Deepest path
critic_loops count(critic iterations) [0, ∞) Critic cycles
trajectory_efficiency useful_steps / total_steps [0, 1] Step efficiency
trajectory_summary Pattern string str Execution pattern

Common trajectory patterns:

  • UNDERSTAND→PLAN→EXECUTE→VERIFY→SUMMARIZE
  • UNDERSTAND→PLAN→EXECUTE→VERIFY→RETRY→SUMMARIZE
  • UNDERSTAND→PLAN→REPLAN→EXECUTE→VERIFY→SUMMARIZE

9. Failure Taxonomy

Metric Type Description
failure_category enum PLANNING / DATA_UNDERSTANDING / TOOL_EXECUTION / etc.
root_cause str Specific error cause
severity enum CRITICAL / HIGH / MEDIUM / LOW
recoverable_failure bool Can be recovered

Categories:

  • PLANNING_FAILURE - Bad plan generation
  • DATA_UNDERSTANDING_FAILURE - Schema/data misunderstanding
  • TOOL_SELECTION_FAILURE - Wrong tool chosen
  • TOOL_EXECUTION_FAILURE - Tool crash/error
  • REASONING_FAILURE - Logic errors
  • VERIFICATION_FAILURE - Verifier malfunction
  • AGGREGATION_FAILURE - Data aggregation errors
  • ANSWER_FORMAT_FAILURE - Wrong output format
  • SCHEMA_MISMATCH_FAILURE - Schema mismatch
  • UNKNOWN_FAILURE - Needs manual inspection

10. Confidence & Calibration Metrics

Metric Formula Range Description
confidence_score From agent output [0, 1] Agent confidence
confidence_correct conf ≥ 0.8 ∧ correct {0, 1} High conf + correct
confidence_error |conf - correctness| [0, 1] Calibration error
calibration_bucket Binned by confidence str Calibration bin

Calibration buckets: very_low (0-0.2), low (0.2-0.4), medium (0.4-0.6), high (0.6-0.8), very_high (0.8-1.0)

11. Composite Scores (For Ranking)

Metric Formula Range Description
analyst_score w₁·autonomy + w₂·efficiency + w₃·verification [0, 1] Weighted composite

Analyst Score Formula:

analyst_score = 0.4·autonomy_score + 0.3·tool_efficiency + 0.3·verification_score

This composite metric balances:

  • 40% Autonomy: Independence and minimal human intervention
  • 30% Efficiency: Effective tool usage and resource management
  • 30% Verification: Quality assurance and self-checking

Execution Metrics

Metric Type Description
execution_success bool Code ran without crashes
execution_time float Wall clock time (seconds)
total_tokens int Total LLM tokens used
llm_calls int Number of LLM invocations
tool_calls int Number of tool invocations
tool_failures int Failed tool calls
trajectory_length int Number of execution steps

AAT Architecture Metrics

Metric Type Description
coordinator_calls int Strategic coordinator invocations
coordinator_final_decision str APPROVE_FINAL / RETRY_EXECUTION / REPLAN
aat_verification_passed bool AAT verifier approval
schema_agent_called bool Schema specialist invoked
domain_agent_called bool Domain specialist invoked
document_agent_called bool Document specialist invoked
specialist_participation float Proportion of specialists used

Research Metrics (KDD Paper)

Metric Formula Range Description
cross_source_reasoning_success From task metadata {0, 1} Multi-source reasoning
explanation_quality_score From output [0, 1] Explanation quality
reproducibility_score Deterministic replay [0, 1] Result stability

� Statistical Analysis for Papers

Recommended Metrics for Publication

Primary Metrics (Table 1 - Main Results):

  • final_score (correctness) - Mean ± Std
  • analyst_score (composite) - Mean ± Std
  • answer_f1 (cell-level) - Mean ± Std
  • autonomy_score - Mean ± Std
  • tool_efficiency - Mean ± Std

Breakdown by Difficulty (Table 2):

Report now distinguishes execution success from answer quality:

  • Execution Success Rate: Tasks that ran without crashes (execution_success=1)
  • Answer Accuracy: Tasks with correct answers (final_score ≥ 0.8)
  • Mean Final Score: Average correctness score
# Compute breakdown
SUCCESS_THRESHOLD = 0.8
for difficulty in ['Easy', 'Medium', 'Hard', 'Extreme']:
    subset = df[df['difficulty'] == difficulty]
    print(f"{difficulty}:")
    print(f"  Execution success: {(subset['execution_success']==1).mean():.1%}")
    print(f"  Answer accuracy: {(subset['final_score']>=SUCCESS_THRESHOLD).mean():.1%}")
    print(f"  Mean score: {subset['final_score'].mean():.3f}")

Breakdown by Task Type (Table 3):

df.groupby('task_type')[['final_score', 'analyst_score']].agg(['mean', 'std', 'count'])

Multi-Agent Performance (Table 4):

# Compare specialist participation
df.groupby('specialist_participation')[['final_score', 'autonomy_score']].mean()

Ablation Studies

Ablation 1: Impact of Verification

with_verification = df[df['verification_triggered'] == 1]
without_verification = df[df['verification_triggered'] == 0]

print("With verification:", with_verification['final_score'].mean())
print("Without verification:", without_verification['final_score'].mean())
# Statistical test
from scipy.stats import mannwhitneyu
stat, p_value = mannwhitneyu(with_verification['final_score'], 
                               without_verification['final_score'])

Ablation 2: Impact of Replanning

first_try = df[df['first_try_success'] == 1]
with_replans = df[df['replan_count'] > 0]

print("First try success rate:", first_try['final_score'].mean())
print("After replanning:", with_replans['final_score'].mean())

Ablation 3: Impact of Tool Efficiency

# Quartile analysis
df['efficiency_quartile'] = pd.qcut(df['tool_efficiency'], q=4, labels=['Q1', 'Q2', 'Q3', 'Q4'])
df.groupby('efficiency_quartile')['final_score'].agg(['mean', 'std', 'count'])

Correlation Analysis

import seaborn as sns
import matplotlib.pyplot as plt

# Select key metrics for correlation
metrics = ['final_score', 'analyst_score', 'autonomy_score', 
           'tool_efficiency', 'verification_score', 'data_understanding_score']

# Compute correlation matrix
corr = df[metrics].corr()

# Visualize
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, cmap='coolwarm', center=0, 
            square=True, linewidths=1)
plt.title('Metric Correlation Matrix')
plt.tight_layout()
plt.savefig('correlation_matrix.png', dpi=300)

Statistical Significance Testing

from scipy.stats import wilcoxon, mannwhitneyu

# Compare two systems (e.g., baseline vs. proposed)
baseline_df = pd.read_csv('baseline_run/task_metrics.csv')
proposed_df = pd.read_csv('proposed_run/task_metrics.csv')

# Paired test (same tasks)
merged = baseline_df.merge(proposed_df, on='task_id', suffixes=('_baseline', '_proposed'))
stat, p_value = wilcoxon(merged['final_score_baseline'], merged['final_score_proposed'])
print(f"Wilcoxon signed-rank test: p={p_value:.4f}")

# Effect size (Cohen's d)
mean_diff = merged['final_score_proposed'].mean() - merged['final_score_baseline'].mean()
pooled_std = np.sqrt((merged['final_score_proposed'].std()**2 + 
                       merged['final_score_baseline'].std()**2) / 2)
cohens_d = mean_diff / pooled_std
print(f"Effect size (Cohen's d): {cohens_d:.3f}")

Failure Analysis for Papers

# Failure distribution (Figure 2)
failure_dist = df[df['final_score'] < 0.8]['failure_category'].value_counts()
plt.figure(figsize=(10, 6))
failure_dist.plot(kind='bar')
plt.xlabel('Failure Category')
plt.ylabel('Count')
plt.title('Failure Distribution')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.savefig('failure_distribution.png', dpi=300)

# Root cause analysis (Table 5)
root_causes = df[df['final_score'] < 0.8]['root_cause'].value_counts().head(10)
print(root_causes)

Reporting Template

Results Section:

We evaluate our system on the DABench benchmark containing 50 tasks of varying 
difficulty (Easy: 15, Medium: 20, Hard: 15). Our system achieves a mean final 
score of X.XX ± Y.YY (mean ± std), significantly outperforming the baseline 
(p < 0.001, Wilcoxon signed-rank test). The analyst score, a composite metric 
combining autonomy (weight=0.4), tool efficiency (weight=0.3), and verification 
quality (weight=0.3), reaches Z.ZZ ± W.WW.

Breakdown by difficulty reveals consistent performance across all levels:
- Easy: X1 ± Y1 (n=15)
- Medium: X2 ± Y2 (n=20)
- Hard: X3 ± Y3 (n=15)

Our multi-agent architecture demonstrates strong autonomy with AA% first-try 
success rate and an average of B.B replanning operations per task. Tool 
efficiency reaches C.C ± D.D, indicating effective tool selection. Verification 
mechanisms trigger in VV% of executions and detect EE failures, contributing to 
improved final scores.

Failure analysis (Figure 2) shows the primary failure categories are:
1. TOOL_EXECUTION_FAILURE (XX%)
2. DATA_UNDERSTANDING_FAILURE (YY%)
3. SCHEMA_MISMATCH_FAILURE (ZZ%)

🎓 Experimental Methodology

Dataset Preparation

  1. Task Selection: Use stratified sampling by difficulty
  2. Data Splits: Train/Val/Test or K-fold cross-validation
  3. Seed Control: Fix random seeds for reproducibility
# Stratified sampling
from sklearn.model_selection import train_test_split

df = pd.read_csv('all_tasks.csv')
train, test = train_test_split(df, test_size=0.3, 
                                 stratify=df['difficulty'], 
                                 random_state=42)

Baseline Comparisons

Recommended Baselines:

  1. Random tool selection
  2. Fixed planning strategy
  3. No verification
  4. Single-agent (no specialists)
  5. Prior work (if available)

Reproducibility

Report:

  • Hardware (GPU type, RAM)
  • Software versions (Python, LLM API version)
  • Random seeds
  • Hyperparameters
  • Number of runs (recommend 3-5 for variance)

Provide:

  • Code repository
  • Trained model weights (if applicable)
  • Full evaluation CSVs (task_metrics.csv, trajectory.csv)
  • Configuration files

Ethical Considerations

  • Data privacy: Ensure benchmark tasks don't contain PII
  • Computational cost: Report total compute time and carbon footprint
  • Failure modes: Document dangerous failure patterns
  • Limitations: Clearly state what the system cannot do

�🔍 Quick Analysis Examples

Load and Analyze

import pandas as pd

# Load metrics
df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")

# Success rate
success_rate = (df['final_score'] >= 0.8).mean()
print(f"Success rate: {success_rate:.1%}")

# By difficulty
print("\nScores by difficulty:")
print(df.groupby("difficulty")[["final_score", "analyst_score"]].mean())

# Failed tasks
failed = df[df['final_score'] < 0.8]
print(f"\nFailed: {len(failed)} tasks")
print(failed[['task_id', 'final_score', 'failure_category', 'root_cause']])

Debug Failed Task

import json

# Load replay artifact
with open("artifacts/runs/<RUN_ID>/task_38/task_replay.json") as f:
    replay = json.load(f)

# Check failure
if replay['failure_attribution']:
    fa = replay['failure_attribution']
    print(f"Category: {fa['failure_category']}")
    print(f"Root cause: {fa['root_cause']}")
    print(f"Stage: {fa['failure_stage']}")
    print(f"Reason: {fa['failure_reason']}")
    print(f"Suggested fix: {fa['suggested_fix']}")

# Review execution
print(f"\nFinal score: {replay['evaluation_result']['final_score']}")
print(f"Coordinator decision: {replay['coordinator_final_decision']}")
print(f"Verification: {replay['verification_passed']}")

View Trajectory

import pandas as pd

# Load trajectory for specific task
traj = pd.read_csv("artifacts/runs/<RUN_ID>/trajectory.csv")
task_traj = traj[traj['task_id'] == 'task_38']

# View execution flow
print(task_traj[['step_id', 'phase', 'agent', 'tool', 'success', 'tokens']])

# Analyze failures
failures = task_traj[~task_traj['success']]
print(f"\nFailures: {len(failures)}")
print(failures[['step_id', 'tool', 'observation']])

🛠️ Hardening Features (Integrated)

1. Artifact Reconciliation

Validates:

  • ✅ Tool call counts match across artifacts
  • ✅ Token counts reconcile (trajectory vs metrics)
  • ✅ Verification semantics consistent (verification_passed ↔ coordinator_final_decision)
  • ✅ Time accounting (wall clock ≥ component time)
  • ✅ Trajectory completeness

Report: artifact_reconciliation_report.txt

2. Replay Artifacts

Complete debug snapshots per task:

  • Question & context
  • All execution attempts (plan, code, stdout, stderr)
  • Agent executions (MAS observability)
  • Tool calls with success/failure
  • Coordinator decisions
  • Verifier output
  • Final answer
  • Evaluation result
  • Structured failure attribution

Location: task_*/task_replay.json

3. Engineering Health Report

System diagnostics:

  • Reconciliation pass/fail status
  • Invariant violations (verification, tool calls, tokens)
  • Time accounting gaps (overhead analysis)
  • Failure taxonomy distribution
  • Top recurring root causes
  • System health indicators
  • Health score (0-10)

Report: engineering_health_report.txt


🏥 Health Score Interpretation

Score Status Action
9-10 ✅ HEALTHY Ready to use
7-8 ⚠️ GOOD Review warnings
5-6 ⚠️ FAIR Fix issues before publication
3-4 ❌ POOR Investigation required
0-2 ❌ UNHEALTHY Do not use

🔧 Implementation Details

Evaluation Pipeline Architecture

┌─────────────────────────────────────────────────────────────┐
│                     Evaluation Harness V2                   │
└─────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│                      Data Collection                         │
│  • Load trace.json (agent execution trace)                  │
│  • Load prediction.csv (agent output)                       │
│  • Load gold.csv (ground truth)                             │
│  • Load task.json (metadata)                                │
└─────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│                    Metric Computation                        │
│  • Correctness (multi-level F1: answer/column/row)          │
│  • Autonomy (first-try success, replans, autonomy score)    │
│  • Planning (revisions, dead ends, alignment)               │
│  • Tool Usage (diversity, efficiency, selection accuracy)   │
│  • Data Understanding (tables/columns, exploration)         │
│  • Verification (triggered, passed, failures detected)      │
│  • Recovery (detected, stage, success, depth)               │
│  • Trajectory (length, branches, efficiency, patterns)      │
│  • Failure Taxonomy (category, root cause, severity)        │
│  • Confidence & Calibration (score, error, bucket)          │
│  • Composite Scores (analyst_score = weighted blend)        │
└─────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│                   Normalized Storage (3 CSV Files)           │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ task_metrics.csv (Primary Evaluation Table)           │  │
│  │ • One row per task execution                          │  │
│  │ • 100+ columns covering all metric dimensions         │  │
│  │ • Granularity: Task-level                             │  │
│  └───────────────────────────────────────────────────────┘  │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ trajectory.csv (Trajectory Trace)                     │  │
│  │ • One row per trajectory step                         │  │
│  │ • Enables process mining and step-level debugging     │  │
│  │ • Granularity: Step-level                             │  │
│  └───────────────────────────────────────────────────────┘  │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ tool_calls.csv (Tool Usage Analysis)                  │  │
│  │ • One row per tool invocation                         │  │
│  │ • Tracks latency, tokens, retries, errors             │  │
│  │ • Granularity: Tool-call-level                        │  │
│  └───────────────────────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│              Evaluation Hardening Suite (Integrated)         │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ 1. Artifact Reconciliation                            │  │
│  │    • Validates CSV consistency                        │  │
│  │    • Enforces invariants                              │  │
│  │    • Generates validation report                      │  │
│  └───────────────────────────────────────────────────────┘  │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ 2. Replay Artifact Generation                         │  │
│  │    • Complete debug snapshots per task                │  │
│  │    • Structured failure attribution                   │  │
│  │    • MAS observability tracking                       │  │
│  └───────────────────────────────────────────────────────┘  │
│  ┌───────────────────────────────────────────────────────┐  │
│  │ 3. Engineering Health Report                          │  │
│  │    • System health diagnostics                        │  │
│  │    • Time accounting analysis                         │  │
│  │    • Health score (0-10)                              │  │
│  └───────────────────────────────────────────────────────┘  │
└─────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────┐
│              Terminal Visualization (3 Modes)                │
│  • Standard: Core metrics + summary                         │
│  • Verbose: + Agent behavior analysis                       │
│  • Research: + All metrics for papers (mean, std)           │
└─────────────────────────────────────────────────────────────┘

Data Model Schema

1. task_metrics.csv (Primary Evaluation Table)

Purpose: Comprehensive per-task evaluation metrics for academic publication.

Granularity: One row per task execution (50-500 tasks typical).

Column Count: 100+ columns organized into 14 categories.

Schema Categories:

  1. Identification (5 cols): run_id, task_id, trace_id, difficulty, timestamp
  2. Task Metadata (5 cols): task_type, source_count, source_types, requires_cross_source_reasoning, ground_truth_available
  3. Correctness (15 cols): final_score, answer_precision/recall/f1, column_precision/recall/f1, row_precision/recall/f1, matched_columns, pred_rows/cols, gold_rows/cols
  4. Execution (6 cols): execution_success, execution_time, total_tokens, llm_calls, tool_calls, tool_failures
  5. Autonomy (5 cols): first_try_success, replan_count, user_intervention_count, autonomy_score, coordinator_interventions
  6. Planning (5 cols): planner_steps, planner_revisions, planner_dead_ends, plan_execution_alignment, plan_attempts
  7. Tool Usage (10 cols): unique_tools_used, tool_diversity, useful_tool_calls, wasted_tool_calls, tool_efficiency, tool_retry_count, tool_selection_accuracy
  8. Data Understanding (7 cols): tables_discovered, columns_discovered, relevant_tables_found, relevant_columns_found, schema_exploration_steps, data_understanding_score
  9. Verification (6 cols): verification_triggered, verification_steps, verification_passed, verification_failures_detected, verification_score, aat_verification_passed
  10. Recovery (6 cols): failure_detected, failure_stage, recovery_success, recovery_depth, recovery_success_rate
  11. Trajectory (7 cols): trajectory_length, trajectory_summary, branching_factor, max_execution_depth, critic_loops, trajectory_efficiency
  12. Failure Taxonomy (6 cols): failure_category, root_cause, recoverable_failure, severity, failure_reason, failure_agent
  13. Confidence (5 cols): confidence_score, confidence_correct, confidence_error, calibration_bucket
  14. Composite (1 col): analyst_score
  15. AAT Architecture (10 cols): coordinator_calls, coordinator_final_decision, coordinator_checkpoints, schema_agent_called, domain_agent_called, document_agent_called, specialist_participation
  16. Per-Stage Metrics (25 cols): understanding_time/calls/tokens, planning_time/calls/tokens, execution_time/calls/tokens, verification_time/calls/tokens, summary_time/calls/tokens
  17. Per-Action Metrics (28 cols): list_context_calls/time, read_json_calls/time, read_knowledge_calls/time, execute_python_calls/time, etc.

Total: 121 columns

2. trajectory.csv (Step-by-Step Trace)

Purpose: Detailed execution trace for process mining and debugging.

Granularity: One row per trajectory step (10-50 steps per task typical).

Columns (24):

  • Identification: run_id, task_id, trace_id, step_id, parent_step_id
  • Execution Context: stage, legacy_stage, phase, agent, trajectory_agent, agent_type
  • Action: tool, action, coordinator_checkpoint
  • Decision: confidence, decision, review_type, verification_status, specialist_selected
  • Outcome: success, duration_seconds, tokens, retries
  • Timing: timestamp, elapsed_seconds
  • Details: thought, action_input, observation, raw_response, metadata_json

Use Cases:

  • Process mining (find common execution patterns)
  • Step-level debugging (identify exact failure point)
  • Agent behavior analysis (tool selection patterns)
  • Performance profiling (step latencies)

3. tool_calls.csv (Tool-Level Analysis)

Purpose: Per-tool-invocation metrics for optimization.

Granularity: One row per tool call (matches trajectory tool steps).

Columns (15):

  • Identification: run_id, task_id, trace_id, step_id, call_id
  • Tool: tool_name, tool_category
  • Performance: success, latency_seconds, tokens, retry_count
  • Data: input_size, output_size
  • Errors: error_type, error_message
  • Metadata: timestamp, metadata_json

Use Cases:

  • Tool efficiency analysis (which tools are slow?)
  • Error rate tracking (which tools fail most?)
  • Token cost analysis (which tools are expensive?)
  • Tool selection optimization (which tools work best for what?)

4. task_replay.json (Debug Snapshot - Per Task)

Purpose: Complete execution context for offline debugging without re-running.

Granularity: One JSON file per task.

Structure:

{
  "task_id": "task_38",
  "trace_id": "...",
  "run_id": "...",
  "difficulty": "Medium",
  "question": "...",
  "available_sources": [...],
  
  "execution_attempts": [
    {
      "attempt_number": 1,
      "plan": {...},
      "code_executed": "...",
      "stdout": "...",
      "stderr": "...",
      "success": false
    }
  ],
  
  "agent_executions": [
    {
      "agent_name": "StrategicCoordinator",
      "checkpoint": "UNDERSTANDING",
      "decision": "PROCEED",
      "confidence": 0.85,
      "duration": 5.2,
      "tokens": 1500
    }
  ],
  
  "coordinator_decisions": [...],
  "verifier_output": {...},
  "final_answer_csv": "...",
  "evaluation_result": {
    "final_score": 0.0,
    "answer_f1": 0.0
  },
  
  "failure_attribution": {
    "failure_category": "TOOL_EXECUTION_FAILURE",
    "root_cause": "KeyError: 'column_name'",
    "failure_stage": "execution",
    "failure_agent": "Executor",
    "evidence_trace_step_ids": [12, 14],
    "suggested_fix": "Validate column existence"
  }
}

Module Structure

src/data_agent_baseline/langgraph_agent/
├── eval_v2.py                       # Main orchestrator (650 lines)
│   ├── evaluate_task_v2()           # Single task evaluation
│   ├── evaluate_run_v2()            # Full run evaluation
│   ├── write_evaluation_v2()        # CSV output
│   └── extract_trajectory/tools()   # Trace extraction
│
├── eval_v2_metrics.py               # Metric computations (1280 lines)
│   ├── compute_correctness()        # Multi-level F1
│   ├── compute_autonomy()           # Independence metrics
│   ├── compute_planning()           # Planning quality
│   ├── compute_tool_usage()         # Tool efficiency
│   ├── compute_data_understanding() # Schema understanding
│   ├── compute_verification()       # Quality assurance
│   ├── compute_recovery()           # Error recovery
│   ├── compute_trajectory()         # Execution path analysis
│   ├── compute_failure_taxonomy()   # Error classification
│   ├── compute_confidence()         # Calibration metrics
│   └── compute_analyst_score()      # Composite score
│
├── eval_v2_viz.py                   # Visualization (580 lines)
│   ├── render_evaluation_report()   # Main report
│   ├── render_task_table()          # Per-task table
│   ├── render_summary_sections()    # Analysis sections
│   └── render_verbose_task_detail() # Task drill-down
│
├── eval_artifact_reconciliation.py # Validation (435 lines)
│   ├── ArtifactReconciliator        # Cross-artifact validation
│   ├── validate_run()               # Run-level validation
│   └── format_report()              # Validation report
│
├── eval_replay_artifacts.py        # Debug snapshots (630 lines)
│   ├── ReplayArtifactGenerator      # Snapshot generation
│   ├── FailureAttributor            # Root cause analysis
│   └── generate_all_replays()       # Batch generation
│
└── eval_health_report.py            # Health diagnostics (490 lines)
    ├── EngineeringHealthReport      # Health metrics
    ├── TimeAccountingGap            # Overhead analysis
    └── generate_health_report()     # Report generation

Key Files

Core Evaluation:

  • eval_v2.py - Main evaluation engine (scoring, metrics)
  • eval_v2_metrics.py - Individual metric calculators
  • eval_v2_viz.py - Rendering & visualization

Hardening Suite:

  • eval_artifact_reconciliation.py - Cross-artifact validation
  • eval_replay_artifacts.py - Debug snapshot generation
  • eval_health_report.py - System health diagnostics

CLI Integration:

  • cli.py:eval_v2_command() - Orchestrates entire pipeline

Critical Invariants

# Verification consistency
aat_verification_passed = 1 ↔ coordinator_final_decision = "APPROVE_FINAL"
# Note: critic_verification_passed is independent of AAT coordinator decision

# Tool call reconciliation
len(tool_calls_df) == task_metrics['tool_calls']
tool_failures_count == task_metrics['tool_failures']

# Token conservation
abs(trajectory_tokens - metrics_tokens) < 100

# Time accounting (with overhead)
wall_clock_time = execution_time  # End-to-end elapsed
component_compute_time = sum(agent/tool tracked durations)
unaccounted_overhead = wall_clock_time - component_compute_time
time_accounting_ratio = component_compute_time / wall_clock_time

# Expected: component_time < wall_clock_time (overhead exists)
# Framework overhead, I/O wait, concurrency gaps, logging typically 20-50%
# Warning only if time_accounting_ratio < 0.05 (95% unaccounted)
# Info if component_time > wall_clock_time (indicates concurrent execution)

Comparison with Previous Systems

Feature Traditional Eval DABench V1 DABench V2 (Ours)
Storage Model Single CSV Single CSV 3 normalized CSVs
Metrics 5-10 basic ~30 metrics 100+ comprehensive
Correctness Overall F1 Overall F1 Multi-level F1 (answer/column/row)
Autonomy Not measured Retry count Composite autonomy score
Planning Not measured Step count Revisions, alignment, dead ends
Tool Analysis Call count Call count Efficiency, diversity, selection accuracy
Trajectory Not captured Basic steps Full trace with branching, patterns
Failure Analysis Error message Simple bucket Structured taxonomy + root cause
Verification Not measured Basic check Multi-stage verification score
Confidence Not measured Not measured Calibration metrics
Composite Scores None None Analyst score (weighted)
Debugging Manual Manual Automated replay artifacts
Validation None Basic Comprehensive reconciliation
Process Mining Not supported Not supported Full trajectory CSV
MAS Observability Not supported Not supported Per-agent tracking
Time Accounting Wall clock only Wall clock only Component + overhead

Key Innovations:

  1. Multi-level correctness: Separate precision/recall/F1 at answer/column/row levels
  2. Autonomy quantification: First-try success, replan count, composite autonomy score
  3. Composite analyst score: Weighted blend of autonomy, efficiency, verification (suitable for ranking)
  4. Structured failure taxonomy: 10 failure categories with root cause attribution
  5. Integrated validation: Automatic artifact reconciliation with invariant enforcement
  6. Debug-ready artifacts: Complete replay snapshots for offline debugging
  7. Process mining support: Full trajectory CSV for pattern discovery
  8. MAS observability: Per-agent execution tracking in multi-agent systems

🎯 Failure Attribution

Failure Categories

PLANNING_FAILURE             # Bad plan generation
DATA_UNDERSTANDING_FAILURE   # Misunderstood data/schema
TOOL_SELECTION_FAILURE       # Wrong tool chosen
TOOL_EXECUTION_FAILURE       # Tool crashed/errored
REASONING_FAILURE            # Logic errors
VERIFICATION_FAILURE         # Verifier malfunction
AGGREGATION_FAILURE          # Data aggregation errors
ANSWER_FORMAT_FAILURE        # Wrong output format
SCHEMA_MISMATCH_FAILURE      # Schema mismatch
UNKNOWN_FAILURE              # Needs manual inspection

Attribution Structure

{
  "failure_category": "TOOL_EXECUTION_FAILURE",
  "root_cause": "KeyError: 'column_name'",
  "failure_stage": "execution",
  "failure_agent": "Executor",
  "failure_reason": "Attempted to access non-existent column",
  "evidence_trace_step_ids": [12, 14],
  "suggested_fix": "Validate column existence before access",
  "evidence_summary": "Step 12: execute_python failed with KeyError",
  "execution_success": false,
  "tool_failure_count": 1
}

📖 Common Workflows

1. Evaluate New Run

# Run prediction first (if not already done)
dabench run input_full --agent aat

# Evaluate
dabench eval-v2 <RUN_ID> --mode research

# Review health
cat artifacts/runs/<RUN_ID>/engineering_health_report.txt

2. Debug Failed Tasks

# Find failed tasks
python3 -c "
import pandas as pd
df = pd.read_csv('artifacts/runs/<RUN_ID>/task_metrics.csv')
failed = df[df['final_score'] < 0.8]
print(failed[['task_id', 'failure_category', 'root_cause']])
"

# Debug specific task
cat artifacts/runs/<RUN_ID>/task_38/task_replay.json | jq '.failure_attribution'

3. Compare Runs

import pandas as pd

# Load two runs
run1 = pd.read_csv("artifacts/runs/RUN_A/task_metrics.csv")
run2 = pd.read_csv("artifacts/runs/RUN_B/task_metrics.csv")

# Merge on task_id
merged = run1.merge(run2, on='task_id', suffixes=('_A', '_B'))

# Compare
print(f"Run A mean: {merged['final_score_A'].mean():.3f}")
print(f"Run B mean: {merged['final_score_B'].mean():.3f}")

# Tasks improved in Run B
improved = merged[merged['final_score_B'] > merged['final_score_A']]
print(f"\nImproved: {len(improved)} tasks")

4. Generate Paper Figures

import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")

# Score distribution
plt.figure(figsize=(10, 6))
plt.hist(df['final_score'], bins=20, edgecolor='black')
plt.xlabel('Final Score')
plt.ylabel('Count')
plt.title('Score Distribution')
plt.savefig('score_distribution.png')

# Autonomy vs Efficiency
plt.figure(figsize=(10, 6))
plt.scatter(df['autonomy_score'], df['tool_efficiency'], 
            c=df['final_score'], cmap='viridis')
plt.xlabel('Autonomy Score')
plt.ylabel('Tool Efficiency')
plt.colorbar(label='Final Score')
plt.savefig('autonomy_vs_efficiency.png')

🧪 Testing

Run Test Suite

cd /workspace/ainn-cm-poc-data-agent

# Test evaluation harness
pytest tests/test_eval_harness.py -v

# Test specific validation
pytest tests/test_eval_harness.py::TestVerificationConsistency -v

Manual Validation

# Re-validate existing run
python3 -c "
from pathlib import Path
from src.data_agent_baseline.langgraph_agent.eval_artifact_reconciliation import validate_evaluation_run

passed, report = validate_evaluation_run(Path('artifacts/runs/<RUN_ID>'))
print(report)
print(f'\nPassed: {passed}')
"

🐛 Troubleshooting

Issue: "Evaluation produced inconsistent metrics"

Cause: Old consistency validator found errors

Solution: Check artifact_reconciliation_report.txt for details:

cat artifacts/runs/<RUN_ID>/artifact_reconciliation_report.txt

Common issues:

  • verification_outcome_mismatch: Verification flag doesn't match coordinator decision
  • tool_call_count_mismatch: Tool calls CSV doesn't match metrics
  • data_understanding_inflation: Score exceeds theoretical maximum

Issue: Health score < 7

Cause: System detected quality issues

Solution: Review engineering_health_report.txt:

cat artifacts/runs/<RUN_ID>/engineering_health_report.txt

Look for:

  • Time accounting gaps > 50%
  • High tool failure rate > 10%
  • Missing failure attribution

Issue: Missing trajectory duration

Symptom: Warnings about "missing time data"

Cause: Trajectory extraction didn't populate duration_seconds column

Impact: Time validation skipped (not critical)


📚 Additional Documentation

Architecture Details:

  • See src/data_agent_baseline/langgraph_agent/eval_v2.py for scoring logic
  • See src/data_agent_baseline/langgraph_agent/eval_v2_metrics.py for metric definitions

Test Coverage:

  • See tests/test_eval_harness.py for validation tests

CLI Integration:

  • See src/data_agent_baseline/cli.py:eval_v2_command() for integration

🔄 Version History

V2.1 (Current) - June 14, 2026

Phase 2: MAS Debugging Enhancements

Focused improvements for multi-agent system debugging and actionable diagnostics:

  1. Deterministic Failure Attribution: Maps evaluation buckets (low_recall, wrong_schema, etc.) to structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, AGGREGATION_FAILURE)

    • Eliminates UNKNOWN_FAILURE when bucket exists
    • Infers failure_stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
    • Identifies failure_agent (Schema Agent, Planner, Executor, etc.)
  2. Separated Health Assessments: Clear distinction between infrastructure health and run quality

    • Harness Health: Infrastructure metrics (reconciliation pass rate, validator errors, attribution coverage)
    • Run Quality: Outcome metrics (answer accuracy, execution success rate)
    • Status thresholds: HEALTHY (9.0+, no errors), DEGRADED (7.0+), UNHEALTHY (else)
  3. Enhanced Visualization:

    • Per-task table: Renamed "Succ" → "Exec" to clarify execution vs. correctness
    • Added "Root Cause" column showing failure diagnostics (filter_logic_error, schema_misunderstanding, etc.)
    • New "MAS Failure Analysis" section with:
      • MAS Failure Categories table (structured categories: REASONING_FAILURE, etc.)
      • Failure Distribution by Stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
      • Outcome Error Types (evaluation buckets)
  4. Partial-Correct Task Handling: New outcome_status field distinguishes:

    • "correct": final_score ≥ 0.8
    • "partial": succeeded but 0 < final_score < 0.8
    • "failed": final_score < 0.8 or execution failure
  5. Improved Terminology: Renamed "time gaps" → "Unaccounted Overhead Time" with clear explanation (framework overhead, I/O wait, async queuing)

  6. Complete Replay Artifacts: Every task_replay.json includes:

    • Full failure diagnostics (failure_category, root_cause, failure_stage, failure_agent)
    • Run context (harness_health_status, run_quality_status, outcome_status)
    • Time accounting (unaccounted_overhead_time_seconds, time_accounting_ratio)

Design Philosophy: Optimized for debugging and MAS improvement, not paper metrics. All changes maintain backward compatibility.

Phase 2.1 Cleanup (Jan 2026):

  • Separated harness health (infrastructure) from run quality (outcomes)
  • Fixed MAS Failure Categories display to show structured categories instead of buckets
  • Added outcome_status field for partial-correct handling
  • Enhanced replay artifacts with health/quality context
  • Improved time accounting terminology

Phase 0: MAS Debugging Enhancements - Final Round (Jan 2026)

Goal: Make evaluation harness maximally useful for all future phases (Baseline ReAct, MAS, DAG Visualization, Replay/Time Travel, Confidence & Verification, Research/Ablation Studies).

Key Improvements:

  1. MAS Recovery Effectiveness Metrics (Task 2):

    • New fields: initial_answer_correct, final_answer_correct, recovered_after_replan, recovered_after_retry
    • Display: "MAS Effectiveness" section showing:
      • First Attempt Accuracy: Initial correct rate
      • Final Accuracy: Final correct rate
      • Recovered Tasks: Count of tasks improved through MAS interventions
      • MAS Recovery Gain: Percentage improvement (e.g., +14%)
    • Impact: Directly answers "Did MAS actually improve answers?"
  2. Replan/Retry Effectiveness Tracking (Task 3):

    • New fields: replan_requested, replan_successful, retry_requested, retry_successful
    • Display: "Coordinator Intervention Effectiveness" section showing:
      • Replans: Requested count, successful count, success rate
      • Retries: Requested count, successful count, success rate
    • Impact: Shows which coordinator interventions actually help
  3. Specialist Agent Value Analysis (Task 4):

    • Existing fields: schema_agent_used, domain_agent_used, document_agent_used
    • Display: "Specialist Agent Value Analysis" section showing for each agent:
      • Tasks Used
      • Accuracy With Agent
      • Accuracy Without Agent
      • Impact (delta %)
    • Impact: Automatic ablation showing which specialists add value
  4. Expanded Failure Stage Taxonomy (Task 5):

    • Updated FailureStage enum: UNDERSTAND, PLAN, EXECUTE, VERIFY, AGGREGATE
    • Replaced coarse stages (EXPLORATION, PLANNING, EXECUTION) with AAT-aligned taxonomy
    • Impact: Finer-grained debugging for AAT phase-specific failures
  5. Cost by Difficulty (Task 6):

    • Difficulty breakdown now includes Mean Runtime and Mean Tokens columns
    • Impact: Required for Baseline vs MAS vs Future comparisons
  6. Verification Timeline Clarity (Task 7):

    • Separated "Execution Approval" (coordinator decision) from "Ground Truth Result" (evaluation correctness)
    • Added explanatory note distinguishing verification from correctness
    • Impact: Eliminates confusion between process approval and actual correctness
  7. Comprehensive CSV Storage (Task 8):

    • All new fields stored in task_metrics.csv: outcome_error_type, MAS recovery fields, specialist usage
    • Impact: Future phases can run aggregations without parsing replay artifacts
  8. Removed Duplicate Reporting (Task 1):

    • Eliminated redundant "Failure Categories" section
    • Kept: "MAS Failure Categories" (structured) and "Outcome Error Types" (buckets)

Design Philosophy: Every change focused on making the evaluation harness more actionable for debugging and improving the MAS, not for paper-writing. Provides automatic ablation studies and directly answers key questions about MAS effectiveness.

V2.0 - June 2026

  • ✅ Integrated hardening suite (automatic reconciliation, replay, health)
  • ✅ 100+ comprehensive metrics for KDD Creative Track
  • ✅ AAT architecture observability
  • ✅ Structured failure attribution
  • ✅ MAS-aware trajectory extraction
  • ✅ Time accounting with overhead tracking

V1 (Legacy)

  • Basic metrics (precision, recall, F1)
  • Manual validation required
  • Limited debugging support

📝 Summary

One command does it all:

dabench eval-v2 <RUN_ID> --mode standard

Automatically provides:

  • ✅ Comprehensive metrics (100+ columns)
  • ✅ Complete validation & reconciliation
  • ✅ Full debug snapshots (replay artifacts)
  • ✅ System health diagnostics
  • ✅ Failure attribution & root cause analysis

No manual steps required.