DAO_kdd26 / docs /implementation /EVALUATION.md
sipe5001's picture
Add Hugging Face Docker Space configuration
d3d0e0e
|
Raw
History Blame Contribute Delete
55.9 kB
# DABench Evaluation System
Complete guide to the evaluation harness with integrated hardening features.
---
## ⚠️ Recent Fixes (June 2026)
**Validation and Reporting Improvements:**
- Fixed `compute_verification_passed()` bug: returns boolean (0/1), resolving `verification_passed=2` errors.
- Coordinator consistency reconciliation (`replan_count_consistent`, `retry_count_consistent`, `coordinator_metrics_consistent`) is now diagnostic-only warning output, not a hard eval-v2 failure condition.
- Meaningful disagreement reporting is hardened to reduce confidence-only inflation.
- **Result:** eval-v2 completes across modes with reconciliation diagnostics reported as warnings when present.
---
## 🚀 Quick Start
### Run Evaluation (with automatic hardening)
```bash
cd /data3/dataFAIR/kdd-dev/public
# Standard mode (quick overview)
dabench eval-v2 <RUN_ID> --mode standard
# Verbose mode (detailed analysis)
dabench eval-v2 <RUN_ID> --mode verbose
# Research mode (all metrics for papers)
dabench eval-v2 <RUN_ID> --mode research
```
**What runs automatically:**
1. ✅ Standard evaluation (task_metrics, trajectory, tool_calls CSVs)
2. ✅ Artifact reconciliation validation
3. ✅ Replay artifact generation (debug snapshots)
4. ✅ Engineering health report
**Report Display Features:**
- **Separated Health Assessments**:
- **Harness Health** (9.7/10 ✓ HEALTHY): Infrastructure quality (reconciliation, validators, attribution coverage)
- **Run Quality** (⚠ DEGRADED 64% accuracy): Outcome metrics (answer accuracy, execution success)
- **Per-Task Results Table**: Shows all tasks with execution success, answer quality, timing, and trajectory
- **Exec** column: Execution success (✓ = code ran, ✗ = crash)
- **Root Cause** column: Displays failure diagnostics for failed tasks (e.g., filter_logic_error, schema_misunderstanding)
- **Overall Summary**: Distinguishes execution success (code ran) from answer accuracy (correct results)
- **Difficulty Breakdown**: Three-column view:
- *Execution Success Rate*: % tasks that ran without crashes
- *Answer Accuracy*: % tasks with final_score ≥ 0.8
- *Mean Final Score*: Average correctness
- *Mean Runtime*: Average execution time per difficulty
- *Mean Tokens*: Average token usage per difficulty
- **MAS Effectiveness**: Shows first attempt vs final accuracy and recovery gain
- **Coordinator Intervention Effectiveness**: Replan and retry success rates
- **Specialist Agent Value Analysis**: Automatic ablation showing agent impact
- **MAS Failure Analysis**: Debugging-focused breakdown showing:
- **MAS Failure Categories**: Structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, etc.)
- **Failure Distribution by Stage**: AAT phases (UNDERSTAND/PLAN/EXECUTE/VERIFY/AGGREGATE)
- **Outcome Error Types**: Evaluation buckets (low_recall, wrong_schema, etc.)
- **AAT Metrics**: Coordinator decisions, specialist activation, verification outcomes
- **Analyst Team Summary (verbose/research)**:
- mean agreement score
- tasks with meaningful disagreement
- disagreement type distribution
- coordinator override count
- verifier disagreement count
- critical disagreement score distribution (0-3)
- auditor trigger totals + precision/recall (warning->failure, failure->warning)
- **Verification Timeline**: Separates execution approval from ground truth correctness
- **Phase Timing**: UNDERSTAND → PLAN → EXECUTE → VERIFY → SUMMARIZE with reconciliation view
---
## 📁 Generated Artifacts
Every `eval-v2` run produces:
```
artifacts/runs/<RUN_ID>/
├── task_metrics.csv # Per-task metrics (100+ columns)
├── trajectory.csv # Step-by-step execution trace
├── tool_calls.csv # Per-tool-call analysis
├── comprehensive_evaluation.csv # Backward compatibility
├── artifact_reconciliation_report.txt # Validation results
├── engineering_health_report.txt # System health diagnostics (includes answer accuracy, attribution coverage)
├── auditor_validation_report.md # Auditor effectiveness and failure-prevention diagnostics
└── task_*/
├── trace.json # Raw execution trace
├── answer.csv # Generated answer
└── task_replay.json # Complete debug context with failure attribution
```
---
## 📊 Key Metrics (100+ Total)
The evaluation system provides comprehensive metrics across 11 dimensions suitable for academic publication.
### 1. Correctness Metrics (Multi-Level F1)
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `answer_precision` | matched_cells / pred_cells | [0, 1] | Cell-level precision |
| `answer_recall` | matched_cells / gold_cells | [0, 1] | Cell-level recall |
| `answer_f1` | 2·P·R/(P+R) | [0, 1] | Cell-level F1 score |
| `column_precision` | matched_cols / pred_cols | [0, 1] | Column-level precision |
| `column_recall` | matched_cols / gold_cols | [0, 1] | Column-level recall |
| `column_f1` | 2·P·R/(P+R) | [0, 1] | Column-level F1 score |
| `row_precision` | min(pred, gold) / pred | [0, 1] | Row-level precision |
| `row_recall` | min(pred, gold) / gold | [0, 1] | Row-level recall |
| `row_f1` | 2·P·R/(P+R) | [0, 1] | Row-level F1 score |
| `final_score` | From evaluator | [0, 1] | Legacy overall score |
### 2. Autonomy Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `first_try_success` | succeeded ∧ attempts=1 | {0, 1} | Success without retry |
| `replan_count` | max(0, plan_attempts - 1) | [0, ∞) | Number of replans |
| `autonomy_score` | 1.0 - replans/max_replans | [0, 1] | Independence measure |
| `coordinator_interventions` | coordinator_calls - 3 | [0, ∞) | Beyond baseline |
| `user_intervention_count` | 0 (autonomous) | 0 | Manual interventions |
### 3. Planning Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `planner_steps` | count(action="planner") | [0, ∞) | Planning invocations |
| `planner_revisions` | max(0, plan_attempts - 1) | [0, ∞) | Plan revisions |
| `planner_dead_ends` | execution_attempts - 1 | [0, ∞) | Failed plans |
| `plan_execution_alignment` | 1.0 if first_try else decay | [0, 1] | Plan-execution match |
### 4. Tool Usage Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `unique_tools_used` | \|{tools}\| | [0, ∞) | Tool diversity count |
| `tool_diversity` | unique_tools / total_calls | [0, 1] | Diversity ratio |
| `useful_tool_calls` | explore + final_exec | [0, ∞) | Contributory calls |
| `wasted_tool_calls` | total - useful | [0, ∞) | Non-contributory |
| `tool_efficiency` | useful / total | [0, 1] | Efficiency ratio |
| `tool_selection_accuracy` | (useful - failures) / total | [0, 1] | Selection quality |
| `tool_calls` | From trace | [0, ∞) | Total invocations |
| `tool_failures` | From trace | [0, ∞) | Failed invocations |
### 5. Data Understanding Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `tables_discovered` | From explore phase | [0, ∞) | Tables found |
| `columns_discovered` | From explore phase | [0, ∞) | Columns found |
| `relevant_tables_found` | Heuristic: all discovered | [0, ∞) | Relevant tables |
| `relevant_columns_found` | matched_columns | [0, ∞) | Relevant columns |
| `schema_exploration_steps` | count(phase="explore") | [0, ∞) | Exploration steps |
| `data_understanding_score` | (table_score + col_score) / 2 | [0, 1] | Composite score |
Formula for `data_understanding_score` (weighted composite):
```
# Specialist activation component
required_specialists = 2 + (1 if documents_required else 0) # schema, domain, [document]
activated_specialists = schema_used + domain_used + (document_used if docs_required else 0)
specialist_activation_score = activated_specialists / required_specialists
# Discovery component (from exploration phase)
table_score = relevant_tables_found / tables_discovered
column_score = relevant_columns_found / columns_discovered
discovery_score = (table_score + column_score) / 2
# Weighted formula (can exceed pure specialist score)
data_understanding_score = 0.5 * specialist_activation_score + 0.5 * discovery_score
# Bounds enforced: [0, 1]
```
**Note:** Score may exceed simple specialist calculation (e.g., 0.833 with 2/3 specialists if discovery_score is high).
### 6. Verification Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `verification_triggered` | 1 if critic steps > 0 | {0, 1} | Verification used |
| `verification_steps` | count(critic actions) | [0, ∞) | Verification count |
| `critic_verification_passed` | From critic trace | {0, 1} | **Critic** passed check |
| `critic_verification_score` | critic_passed / critic_checks | [0, 1] | Critic quality |
| `critic_failures_detected` | From critic | [0, ∞) | Critic issues detected |
| `aat_verification_triggered` | 1 if AAT verifier ran | {0, 1} | AAT verifier used |
| `aat_verification_passed` | From coordinator | {0, 1} | **AAT** verifier result |
| `aat_verification_score` | AAT verifier confidence | [0, 1] | AAT verification quality |
| `verification_passed` | **Legacy** (= critic_passed) | {0, 1} | Backward compat (critic) |
| `verification_score` | **Legacy** (= critic_score) | [0, 1] | Backward compat (critic) |
**Critical Distinction:**
- `critic_verification_passed`: Step-level critic checks (legacy ReAct agent)
- `aat_verification_passed`: Final AAT coordinator approval (multi-agent system)
- **They are independent**: Critic may pass but coordinator may still request retry
**Invariant:** `aat_verification_passed=1 ↔ coordinator_final_decision="APPROVE_FINAL"`
**Backward Compatibility:** `verification_passed` and `verification_score` maintain critic semantics for legacy comparisons.
### 7. Recovery Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `failure_detected` | From trace | {0, 1} | Failure occurred |
| `failure_stage` | First failure phase | str | Stage of failure |
| `recovery_success` | recovered after failure | {0, 1} | Recovery outcome |
| `recovery_depth` | attempts until success | [0, ∞) | Recovery iterations |
| `recovery_success_rate` | successes / attempts | [0, 1] | Recovery rate |
### 8. Trajectory Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `trajectory_length` | count(steps) | [0, ∞) | Total steps |
| `branching_factor` | avg(children per node) | [1, ∞) | Execution branches |
| `max_execution_depth` | max(depth in tree) | [0, ∞) | Deepest path |
| `critic_loops` | count(critic iterations) | [0, ∞) | Critic cycles |
| `trajectory_efficiency` | useful_steps / total_steps | [0, 1] | Step efficiency |
| `trajectory_summary` | Pattern string | str | Execution pattern |
Common trajectory patterns:
- `UNDERSTAND→PLAN→EXECUTE→VERIFY→SUMMARIZE`
- `UNDERSTAND→PLAN→EXECUTE→VERIFY→RETRY→SUMMARIZE`
- `UNDERSTAND→PLAN→REPLAN→EXECUTE→VERIFY→SUMMARIZE`
### 9. Failure Taxonomy
| Metric | Type | Description |
|--------|------|-------------|
| `failure_category` | enum | PLANNING / DATA_UNDERSTANDING / TOOL_EXECUTION / etc. |
| `root_cause` | str | Specific error cause |
| `severity` | enum | CRITICAL / HIGH / MEDIUM / LOW |
| `recoverable_failure` | bool | Can be recovered |
Categories:
- `PLANNING_FAILURE` - Bad plan generation
- `DATA_UNDERSTANDING_FAILURE` - Schema/data misunderstanding
- `TOOL_SELECTION_FAILURE` - Wrong tool chosen
- `TOOL_EXECUTION_FAILURE` - Tool crash/error
- `REASONING_FAILURE` - Logic errors
- `VERIFICATION_FAILURE` - Verifier malfunction
- `AGGREGATION_FAILURE` - Data aggregation errors
- `ANSWER_FORMAT_FAILURE` - Wrong output format
- `SCHEMA_MISMATCH_FAILURE` - Schema mismatch
- `UNKNOWN_FAILURE` - Needs manual inspection
### 10. Confidence & Calibration Metrics
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `confidence_score` | From agent output | [0, 1] | Agent confidence |
| `confidence_correct` | conf ≥ 0.8 ∧ correct | {0, 1} | High conf + correct |
| `confidence_error` | \|conf - correctness\| | [0, 1] | Calibration error |
| `calibration_bucket` | Binned by confidence | str | Calibration bin |
Calibration buckets: `very_low` (0-0.2), `low` (0.2-0.4), `medium` (0.4-0.6), `high` (0.6-0.8), `very_high` (0.8-1.0)
### 11. Composite Scores (For Ranking)
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `analyst_score` | w₁·autonomy + w₂·efficiency + w₃·verification | [0, 1] | Weighted composite |
**Analyst Score Formula:**
```
analyst_score = 0.4·autonomy_score + 0.3·tool_efficiency + 0.3·verification_score
```
This composite metric balances:
- **40% Autonomy**: Independence and minimal human intervention
- **30% Efficiency**: Effective tool usage and resource management
- **30% Verification**: Quality assurance and self-checking
### Execution Metrics
| Metric | Type | Description |
|--------|------|-------------|
| `execution_success` | bool | Code ran without crashes |
| `execution_time` | float | Wall clock time (seconds) |
| `total_tokens` | int | Total LLM tokens used |
| `llm_calls` | int | Number of LLM invocations |
| `tool_calls` | int | Number of tool invocations |
| `tool_failures` | int | Failed tool calls |
| `trajectory_length` | int | Number of execution steps |
### AAT Architecture Metrics
| Metric | Type | Description |
|--------|------|-------------|
| `coordinator_calls` | int | Strategic coordinator invocations |
| `coordinator_final_decision` | str | APPROVE_FINAL / RETRY_EXECUTION / REPLAN |
| `aat_verification_passed` | bool | AAT verifier approval |
| `schema_agent_called` | bool | Schema specialist invoked |
| `domain_agent_called` | bool | Domain specialist invoked |
| `document_agent_called` | bool | Document specialist invoked |
| `specialist_participation` | float | Proportion of specialists used |
### Research Metrics (KDD Paper)
| Metric | Formula | Range | Description |
|--------|---------|-------|-------------|
| `cross_source_reasoning_success` | From task metadata | {0, 1} | Multi-source reasoning |
| `explanation_quality_score` | From output | [0, 1] | Explanation quality |
| `reproducibility_score` | Deterministic replay | [0, 1] | Result stability |
---
## � Statistical Analysis for Papers
### Recommended Metrics for Publication
**Primary Metrics (Table 1 - Main Results):**
- `final_score` (correctness) - Mean ± Std
- `analyst_score` (composite) - Mean ± Std
- `answer_f1` (cell-level) - Mean ± Std
- `autonomy_score` - Mean ± Std
- `tool_efficiency` - Mean ± Std
**Breakdown by Difficulty (Table 2):**
*Report now distinguishes execution success from answer quality:*
- **Execution Success Rate**: Tasks that ran without crashes (execution_success=1)
- **Answer Accuracy**: Tasks with correct answers (final_score ≥ 0.8)
- **Mean Final Score**: Average correctness score
```python
# Compute breakdown
SUCCESS_THRESHOLD = 0.8
for difficulty in ['Easy', 'Medium', 'Hard', 'Extreme']:
subset = df[df['difficulty'] == difficulty]
print(f"{difficulty}:")
print(f" Execution success: {(subset['execution_success']==1).mean():.1%}")
print(f" Answer accuracy: {(subset['final_score']>=SUCCESS_THRESHOLD).mean():.1%}")
print(f" Mean score: {subset['final_score'].mean():.3f}")
```
**Breakdown by Task Type (Table 3):**
```python
df.groupby('task_type')[['final_score', 'analyst_score']].agg(['mean', 'std', 'count'])
```
**Multi-Agent Performance (Table 4):**
```python
# Compare specialist participation
df.groupby('specialist_participation')[['final_score', 'autonomy_score']].mean()
```
### Ablation Studies
**Ablation 1: Impact of Verification**
```python
with_verification = df[df['verification_triggered'] == 1]
without_verification = df[df['verification_triggered'] == 0]
print("With verification:", with_verification['final_score'].mean())
print("Without verification:", without_verification['final_score'].mean())
# Statistical test
from scipy.stats import mannwhitneyu
stat, p_value = mannwhitneyu(with_verification['final_score'],
without_verification['final_score'])
```
**Ablation 2: Impact of Replanning**
```python
first_try = df[df['first_try_success'] == 1]
with_replans = df[df['replan_count'] > 0]
print("First try success rate:", first_try['final_score'].mean())
print("After replanning:", with_replans['final_score'].mean())
```
**Ablation 3: Impact of Tool Efficiency**
```python
# Quartile analysis
df['efficiency_quartile'] = pd.qcut(df['tool_efficiency'], q=4, labels=['Q1', 'Q2', 'Q3', 'Q4'])
df.groupby('efficiency_quartile')['final_score'].agg(['mean', 'std', 'count'])
```
### Correlation Analysis
```python
import seaborn as sns
import matplotlib.pyplot as plt
# Select key metrics for correlation
metrics = ['final_score', 'analyst_score', 'autonomy_score',
'tool_efficiency', 'verification_score', 'data_understanding_score']
# Compute correlation matrix
corr = df[metrics].corr()
# Visualize
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, cmap='coolwarm', center=0,
square=True, linewidths=1)
plt.title('Metric Correlation Matrix')
plt.tight_layout()
plt.savefig('correlation_matrix.png', dpi=300)
```
### Statistical Significance Testing
```python
from scipy.stats import wilcoxon, mannwhitneyu
# Compare two systems (e.g., baseline vs. proposed)
baseline_df = pd.read_csv('baseline_run/task_metrics.csv')
proposed_df = pd.read_csv('proposed_run/task_metrics.csv')
# Paired test (same tasks)
merged = baseline_df.merge(proposed_df, on='task_id', suffixes=('_baseline', '_proposed'))
stat, p_value = wilcoxon(merged['final_score_baseline'], merged['final_score_proposed'])
print(f"Wilcoxon signed-rank test: p={p_value:.4f}")
# Effect size (Cohen's d)
mean_diff = merged['final_score_proposed'].mean() - merged['final_score_baseline'].mean()
pooled_std = np.sqrt((merged['final_score_proposed'].std()**2 +
merged['final_score_baseline'].std()**2) / 2)
cohens_d = mean_diff / pooled_std
print(f"Effect size (Cohen's d): {cohens_d:.3f}")
```
### Failure Analysis for Papers
```python
# Failure distribution (Figure 2)
failure_dist = df[df['final_score'] < 0.8]['failure_category'].value_counts()
plt.figure(figsize=(10, 6))
failure_dist.plot(kind='bar')
plt.xlabel('Failure Category')
plt.ylabel('Count')
plt.title('Failure Distribution')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.savefig('failure_distribution.png', dpi=300)
# Root cause analysis (Table 5)
root_causes = df[df['final_score'] < 0.8]['root_cause'].value_counts().head(10)
print(root_causes)
```
### Reporting Template
**Results Section:**
```
We evaluate our system on the DABench benchmark containing 50 tasks of varying
difficulty (Easy: 15, Medium: 20, Hard: 15). Our system achieves a mean final
score of X.XX ± Y.YY (mean ± std), significantly outperforming the baseline
(p < 0.001, Wilcoxon signed-rank test). The analyst score, a composite metric
combining autonomy (weight=0.4), tool efficiency (weight=0.3), and verification
quality (weight=0.3), reaches Z.ZZ ± W.WW.
Breakdown by difficulty reveals consistent performance across all levels:
- Easy: X1 ± Y1 (n=15)
- Medium: X2 ± Y2 (n=20)
- Hard: X3 ± Y3 (n=15)
Our multi-agent architecture demonstrates strong autonomy with AA% first-try
success rate and an average of B.B replanning operations per task. Tool
efficiency reaches C.C ± D.D, indicating effective tool selection. Verification
mechanisms trigger in VV% of executions and detect EE failures, contributing to
improved final scores.
Failure analysis (Figure 2) shows the primary failure categories are:
1. TOOL_EXECUTION_FAILURE (XX%)
2. DATA_UNDERSTANDING_FAILURE (YY%)
3. SCHEMA_MISMATCH_FAILURE (ZZ%)
```
---
## 🎓 Experimental Methodology
### Dataset Preparation
1. **Task Selection**: Use stratified sampling by difficulty
2. **Data Splits**: Train/Val/Test or K-fold cross-validation
3. **Seed Control**: Fix random seeds for reproducibility
```python
# Stratified sampling
from sklearn.model_selection import train_test_split
df = pd.read_csv('all_tasks.csv')
train, test = train_test_split(df, test_size=0.3,
stratify=df['difficulty'],
random_state=42)
```
### Baseline Comparisons
**Recommended Baselines:**
1. Random tool selection
2. Fixed planning strategy
3. No verification
4. Single-agent (no specialists)
5. Prior work (if available)
### Reproducibility
**Report:**
- Hardware (GPU type, RAM)
- Software versions (Python, LLM API version)
- Random seeds
- Hyperparameters
- Number of runs (recommend 3-5 for variance)
**Provide:**
- Code repository
- Trained model weights (if applicable)
- Full evaluation CSVs (task_metrics.csv, trajectory.csv)
- Configuration files
### Ethical Considerations
- Data privacy: Ensure benchmark tasks don't contain PII
- Computational cost: Report total compute time and carbon footprint
- Failure modes: Document dangerous failure patterns
- Limitations: Clearly state what the system cannot do
---
## �🔍 Quick Analysis Examples
### Load and Analyze
```python
import pandas as pd
# Load metrics
df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")
# Success rate
success_rate = (df['final_score'] >= 0.8).mean()
print(f"Success rate: {success_rate:.1%}")
# By difficulty
print("\nScores by difficulty:")
print(df.groupby("difficulty")[["final_score", "analyst_score"]].mean())
# Failed tasks
failed = df[df['final_score'] < 0.8]
print(f"\nFailed: {len(failed)} tasks")
print(failed[['task_id', 'final_score', 'failure_category', 'root_cause']])
```
### Debug Failed Task
```python
import json
# Load replay artifact
with open("artifacts/runs/<RUN_ID>/task_38/task_replay.json") as f:
replay = json.load(f)
# Check failure
if replay['failure_attribution']:
fa = replay['failure_attribution']
print(f"Category: {fa['failure_category']}")
print(f"Root cause: {fa['root_cause']}")
print(f"Stage: {fa['failure_stage']}")
print(f"Reason: {fa['failure_reason']}")
print(f"Suggested fix: {fa['suggested_fix']}")
# Review execution
print(f"\nFinal score: {replay['evaluation_result']['final_score']}")
print(f"Coordinator decision: {replay['coordinator_final_decision']}")
print(f"Verification: {replay['verification_passed']}")
```
### View Trajectory
```python
import pandas as pd
# Load trajectory for specific task
traj = pd.read_csv("artifacts/runs/<RUN_ID>/trajectory.csv")
task_traj = traj[traj['task_id'] == 'task_38']
# View execution flow
print(task_traj[['step_id', 'phase', 'agent', 'tool', 'success', 'tokens']])
# Analyze failures
failures = task_traj[~task_traj['success']]
print(f"\nFailures: {len(failures)}")
print(failures[['step_id', 'tool', 'observation']])
```
---
## 🛠️ Hardening Features (Integrated)
### 1. Artifact Reconciliation
**Validates:**
- ✅ Tool call counts match across artifacts
- ✅ Token counts reconcile (trajectory vs metrics)
- ✅ Verification semantics consistent (`verification_passed ↔ coordinator_final_decision`)
- ✅ Time accounting (wall clock ≥ component time)
- ✅ Trajectory completeness
**Report:** `artifact_reconciliation_report.txt`
### 2. Replay Artifacts
**Complete debug snapshots per task:**
- Question & context
- All execution attempts (plan, code, stdout, stderr)
- Agent executions (MAS observability)
- Tool calls with success/failure
- Coordinator decisions
- Verifier output
- Final answer
- Evaluation result
- **Structured failure attribution**
**Location:** `task_*/task_replay.json`
### 3. Engineering Health Report
**System diagnostics:**
- Reconciliation pass/fail status
- Invariant violations (verification, tool calls, tokens)
- Time accounting gaps (overhead analysis)
- Failure taxonomy distribution
- Top recurring root causes
- System health indicators
- **Health score (0-10)**
**Report:** `engineering_health_report.txt`
---
## 🏥 Health Score Interpretation
| Score | Status | Action |
|-------|--------|--------|
| 9-10 | ✅ HEALTHY | Ready to use |
| 7-8 | ⚠️ GOOD | Review warnings |
| 5-6 | ⚠️ FAIR | Fix issues before publication |
| 3-4 | ❌ POOR | Investigation required |
| 0-2 | ❌ UNHEALTHY | Do not use |
---
## 🔧 Implementation Details
### Evaluation Pipeline Architecture
```
┌─────────────────────────────────────────────────────────────┐
│ Evaluation Harness V2 │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Data Collection │
│ • Load trace.json (agent execution trace) │
│ • Load prediction.csv (agent output) │
│ • Load gold.csv (ground truth) │
│ • Load task.json (metadata) │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Metric Computation │
│ • Correctness (multi-level F1: answer/column/row) │
│ • Autonomy (first-try success, replans, autonomy score) │
│ • Planning (revisions, dead ends, alignment) │
│ • Tool Usage (diversity, efficiency, selection accuracy) │
│ • Data Understanding (tables/columns, exploration) │
│ • Verification (triggered, passed, failures detected) │
│ • Recovery (detected, stage, success, depth) │
│ • Trajectory (length, branches, efficiency, patterns) │
│ • Failure Taxonomy (category, root cause, severity) │
│ • Confidence & Calibration (score, error, bucket) │
│ • Composite Scores (analyst_score = weighted blend) │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Normalized Storage (3 CSV Files) │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ task_metrics.csv (Primary Evaluation Table) │ │
│ │ • One row per task execution │ │
│ │ • 100+ columns covering all metric dimensions │ │
│ │ • Granularity: Task-level │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ trajectory.csv (Trajectory Trace) │ │
│ │ • One row per trajectory step │ │
│ │ • Enables process mining and step-level debugging │ │
│ │ • Granularity: Step-level │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ tool_calls.csv (Tool Usage Analysis) │ │
│ │ • One row per tool invocation │ │
│ │ • Tracks latency, tokens, retries, errors │ │
│ │ • Granularity: Tool-call-level │ │
│ └───────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Evaluation Hardening Suite (Integrated) │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 1. Artifact Reconciliation │ │
│ │ • Validates CSV consistency │ │
│ │ • Enforces invariants │ │
│ │ • Generates validation report │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 2. Replay Artifact Generation │ │
│ │ • Complete debug snapshots per task │ │
│ │ • Structured failure attribution │ │
│ │ • MAS observability tracking │ │
│ └───────────────────────────────────────────────────────┘ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ 3. Engineering Health Report │ │
│ │ • System health diagnostics │ │
│ │ • Time accounting analysis │ │
│ │ • Health score (0-10) │ │
│ └───────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Terminal Visualization (3 Modes) │
│ • Standard: Core metrics + summary │
│ • Verbose: + Agent behavior analysis │
│ • Research: + All metrics for papers (mean, std) │
└─────────────────────────────────────────────────────────────┘
```
### Data Model Schema
#### 1. task_metrics.csv (Primary Evaluation Table)
**Purpose:** Comprehensive per-task evaluation metrics for academic publication.
**Granularity:** One row per task execution (50-500 tasks typical).
**Column Count:** 100+ columns organized into 14 categories.
**Schema Categories:**
1. **Identification** (5 cols): run_id, task_id, trace_id, difficulty, timestamp
2. **Task Metadata** (5 cols): task_type, source_count, source_types, requires_cross_source_reasoning, ground_truth_available
3. **Correctness** (15 cols): final_score, answer_precision/recall/f1, column_precision/recall/f1, row_precision/recall/f1, matched_columns, pred_rows/cols, gold_rows/cols
4. **Execution** (6 cols): execution_success, execution_time, total_tokens, llm_calls, tool_calls, tool_failures
5. **Autonomy** (5 cols): first_try_success, replan_count, user_intervention_count, autonomy_score, coordinator_interventions
6. **Planning** (5 cols): planner_steps, planner_revisions, planner_dead_ends, plan_execution_alignment, plan_attempts
7. **Tool Usage** (10 cols): unique_tools_used, tool_diversity, useful_tool_calls, wasted_tool_calls, tool_efficiency, tool_retry_count, tool_selection_accuracy
8. **Data Understanding** (7 cols): tables_discovered, columns_discovered, relevant_tables_found, relevant_columns_found, schema_exploration_steps, data_understanding_score
9. **Verification** (6 cols): verification_triggered, verification_steps, verification_passed, verification_failures_detected, verification_score, aat_verification_passed
10. **Recovery** (6 cols): failure_detected, failure_stage, recovery_success, recovery_depth, recovery_success_rate
11. **Trajectory** (7 cols): trajectory_length, trajectory_summary, branching_factor, max_execution_depth, critic_loops, trajectory_efficiency
12. **Failure Taxonomy** (6 cols): failure_category, root_cause, recoverable_failure, severity, failure_reason, failure_agent
13. **Confidence** (5 cols): confidence_score, confidence_correct, confidence_error, calibration_bucket
14. **Composite** (1 col): analyst_score
15. **AAT Architecture** (10 cols): coordinator_calls, coordinator_final_decision, coordinator_checkpoints, schema_agent_called, domain_agent_called, document_agent_called, specialist_participation
16. **Per-Stage Metrics** (25 cols): understanding_time/calls/tokens, planning_time/calls/tokens, execution_time/calls/tokens, verification_time/calls/tokens, summary_time/calls/tokens
17. **Per-Action Metrics** (28 cols): list_context_calls/time, read_json_calls/time, read_knowledge_calls/time, execute_python_calls/time, etc.
**Total:** 121 columns
#### 2. trajectory.csv (Step-by-Step Trace)
**Purpose:** Detailed execution trace for process mining and debugging.
**Granularity:** One row per trajectory step (10-50 steps per task typical).
**Columns (24):**
- **Identification**: run_id, task_id, trace_id, step_id, parent_step_id
- **Execution Context**: stage, legacy_stage, phase, agent, trajectory_agent, agent_type
- **Action**: tool, action, coordinator_checkpoint
- **Decision**: confidence, decision, review_type, verification_status, specialist_selected
- **Outcome**: success, duration_seconds, tokens, retries
- **Timing**: timestamp, elapsed_seconds
- **Details**: thought, action_input, observation, raw_response, metadata_json
**Use Cases:**
- Process mining (find common execution patterns)
- Step-level debugging (identify exact failure point)
- Agent behavior analysis (tool selection patterns)
- Performance profiling (step latencies)
#### 3. tool_calls.csv (Tool-Level Analysis)
**Purpose:** Per-tool-invocation metrics for optimization.
**Granularity:** One row per tool call (matches trajectory tool steps).
**Columns (15):**
- **Identification**: run_id, task_id, trace_id, step_id, call_id
- **Tool**: tool_name, tool_category
- **Performance**: success, latency_seconds, tokens, retry_count
- **Data**: input_size, output_size
- **Errors**: error_type, error_message
- **Metadata**: timestamp, metadata_json
**Use Cases:**
- Tool efficiency analysis (which tools are slow?)
- Error rate tracking (which tools fail most?)
- Token cost analysis (which tools are expensive?)
- Tool selection optimization (which tools work best for what?)
#### 4. task_replay.json (Debug Snapshot - Per Task)
**Purpose:** Complete execution context for offline debugging without re-running.
**Granularity:** One JSON file per task.
**Structure:**
```json
{
"task_id": "task_38",
"trace_id": "...",
"run_id": "...",
"difficulty": "Medium",
"question": "...",
"available_sources": [...],
"execution_attempts": [
{
"attempt_number": 1,
"plan": {...},
"code_executed": "...",
"stdout": "...",
"stderr": "...",
"success": false
}
],
"agent_executions": [
{
"agent_name": "StrategicCoordinator",
"checkpoint": "UNDERSTANDING",
"decision": "PROCEED",
"confidence": 0.85,
"duration": 5.2,
"tokens": 1500
}
],
"coordinator_decisions": [...],
"verifier_output": {...},
"final_answer_csv": "...",
"evaluation_result": {
"final_score": 0.0,
"answer_f1": 0.0
},
"failure_attribution": {
"failure_category": "TOOL_EXECUTION_FAILURE",
"root_cause": "KeyError: 'column_name'",
"failure_stage": "execution",
"failure_agent": "Executor",
"evidence_trace_step_ids": [12, 14],
"suggested_fix": "Validate column existence"
}
}
```
### Module Structure
```
src/data_agent_baseline/langgraph_agent/
├── eval_v2.py # Main orchestrator (650 lines)
│ ├── evaluate_task_v2() # Single task evaluation
│ ├── evaluate_run_v2() # Full run evaluation
│ ├── write_evaluation_v2() # CSV output
│ └── extract_trajectory/tools() # Trace extraction
├── eval_v2_metrics.py # Metric computations (1280 lines)
│ ├── compute_correctness() # Multi-level F1
│ ├── compute_autonomy() # Independence metrics
│ ├── compute_planning() # Planning quality
│ ├── compute_tool_usage() # Tool efficiency
│ ├── compute_data_understanding() # Schema understanding
│ ├── compute_verification() # Quality assurance
│ ├── compute_recovery() # Error recovery
│ ├── compute_trajectory() # Execution path analysis
│ ├── compute_failure_taxonomy() # Error classification
│ ├── compute_confidence() # Calibration metrics
│ └── compute_analyst_score() # Composite score
├── eval_v2_viz.py # Visualization (580 lines)
│ ├── render_evaluation_report() # Main report
│ ├── render_task_table() # Per-task table
│ ├── render_summary_sections() # Analysis sections
│ └── render_verbose_task_detail() # Task drill-down
├── eval_artifact_reconciliation.py # Validation (435 lines)
│ ├── ArtifactReconciliator # Cross-artifact validation
│ ├── validate_run() # Run-level validation
│ └── format_report() # Validation report
├── eval_replay_artifacts.py # Debug snapshots (630 lines)
│ ├── ReplayArtifactGenerator # Snapshot generation
│ ├── FailureAttributor # Root cause analysis
│ └── generate_all_replays() # Batch generation
└── eval_health_report.py # Health diagnostics (490 lines)
├── EngineeringHealthReport # Health metrics
├── TimeAccountingGap # Overhead analysis
└── generate_health_report() # Report generation
```
### Key Files
**Core Evaluation:**
- `eval_v2.py` - Main evaluation engine (scoring, metrics)
- `eval_v2_metrics.py` - Individual metric calculators
- `eval_v2_viz.py` - Rendering & visualization
**Hardening Suite:**
- `eval_artifact_reconciliation.py` - Cross-artifact validation
- `eval_replay_artifacts.py` - Debug snapshot generation
- `eval_health_report.py` - System health diagnostics
**CLI Integration:**
- `cli.py:eval_v2_command()` - Orchestrates entire pipeline
### Critical Invariants
```python
# Verification consistency
aat_verification_passed = 1 ↔ coordinator_final_decision = "APPROVE_FINAL"
# Note: critic_verification_passed is independent of AAT coordinator decision
# Tool call reconciliation
len(tool_calls_df) == task_metrics['tool_calls']
tool_failures_count == task_metrics['tool_failures']
# Token conservation
abs(trajectory_tokens - metrics_tokens) < 100
# Time accounting (with overhead)
wall_clock_time = execution_time # End-to-end elapsed
component_compute_time = sum(agent/tool tracked durations)
unaccounted_overhead = wall_clock_time - component_compute_time
time_accounting_ratio = component_compute_time / wall_clock_time
# Expected: component_time < wall_clock_time (overhead exists)
# Framework overhead, I/O wait, concurrency gaps, logging typically 20-50%
# Warning only if time_accounting_ratio < 0.05 (95% unaccounted)
# Info if component_time > wall_clock_time (indicates concurrent execution)
```
### Comparison with Previous Systems
| Feature | Traditional Eval | DABench V1 | DABench V2 (Ours) |
|---------|-----------------|------------|-------------------|
| **Storage Model** | Single CSV | Single CSV | 3 normalized CSVs |
| **Metrics** | 5-10 basic | ~30 metrics | 100+ comprehensive |
| **Correctness** | Overall F1 | Overall F1 | Multi-level F1 (answer/column/row) |
| **Autonomy** | Not measured | Retry count | Composite autonomy score |
| **Planning** | Not measured | Step count | Revisions, alignment, dead ends |
| **Tool Analysis** | Call count | Call count | Efficiency, diversity, selection accuracy |
| **Trajectory** | Not captured | Basic steps | Full trace with branching, patterns |
| **Failure Analysis** | Error message | Simple bucket | Structured taxonomy + root cause |
| **Verification** | Not measured | Basic check | Multi-stage verification score |
| **Confidence** | Not measured | Not measured | Calibration metrics |
| **Composite Scores** | None | None | Analyst score (weighted) |
| **Debugging** | Manual | Manual | Automated replay artifacts |
| **Validation** | None | Basic | Comprehensive reconciliation |
| **Process Mining** | Not supported | Not supported | Full trajectory CSV |
| **MAS Observability** | Not supported | Not supported | Per-agent tracking |
| **Time Accounting** | Wall clock only | Wall clock only | Component + overhead |
**Key Innovations:**
1. **Multi-level correctness**: Separate precision/recall/F1 at answer/column/row levels
2. **Autonomy quantification**: First-try success, replan count, composite autonomy score
3. **Composite analyst score**: Weighted blend of autonomy, efficiency, verification (suitable for ranking)
4. **Structured failure taxonomy**: 10 failure categories with root cause attribution
5. **Integrated validation**: Automatic artifact reconciliation with invariant enforcement
6. **Debug-ready artifacts**: Complete replay snapshots for offline debugging
7. **Process mining support**: Full trajectory CSV for pattern discovery
8. **MAS observability**: Per-agent execution tracking in multi-agent systems
---
## 🎯 Failure Attribution
### Failure Categories
```python
PLANNING_FAILURE # Bad plan generation
DATA_UNDERSTANDING_FAILURE # Misunderstood data/schema
TOOL_SELECTION_FAILURE # Wrong tool chosen
TOOL_EXECUTION_FAILURE # Tool crashed/errored
REASONING_FAILURE # Logic errors
VERIFICATION_FAILURE # Verifier malfunction
AGGREGATION_FAILURE # Data aggregation errors
ANSWER_FORMAT_FAILURE # Wrong output format
SCHEMA_MISMATCH_FAILURE # Schema mismatch
UNKNOWN_FAILURE # Needs manual inspection
```
### Attribution Structure
```python
{
"failure_category": "TOOL_EXECUTION_FAILURE",
"root_cause": "KeyError: 'column_name'",
"failure_stage": "execution",
"failure_agent": "Executor",
"failure_reason": "Attempted to access non-existent column",
"evidence_trace_step_ids": [12, 14],
"suggested_fix": "Validate column existence before access",
"evidence_summary": "Step 12: execute_python failed with KeyError",
"execution_success": false,
"tool_failure_count": 1
}
```
---
## 📖 Common Workflows
### 1. Evaluate New Run
```bash
# Run prediction first (if not already done)
dabench run input_full --agent aat
# Evaluate
dabench eval-v2 <RUN_ID> --mode research
# Review health
cat artifacts/runs/<RUN_ID>/engineering_health_report.txt
```
### 2. Debug Failed Tasks
```bash
# Find failed tasks
python3 -c "
import pandas as pd
df = pd.read_csv('artifacts/runs/<RUN_ID>/task_metrics.csv')
failed = df[df['final_score'] < 0.8]
print(failed[['task_id', 'failure_category', 'root_cause']])
"
# Debug specific task
cat artifacts/runs/<RUN_ID>/task_38/task_replay.json | jq '.failure_attribution'
```
### 3. Compare Runs
```python
import pandas as pd
# Load two runs
run1 = pd.read_csv("artifacts/runs/RUN_A/task_metrics.csv")
run2 = pd.read_csv("artifacts/runs/RUN_B/task_metrics.csv")
# Merge on task_id
merged = run1.merge(run2, on='task_id', suffixes=('_A', '_B'))
# Compare
print(f"Run A mean: {merged['final_score_A'].mean():.3f}")
print(f"Run B mean: {merged['final_score_B'].mean():.3f}")
# Tasks improved in Run B
improved = merged[merged['final_score_B'] > merged['final_score_A']]
print(f"\nImproved: {len(improved)} tasks")
```
### 4. Generate Paper Figures
```python
import pandas as pd
import matplotlib.pyplot as plt
df = pd.read_csv("artifacts/runs/<RUN_ID>/task_metrics.csv")
# Score distribution
plt.figure(figsize=(10, 6))
plt.hist(df['final_score'], bins=20, edgecolor='black')
plt.xlabel('Final Score')
plt.ylabel('Count')
plt.title('Score Distribution')
plt.savefig('score_distribution.png')
# Autonomy vs Efficiency
plt.figure(figsize=(10, 6))
plt.scatter(df['autonomy_score'], df['tool_efficiency'],
c=df['final_score'], cmap='viridis')
plt.xlabel('Autonomy Score')
plt.ylabel('Tool Efficiency')
plt.colorbar(label='Final Score')
plt.savefig('autonomy_vs_efficiency.png')
```
---
## 🧪 Testing
### Run Test Suite
```bash
cd /workspace/ainn-cm-poc-data-agent
# Test evaluation harness
pytest tests/test_eval_harness.py -v
# Test specific validation
pytest tests/test_eval_harness.py::TestVerificationConsistency -v
```
### Manual Validation
```bash
# Re-validate existing run
python3 -c "
from pathlib import Path
from src.data_agent_baseline.langgraph_agent.eval_artifact_reconciliation import validate_evaluation_run
passed, report = validate_evaluation_run(Path('artifacts/runs/<RUN_ID>'))
print(report)
print(f'\nPassed: {passed}')
"
```
---
## 🐛 Troubleshooting
### Issue: "Evaluation produced inconsistent metrics"
**Cause:** Old consistency validator found errors
**Solution:** Check `artifact_reconciliation_report.txt` for details:
```bash
cat artifacts/runs/<RUN_ID>/artifact_reconciliation_report.txt
```
Common issues:
- `verification_outcome_mismatch`: Verification flag doesn't match coordinator decision
- `tool_call_count_mismatch`: Tool calls CSV doesn't match metrics
- `data_understanding_inflation`: Score exceeds theoretical maximum
### Issue: Health score < 7
**Cause:** System detected quality issues
**Solution:** Review `engineering_health_report.txt`:
```bash
cat artifacts/runs/<RUN_ID>/engineering_health_report.txt
```
Look for:
- Time accounting gaps > 50%
- High tool failure rate > 10%
- Missing failure attribution
### Issue: Missing trajectory duration
**Symptom:** Warnings about "missing time data"
**Cause:** Trajectory extraction didn't populate `duration_seconds` column
**Impact:** Time validation skipped (not critical)
---
## 📚 Additional Documentation
**Architecture Details:**
- See `src/data_agent_baseline/langgraph_agent/eval_v2.py` for scoring logic
- See `src/data_agent_baseline/langgraph_agent/eval_v2_metrics.py` for metric definitions
**Test Coverage:**
- See `tests/test_eval_harness.py` for validation tests
**CLI Integration:**
- See `src/data_agent_baseline/cli.py:eval_v2_command()` for integration
---
## 🔄 Version History
### V2.1 (Current) - June 14, 2026
**Phase 2: MAS Debugging Enhancements**
Focused improvements for multi-agent system debugging and actionable diagnostics:
1. **Deterministic Failure Attribution**: Maps evaluation buckets (low_recall, wrong_schema, etc.) to structured categories (REASONING_FAILURE, DATA_UNDERSTANDING_FAILURE, AGGREGATION_FAILURE)
- Eliminates UNKNOWN_FAILURE when bucket exists
- Infers failure_stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
- Identifies failure_agent (Schema Agent, Planner, Executor, etc.)
2. **Separated Health Assessments**: Clear distinction between infrastructure health and run quality
- **Harness Health**: Infrastructure metrics (reconciliation pass rate, validator errors, attribution coverage)
- **Run Quality**: Outcome metrics (answer accuracy, execution success rate)
- Status thresholds: HEALTHY (9.0+, no errors), DEGRADED (7.0+), UNHEALTHY (else)
3. **Enhanced Visualization**:
- Per-task table: Renamed "Succ" → "Exec" to clarify execution vs. correctness
- Added "Root Cause" column showing failure diagnostics (filter_logic_error, schema_misunderstanding, etc.)
- New "MAS Failure Analysis" section with:
- MAS Failure Categories table (structured categories: REASONING_FAILURE, etc.)
- Failure Distribution by Stage (UNDERSTAND/PLAN/EXECUTE/VERIFY)
- Outcome Error Types (evaluation buckets)
4. **Partial-Correct Task Handling**: New `outcome_status` field distinguishes:
- "correct": final_score ≥ 0.8
- "partial": succeeded but 0 < final_score < 0.8
- "failed": final_score < 0.8 or execution failure
5. **Improved Terminology**: Renamed "time gaps" → "Unaccounted Overhead Time" with clear explanation (framework overhead, I/O wait, async queuing)
6. **Complete Replay Artifacts**: Every task_replay.json includes:
- Full failure diagnostics (failure_category, root_cause, failure_stage, failure_agent)
- Run context (harness_health_status, run_quality_status, outcome_status)
- Time accounting (unaccounted_overhead_time_seconds, time_accounting_ratio)
**Design Philosophy**: Optimized for debugging and MAS improvement, not paper metrics. All changes maintain backward compatibility.
**Phase 2.1 Cleanup (Jan 2026)**:
- Separated harness health (infrastructure) from run quality (outcomes)
- Fixed MAS Failure Categories display to show structured categories instead of buckets
- Added outcome_status field for partial-correct handling
- Enhanced replay artifacts with health/quality context
- Improved time accounting terminology
### Phase 0: MAS Debugging Enhancements - Final Round (Jan 2026)
**Goal**: Make evaluation harness maximally useful for all future phases (Baseline ReAct, MAS, DAG Visualization, Replay/Time Travel, Confidence & Verification, Research/Ablation Studies).
**Key Improvements**:
1. **MAS Recovery Effectiveness Metrics** (Task 2):
- New fields: `initial_answer_correct`, `final_answer_correct`, `recovered_after_replan`, `recovered_after_retry`
- Display: "MAS Effectiveness" section showing:
- First Attempt Accuracy: Initial correct rate
- Final Accuracy: Final correct rate
- Recovered Tasks: Count of tasks improved through MAS interventions
- MAS Recovery Gain: Percentage improvement (e.g., +14%)
- **Impact**: Directly answers "Did MAS actually improve answers?"
2. **Replan/Retry Effectiveness Tracking** (Task 3):
- New fields: `replan_requested`, `replan_successful`, `retry_requested`, `retry_successful`
- Display: "Coordinator Intervention Effectiveness" section showing:
- Replans: Requested count, successful count, success rate
- Retries: Requested count, successful count, success rate
- **Impact**: Shows which coordinator interventions actually help
3. **Specialist Agent Value Analysis** (Task 4):
- Existing fields: `schema_agent_used`, `domain_agent_used`, `document_agent_used`
- Display: "Specialist Agent Value Analysis" section showing for each agent:
- Tasks Used
- Accuracy With Agent
- Accuracy Without Agent
- Impact (delta %)
- **Impact**: Automatic ablation showing which specialists add value
4. **Expanded Failure Stage Taxonomy** (Task 5):
- Updated `FailureStage` enum: UNDERSTAND, PLAN, EXECUTE, VERIFY, AGGREGATE
- Replaced coarse stages (EXPLORATION, PLANNING, EXECUTION) with AAT-aligned taxonomy
- **Impact**: Finer-grained debugging for AAT phase-specific failures
5. **Cost by Difficulty** (Task 6):
- Difficulty breakdown now includes Mean Runtime and Mean Tokens columns
- **Impact**: Required for Baseline vs MAS vs Future comparisons
6. **Verification Timeline Clarity** (Task 7):
- Separated "Execution Approval" (coordinator decision) from "Ground Truth Result" (evaluation correctness)
- Added explanatory note distinguishing verification from correctness
- **Impact**: Eliminates confusion between process approval and actual correctness
7. **Comprehensive CSV Storage** (Task 8):
- All new fields stored in task_metrics.csv: `outcome_error_type`, MAS recovery fields, specialist usage
- **Impact**: Future phases can run aggregations without parsing replay artifacts
8. **Removed Duplicate Reporting** (Task 1):
- Eliminated redundant "Failure Categories" section
- Kept: "MAS Failure Categories" (structured) and "Outcome Error Types" (buckets)
**Design Philosophy**: Every change focused on making the evaluation harness more actionable for debugging and improving the MAS, not for paper-writing. Provides automatic ablation studies and directly answers key questions about MAS effectiveness.
### V2.0 - June 2026
- ✅ Integrated hardening suite (automatic reconciliation, replay, health)
- ✅ 100+ comprehensive metrics for KDD Creative Track
- ✅ AAT architecture observability
- ✅ Structured failure attribution
- ✅ MAS-aware trajectory extraction
- ✅ Time accounting with overhead tracking
### V1 (Legacy)
- Basic metrics (precision, recall, F1)
- Manual validation required
- Limited debugging support
---
## 📝 Summary
**One command does it all:**
```bash
dabench eval-v2 <RUN_ID> --mode standard
```
**Automatically provides:**
- ✅ Comprehensive metrics (100+ columns)
- ✅ Complete validation & reconciliation
- ✅ Full debug snapshots (replay artifacts)
- ✅ System health diagnostics
- ✅ Failure attribution & root cause analysis
**No manual steps required.**