Harbor Eval Visualizations
Explore model pass@k results with interactive dashboard
Explore model pass@k results with interactive dashboard
Note Canonical TRAINING suite: 2238 Harbor data-analysis tasks over real Kaggle datasets. Single correctness reward.
Note Held-out EVAL suite: 366 verified Harbor tasks. For benchmarking, not training.
Note Source-of-truth row-level split manifest (parquet, eval/train, ~30k rows). The Harbor task suites are built from this.
Note Shared Kaggle DATA BUCKET every task pulls from at container start (keyed by BUCKET_PREFIX).
Note Fast-iteration SUBSET: 100 easy (L1) numeric-reward tasks sampled from the train suite.
Note Train suite (2238) with a 3-REWARD verifier: correctness + submission + tool_efficiency (reward.json).
Note Train suite (2238) ORDERED easy->hard by empirical pass@4 difficulty; registry carries rank/difficulty/solve_frac. Curriculum-ready + multireward.