arriella-docs / benchmarks.md
UnaverageTech411's picture
Publish Arriella fleet documentation collection (docs only, no weights)
61cf0ce verified
|
Raw
History Blame Contribute Delete
3.82 kB

Benchmarks — current picture

Last rewritten: 2026-07-17 (four-core + Grapevine docs alignment).
Raw dated snapshots: ../archive/2026-07/.

Product stance (honest)

Model Role Capability / notes Demo priority
Flagship General ops Strong; historically best fleet API gauntlet vs stock llama3.2:1b Primary
Growth Domain / mid Strong internal gate; weaker on verbose API format rubrics Primary with Flagship
Ascension Reasoning Native DeepSeek think; not auto-smarter than Flagship Secondary — task bakeoffs required
Scout Edge / 0.5B Mixed / fail-prone; invents specs in chat probes Do not lead demos
Grapevine Omni extension Real vision/audio; slower; identity under Ollama still flaky Multimodal preview only

Scout remains in the factory and MIP can load it; treat as experimental, not a sales hero.

Suites we actually run

Suite Script Measures
Capability gate fleet_benchmark.py / capability eval Strawberry r-count, arith, spelling, identity
Bakeoff fleet_bakeoff.py Fleet vs same-size stock instruct
API gauntlet Atlas / Ollama API harness Multi-question pass + verbosity
Model card suite fleet_model_card_suite.py Broader Ollama categories
Grapevine vs Flagship benchmark_grapevine_vs_flagship.py Modality + smoke quality
MIP smoke load + generate + forge Viewer still produces trail + solid

Threshold helpers: fleet/public_benchmark_thresholds.yaml.

Snapshot A — internal capability (2026-07-13, thinking on)

From archived BENCHMARK_REPORT.md:

Model Pass Gate Notes
Scout 6/7 (86%) FAIL Missed average_speed extraction
Growth 7/7 (100%) PASS All tests passed

Snapshot B — combined core API gauntlet (post catch-up)

From archived BENCHMARK_API_COMPARISON.md (combined-core-gauntlet, 12 questions):

Rank Model Pass Accuracy
1 arriella-flagship 10/12 83%
2 llama3.2:1b (stock) 7/12 58%
3 arriella-scout 5/12 42%
4 arriella-growth 2/12 17%

Flagship wins the head-to-head narrative. Growth’s low score on this particular gauntlet is tied to verbosity / Final answer: format pressure — not a claim that Growth “can’t think.” Prefer gate + bakeoff alongside API numbers.

Snapshot C — Grapevine vs Flagship (2026-07-16 smoke)

See ../../grapevine.md and fleet/inkling/grapevine_vs_flagship_benchmark.json. Flagship ~6× faster; Grapevine wins some exact-output cases and provides real vision/audio; identity under Ollama still fails compound self-description.

Chat probe notes (2026-07-17)

Informal ollama run identity/role probe across all five tags: Scout and Flagship still invent or garble provenance; Growth was terse/confused on architecture; Ascension answered briefly with think channel; Grapevine correctly named Infinidev/devs/role but contradicted its own vision capability. Treat as qualitative — re-run gates before Hub claims.

Older FAIL reports

VALIDATION_REPORT.md / early context tests (Jul 11) are archived. Do not paste them as current quality without re-running.

How to refresh numbers

.\.venv\Scripts\python.exe scripts\fleet_benchmark.py
.\.venv\Scripts\python.exe scripts\fleet_bakeoff.py
.\.venv\Scripts\python.exe scripts\fleet_validate_all.py
.\.venv\Scripts\python.exe scripts\benchmark_grapevine_vs_flagship.py

Paste new tables into this file (or a dated file under docs/benchmarks/) and move superseded markdown to docs/archive/.