--- license: other license_name: research-and-demo base_model: Qwen/Qwen3.5-4B library_name: peft tags: [decision, classification, calibration, lora] --- # Hopper (G) A general-purpose version of [Hopper](https://huggingface.co/HopitAI/hopper): a LoRA adapter for Qwen3.5-4B that answers typed decision questions in one forward pass by reading the probability of each option letter, with a per-kind calibration map. Served with the Hopper code at https://github.com/hopit-ai/hopper. **Current version: Hopper (G) 1.3** (revision `8b4cd7c`, tag `g-1.3.0`). Hopper (G) 1.2 stays available at revision `d60a1d6` and is what the Decision Index row below measures. **Research and demo use only.** These adapters continue training from Hopper 1.0's adapter, whose training data included passages from RACE (non-commercial research only), and their training data also includes material made with LLM-based generation. Do not use them commercially. ## Leaderboards (official) - **[Jev Decision Index](https://huggingface.co/spaces/multimodalart/jev-decision-index)** (edition 0.2.1, 27 Sep 2026): Hopper (G) 1.2 scores **40.77, #16 of 68**, the highest of the 4B models (Decider 4B: 40.70), from a complete self-scored run ([results](https://huggingface.co/datasets/HopitAI/hopper-g-decision-index-results)). Hopper (G) 1.3 has not been scored on the Index yet. - **[JevBench](https://benchmarkheaven.com/jev-models)**: not yet measured. We have asked for Hopper (G) 1.3 to be measured as a separate row ([issue #112](https://github.com/fstandhartinger/jevbench/issues/112)). Hopper 1.0's official result is 59.43 (v1.4.2.2). ## What changed in 1.3 - **Weights only.** Continued from Hopper (G) 1.2 on 6,400 new code-generated decision items (dated arithmetic, long policies with exceptions and precedence, multi-table lookups, ambiguity, probability, safety and judging checklists, trade-offs, paraphrase pairs, injected-instruction traps), every answer computed and checked by code, plus maintenance and replay of earlier training data under the same retention constraint as 1.2. - **Serving: unchanged.** Same code as Hopper 1.1.1 / Hopper (G) 1.2, same calibration map (byte-identical), eager by default. ## Evaluation (our runs, not official scores) On our private held-out decision set (2,700 items, nine decision families, built for this purpose and never trained on), 1.3 scores **+1.7 points over 1.2** (template bootstrap 95% interval +0.5 to +2.9). **This is below the +3.0 we pre-registered as our own bar for this build**, and a smaller run with a quarter of the new data did about as well (+2.1), so more of the same data did not help. We release it anyway as a disclosed, qualified release: on our regression checks against Hopper 1.0 (document reading, numeric, paraphrase, abstention, routing, calibration and a general-knowledge check) it passes every registered floor, with a pooled gain of +2.3 points (one-sided 95% lower bound +1.9) and hard-tier calibration 83.5 under the shipped map. None of this predicts a JevBench result. ## Limitations - English, 4B parameters; it reads options, it does not generate reasoning. - It rarely concludes "no match" or "cannot be determined" when an explicitly incomplete record leaves a multi-step lookup unresolved; 1.3 did not fix this. - Research and demo use only (see above). ## Revisions - `8b4cd7c` (tag `g-1.3.0`, full `8b4cd7cc87ef3ec6990245974a91ce6acdd73e4d`): Hopper (G) 1.3, the evaluated adapter bytes. - Later commit: `adapter_config.json` sets `"task_type": "CAUSAL_LM"` (it was `null`) for the Hub's metadata parser only; weights, calibration map and outputs unchanged. - `d60a1d6`: Hopper (G) 1.2 (the Decision Index row). `060b1bd`: its config-only fix.