Hopper (G)

A general-purpose version of Hopper: a LoRA adapter for Qwen3.5-4B that answers typed decision questions in one forward pass by reading the probability of each option letter, with a per-kind calibration map. Served with the Hopper code at https://github.com/hopit-ai/hopper.

Current version: Hopper (G) 1.3 (revision 8b4cd7c, tag g-1.3.0). Hopper (G) 1.2 stays available at revision d60a1d6 and is what the Decision Index row below measures.

Research and demo use only. These adapters continue training from Hopper 1.0's adapter, whose training data included passages from RACE (non-commercial research only), and their training data also includes material made with LLM-based generation. Do not use them commercially.

Leaderboards (official)

  • Jev Decision Index (edition 0.2.1, 27 Sep 2026): Hopper (G) 1.2 scores 40.77, #16 of 68, the highest of the 4B models (Decider 4B: 40.70), from a complete self-scored run (results). Hopper (G) 1.3 has not been scored on the Index yet.
  • JevBench: not yet measured. We have asked for Hopper (G) 1.3 to be measured as a separate row (issue #112). Hopper 1.0's official result is 59.43 (v1.4.2.2).

What changed in 1.3

  • Weights only. Continued from Hopper (G) 1.2 on 6,400 new code-generated decision items (dated arithmetic, long policies with exceptions and precedence, multi-table lookups, ambiguity, probability, safety and judging checklists, trade-offs, paraphrase pairs, injected-instruction traps), every answer computed and checked by code, plus maintenance and replay of earlier training data under the same retention constraint as 1.2.
  • Serving: unchanged. Same code as Hopper 1.1.1 / Hopper (G) 1.2, same calibration map (byte-identical), eager by default.

Evaluation (our runs, not official scores)

On our private held-out decision set (2,700 items, nine decision families, built for this purpose and never trained on), 1.3 scores +1.7 points over 1.2 (template bootstrap 95% interval +0.5 to +2.9). This is below the +3.0 we pre-registered as our own bar for this build, and a smaller run with a quarter of the new data did about as well (+2.1), so more of the same data did not help. We release it anyway as a disclosed, qualified release: on our regression checks against Hopper 1.0 (document reading, numeric, paraphrase, abstention, routing, calibration and a general-knowledge check) it passes every registered floor, with a pooled gain of +2.3 points (one-sided 95% lower bound +1.9) and hard-tier calibration 83.5 under the shipped map. None of this predicts a JevBench result.

Limitations

  • English, 4B parameters; it reads options, it does not generate reasoning.
  • It rarely concludes "no match" or "cannot be determined" when an explicitly incomplete record leaves a multi-step lookup unresolved; 1.3 did not fix this.
  • Research and demo use only (see above).

Revisions

  • 8b4cd7c (tag g-1.3.0, full 8b4cd7cc87ef3ec6990245974a91ce6acdd73e4d): Hopper (G) 1.3, the evaluated adapter bytes.
  • Later commit: adapter_config.json sets "task_type": "CAUSAL_LM" (it was null) for the Hub's metadata parser only; weights, calibration map and outputs unchanged.
  • d60a1d6: Hopper (G) 1.2 (the Decision Index row). 060b1bd: its config-only fix.
Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for HopitAI/hopper-g

Finetuned
Qwen/Qwen3.5-4B
Adapter
(656)
this model

Spaces using HopitAI/hopper-g 2