AI that gets better at its job by doing it.
hopit.ai · Models · Benchmarks · Research notes
Hopit is a continual-learning lab. We build models and harnesses that learn new tasks, keep what they already know, and improve inside the enterprises that use them.
Small models that answer a typed decision question in one forward pass, with a calibrated probability for each option.
| Model | What it is | Where it stands |
|---|---|---|
| Hopper | 4B decision model, LoRA on Qwen3.5-4B | #2 in JevBench's Jev-class capability ranking¹ |
| Hopper (G) | General-purpose version, 4.66B served | Top five of 46 under 5B on the Jev Decision Index² |
Released for research and demonstration only; see each model card for its licence and training data.
Our update method lifted an internal tool-use evaluation from 57.9% to 66.1% without losing earlier abilities on the retention suites we track. One seed and an internal measurement — we will publish the protocol and artifacts before treating it as established.
Before continual learning, we built open models and public benchmark suites for fashion retrieval and attribute extraction, and held ourselves to them in public. They remain the standard our newer work has to clear.
- hopper — the decision server behind Hopper and Hopper (G).
- Moda — open retrieval models and their benchmark: harness, evaluation code, and the experiments that failed.
- Moda_ner — an open attribute-extraction suite: four frozen tracks, scorers, prediction files and their hashes.
- india-trade-cli — agentic research over Indian equities. Different domain, same conviction: publish the method, measure the result.
- Every rank is quoted with its qualifier. A leaderboard position without its scope is a claim nobody can check.
- Frozen before inference. Protocols are fixed and predictions hashed before labels open; scorers fail closed.
- Losses shown. The runs we lose are published beside the runs we win.
We deploy with a small forward-deployed team inside your environment. Your data and your deployed models stay yours. The approach is ideal for regulated enterprises. → hopit.ai
¹ JevBench v1.4.2, 24 September 2026 snapshot, scored by an independent maintainer. Second on capability; fifth on the composite score, which also weighs speed and cost. ² Jev Decision Index 0.2.1, 28 September 2026: Hopper (G) 1.2 is third of 46 systems under 5B on the chance-corrected headline score (40.77), within 0.1 of fourth, and 18th of 70 overall. The edition is 76% scored.