AI & ML interests

Open models built to place at the top of their class on public leaderboards, from fine-tunes the open-source community has already published. Every result is measured against the base model and published with its evidence.

Recent Activity

Optitransferย  updated a Space 7 days ago
OptiTransferData/README
Optitransferย  updated a Space 6 months ago
OptiTransferData/README
Optitransferย  published a Space 7 months ago
OptiTransferData/README
View all activity

Organization Card

Optitransfer

Open models built to place at the top of their class on public leaderboards, from the work the open-source community has already published.

Models Paper Paper crdt-merge

At a Glance

7B
Class in progress
2
Public releases
12
Board benchmarks
2
Papers

The Thesis

Thousands of fine-tuned models sit on public repositories. Each adds a skill to a shared base model. Their training is already paid for.

Hypothesis: that capability can be consolidated into one model per class without pre-training, with trade-offs between skills controlled and capability then maximised across the board through iteration. Converge is the programme that tests it, one release at a time, in public.

Where We Are

  • v2 against its base. 12 benchmarks: 2 improved, 2 regressed, 8 with no measurable difference.
  • Improved. GSM8K under HELM's protocol: 87.0 against 83.3 (+3.7, calibrated). MMLU under HELM's protocol: +0.47 (calibrated).
  • Regressed. MATH Level 5: -5.74. MMLU-Pro: -1.89.
  • Defect found. In the merge that produced v1, inherited by v2. Corrected. The corrected v1 measures at its base model's level on MMLU-Pro (TIGER-Lab).
  • Not yet shown. The end state: a release that improves on its base on every axis. Each iteration is built to recover the regressions and raise the rest.

Why It Matters

  • Science. Whether trade-offs can be controlled, and then removed through iteration, is an open, testable question.
  • Economics. If the hypothesis holds, the marginal cost of a stronger open model is evaluation compute, not pre-training.
  • Adoption. Capability without dependence on one vendor, with a record of every model change.

The Goal

Top placements in class, across the public leaderboards.

  • Scope. 7B first. Then each larger size class and other model families.
  • Method. Each release controls the trade-offs between skills. Iteration then raises every axis: reasoning, mathematics, code, instruction following and knowledge.
  • End state. Improvement on every axis, with no statistically significant regression against the base model.
  • Done. A class is complete when no public fine-tune improves any axis further.

The Process

  • Start from a strong open base model in the class
  • Discover the compatible fine-tunes the community has published for it
  • Assimilate them into one model
  • Measure every axis against the base model, on identical items and public protocols
  • Release with every gain and every regression published, against the base model and the previous release
  • Iterate to recover the regressions and raise the weaker axes, until nothing more improves. Then the next class

The construction method is proprietary. The evaluation is public.

Releases

Published under this organisation from converge v3.

converge v3
Next iteration. Built to recover v2's regressions and raise the remaining axes
converge v4
Planned. The iteration after v3
Larger classes
Planned. The same process, at each larger size class and other model families

Research

The research behind the releases is on the founder's page, @Optitransfer.

The corrected models' boards are published when complete, whichever way they fall.

Evidence Standard

  • Paired. Every score is the model minus its base on the same items, with a 95% confidence interval and an exact significance test
  • Regressions reported. A card lists what went down as prominently as what went up
  • Calibration labelled. Calibrated means the harness first reproduced the base model's published score. Every result is marked either way
  • Public protocols. HELM-protocol GSM8K and MMLU, MMLU-Pro (TIGER-Lab), ZeroEval, EvalPlus HumanEval+ and MBPP+, MATH, AIME, ARC, IFEval
  • Corrected in public. When a defect is found, the affected cards say so first

Open Questions

  • Can trade-offs be controlled, and every axis then raised through iteration, across a full board?
  • Does it beat ensembling, routing and best-of-n sampling at matched compute?
  • Does it carry to larger classes and other model families?
  • Does it hold on calibration, hallucination and instruction following?

Each answer is published as it is measured.

Who It Is For

  • Open community. A stronger open model per class, with the evidence to check it
  • Regulated industries. Healthcare and finance teams whose fine-tunes cannot leave their boundary, and who must show how each model change was made
  • National programmes. Capability that grows by contribution, without dependence on a few vendors

Open model, paid guarantees. The public model stays open. The planned commercial offering is private consolidation of an organisation's own fine-tunes, inside its boundary, with the same measurement and evidence.

models 0

None public yet

datasets 0

None public yet