BestTerm-440M

BestTerm-440M improves short-sequence puzzle completion over NextTerm-440M while retaining similar OEIS and polynomial continuation performance. It is a 440.5M-parameter causal transformer specialized for comma-separated integer sequences, produced by continued pretraining on OEIS b-files and a weight-space merge with the released base model.

On the matched Ryskina & Knight comparison, greedy accuracy rises from 30/57 (52.63%) to 38/57 (66.67%). Beam-4 rises from 32/57 (56.14%) to 40/57 (70.18%). The improvements on this small puzzle benchmark come with much smaller changes on the broader evaluations below.

Benchmark comparison: Ryskina, OEIS-Eval-Neo, M1 forecasting, and polynomial continuation

This repository contains the selected SLERP t=0.80 checkpoint. The original July 2026 results and a separate DGX Spark reproduction are reported independently. The documentation and supporting summaries were recovered and checked in September 2026; the scores are from those historical runs.

Direct comparison with NextTerm-440M

Sequence completion and forecasting

Benchmark NextTerm-440M control BestTerm-440M: original evaluation BestTerm-440M: DGX Spark reproduction
Ryskina & Knight, greedy โ†‘ 30/57 ยท 52.63% 38/57 ยท 66.67% 38/57 ยท 66.67%
Ryskina & Knight, beam-4 โ†‘ 32/57 ยท 56.14% 40/57 ยท 70.18% โ€”
OEIS-Eval-Neo โ†‘ 6,555/19,034 ยท 34.438% 6,532/19,034 ยท 34.318% 6,536/19,034 ยท 34.339%
M1 Competition 111, macro MAPE โ†“ 17.623927 17.582548 17.813680

Interpretation: the matched greedy Ryskina gain is 8 questions / 14.04 percentage points. OEIS accuracy remains within 0.13 percentage points of the base control. The original M1 gain is only 0.0414 MAPE points and does not persist in the Spark reproduction. Naive2 scores 17.798674 macro MAPE in the saved M1 evaluation, slightly better than the reproduced BestTerm result.

The Ryskina controls use the matched Transformers greedy/beam runs. The initial MLX merge sweep instead scored the base at 29/57, with BestTerm still at 38/57; the two base results should not be interchanged. A separate later decoding sweep also reached 40/57 with beam-4 and beam-8, but is outside the main reproduction summary.

Sources: matched Ryskina results, original merge sweep, base OEIS, base M1, and Spark reproduction.

Polynomial continuation

Sequence family NextTerm-440M control BestTerm: original evaluation BestTerm: Spark reproduction
Arithmetic 94.2917% 94.5625% 94.5417% ยท 4,538/4,800
Quadratic 86.3696% 86.3043% 86.2174% ยท 3,966/4,600
Cubic 74.8409% 74.5682% 74.5682% ยท 3,281/4,400
Quartic 68.1190% 67.9524% 68.1905% ยท 2,864/4,200

These results show preservation of the base model's polynomial continuation ability, with small gains and losses across families. They aggregate over prefix lengths; they are not scores at one fixed prompt length. The recovered experiment controls differ slightly from the older published NextTerm-440M model-card values.

Sources: base polynomial results, original merge sweep, and Spark polynomial results by prefix length.

Context among other models

The following historical comparison comes from the NextTerm-440M model card. BestTerm's row uses its Spark greedy reproduction. The other rows are previously reported results, not fresh matched reruns; the direct control comparisons above are preferable for measuring this intervention.

Model Ryskina & Knight, greedy โ†‘ OEIS-Eval-Neo โ†‘ M1 macro MAPE โ†“
BestTerm-440M 66.67% 34.34% 17.8137
NextTerm-440M, published 52.63% 34.43% 17.6239
NextTerm-47M 70.18% 29.49% 18.7621
Qwen3-0.6B 33.33% 18.44% 22.7984
Qwen3-1.7B 49.12% 20.77% 22.2411
Qwen3-4B 63.16% 23.74% 19.1731
Qwen3-8B 57.89% 24.62% 18.4027
Qwen3-14B 59.65% 26.00% 17.9837

NextTerm-47M remains ahead on greedy Ryskina. BestTerm reaches the same 70.18% score with beam-4, which uses additional decoding compute. OEIS-Eval-Neo is also distinct from the older OEIS-Eval benchmark; their scores should not be mixed.

How BestTerm was made

  1. Base pretraining: NextTerm-440M was trained for 13,999,999,995 tokens on extended OEIS sequences plus synthetic augmentation, using preserved sequence prefixes and a 4,096-token training cap.
  2. B-file-only continued pretraining: the synthetic augmentation was removed for this phase, using oeis_train_bfile_prefix4096.packed. OEIS b-files provide extended lists of sequence terms. The selected intermediate checkpoint had seen 500,020,611 additional tokens over 20,486 steps (training metadata), from a planned 754,629,901-token run. The internal name hot500 refers to the higher learning-rate schedule and approximately 500M tokens, not sampling temperature. The recipe used a Muon/AdamW hybrid with peak learning rates 0.01/0.0001, length-bucketed batches, and BF16 live weights with FP32 master weights.
  3. Global SLERP merge: the base and hot500 parameter vectors were combined at t=0.80, toward hot500. The recorded coefficients are 0.2010891797 ร— base + 0.8016321071 ร— hot500. These coefficients come from spherical interpolation, so this is not an exact 20/80 linear average. See merge metadata.

Why this merge?

Selected candidate Ryskina greedy OEIS-Eval-Neo M1 macro MAPE โ†“
t=0.45 35/57 34.517% 18.469475
t=0.60 36/57 34.543% 17.775277
t=0.70 37/57 34.428% 17.821592
t=0.80 โ€” this repository 38/57 34.318% 17.582548
t=0.90 38/57 34.254% 17.851500

t=0.80 was selected for its Ryskina/M1 tradeoff; t=0.60 was the more conservative OEIS-preservation point. These benchmarks were used during checkpoint selection. The sweep therefore does not constitute an untouched evaluation of a preselected model. All rows in this table are original sweep results.

Model and input format

Property Value
Parameters 440,500,224
Architecture Qwen3-style causal LM
Layers / hidden size / FFN size 28 / 1,024 / 3,072
Attention heads / KV heads 16 / 8
Vocabulary 16 tokens
Training sequence cap 4,096 tokens
Export dtype BF16
Weight file 881,035,528 bytes, approximately 840 MiB

Use comma-separated integers, for example 1,2,4,8, or 1,-2,3,-4,. Decimal digits are tokenized individually; large integers consume more tokens. The tokenizer ignores characters outside digits, comma, and minus. Natural-language instructions and chat templates are inappropriate for this vocabulary. Leading-zero forms such as 01,02,03, were outside the training format.

The exported configuration permits 40,960 positions, but this is not evidence of training or evaluation at that context length.

Usage

Transformers: predict one integer

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "N8Programs/BestTerm-440M-Checkpts"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    torch_dtype=torch.bfloat16,
    device_map="auto",
).eval()

inputs = tokenizer("1,2,4,8,", return_tensors="pt", return_token_type_ids=False)
inputs = {name: value.to(model.device) for name, value in inputs.items()}
comma_id = tokenizer.convert_tokens_to_ids(",")

with torch.inference_mode():
    output = model.generate(
        **inputs,
        do_sample=False,
        num_beams=1,
        max_new_tokens=64,
        eos_token_id=[comma_id, tokenizer.eos_token_id],
        pad_token_id=tokenizer.pad_token_id,
    )

completion = tokenizer.decode(
    output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)
print(completion.split(",", 1)[0].strip())

For the reported beam-4 decoding convention, use num_beams=4, eos_token_id=comma_id, and suppress_tokens=[tokenizer.eos_token_id, tokenizer.pad_token_id]. This stops at the next comma while suppressing premature EOS/PAD. The snippet is a single-prompt usage example; the reported benchmarks use the protocols below.

MLX: generate a continuation

mlx_lm.generate \
  --model N8Programs/BestTerm-440M-Checkpts \
  --prompt "1,2,4,8," \
  --temp 0 \
  --max-tokens 64

Evaluation protocols and reproducibility

  • Ryskina & Knight: 57 short sequence-completion puzzles, exact next-term accuracy. Spark greedy evaluation stops at comma/EOS/PAD with a 20-token generation cap. The original beam-4 comparison uses comma stopping and suppresses EOS/PAD.
  • OEIS-Eval-Neo: 19,034 held-out next-term examples, described as decontaminated in the base model release. The Spark run uses a 4,096-token context cap and 196 new-token cap; all examples were evaluated and parsed, with none skipped. The recovered predictions were checked against their saved answers, confirming 6,536 exact matches. Decontamination was not independently re-audited during this documentation recovery.
  • Polynomials: 200 generated sequences per prefix length, seed 0; lengths 2โ€“25 for arithmetic, 3โ€“25 for quadratic, 4โ€“25 for cubic and 5โ€“25 for quartic. Exact next-term match gives 4,800 / 4,600 / 4,400 / 4,200 examples. See the base release's evaluation script for coefficient sampling.
  • M1 Competition 111: 68 monthly series with horizon 18, 23 quarterly with horizon 8, and 20 yearly with horizon 6. Macro MAPE weights each series equally. The reproduced score is 17.813680; point-weighted MAPE is 18.409850. All 111 series parsed fully, with no missing terms or malformed series.

The Spark reproduction used MLX 0.31.1, mlx-cuda-13 0.31.1 and mlx-lm 0.31.2, with CUDA graphs disabled. OEIS batch/queue sizes were 32/512, M1 queue size was 16, and polynomial queue size was 64. The original evaluation had different execution settings. Small numerical and generation differences are visible across runs; their exact cause was not isolated.

The benchmark summaries, comparison CSV, and provenance manifest are included here. Absolute machine-specific paths in copied summaries were reduced to filenames; metrics and settings were preserved. Benchmark datasets and original evaluator implementations are available in the NextTerm-440M repository.

Weight SHA-256, matching the recovered and published checkpoint:

a6a4244107d621c7138060d37168345b9e83a2e8036ccf4ccbcd465b048068f2

Limitations and attribution

This is a specialized integer-sequence model. A finite prefix can admit multiple plausible rules, so a generated continuation need not be the intended one. Improvements on 57 selected puzzles do not establish universal rule induction, and the merge does not improve every benchmark. The reported M1 advantage is sensitive to the evaluation run.

The checkpoint inherits training on the On-Line Encyclopedia of Integer Sequences through NextTerm-440M and its b-file-only continuation phase. See the base release for dataset attribution and licensing information.

Downloads last month
563
Safetensors
Model size
0.4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for N8Programs/BestTerm-440M-Checkpts

Finetuned
(1)
this model

Dataset used to train N8Programs/BestTerm-440M-Checkpts