Instructions to use N8Programs/BestTerm-440M-Checkpts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use N8Programs/BestTerm-440M-Checkpts with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="N8Programs/BestTerm-440M-Checkpts")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("N8Programs/BestTerm-440M-Checkpts") model = AutoModelForCausalLM.from_pretrained("N8Programs/BestTerm-440M-Checkpts", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use N8Programs/BestTerm-440M-Checkpts with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "N8Programs/BestTerm-440M-Checkpts" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "N8Programs/BestTerm-440M-Checkpts", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/N8Programs/BestTerm-440M-Checkpts
- SGLang
How to use N8Programs/BestTerm-440M-Checkpts with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "N8Programs/BestTerm-440M-Checkpts" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "N8Programs/BestTerm-440M-Checkpts", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "N8Programs/BestTerm-440M-Checkpts" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "N8Programs/BestTerm-440M-Checkpts", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use N8Programs/BestTerm-440M-Checkpts with Docker Model Runner:
docker model run hf.co/N8Programs/BestTerm-440M-Checkpts
BestTerm-440M
BestTerm-440M improves short-sequence puzzle completion over NextTerm-440M while retaining similar OEIS and polynomial continuation performance. It is a 440.5M-parameter causal transformer specialized for comma-separated integer sequences, produced by continued pretraining on OEIS b-files and a weight-space merge with the released base model.
On the matched Ryskina & Knight comparison, greedy accuracy rises from 30/57 (52.63%) to 38/57 (66.67%). Beam-4 rises from 32/57 (56.14%) to 40/57 (70.18%). The improvements on this small puzzle benchmark come with much smaller changes on the broader evaluations below.
This repository contains the selected SLERP t=0.80 checkpoint. The original July 2026 results and a separate DGX Spark reproduction are reported independently. The documentation and supporting summaries were recovered and checked in September 2026; the scores are from those historical runs.
Direct comparison with NextTerm-440M
Sequence completion and forecasting
| Benchmark | NextTerm-440M control | BestTerm-440M: original evaluation | BestTerm-440M: DGX Spark reproduction |
|---|---|---|---|
| Ryskina & Knight, greedy โ | 30/57 ยท 52.63% | 38/57 ยท 66.67% | 38/57 ยท 66.67% |
| Ryskina & Knight, beam-4 โ | 32/57 ยท 56.14% | 40/57 ยท 70.18% | โ |
| OEIS-Eval-Neo โ | 6,555/19,034 ยท 34.438% | 6,532/19,034 ยท 34.318% | 6,536/19,034 ยท 34.339% |
| M1 Competition 111, macro MAPE โ | 17.623927 | 17.582548 | 17.813680 |
Interpretation: the matched greedy Ryskina gain is 8 questions / 14.04 percentage points. OEIS accuracy remains within 0.13 percentage points of the base control. The original M1 gain is only 0.0414 MAPE points and does not persist in the Spark reproduction. Naive2 scores 17.798674 macro MAPE in the saved M1 evaluation, slightly better than the reproduced BestTerm result.
The Ryskina controls use the matched Transformers greedy/beam runs. The initial MLX merge sweep instead scored the base at 29/57, with BestTerm still at 38/57; the two base results should not be interchanged. A separate later decoding sweep also reached 40/57 with beam-4 and beam-8, but is outside the main reproduction summary.
Sources: matched Ryskina results, original merge sweep, base OEIS, base M1, and Spark reproduction.
Polynomial continuation
| Sequence family | NextTerm-440M control | BestTerm: original evaluation | BestTerm: Spark reproduction |
|---|---|---|---|
| Arithmetic | 94.2917% | 94.5625% | 94.5417% ยท 4,538/4,800 |
| Quadratic | 86.3696% | 86.3043% | 86.2174% ยท 3,966/4,600 |
| Cubic | 74.8409% | 74.5682% | 74.5682% ยท 3,281/4,400 |
| Quartic | 68.1190% | 67.9524% | 68.1905% ยท 2,864/4,200 |
These results show preservation of the base model's polynomial continuation ability, with small gains and losses across families. They aggregate over prefix lengths; they are not scores at one fixed prompt length. The recovered experiment controls differ slightly from the older published NextTerm-440M model-card values.
Sources: base polynomial results, original merge sweep, and Spark polynomial results by prefix length.
Context among other models
The following historical comparison comes from the NextTerm-440M model card. BestTerm's row uses its Spark greedy reproduction. The other rows are previously reported results, not fresh matched reruns; the direct control comparisons above are preferable for measuring this intervention.
| Model | Ryskina & Knight, greedy โ | OEIS-Eval-Neo โ | M1 macro MAPE โ |
|---|---|---|---|
| BestTerm-440M | 66.67% | 34.34% | 17.8137 |
| NextTerm-440M, published | 52.63% | 34.43% | 17.6239 |
| NextTerm-47M | 70.18% | 29.49% | 18.7621 |
| Qwen3-0.6B | 33.33% | 18.44% | 22.7984 |
| Qwen3-1.7B | 49.12% | 20.77% | 22.2411 |
| Qwen3-4B | 63.16% | 23.74% | 19.1731 |
| Qwen3-8B | 57.89% | 24.62% | 18.4027 |
| Qwen3-14B | 59.65% | 26.00% | 17.9837 |
NextTerm-47M remains ahead on greedy Ryskina. BestTerm reaches the same 70.18% score with beam-4, which uses additional decoding compute. OEIS-Eval-Neo is also distinct from the older OEIS-Eval benchmark; their scores should not be mixed.
How BestTerm was made
- Base pretraining: NextTerm-440M was trained for 13,999,999,995 tokens on extended OEIS sequences plus synthetic augmentation, using preserved sequence prefixes and a 4,096-token training cap.
- B-file-only continued pretraining: the synthetic augmentation was removed for this phase, using
oeis_train_bfile_prefix4096.packed. OEIS b-files provide extended lists of sequence terms. The selected intermediate checkpoint had seen 500,020,611 additional tokens over 20,486 steps (training metadata), from a planned 754,629,901-token run. The internal name hot500 refers to the higher learning-rate schedule and approximately 500M tokens, not sampling temperature. The recipe used a Muon/AdamW hybrid with peak learning rates 0.01/0.0001, length-bucketed batches, and BF16 live weights with FP32 master weights. - Global SLERP merge: the base and hot500 parameter vectors were combined at
t=0.80, toward hot500. The recorded coefficients are 0.2010891797 ร base + 0.8016321071 ร hot500. These coefficients come from spherical interpolation, so this is not an exact 20/80 linear average. See merge metadata.
Why this merge?
| Selected candidate | Ryskina greedy | OEIS-Eval-Neo | M1 macro MAPE โ |
|---|---|---|---|
t=0.45 |
35/57 | 34.517% | 18.469475 |
t=0.60 |
36/57 | 34.543% | 17.775277 |
t=0.70 |
37/57 | 34.428% | 17.821592 |
t=0.80 โ this repository |
38/57 | 34.318% | 17.582548 |
t=0.90 |
38/57 | 34.254% | 17.851500 |
t=0.80 was selected for its Ryskina/M1 tradeoff; t=0.60 was the more conservative OEIS-preservation point. These benchmarks were used during checkpoint selection. The sweep therefore does not constitute an untouched evaluation of a preselected model. All rows in this table are original sweep results.
Model and input format
| Property | Value |
|---|---|
| Parameters | 440,500,224 |
| Architecture | Qwen3-style causal LM |
| Layers / hidden size / FFN size | 28 / 1,024 / 3,072 |
| Attention heads / KV heads | 16 / 8 |
| Vocabulary | 16 tokens |
| Training sequence cap | 4,096 tokens |
| Export dtype | BF16 |
| Weight file | 881,035,528 bytes, approximately 840 MiB |
Use comma-separated integers, for example 1,2,4,8, or 1,-2,3,-4,. Decimal digits are tokenized individually; large integers consume more tokens. The tokenizer ignores characters outside digits, comma, and minus. Natural-language instructions and chat templates are inappropriate for this vocabulary. Leading-zero forms such as 01,02,03, were outside the training format.
The exported configuration permits 40,960 positions, but this is not evidence of training or evaluation at that context length.
Usage
Transformers: predict one integer
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "N8Programs/BestTerm-440M-Checkpts"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo,
torch_dtype=torch.bfloat16,
device_map="auto",
).eval()
inputs = tokenizer("1,2,4,8,", return_tensors="pt", return_token_type_ids=False)
inputs = {name: value.to(model.device) for name, value in inputs.items()}
comma_id = tokenizer.convert_tokens_to_ids(",")
with torch.inference_mode():
output = model.generate(
**inputs,
do_sample=False,
num_beams=1,
max_new_tokens=64,
eos_token_id=[comma_id, tokenizer.eos_token_id],
pad_token_id=tokenizer.pad_token_id,
)
completion = tokenizer.decode(
output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)
print(completion.split(",", 1)[0].strip())
For the reported beam-4 decoding convention, use num_beams=4, eos_token_id=comma_id, and suppress_tokens=[tokenizer.eos_token_id, tokenizer.pad_token_id]. This stops at the next comma while suppressing premature EOS/PAD. The snippet is a single-prompt usage example; the reported benchmarks use the protocols below.
MLX: generate a continuation
mlx_lm.generate \
--model N8Programs/BestTerm-440M-Checkpts \
--prompt "1,2,4,8," \
--temp 0 \
--max-tokens 64
Evaluation protocols and reproducibility
- Ryskina & Knight: 57 short sequence-completion puzzles, exact next-term accuracy. Spark greedy evaluation stops at comma/EOS/PAD with a 20-token generation cap. The original beam-4 comparison uses comma stopping and suppresses EOS/PAD.
- OEIS-Eval-Neo: 19,034 held-out next-term examples, described as decontaminated in the base model release. The Spark run uses a 4,096-token context cap and 196 new-token cap; all examples were evaluated and parsed, with none skipped. The recovered predictions were checked against their saved answers, confirming 6,536 exact matches. Decontamination was not independently re-audited during this documentation recovery.
- Polynomials: 200 generated sequences per prefix length, seed 0; lengths 2โ25 for arithmetic, 3โ25 for quadratic, 4โ25 for cubic and 5โ25 for quartic. Exact next-term match gives 4,800 / 4,600 / 4,400 / 4,200 examples. See the base release's evaluation script for coefficient sampling.
- M1 Competition 111: 68 monthly series with horizon 18, 23 quarterly with horizon 8, and 20 yearly with horizon 6. Macro MAPE weights each series equally. The reproduced score is 17.813680; point-weighted MAPE is 18.409850. All 111 series parsed fully, with no missing terms or malformed series.
The Spark reproduction used MLX 0.31.1, mlx-cuda-13 0.31.1 and mlx-lm 0.31.2, with CUDA graphs disabled. OEIS batch/queue sizes were 32/512, M1 queue size was 16, and polynomial queue size was 64. The original evaluation had different execution settings. Small numerical and generation differences are visible across runs; their exact cause was not isolated.
The benchmark summaries, comparison CSV, and provenance manifest are included here. Absolute machine-specific paths in copied summaries were reduced to filenames; metrics and settings were preserved. Benchmark datasets and original evaluator implementations are available in the NextTerm-440M repository.
Weight SHA-256, matching the recovered and published checkpoint:
a6a4244107d621c7138060d37168345b9e83a2e8036ccf4ccbcd465b048068f2
Limitations and attribution
This is a specialized integer-sequence model. A finite prefix can admit multiple plausible rules, so a generated continuation need not be the intended one. Improvements on 57 selected puzzles do not establish universal rule induction, and the merge does not improve every benchmark. The reported M1 advantage is sensitive to the evaluation run.
The checkpoint inherits training on the On-Line Encyclopedia of Integer Sequences through NextTerm-440M and its b-file-only continuation phase. See the base release for dataset attribution and licensing information.
- Downloads last month
- 563
Model tree for N8Programs/BestTerm-440M-Checkpts
Base model
N8Programs/NextTerm-440M