CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Abstract
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.
Community
A very interesting take on agent data synthesis: instead of just generating and validating executable tasks, CalibForge uses solver behavior to actively reshape tasks into a solver-relative “learnable zone.” The idea of calibrating task difficulty through multi-solver disagreement or strong-vs-weak contrast feels simple but powerful, and the downstream gains suggest that which tasks we train on matters as much as how many we generate.
HF: https://huggingface.co/datasets/AweAI-Team/CalibForge
Model: https://huggingface.co/collections/AweAI-Team/calibforge
Github: https://github.com/AweAI-Team/CalibForge
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents (2026)
- SETA: Scaling Environments for Terminal Agents (2026)
- NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs (2026)
- Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading (2026)
- VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct (2026)
- KAT-Coder-V2.5 Technical Report (2026)
- Cross-Benchmark Generalization in Long-Horizon Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.06352 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
AweAI-Team/CalibForge-35B-A3B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper