DesignCoder

Checkpoint collection for DesignCoder, a family of full-parameter SFT models for UI design research and end-to-end HTML/CSS/JavaScript implementation.

Each subfolder in this repository is a self-contained, directly loadable checkpoint.

Naming convention

designcoder_{basemodel}_{size}_{optimizer}_bs{global_batch}[_{extra_axes}]_step{global_step}
  • basemodel / size: base model family and parameter scale
  • optimizer: muon or adamw
  • bs: global batch size (per_device Γ— grad_accum Γ— world_size)
  • extra_axes: any hyper-parameter that deviates from the default recipe, e.g. wd0.05 (weight decay, default 0.0), ep20 (epochs, default 2), or data41287 (dataset revision)
  • step: trainer global_step of the exported weights

Dataset revisions

Checkpoints in this repository come from two different dataset revisions. Scores and loss values are only comparable within the same revision.

Tag Samples Used by
(untagged) data37865 37,865 *_step1900, *_step3800
data41287 41,287 *_data41287_step200, *_data41287_step400

Checkpoints

Subfolder Base model Optimizer LR Global batch Dataset Step bench-200 (full, n=200) Notes
designcoder_qwen3.5_4b_muon_bs32_step1900 Qwen3.5-4B Muon 1e-5 32 37,865 1900 – smallest of the first release
designcoder_qwen3.5_9b_muon_bs16_step3800 Qwen3.5-9B Muon 1e-5 16 37,865 3800 – optimizer ablation (Muon arm)
designcoder_qwen3.5_9b_adamw_bs16_step3800 Qwen3.5-9B AdamW 2e-5 16 37,865 3800 – optimizer ablation (AdamW arm)
designcoder_qwen3.6_27b_adamw_bs32_step1900 Qwen3.6-27B AdamW 1e-5 32 37,865 1900 – largest of the first release
designcoder_qwen3.5_4b_adamw_bs256_data41287_step200 Qwen3.5-4B AdamW 2e-5 256 41,287 200 82.12 best 4B / AdamW
designcoder_qwen3.5_4b_muon_bs256_data41287_step200 Qwen3.5-4B Muon 2e-5 256 41,287 200 77.62 best 4B / Muon
designcoder_qwen3.5_9b_adamw_bs256_data41287_step200 Qwen3.5-9B AdamW 2e-5 256 41,287 200 84.40 best 9B
designcoder_qwen3.8_27b_adamw_bs128_data41287_step400 Qwen3.8-27B AdamW 1e-5 128 41,287 400 87.89 strongest checkpoint in the collection

All four data41287 scores are final full-benchmark runs: 200/200 rollouts, 200/200 screenshot captures, 200/200 judge evaluations per model (no subsetting).

Benchmark

bench-200 is the frozen 200-case DesignCoder benchmark (100 Track A landing, 40 Track A dashboard, 30 Track B landing, 30 Track B dashboard; Track A cases specify a style, Track B cases are style-free).

Rubric composition. Every prompt ships with its own reference rubric of 23–25 binary screenshot checks (184 prompts carry 25 checks, 15 carry 24, 1 carries 23 β€” 4,983 frozen checks in total), all evaluated with check_with=screenshot by a vision judge over the full-page render. Check distribution across dimensions:

Dimension Checks Share
Components 1,517 30.4%
Layout 842 16.9%
Aesthetics 782 15.7%
Typography 642 12.9%
Alignment 616 12.4%
Assets 584 11.7%

On top of the frozen checks, the judge scores 5 surface-specific Prompt-Fit items (0–2 each) per case. The reported overall_score (0–100) is the unweighted mean of Prompt Fit and the six rubric dimensions. Judge: gpt-5.6-sol (vision) with structured JSON output.

Full-run results (n=200 per model)

Model Overall Landing Dashboard Track A Track B Prompt Fit Frozen pass rate Render fails
27B AdamW step400 87.89 88.55 86.67 86.90 90.20 84.10 88.6% 0/200
9B AdamW step200 84.40 86.31 80.85 84.10 85.09 78.35 85.4% 1/200
4B AdamW step200 82.12 85.07 76.65 82.39 81.50 75.25 83.2% 1/200
4B Muon step200 77.62 81.65 70.14 77.64 77.58 64.60 79.4% 3/200

Scores increase strictly monotonically with scale (all 6 pairwise differences significant, paired bootstrap 10k-resample 95% CI excludes 0 and Wilcoxon p < 0.013 β€” see eval/significance_tests.json). The gap is far larger on dashboards (+16.5 from 4B Muon to 27B) than on landings (+6.9), and Assets is the weakest dimension for every scale (55–67% pass rate), indicating a data-level bottleneck rather than a capability ceiling.

Evaluation artifacts (eval/)

File Content
eval/benchmark_summary.csv per-model aggregates: overall, Track/Surface splits, six dimensions, Prompt Fit, frozen pass rate
eval/benchmark_per_case.csv long-form per-case scores for all 4 models Γ— 200 cases
eval/significance_tests.json paired bootstrap (10k resamples) + Wilcoxon signed-rank for all 6 model pairs
eval/rubric_stats.json rubric composition statistics (checks per prompt, per dimension, per track)
eval/reports.html self-contained interactive HTML report: model comparison, dimension heatmap, score distributions, per-case tables

Checkpoint selection

The data41287 checkpoints were selected by running the benchmark, not by taking the lowest training loss. In all four runs the best checkpoint sits at roughly 75% of training, and loss kept improving while benchmark scores fell. The table below shows the 8-case selection subset (used only to rank checkpoints, not comparable to the final full-run numbers in the tables above):

Run Step Train loss subset bench (n=8, selection only) final full bench (n=200)
4B AdamW 200 0.2696 84.22 82.12
4B AdamW 266 0.2682 68.35 –
4B Muon 200 0.3339 83.36 77.62
4B Muon 266 0.3349 81.27 –
9B AdamW 200 0.2518 84.40 84.40
9B AdamW 266 0.2504 lowest of the three –
27B AdamW 400 0.2067 91.19 87.89
27B AdamW 530 0.2059 86.37 –

The 4B AdamW pair is the clearest example: loss improved from 0.2696 to 0.2682 while the subset score collapsed from 84.22 to 68.35. Do not pick checkpoints from this family by loss. Note also that small subsets systematically overestimate: the subset ranks checkpoints correctly but runs several points above the full 200-case benchmark.

Shared training setup

  • Objective: full-parameter supervised fine-tuning (no LoRA / adapters)
  • Dataset: designcoder_sft_v2_train in ShareGPT format (see revision table above)
  • Chat template: qwen3_5 with thinking enabled
  • Context length: 32,768
  • Sequence packing: enabled, with neat packing (no cross-sample attention)
  • LR schedule: cosine, warmup ratio 0.1

Usage

from transformers import AutoModelForCausalLM, AutoProcessor

repo = "xingxm/DesignCoder"
subfolder = "designcoder_qwen3.8_27b_adamw_bs128_data41287_step400"

model = AutoModelForCausalLM.from_pretrained(repo, subfolder=subfolder, dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(repo, subfolder=subfolder)

To download a single checkpoint only:

hf download xingxm/DesignCoder --include "designcoder_qwen3.8_27b_adamw_bs128_data41287_step400/*" --local-dir ./DesignCoder

Inference contract

These models are trained as tool-using agents, not single-turn generators. A case runs design_search β†’ (websearch, landing only) β†’ a final answer containing exactly three code blocks in the order html, css, js. Reproduce the system prompts and tool observation format from examples/designcoder/runtime/infer_designcoder.py; prompting with a bare instruction and no tool turns does not match the training distribution and will score far below the numbers above.

Provenance

Each subfolder additionally ships trainer_state.json / trainer_log.jsonl (and training_loss.png where available) so that the loss curve and exact step schedule of the run can be recovered from the checkpoint itself.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support