AI & ML interests
None defined yet.
Recent Activity
Alignment ā and the assistant identity itself ā is normally introduced only after
pretraining, once behavioral priors are already set. SPP installs the desired persona from
token zero instead: we define it through normative values in a constitution, generate
first-person moral reflections grounded in that constitution, and insert them throughout the
pretraining corpus behind an <assistant> token. Post-training then binds the chat assistant
identity to the installed persona. Pretraining up to 3B on 500B tokens, SPP improves
constitution following and jailbreak robustness while preserving capabilities ā and when the
data arrives matters: models trained with reflections from token zero prioritize values
differently and take fewer risky actions in out-of-distribution moral dilemmas than models
given the exact same data only at the end of pretraining, an advantage that grows with scale.
Collections
š¦ Pretraining Datasets ā the reflection data, the corpus selection manifest, safety scores, and verification files.
š¤ Models ā 3B Ā· Models ā 1.7B ā all five recipes, base and instruct, at both scales.
š¬ Post-training Dataset ā SP-SFT, the mixture that performs persona binding.
š Evals ā ConstitutionEval and an audited AIRiskDilemmas.
From EPFL DLAB.