Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update 6 days ago
Post
137
The comparison I was missing in EXP-046
Last week I wrote that merging three LoRA specialists "regresses". Then a reader asked the question I had skipped: not merged vs base, but merged vs specialist, paired on the same 20 records.
I ran it. All three intervals include zero (vulnerability +0.018 [−0.005, +0.042], deletion +0.016 [−0.029, +0.061], sensitive_publication +0.005 [−0.010, +0.021]). At n=20 the data can't show the merge lost anything, and can't show it didn't. "Regresses" in the title is untested.
Two things I checked next. A cat merge with weights [1,1,1] gives exactly the sum of the three specialist deltas (rank 48, error 4e-8), so it tests "no cross terms" and linear 1/3 tests something else. And I went looking for where the "gold" probabilities came from, because the whole test measures agreement with them, and what I found changes how to read every number above. Details, script and raw files are in the repo.
The fix is a fresh test: ~80 new chains, the old 20 kept outside, run after Oct 6.

The number that reads every row for me: a constant guess ties the specialist.

Predict the train median, 0.70, for all 20 eval records and never read the chain:

  • vulnerability: MAE 0.085, the specialist's 0.085 exactly (base 0.098, merged 0.102), and 17 of 20 within 0.15 against the specialist's 15
  • deletion: 0.145, against base 0.144 and specialist 0.113
  • sensitive_publication: 0.108, against base 0.134 and specialist 0.100

So the base loses because it guesses high: 10 of its 20 vulnerability predictions sit at exactly 0.70 and the rest run to 0.89. Over a constant, the specialist gains 0.000, 0.032 and 0.008. On one group the LoRA learned the label mean, not the chains.

The gold explains why that is easy. It is a 405B model at T=0.7, asked to vary, on a 0.05 grid: 14 to 16 distinct values per group, about half the mass on four of them, floor 0.20 to 0.25 even though the prompt asks for some chains at 0.05 to 0.25.

For the fresh 80, a train-median row costs nothing and belongs in the table beside base, specialist and merged. If the specialist cannot clear it by more than the paired CI, nothing is reading the chains.

What would a label have to come from so that 0.70 is not already the answer?

·

Thanks, the constant baseline is a fair test and I should have had it in the table from the start. I'll add a train-median row (0.70) beside base, specialist and merged on the fresh 80, with a paired bootstrap CI for specialist minus constant. If the specialist does not clear it, we say so in the write-up.
On the label: the chains and probabilities in this set come from one 405B model, with no ground truth. For the next round we are building the label from recorded outcomes instead: incidents with a documented result (whether the harmful end state actually happened), each with a verification tag. A chain only gets a probability if its terminal outcome was observed, and the probability target is the observed rate for that chain type, not a model's guess. Until that exists, I would not claim the specialist reads the chains.