← Kurt Pieper
Recommendations & Commentary
This contains interesting articles and sketched thoughts,
which I hope will be helpful. All polished writing is on LW.
Can activation verbalizers surface an internal chain of thought?
A bunch of negative results, basically. Loss not good, Claude can few-shot
comparable explanations. Explanations are meh, and more speculatively,
it might be due the AV reading shallow token representations
(“ah, this activation vector represents this and that input”), and
reconstructing what the model “ought to be thinking”, instead of
going deeply into the activation vector to see what is actually being
thought. How the hell would we even know? It must be said that the
math problems are hard though, and NLA’s are bad with the numbers in
general.
NLA explanations can be shortened without harming reconstruction
The bag of NLA-selection forces contains the warm-start (Claude’s
guess as to what the model might think), reconstruction loss and the KL penalty
of GRPO (which might improve legibility via “keep it close to
pretraining”), parsimony is a nice addition, yet it remains to be
demonstrated that the explanations are not just shorter, but also
better (see also my draft).
Sleeper Agent Backdoor Results Are Messy
Making AI’s talk like a pirate is competetive to HHH. One could spin
this against the PSM: Behaviour that is fine-tuned in can be removed by
knocking the model off-distribution.
Furthermore: One could ambitiously
resample different models, i.e.
make them talk in a bunch of different distributions (i.e. talk like a pirate,
talk in French …), and just hope they’re not all scheming in the
same way. Unfortunately, since this is likely pretty capabilities-costly,
compared to e.g. monitors, Pirate resampling may never see the light of day at
frontier labs.
Reward Hacking Without Egregious Misalignment in an RL-Only Setting
We can RL-fine-tune models in a specific, bad way, and they don’t
become bad in general. To me, it is still confusing why mechanistically
small weight changes via supervised fine-tuning seem to be different from small
weight changes via GRPO. Perhaps if one make the LoRA dimensions really small
for both things, i.e. updating just a few parameters, one could mechanistically
locate the “persona selection circuitry”, see
Persona Cartography: Charting Language Model Personality Traits in Weight Space
Might ASI’s just be good by default (i.e. via strong moral
internalism: any sufficiently intelligent agent becomes ‘good’
upon learning what that is)? If the human utility function is an arbitrary
product of inclusive-genetic-fitness maximization, how can alignment be
justified under a universalistic framework? Read further below, or
my draft :)
Relatedly, what have we learned about
the idea of expected utility,
or should we accept the VNM-model of intelligent agency a priori?
Exercise: Try to think of falsifiable predictions the EUM-model made of humans
or LLMs.
Does the
information bottleneck of RLVR
(Toby Ord) represent a natural Occam’s razor, a blessing in disguise?
The
2021 MIRI Conversation consensus
seems to be that humans had a similar helper,
Humans face the genomic bottleneck which means that each individual has to
rederive all the knowledge about the world that their parents already had. If
this genetic bottleneck hadn’t been so tight, then individual humans
would have been significantly less capable of performing novel tasks.
Furthermore, we recently learned that humans got smarter a bunch only
recently, i.e. with just a few bits, see:
Ancient DNA reveals pervasive directional selection across West Eurasia
(Nature)
Similarly, it is possible to train reasoning models by
updating just 13 Parameters.
We can collect evidence by asking how strong the
model’s positive manifold (Epoch
Capabilities Index) is, i.e. their anti-jagged frontier, comparing between
reasoning and non-reasoning models. That is, how strongly predictive of
downstream benchmark scores is the general factor, i.e. the Epoch Capabilities
Index. There seems to be no difference between reasoning and non-reasoning
models here (detailed evaluation will be added here soon), and holistically I
would say RLVR generalizes more weakly, despite all the fancy arguments
above.
Intelligence – A Very Short Introduction
(Oxford University Press)
Basically all of Intelligence Psychometrics. But probably you could still
boil it down to ten Pearson’s r’s.
Esketamine for treatment-resistant depression: seven concerns about efficacy and FDA approval
(Lancet Psychiatry)
On Johnson & Johnson’s polish outlier sides. But ketamine has
other things than four MADRS points going for it :).
Behavioural therapies versus other psychological therapies for depression
(Cochrane)
As ‘behavioural activation’ (erm, doing more stuff?) is part of
CBT, with seemingly equal effectiveness, one can get the favourite razor
out.
Genetics of attention deficit disorder
(Nature)
Most of the variance in teacher-annoying among all school-children
explained by genes,
Falconer equations say
Geeks, Mops, Sociopaths
Absolute post-rat classic, though not to be taken literally.
Exercise: Come up with different words for Geeks and Sociopaths, there are
many. For example, Genius in it’s meaning before Galton’s
Hereditary Genius, i.e. as creator of culture, and Opportunist, which captures
the fact that people in question lack clinically-sociopathic traits, yet
doesn’t represent the competence with which they swing up in social
networks (e.g. if they were on an island …)
Crony Beliefs
Meltingasphalt seems less known than he should be. The most useful part is
probably this part on how to indentify the crony beliefs.
- Abstract and impractical. Merit beliefs have value
only insofar as we’re able to make use of them for choosing actions;
we need some “skin in the game.” If a belief isn’t
actionable, or if the actions we might take based on the belief (e.g.,
voting) don’t provide material benefits one way or the other, then
it’s more likely to be a crony.
- Benefit of the doubt. When we have social incentives
to believe something, we stack the deck in its favor. Or to use another
metaphor, we put our thumbs on the scale as we weigh the evidence. Blind
faith — religious, political, or otherwise — is simply
“benefit of the doubt” taken to its logical extreme.
- Conspicuousness. The whole point of a crony belief is
to reap social and political rewards, but in order to get these rewards, we
need to advertise the belief in question. So the greater our urge to talk
about a belief, to wear it like a badge, the more likely it is to be a
crony.
- Overconfidence. Related to the above, crony beliefs
will typically provide more social value the more confident we seem in
them. (If Acme hires the mayor’s nephew, but seems constantly on the
verge of firing him, the mayor isn’t going to be happy.)
Overconfidence also acts as a form of protection for beliefs that
can’t survive on their own within the meritocracy.
- Reluctance to bet. Betting on a belief is just as good
as acting on it; both mechanisms create incentives for accuracy. If
we’re reluctant to bet on a belief, then, it’s often because
some parts of our psyche know that the belief is unlikely to be true. Hence
the challenge: “Put up or shut up.