← Kurt Pieper

Recommendations & Commentary

This contains interesting articles and sketched thoughts, which I hope will be helpful. All polished writing is on LW.

Can activation verbalizers surface an internal chain of thought?

A bunch of negative results, basically. Loss not good, Claude can few-shot comparable explanations. Explanations are meh, and more speculatively, it might be due the AV reading shallow token representations (“ah, this activation vector represents this and that input”), and reconstructing what the model “ought to be thinking”, instead of going deeply into the activation vector to see what is actually being thought. How the hell would we even know? It must be said that the math problems are hard though, and NLA’s are bad with the numbers in general.

NLA explanations can be shortened without harming reconstruction

The bag of NLA-selection forces contains the warm-start (Claude’s guess as to what the model might think), reconstruction loss and the KL penalty of GRPO (which might improve legibility via “keep it close to pretraining”), parsimony is a nice addition, yet it remains to be demonstrated that the explanations are not just shorter, but also better (see also my draft).


Sleeper Agent Backdoor Results Are Messy

Making AI’s talk like a pirate is competetive to HHH. One could spin this against the PSM: Behaviour that is fine-tuned in can be removed by knocking the model off-distribution. Furthermore: One could ambitiously resample different models, i.e. make them talk in a bunch of different distributions (i.e. talk like a pirate, talk in French …), and just hope they’re not all scheming in the same way. Unfortunately, since this is likely pretty capabilities-costly, compared to e.g. monitors, Pirate resampling may never see the light of day at frontier labs.

Reward Hacking Without Egregious Misalignment in an RL-Only Setting

We can RL-fine-tune models in a specific, bad way, and they don’t become bad in general. To me, it is still confusing why mechanistically small weight changes via supervised fine-tuning seem to be different from small weight changes via GRPO. Perhaps if one make the LoRA dimensions really small for both things, i.e. updating just a few parameters, one could mechanistically locate the “persona selection circuitry”, see Persona Cartography: Charting Language Model Personality Traits in Weight Space


Might ASI’s just be good by default (i.e. via strong moral internalism: any sufficiently intelligent agent becomes ‘good’ upon learning what that is)? If the human utility function is an arbitrary product of inclusive-genetic-fitness maximization, how can alignment be justified under a universalistic framework? Read further below, or my draft :)

Relatedly, what have we learned about the idea of expected utility, or should we accept the VNM-model of intelligent agency a priori? Exercise: Try to think of falsifiable predictions the EUM-model made of humans or LLMs.


Does the information bottleneck of RLVR (Toby Ord) represent a natural Occam’s razor, a blessing in disguise?

The 2021 MIRI Conversation consensus seems to be that humans had a similar helper,

Humans face the genomic bottleneck which means that each individual has to rederive all the knowledge about the world that their parents already had. If this genetic bottleneck hadn’t been so tight, then individual humans would have been significantly less capable of performing novel tasks.

Furthermore, we recently learned that humans got smarter a bunch only recently, i.e. with just a few bits, see:

Ancient DNA reveals pervasive directional selection across West Eurasia (Nature)

Similarly, it is possible to train reasoning models by updating just 13 Parameters.

We can collect evidence by asking how strong the model’s positive manifold (Epoch Capabilities Index) is, i.e. their anti-jagged frontier, comparing between reasoning and non-reasoning models. That is, how strongly predictive of downstream benchmark scores is the general factor, i.e. the Epoch Capabilities Index. There seems to be no difference between reasoning and non-reasoning models here (detailed evaluation will be added here soon), and holistically I would say RLVR generalizes more weakly, despite all the fancy arguments above.

Intelligence – A Very Short Introduction (Oxford University Press)

Basically all of Intelligence Psychometrics. But probably you could still boil it down to ten Pearson’s r’s.


Esketamine for treatment-resistant depression: seven concerns about efficacy and FDA approval (Lancet Psychiatry)

On Johnson & Johnson’s polish outlier sides. But ketamine has other things than four MADRS points going for it :).

Behavioural therapies versus other psychological therapies for depression (Cochrane)

As ‘behavioural activation’ (erm, doing more stuff?) is part of CBT, with seemingly equal effectiveness, one can get the favourite razor out.

Genetics of attention deficit disorder (Nature)

Most of the variance in teacher-annoying among all school-children explained by genes, Falconer equations say

Geeks, Mops, Sociopaths

Absolute post-rat classic, though not to be taken literally. Exercise: Come up with different words for Geeks and Sociopaths, there are many. For example, Genius in it’s meaning before Galton’s Hereditary Genius, i.e. as creator of culture, and Opportunist, which captures the fact that people in question lack clinically-sociopathic traits, yet doesn’t represent the competence with which they swing up in social networks (e.g. if they were on an island …)

Crony Beliefs

Meltingasphalt seems less known than he should be. The most useful part is probably this part on how to indentify the crony beliefs.