Without specific countermeasures, the easiest path…
You understand the broad concern, but still want to know how ordinary model training could lead there.
Why I chose itWhy this one
The useful question is how a training process could produce a system we cannot safely rely on. This essay connects the feedback a system receives with the strategies that feedback might reward.
I would pick it for the space it gives the intermediate steps. You can separate an objection to the training setup from an objection to what a capable system would do inside that setup. That makes the mechanism easier to discuss than a general claim about intelligence.
- Training rewards behaviour that looks good to the rater, not behaviour that is good.
- The gap only bites once a system can tell the two apart — and act on the difference.
- "Just don't reward the wrong thing" turns out to be the hard part, not the easy one.
Where it falls short
This is an argument about a possible route to failure, not a measurement of how often that route occurs. The specificity helps you inspect the assumptions; it does not settle them.
Also considered
Other ways into the question, and when I would reach for them. These are illustrative comparisons for this design, not a completed review record.
What failure looks like
Makes the broad failure modes easier to grasp in one sitting. It leaves less room for the training details that matter to this category.
Alignment faking in large language models
Offers a particular experimental setup to inspect. Useful after the mechanism, but too narrow to serve as the whole explanation.
How I would choose
I would look for a worked mechanism, a clear account of the training incentives and identifiable points where the argument could fail.
This prototype proposes the editorial structure and comparison criteria. A published page would need Ben’s review, verified source links and a documented shortlist.