Without specific countermeasures, the easiest path…
by Ajeya Cotra · 2022
EssaySource facsimile
The pick

Without specific countermeasures, the easiest path…

Ajeya Cotra · 2022

You understand the broad concern, but still want to know how ordinary model training could lead there.

Why I chose it
01 / The case for it

Why this one

The useful question is how a training process could produce a system we cannot safely rely on. This essay connects the feedback a system receives with the strategies that feedback might reward.

I would pick it for the space it gives the intermediate steps. You can separate an objection to the training setup from an objection to what a capable system would do inside that setup. That makes the mechanism easier to discuss than a general claim about intelligence.

What should stay with you
  1. Training rewards behaviour that looks good to the rater, not behaviour that is good.
  2. The gap only bites once a system can tell the two apart — and act on the difference.
  3. "Just don't reward the wrong thing" turns out to be the hard part, not the easy one.
02 / The tradeoff

Where it falls short

This is an argument about a possible route to failure, not a measurement of how often that route occurs. The specificity helps you inspect the assumptions; it does not settle them.

03 / The rest of the shortlist

Also considered

Other ways into the question, and when I would reach for them. These are illustrative comparisons for this design, not a completed review record.

01
The shorter introduction

What failure looks like

Paul Christiano · 15 min

Makes the broad failure modes easier to grasp in one sitting. It leaves less room for the training details that matter to this category.

02
A narrower empirical companion

Alignment faking in large language models

Greenblatt et al., Anthropic & Redwood · 40 min

Offers a particular experimental setup to inspect. Useful after the mechanism, but too narrow to serve as the whole explanation.

04 / The selection

How I would choose

I would look for a worked mechanism, a clear account of the training incentives and identifiable points where the argument could fail.

This prototype proposes the editorial structure and comparison criteria. A published page would need Ben’s review, verified source links and a documented shortlist.