Skip to content

The Best Ways to Understand AI Alignment Research (and Whether It's on Track)

Alignment faking and Power-Seeking AI

Cofounder of Lightcone Infrastructure; builder of LessWrong & the AI Alignment Forum

An experimental result becomes easier to evaluate when you can see what the model was told and how the authors interpreted its responses. The connection between setup and behavior is the center of this recommendation.

I would start with the examples, then return to the methodology. That order lets you form a provisional view before absorbing the authors’ interpretation. The valuable question is which explanation of the observed behavior survives close inspection.

Where it falls short

Behavior in a constructed experiment does not establish how a deployed system will behave. Reasoning traces also require interpretation; they are not a transparent window into everything a model is doing.

The rest of the short list

Without specific countermeasures, the easiest path…
by Ajeya Cotra · 2022
EssaySource facsimile
For the explanation behind the experiment

Without specific countermeasures, the easiest path…

Ajeya Cotra · 75 min

Provides a broader account of incentives. It supplies a conceptual framework, rather than the experimental record this category is meant to foreground.

Research report
Is Power-Seeking AI an Existential Risk?
Joe Carlsmith · 2022
  1. Introduction
  2. Six premises
  3. Timelines
  4. Incentives
  5. Probability estimate
  6. Objections
ReportSource facsimile
For placing evidence in the larger argument

Is power-seeking AI an existential risk?

Joe Carlsmith · 3 hr

Useful for asking which premise an experiment bears on. The report is a synthesis, so it is a less direct encounter with a single result.

Who is this for?

You want to examine a specific piece of evidence rather than rely on a secondhand account.

Why you should trust me?

For the last nine years I have built LessWrong and the AI Alignment Forum, where professionals and independent researchers discuss this subject. Along with the pieces themselves, I have read rebuttals, failed ideas and more comment threads than I would recommend.

I built and ran the LessWrong Annual Review, which evaluates the writing each year. I have also worked alongside people in the field as an editor, web developer, event host and occasionally an argumentative commenter at 2am.

I have watched arguments and research areas take shape, and I have watched some of them fail.