Skip to content
agi.fyi

The best ways to understand AI alignment research (and whether it's on track)

By Ben PaceCofounder of Lightcone Infrastructure;
builder of LessWrong & the AI Alignment Forum

An experimental result becomes easier to evaluate when you can see what the model was told and how the authors interpreted its responses. The connection between setup and behavior is the center of this recommendation.

I would start with the examples, then return to the methodology. That order lets you form a provisional view before absorbing the authors’ interpretation. The valuable question is which explanation of the observed behavior survives close inspection.

Get an email when a guide is updated.

Picks change as better pieces appear. A confirmation link arrives first; unsubscribe from any email.

Alignment faking in large language models
Greenblatt et al., Anthropic & Redwood
PaperSource facsimile
Top read

Alignment faking in large language models

Greenblatt et al., Anthropic & Redwood · 2024

Read the setup, then the transcripts. Keep the claim as narrow as the evidence.

Read Alignment faking in large language models ↗
Format
Paper
Commitment
40 min

Where it falls short

Behavior in a constructed experiment does not establish how a deployed system will behave. Reasoning traces also require interpretation; they are not a transparent window into everything a model is doing.

The rest of the short list

Without specific countermeasures, the easiest path…
by Ajeya Cotra · 2022
EssaySource facsimile
For the explanation behind the experiment

Without specific countermeasures, the easiest path…

Ajeya Cotra · 75 min

Provides a broader account of incentives. It supplies a conceptual framework, rather than the experimental record this category is meant to foreground.

Research report
Is Power-Seeking AI an Existential Risk?
Joe Carlsmith · 2022
  1. Introduction
  2. Six premises
  3. Timelines
  4. Incentives
  5. Probability estimate
  6. Objections
ReportSource facsimile
For placing evidence in the larger argument

Is Power-Seeking AI an Existential Risk?

Joe Carlsmith · 3 hr

Useful for asking which premise an experiment bears on. The report is a synthesis, so it is a less direct encounter with a single result.

Who is this for?

You want to examine a specific piece of evidence rather than rely on a secondhand account.

Why you should trust me?

For the last nine years I have built LessWrong and the AI Alignment Forum, where professionals and independent researchers discuss this subject. Along with the pieces themselves, I have read rebuttals, failed ideas and more comment threads than I would recommend.

I built and ran the LessWrong Annual Review, which evaluates the writing each year. I have also worked alongside people in the field as an editor, web developer, event host and occasionally an argumentative commenter at 2am.

I have watched arguments and research areas take shape, and I have watched some of them fail.

Thanks for reading. Subscribe to hear when this guide, or any other, changes.

No schedule, no digest: only a short note when a category gets a new pick or a rewrite. Unsubscribe from any email.