A checklist for unpacking a loud claim

The future of AI and AGI: an overview · Lesson 2 / 20

Five questions that remove half the noise

More claims about AI capability appear than anyone can read carefully. What you need is a cheap filter - a set of questions that takes a minute and screens out most of what does not deserve attention. Whatever survives is worth taking apart seriously. Here is that filter, with the reason for each step.

The checklist

  • Who is speaking and what they gain if you believe it. An interest does not make a claim false, but it sets the direction of likely distortion: a vendor stresses capability, a critic stresses failure, a publication stresses drama. Knowing the direction tells you where to double-check.
  • What exactly is being claimed. Restate the claim so that only testable parts remain. The system passed test X with result Y under conditions Z is testable. A breakthrough that changes everything contains no claim at all.
  • What it rests on. A one-off showing, a report by the company itself, an independent check, replication by outside people - these are four different levels, and the gap between them matters more than the gap between two neighbouring numbers in a result.
  • What would count as refutation. Ask yourself what result would make the author admit they were wrong. If there is no answer, the claim is not about facts.
  • What changes for you. Many messages are true and still change nothing in your work. That is a legitimate outcome of the analysis.
Analysis card (one page)
Source and its interest:
Testable core (one sentence):
Rhetoric discarded:
Level of evidence: demo / author report / independent check / replication
What would count as refutation:
Decision: ignore / watch / test myself / change something in my work
Insight. The most useful line on the card is rhetoric discarded. Writing out separately what you removed makes it visible how much of the message was a claim and how much was tone of voice.
Common mistake. Judging a claim by whether you like its conclusion. Pleasant news passes the filter unchecked, unpleasant news is rejected just as unchecked. Equal strictness towards both sides is the only defence.
Pro tip. Keep a separate list of claims marked check later with a return date. Most loud announcements quietly fail to hold up, but only the person who comes back to look ever notices.

Words that most often hide emptiness

Some phrases should make you slow down. A world first is usually true under a narrow enough definition of this. Human-level needs a follow-up question: which human, on which task, under which conditions. Understands and reasons are words everyone reads differently, which is exactly why they are convenient. Soon is not a deadline. Practically solved is not solved. None of these expressions is a sign of dishonesty; they only mean the substantive claim has yet to be extracted.

The other side: criticism can be empty too

Apply the same checklist to debunkings. None of this works is exactly as unfalsifiable as this changes everything. Criticism also has a source, an interest and a level of evidence. A single example of failure says as much about a system as a single example of success: it shows the case is possible and says nothing about how often it happens.

Cheat sheet

  • Five questions: who, what exactly, what it rests on, what would refute it, what changes for me.
  • The level of evidence matters more than the size of the result.
  • Equal strictness towards enthusiasm and towards debunking.
  • A single example is about possibility, not about frequency.
1. Why establish the source's interest when unpacking a claim?
2. What matters more when assessing a system's result?
3. A single example of a system failing proves that...
Task — checked by AI

Find any loud claim about AI capability published in the last few weeks. Take it apart using the card: source and its interest, testable core in one sentence, rhetoric discarded, level of evidence, what would count as refutation, your decision. Separately, write whether your attitude to the claim changed after the analysis and which specific item produced that change.

← Back

🔒 Answer the question correctly to move on to the next lesson.

A checklist for unpacking a loud claim — The future of AI and AGI: an overview — Skilvy