learn.

Curriculum / 01 · How to know / Lesson 2

Theme 01 · How to know · Lesson 2

Falsifiability

A claim is scientific only if something could prove it wrong. Learn to spot claims that can never lose, to design a test that could kill your own favourite idea, and to see why a claim that cannot fail cannot help you act.

  • Socratic
  • Concept Development Study
  • ~50 min with a notebook
Before you begin

How to use this page

As in Lesson 1, keep a notebook beside you. Whenever you reach a Pause and answer box, stop and write your answer — a sentence or two, even a rough guess — before you read on.

Each pause is followed by a folded panel. Open it only after you have written something. The panel is one line of reasoning to compare with yours, not the official answer. Where yours differs, ask which one fits the observation better.

This lesson asks something slightly uncomfortable of you: to aim tests at ideas you like. Pick one belief you hold fondly at the start, and keep it in mind throughout. By the end you will design a test that could kill it.

Step 1 · Foundation

What we assume you already know

We build on Lesson 1 and on ordinary experience. We take these as given:

  • An observation is what reached your senses or instrument; a belief is the story you build to explain it. Errors usually live in the story.
  • A good belief should keep working as you collect more observations, not just the first one.
  • Some claims name an observation that would show them wrong; others have a built-in escape hatch. (If “lucky thread” rings a bell, you have this one.)
  • You have held on to an opinion after someone gave you a good reason to drop it — or watched someone else do so.

In Lesson 1 we met falsifiability briefly. Here we make it the whole subject: what it is, why it matters, and how to use it on yourself.

Step 2 · Questions that arise

If a claim can’t be wrong, is that a strength or a weakness?

Common sense says a claim that is never proven wrong must be a good one. Lesson 1 hinted the opposite. Questions follow:

  1. What exactly makes a claim “able to be wrong”? Is it the claim itself, or how people defend it?
  2. When I test my own idea, am I really testing it — or just collecting support?
  3. Is a bold, precise claim better or worse than a safe, vague one?
  4. How do I design a test that could actually kill an idea I care about?
  5. If a claim can’t be tested, what is it good for — and what is it not good for?
Pause and answer
  1. Write down one belief you hold fondly — about health, food, study habits, luck, people, or how the world works. Then write: “I would change my mind if I saw…” and finish the sentence. If you can’t finish it, write that down too.

Keep this belief close. We will come back for it.

Step 3 · Observation 1

The tiger on the terrace

Your cousin announces, quite seriously, “There is a tiger living on our terrace.” You go up together. You see clothes on the line, a water tank, and some potted tulsi. No tiger.

“It’s an invisible tiger,” she says. You suggest spreading flour on the floor to catch its pawprints. “It floats.” You suggest a thermal camera from the hardware shop. “Its body is the same temperature as the air.” You offer to sit up there overnight and listen. “It never makes a sound.” You suggest leaving meat out. “It doesn’t eat.”

Pause and answer
  1. List every test you proposed and the reply each one got. What do the replies have in common?
  2. After all the replies, what is left of the “tiger”? Name one way the terrace would be different with this tiger compared with no tiger at all.
  3. Has your cousin proved there is a tiger? Have you proved there is not?
  4. Would you change how you use the terrace — hang clothes, water the plants — because of this claim? Why or why not?
After you have written your answers — compare your reasoning

Every reply arrived after a test was proposed, and each one did the same job: it removed exactly the property that the test would have detected. Nothing in the replies came from observing the tiger. They came from protecting the claim.

By the end, the invisible, floating, heatless, silent, never-eating tiger makes no difference to anything. A terrace with it and a terrace without it look, sound, feel, and behave exactly alike.

Neither of you has proved anything. But notice what that means in practice: you go on hanging clothes as before. The claim has no grip on any decision, because it tells you nothing about what will happen.

(This story is adapted from Carl Sagan’s “dragon in my garage”, in The Demon-Haunted World.)

Deduction

A claim can be made unfalsifiable after the fact, by adding a fresh excuse each time a test comes near. Each excuse is called an ad hoc rescue — a patch made for this one occasion, with no independent evidence behind it. A claim that has been rescued from every possible test no longer says anything about the world. A difference that makes no difference is no difference.

Step 3 · Observation 2

The number game

Now turn the lens on yourself. This one works best if you actually play it with a friend.

A friend has a secret rule for sequences of three numbers. She tells you that 2, 4, 6 follows the rule. You may propose any three numbers you like; she will say only “yes, fits” or “no, doesn’t fit”. When you are confident, you announce the rule.

A typical player thinks, “Even numbers going up by two,” and tries 8, 10, 12 — yes. Then 20, 22, 24 — yes. Then 100, 102, 104 — yes. Feeling sure, they announce: “Even numbers, each two more than the last.”

“No,” says the friend. “That’s not my rule.”

Pause and answer
  1. The player got three “yes” answers in a row. Why did these not help them find the rule?
  2. Write down two or three other rules that 2, 4, 6 would also fit.
  3. What sequence should the player have tried — one that their own guess predicts would get a “no”?
  4. Which kind of test feels better to run: one you expect to pass, or one you expect might fail? Which kind teaches you more?
After you have written your answers — compare your reasoning

Every sequence the player tried fitted their own guess and fitted many other rules. A “yes” to 8, 10, 12 is equally expected under “even numbers up by two”, “numbers up by two”, “any increasing numbers”, or “any three different numbers”. A test that every candidate passes can’t pick between them.

The useful moves are the ones the player’s guess says should fail: 1, 3, 5 (odd); 1, 2, 3 (steps of one); 3, 10, 50 (uneven jumps); 6, 4, 2 (going down). If 1, 2, 3 gets a “yes”, the guess “up by two” is dead on the spot — and you have learnt something real.

In the original version of this game, set by the psychologist Peter Wason in 1960, the secret rule was simply “any three increasing numbers”. Many people announced a wrong rule with great confidence, because they only ever tried sequences they expected to fit.

Tests you expect to pass feel pleasant. Tests that could fail feel risky. The risky ones are where the information is.

Deduction

Piling up agreeing cases is not testing. Our minds lean toward looking for support — this is called confirmation bias — and support is cheap, because many ideas fit the same evidence. A real test is one where your idea predicts a result that could fail to happen. To test an idea, aim at the place where it would break.

Step 3 · Observation 3

Two forecasts

So far we have seen claims made safe by excuses, and tests made safe by choice. Now look at the claims themselves.

Forecast A. A printed almanac states that on a certain date, a partial solar eclipse will be visible from your city, beginning at a stated time to the minute and ending about two hours later, with the Moon covering a stated fraction of the Sun at the peak.

Forecast B. A horoscope column says: “This year will bring changes. You may face a challenge at work or at home, but with patience, good things are possible. Someone close to you has something important to say.”

Pause and answer
  1. For each forecast, describe what you would have to observe on the day — or over the year — to say “this forecast was wrong”.
  2. Which forecast is more likely to come true? Which one tells you more?
  3. Suppose both come true. Which success should impress you more, and why?
  4. Could you plan anything — a trip, a photograph, a school demonstration — using Forecast B?
After you have written your answers — compare your reasoning

Forecast A can fail in many ways: the eclipse starts ten minutes late, it isn’t visible from your city, the Moon covers far less of the Sun, or nothing happens at all. A clock, a clear sky, and eclipse glasses are enough to catch it out.

Forecast B can barely fail. Almost every year brings some change, some challenge, and some important conversation. Whatever happens, a reader can find a match. It is very likely to “come true” precisely because it says so little.

So the forecast more likely to come true is the one that tells you less. Forecast A’s success is impressive because it was unlikely by chance: out of all the minutes in the year, it named the right one. That kind of success only comes from a model that has really captured something — here, the motions of the Earth and Moon.

And you can act on Forecast A: book the trip, set up the camera, gather the class at the right time. Forecast B gives you nothing to schedule.

Deduction

Claims can be unfalsifiable not only through excuses but through vagueness. The more outcomes a claim forbids, the more it tells you — and the more easily it could be caught out. Boldness and precision are virtues: a claim that sticks its neck out and survives has earned trust that a safe claim never can.

This is the idea the philosopher Karl Popper made famous: what marks a claim as scientific is not that it has been confirmed, but that it could be refuted by an observation — and has so far survived honest attempts.

Step 3 · Observation 4

The aching knee

Now we design a test ourselves — on an idea that its owner is fond of.

Your grandmother says, “My left knee aches the day before it rains. It never fails.” She is not being silly: she has noticed it many times, and she is sure.

You would like to know whether it is true. You also notice something: on rainy days she often says, “See? My knee told me yesterday.” On dry days, nobody mentions the knee at all.

Pause and answer
  1. Why might she remember the “ache, then rain” days much more than the “ache, then no rain” days?
  2. Turn her idea into a claim that could fail. What exactly would you count, and over how long?
  3. Before collecting any data, write down the result that would mean “the knee does not predict rain”. Why is it important to write this down first?
  4. What could spoil the test? Think about who writes down the ache, and when — and whether they have seen the weather forecast or the clouds.
After you have written your answers — compare your reasoning

The hits are memorable: rain arrives, and the knee gets the credit. The misses — an ache followed by a dry day, or rain with no warning ache — pass without comment. Memory keeps a one-sided scorebook.

One test: for eight weeks, every evening, she marks in a notebook “ache” or “no ache” before looking at the sky or the forecast. Separately, someone records whether it rained the next day. At the end you fill in four boxes: ache-and-rain, ache-and-dry, no-ache-and-rain, no-ache-and-dry.

Write the killing result in advance: “If rain follows an ache day no more often than it follows a no-ache day, the knee does not predict rain.” Fixing this first matters because once you see the numbers, it is very tempting to redraw the line — “well, it works for heavy rain”, “it works in the monsoon only”. Those are the tiger’s excuses again, in a friendlier voice.

Spoilers to guard against: marking the ache after seeing dark clouds; letting the person who knows the forecast decide what counts as an “ache”; stopping the test early on a lucky streak. And the test might come out in the knee’s favour — that is allowed. A test is honest only if both results were possible.

Deduction

To test a favourite idea, turn it into a prediction, name the result that would kill it before you look, record the misses as carefully as the hits, and keep the person with the hope away from the scorebook. If the idea survives that, it has earned some trust. If it doesn’t, you have swapped a comfortable belief for real knowledge — a good trade.

Step 4 · Choosing between models

Occam: the cost of excuses

Return to the terrace. Two models are on offer to explain everything you observed there — clothes, tank, tulsi, no pawprints, no heat, no sound, no missing food.

Model 1. There is a tiger, and it is invisible, and it floats, and its body is exactly air temperature, and it is perfectly silent, and it never eats — and each of these properties happens to match the test you just proposed.

Model 2. There is no tiger on the terrace.

Pause and answer
  1. Do both models fit every observation you made? Check each one.
  2. Count the special assumptions each model needs.
  3. Model 2 makes predictions that could fail: a tiger-sized pawprint in flour, a warm shape on the thermal camera, a missing goat. Can Model 1 be caught out by anything at all?
Occam note

Both models “fit” — but Model 1 only fits because a new assumption was invented for each new observation, and each assumption is there for no reason except to dodge a test. Model 2 needs no extra assumptions, and it stays exposed: one clear pawprint would sink it.

When two models fit the evidence, prefer the one with fewer special assumptions — and be most suspicious of assumptions whose only job is to protect the claim. Occam’s razor and falsifiability work together: excuses add assumptions and remove risk, so each excuse makes a model both less simple and less useful.

Step 5 · Building the model

A working model of testing

Gather what the four observations taught. Write your own summary in three or four lines before you open ours.

Pause and answer
  1. In your own words: what makes a claim falsifiable?
  2. What are two ways a claim can become unfalsifiable?
  3. What are the steps of a test that could kill a favourite idea?
  4. Why does an unfalsifiable claim give you no power to act?
After you have written your summary — compare it with ours
PartWhat it meansThe question to ask
Falsifiable claimA claim that forbids some observable outcome. The more it forbids, the more it tells you.“What would I see if this were wrong?”
Escape by excuseAd hoc rescues invented after each failed or threatened test (the terrace tiger).“Did this excuse exist before the test, and is there any evidence for it apart from saving the claim?”
Escape by vaguenessClaims loose enough to match any outcome (the horoscope).“Which outcomes does this rule out? Name one.”
Real testAim where the idea would break; fix the killing result in advance; record misses as well as hits; keep hopes away from the scorebook.“Could this test have come out against me?”
SurvivalA claim that passes risky tests earns trust — provisional, never final.“How hard has this been tested, and by whom?”
PowerOnly claims that forbid outcomes can guide plans, tools, and decisions.“What does this let me predict, build, or decide?”

The last row is the heart of it. Knowledge is power because it predicts: it tells you what will happen, and so what to do. A claim that fits every outcome predicts nothing, so it can never help you build a bridge, time an eclipse, cure a fever, or decide whether to carry an umbrella. Unfalsifiable claims may be comforting, but they give no grip on the world.

And like everything on this site, this is a model, not a proven theorem. It must survive its own test. So let us push on it.

Step 6 · Further questions → refine

Pushing on the model

Here are some hard cases. Expect sharper thinking, not neat answers.

Pause and answer
  1. Your knee test fails once, but your grandmother had a cold that week and the notebook was filled in late. Should one failed test always kill an idea?
  2. In the 1840s, astronomers found that the planet Uranus was not moving quite where the accepted laws of motion said it should. Some suggested an unseen planet farther out was tugging on it. Is that just the tiger’s excuse in a lab coat? What would make it different?
  3. “Mangoes taste better than apples.” “Be kind to strangers.” “2 + 2 = 4.” None of these can be refuted by an observation in the usual way. Does that make them worthless?
After you have written your answers — one way to refine the model

One failure is a strong signal, not always a verdict. Every test leans on other assumptions — the notebook was filled in honestly, the thermometer works, the rain gauge wasn’t blocked. When a test fails, either the idea is wrong or one of those helpers is. The honest response is to check the helpers openly and repeat the test — not to keep inventing new helpers until the idea is safe.

A rescue is honest when it makes a new, risky prediction of its own. The unseen-planet idea did exactly that: Urbain Le Verrier calculated where the new planet should be, and in 1846 Johann Galle pointed a telescope at that patch of sky and found Neptune. The rescue could have failed — the sky could have been empty there. The tiger’s excuses never risked anything. So the refined rule is: a patch that predicts something new and checkable is a hypothesis; a patch that only dodges is an excuse.

Not every meaningful statement is a claim about observations. Tastes, values, and mathematics are different kinds of statement: mathematics is checked by proof, values by argument and consequences, tastes not at all. Falsifiability marks the boundary of empirical knowledge — claims about what the world will show us. Outside that boundary, a statement may still matter; it just can’t borrow the authority of evidence. The trouble starts when a claim about the world (“this thread brings luck”, “this tiger exists”) shelters behind protections that belong to another kind of statement.

Why this matters now

Fluent, confident claims are becoming cheap. A person, a forwarded message, or an AI system can produce a persuasive explanation for anything in seconds — and a persuasive explanation that fits every outcome is just a better-written horoscope. The defence is a single question: “What would show this wrong — and has anyone looked?”

Apply it hardest to the claims you like most. Machines and markets will happily tell you what you want to hear. The habit of designing a test that could kill your own favourite idea is one that no one else can practise on your behalf — and when paid work no longer decides what you spend your days thinking about, it is how you keep your judgement your own.

Step 7 · Study by reasoning

Review & discussion questions

Answer these by explaining, in full sentences, as if teaching someone who missed the lesson. Try them alone first, then discuss with a friend or study group and compare your reasoning.

  1. Using the terrace tiger, explain how a claim that started out testable can be made unfalsifiable. What is an ad hoc rescue, and why does each one make the claim less useful?
  2. In the number game, why did three “yes” answers in a row fail to reveal the rule? Describe the sequence you would try first, and why.
  3. Explain why the eclipse forecast is more impressive than the horoscope when both “come true”. What does this tell you about the relationship between how likely a claim is and how much it tells you?
  4. Go back to the favourite belief you wrote at the start. Design a test that could kill it: what you would record, for how long, who would record it, and — written in advance — the result that would make you give it up.
  5. Someone says, “You can’t prove my claim is false, so you have to respect it as equally valid.” Explain, politely and clearly, what is wrong with this reasoning.
  6. Compare the unseen-planet rescue of the laws of motion with the tiger’s excuses. What single feature separates an honest hypothesis from an excuse?
  7. Explain in your own words why an unfalsifiable claim gives no power to predict or act. Give an example from health, money, or news where acting on such a claim could cost someone.
  8. Take one confident claim you have recently met — from a person, a message, or a machine. Rewrite it as a falsifiable claim, if you can. If you can’t, say what that tells you.
You own this lesson when…

…you can explain, without looking, why a claim that can’t be wrong tells you nothing; how excuses and vagueness each make a claim unfalsifiable; and how to design a test, fixed in advance, that could kill an idea you care about — and you have actually designed one.

Attribution. Lesson structure — Foundation, questions, observations and deductions, refined models — adapted from John S. Hutchinson, Concept Development Studies in Chemistry (Connexions / Rice University), licensed under Creative Commons Attribution 2.0 (CC BY 2.0). The terrace tiger is adapted from Carl Sagan’s “dragon in my garage” (The Demon-Haunted World, 1995); the number game is after Peter Wason’s 2-4-6 task (1960). The topic, examples, and text of this lesson are otherwise original to learn.curiosta.com.