The Oldest Checklist
Critical Thinking AI Systems

The Oldest Checklist

Ibn Sina wrote seven tests for whether a remedy works. They still catch bad claims.

Ibrahim AbuAlhaol, PhD, P.Eng., SMIEEE

AI Technical Lead

Published: September 14, 2026 | Reading Time: ~5 min read

Ibn Sina had a problem that will sound familiar. People kept telling him a remedy worked, and he had no reliable way to separate a remedy that worked from a good story about one.

So a thousand years ago he wrote down seven rules for settling it. They sit in Book Two of the Canon of Medicine, in a chapter on learning what a drug can do through tajribah, meaning experiment, as opposed to qiyas, meaning reasoning by analogy. He closes the passage by telling the reader these are the rules to observe, and adds two words: take note.

The seven

In plain language, and in his order.

  1. Test the thing itself. A substance that has been heated, chilled, left to spoil or stored against something else is no longer the thing you meant to test.
  2. Test it on one clear problem. If someone has two illnesses at once and improves, you cannot say which one the remedy touched.
  3. Test it on opposite problems. If it appears to help a cold illness and a hot one equally, you have learned nothing about what it does.
  4. Match the strength of the remedy to the strength of the problem. Too weak against something severe and you will see nothing and conclude wrongly. Begin with the smallest dose and work up.
  5. Watch when the effect arrives. Help that comes immediately probably acted on the illness. Help that arrives much later, or arrives and then reverses, is doubtful.
  6. It has to happen again. The effect should hold in every case, or at least in most, because what happens by nature happens regularly and what happens by accident does not.
  7. Test it on people. A substance may behave one way in a horse or a lion and another way in a human being, so a result from a different body can mislead.

That is a clinical trial

Read the list again and notice what is actually in it. Isolate your variable. Exclude confounded cases. Escalate the dose. Watch for delayed and reversing effects. Demand that the result repeat. Check that your test population resembles the one you care about.

The first controlled trial most people can name is James Lind's scurvy experiment aboard the Salisbury in 1747. The randomised controlled trial as we run it now dates to the 1940s. Ibn Sina's rules come roughly seven hundred years before the first of those.

He managed it without statistics, without a control group in the modern sense, and without any theory of bias. He got there by asking one question repeatedly: if I saw an improvement, what else could have caused it?

The same seven catch a sales pitch

Here is the useful part. The rules were written for medicine, but they are not really about medicine. They are about telling a real effect apart from a story. That makes them work on nearly anything being sold to you, software included.

Take the AI tool someone is pitching to your team.

Rule one asks whether you are testing the tool, or testing the tool plus a new process plus the two most enthusiastic people on the team. Rule two asks whether it was tried on one clear task or dropped into a messy quarter in which five other things changed.

Rule three is the sharpest. If a tool is reported to help with writing and planning and debugging and hiring, you have learned nothing about what it does.

A remedy that seems to help everything has told you nothing.

Rule five asks when the benefit showed up. Gains in the first week are probably the tool. Gains in month five, after a reorganisation and two new hires, are probably not. Rule six is the one most demos fail, because you were shown the run that worked and Ibn Sina would want to know how many runs there were.

Rule seven is the one this industry keeps relearning. A benchmark is not your work. A model that scores well on a public test set is, in his terms, a drug that was tested on a lion.

Ibn Sina's seven rules and the modern claim each one catches Each of the seven rules for testing a remedy maps onto a specific way that claims about software and tools go wrong today, from bundled trials and tangled pilots to cherry-picked demos and benchmark scores that do not resemble your actual work. The same seven rules, a thousand years apart HIS RULE WHAT IT CATCHES NOW 1. Test the thing itself Was it bundled with other changes? 2. Use one clear problem Too tangled to credit anything 3. Try it on opposite cases Helps everything, proves nothing 4. Match strength to problem Start small, then raise it 5. Watch when the effect lands Gains that show up months later 6. It must repeat The one impressive demo 7. Test on the real subject A benchmark is not your work Written for plants and minerals. They transfer without modification.
Figure 1. The seven rules as set out in Book Two of the Canon, paired with the modern claim each one is still good at catching. Source: Nasser, Tibi and Savage-Smith (2009).

The rule everyone skips

If you keep only one of the seven, keep the sixth.

Almost every claim that later turns out to be wrong survives because nobody asked how many times it was tried. The supplement that worked for one person. The trading strategy backtested until it fit. The agent demo recorded on the eleventh attempt.

His reasoning here is worth reading slowly. The effect should hold in all cases, or at least in most, because what happens by nature happens regularly and what happens by accident does not. That is a working definition of evidence written in the eleventh century, and it still produces the fastest question you can ask in a meeting. How many times did you run it?

One successful run shown, against a result that would actually count A demo typically shows a single successful run out of many attempts. Ibn Sina's sixth rule asks for the effect to hold in all cases or at least most of them, which looks completely different when the runs are laid out side by side. Rule six, drawn WHAT YOU WERE SHOWN One run in ten, and the nine are not in the slide deck. WHAT WOULD COUNT Holds in all cases, or at least in most of them. Amber is the run you were shown. Blue is a run that happened.
Figure 2. Ibn Sina's sixth rule asks for regularity, because what happens by nature repeats and what happens by accident does not. Source: Nasser, Tibi and Savage-Smith (2009).

What outlasted the pharmacy

Ibn Sina's pharmacology is obsolete. Most of the specific remedies in Book Two were abandoned long ago, and his theory of why any of them worked was wrong. The seven rules survived all of it, because they were never claims about substances. They were a procedure for not fooling yourself.

The claims have changed. A thousand years ago it was a plant that cured a fever. Today it is a tool that doubles your output. The procedure for checking runs exactly the same way.

What to do this week

  1. When someone shows you a result, ask how many times they ran it before that one. That is rule six in eight words.
  2. When you trial a tool, change one thing only. Change the tool, the process and the team together and you have bought a story instead of a test.
  3. Ask where it was tested. If the answer is a public benchmark and your work looks nothing like that benchmark, treat the score as a starting point rather than evidence.
  4. Distrust anything reported to help with everything. Rule three has never had a richer hunting ground than the current AI market.

Related Articles

References & Extended Literature

  1. Nasser, M., Tibi, A., and Savage-Smith, E. (2009). "Ibn Sina's Canon of Medicine: 11th century rules for assessing the effects of drugs." Journal of the Royal Society of Medicine, 102, 78 to 80. https://doi.org/10.1258/jrsm.2008.08k040
  2. The James Lind Library. "Ibn Sina's Canon of Medicine: 11th century rules for assessing the effects of drugs." Full open text of the commentary above. https://www.jameslindlibrary.org/articles/
  3. Avicenna (Ibn Sina). Al-Qanun fi al-Tibb (The Canon of Medicine), Book Two, on knowing the potency of drugs through experiment.
  4. Perel, P., Roberts, I., Sena, E., et al. (2007). "Comparison of treatment effects between animal experiments and clinical trials: systematic review." BMJ, 334, 197. Cited by Nasser and colleagues as the modern echo of the seventh rule. https://doi.org/10.1136/bmj.39048.407928.BE
  5. Internet Encyclopedia of Philosophy. "Avicenna (Ibn Sina)." https://iep.utm.edu/avicenna-ibn-sina/