Occam: Somewhere between a folk proverb and a formal principle of science sits one of the most quoted and most misunderstood ideas in human reasoning. When the dog knocks a glass off the table and you find it shattered on the floor, you do not seriously entertain the theory that an intruder broke in, smashed exactly one glass, and left without touching anything else.

You assume the dog did it. That instinct toward the leaner explanation has a name: Occam’s razor. It is invoked in courtrooms, diagnostic clinics, physics laboratories, and arguments about the supernatural, usually as if it were a law of nature. It is not a law.

It is a remarkably good heuristic with sharp edges, and learning where it cuts cleanly and where it draws blood is one of the more useful skills a thinking person can acquire.

The razor is older than its popular name and more subtle than its popular use. Stripped to its core, it says that when two explanations account for the same evidence, you should prefer the one that assumes less. The trouble begins the moment you ask what “less” means, who decides what counts as an assumption, and whether the simpler story is actually more likely to be true or merely easier to hold in your head.

Those questions turn a tidy slogan into a genuinely deep problem about how we ought to reason under uncertainty.

“Plurality must never be posited without necessity.” - attributed to William of Ockham, paraphrased from his fourteenth-century writings

Key Definitions

Occam’s razor (also spelled Ockham’s razor) is the principle that, among competing explanations that fit the evidence equally well, one should prefer the explanation that introduces the fewest entities, assumptions, or causes. It is a principle of parsimony: do not multiply assumptions beyond what is required to account for what you observe.

It is worth being precise about what the razor does not claim. It does not say the simplest explanation is always true. It does not say complex explanations are forbidden. It says, all else being equal, that simplicity is a reason to prefer one account over another, a tiebreaker that earns its keep only when the competing explanations genuinely explain the same evidence.

The phrase “all else being equal” is doing enormous work, and most misuses of the razor come from quietly dropping it. An explanation can be gloriously simple and still lose to a more complicated rival the instant the rival accounts for some piece of evidence the simple one ignores. Simplicity counts only after the explanations have been put on an even footing.

A Principle Older Than Its Name

The Franciscan friar William of Ockham, working in the early 1300s, never wrote the famous Latin sentence most often attributed to him. He did write things very close to it, repeatedly arguing against needless metaphysical baggage in scholastic philosophy. The crisp formulation entia non sunt multiplicanda praeter necessitatem (entities must not be multiplied beyond necessity) was coined by later writers and pinned to his name centuries afterward.

The underlying instinct predates Ockham by a long way. Aristotle, in his Posterior Analytics, wrote that “we may assume the superiority, all other things being equal, of the demonstration which derives from fewer postulates or hypotheses,” precisely because knowledge is acquired more readily when the premises are fewer. Ptolemy, the astronomer, expressed a version too, holding that one should adopt the simplest hypotheses that fit the observed motions of the heavens.

Ockham’s contribution was to wield the idea so consistently and so famously that his name stuck to the blade. The history matters because it shows the razor was never a single discovery but a recurring conviction, surfacing wherever people tried to reason carefully about a world that offers more theories than facts.

Why Simpler Is Often Better

The deepest reason to prefer simpler explanations is not aesthetic. It is probabilistic. Every independent assumption an explanation requires is another claim that could be false, and the probability that all of several independent claims are simultaneously true can only shrink as you add more of them. A story that needs three improbable coincidences to work is, other things equal, less likely than one that needs none.

Complexity is not free; each added moving part is a fresh opportunity to be wrong.

There is a second, subtler reason that comes from the science of prediction. An explanation with many adjustable parts can be bent to fit almost any data, including the random noise in your particular sample, and a model that fits the noise predicts the future badly. Statisticians call this overfitting.

A simpler model, with fewer parameters to tune, is forced to capture only the genuine pattern, and it tends to generalize better to data it has not seen. This is why parsimony is built directly into modern statistical practice rather than left as a philosopher’s preference.

There is a third reason that is purely practical: simpler explanations are easier to test, to teach, and to act on. A claim with one moving part makes a sharp prediction that a single experiment can confirm or refute. A sprawling claim, hedged with exceptions and special cases, can wriggle out of almost any disconfirming result, which is precisely what makes it scientifically weaker.

Falsifiability and parsimony are cousins; the leaner an explanation, the more it sticks its neck out, and the more it sticks its neck out, the more we learn when we put it to the test.

Reason for ParsimonyMechanismWhere It Shows Up
Fewer chances to be wrongEach assumption can independently failDetective work, medical diagnosis
Better generalizationFewer parameters resist overfitting noiseMachine learning, statistics
Easier to test and falsifySimple claims make sharper predictionsExperimental science
Lower communication costCompressed explanations transfer fasterTeaching, engineering

The Razor in Science

In practice, working scientists treat parsimony as a tool for choosing between theories that the data cannot yet separate. When Copernicus proposed that the planets orbit the sun, his model was not initially more accurate than the elaborate Ptolemaic system of Earth-centered circles within circles. What it offered was economy: it explained the same retrograde motions of the planets without the dense thicket of epicycles the old system required.

The razor favored heliocentrism before the telescope settled the matter.

“We are to admit no more causes of natural things than such as are both true and sufficient to explain their appearances.” - Isaac Newton, Philosophiae Naturalis Principia Mathematica (1687)

Newton placed this rule first among his stated rules of reasoning in natural philosophy, justifying it with the remark that “Nature is pleased with simplicity, and affects not the pomp of superfluous causes.” Physics has honored the rule ever since. Einstein is widely paraphrased as saying that everything should be made as simple as possible, but not simpler, a line that captures the discipline’s attitude precisely even though the exact wording is a later distillation of his views.

Parsimony is prized, but never at the cost of accounting for the evidence. The razor trims; it does not amputate. A theory that is beautifully simple and demonstrably wrong is worth nothing, and no serious scientist prefers elegance to correctness when the two part ways.

Modern statistics has even formalized the trade-off. Tools such as the Akaike Information Criterion and the Bayesian Information Criterion score competing models by rewarding goodness of fit while penalizing the number of parameters, giving a numerical, defensible version of the razor that practitioners use every day to decide how complex a model is justified. Where an old philosopher could only gesture at “fewer assumptions,” a statistician can now put a precise number on the cost of each extra parameter and let the data decide whether that cost is worth paying.

When the Razor Cuts the Wrong Way

The razor’s reputation for reliability hides a real hazard: reality is sometimes complicated, and the simplest explanation is sometimes simply false. The history of medicine is full of patients whose symptoms were attributed to one tidy cause when the truth involved several interacting conditions. The geocentric universe felt parsimonious to people who could see the sun move across the sky; it was wrong.

Continental drift was rejected for decades partly because a stationary Earth seemed the simpler assumption.

There is even a counter-maxim in medicine known as Hickam’s dictum, a deliberate rejoinder to the diagnostic razor. Where the razor (in its clinical form, “a patient’s symptoms should be attributed to a single disease if possible”) urges one unifying diagnosis, Hickam’s dictum dryly observes that a patient “can have as many diseases as they damn well please.” An elderly patient with multiple chronic conditions is often better understood by abandoning the search for a single elegant cause.

The lesson is that simplicity is a prior expectation, not a guarantee, and it must always yield to evidence. The skilled diagnostician holds both maxims at once: reach first for the single cause that explains the whole picture, but stay ready to abandon it the moment the picture refuses to cohere.

DomainTempting Simple StoryReality That Overturned It
AstronomyThe sun moves around a fixed EarthEarth orbits the sun and rotates
GeologyContinents are fixed in placePlate tectonics and continental drift
MedicineOne disease explains all symptomsMultiple coexisting conditions (Hickam’s dictum)
BiologyAcquired traits are inherited directlyGenetic inheritance and natural selection

Simplicity Is in the Eye of the Modeler

A quieter problem undermines naive uses of the razor: there is no neutral, universal way to count assumptions. What looks simple in one framework looks baroque in another. Is a single invisible force acting at a distance simpler than a curved four-dimensional spacetime? Newtonian gravity has fewer concepts; general relativity has fewer unexplained coincidences. Which is “simpler” depends entirely on what you treat as a free assumption versus a derived consequence.

This is why philosophers distinguish between two kinds of parsimony. Syntactic simplicity, or elegance, concerns the number and conciseness of the principles a theory uses. Ontological simplicity, or parsimony proper, concerns the number of kinds of things a theory says exist. A theory can be ontologically lean but syntactically ugly, or vice versa, and the razor gives no automatic verdict on how to weigh one against the other.

The upshot is that “the simplest explanation” is rarely as obvious as the person invoking the razor believes. Often the razor is not settling a dispute so much as smuggling in the speaker’s pre-existing sense of what is plausible, dressed up as a principle. Whenever someone says their position is “the simpler one,” it is worth asking simpler by which measure, because the answer frequently reveals an unstated assumption rather than a neutral count.

The Razor as a Tiebreaker, Not a Truth Detector

The cleanest way to hold the razor is to remember its proper job. It is a method for choosing what to believe and what to test first when the evidence runs out, not a method for determining what is true. When two hypotheses make identical predictions, no experiment can yet separate them, and the simpler one is the better bet and the better starting point for investigation.

But the moment new evidence arrives that the simple theory cannot accommodate, the razor demands you abandon it, because it was never about loving simplicity for its own sake.

“Whenever possible, substitute constructions out of known entities for inferences to unknown entities.” - Bertrand Russell, who called this “the supreme maxim in scientific philosophizing” and explicitly identified it with Occam’s razor

Understood this way, the razor stops being a blunt instrument for dismissing ideas you dislike and becomes a disciplined default. It tells you where the burden of proof sits: the explanation that posits more is the one that owes you more evidence. Adding entities, forces, conspiracies, or hidden causes is allowed, but only when the evidence forces your hand.

Until then, you travel light. This is also why the razor is a tool of intellectual humility rather than arrogance. It does not claim the universe is simple; it claims that you should not credit complications you have not earned.

Bayesian Foundations of the Razor

Modern probability theory gives the razor a rigorous home. In Bayesian reasoning, a theory that can explain many possible outcomes must spread its probability thin across all of them, so it assigns lower probability to the specific outcome you actually observe. A simpler theory that predicts a narrow range of outcomes concentrates its probability, and when one of those outcomes occurs, it is rewarded with a larger boost in credibility.

This is sometimes called the Bayesian Occam’s razor, and it shows that parsimony is not an arbitrary preference bolted onto inference but a consequence of the mathematics of probability itself. The physicist David MacKay devoted an entire chapter of his standard textbook to exactly this point, showing that Bayesian model comparison automatically penalizes models that are needlessly flexible, without any separate rule for simplicity having to be added by hand.

The computer scientist and information theorist Ray Solomonoff pushed this further with a formal theory of prediction in which simpler hypotheses, measured by the length of the shortest program that could generate the data, are assigned higher prior probability. This connects the razor to Kolmogorov complexity and the idea that the best explanation is, in a precise sense, the most compressed one.

The intuition the medieval friar expressed in Latin turns out to anchor some of the deepest results in the theory of induction. Parsimony, it seems, is not merely how careful humans happen to think; it may be close to how any rational agent must think to learn from a noisy world.

The Animal Dimension

The preference for parsimony is not a uniquely human invention; something very like it is wired into how brains and even simpler nervous systems make sense of the world. Perception itself is a relentless exercise in choosing the simplest hypothesis that fits the sensory data. The visual systems of humans and many animals routinely resolve ambiguous images into the simplest stable interpretation, a tendency captured by the Gestalt principle of Pragnanz, the law that we perceive the simplest organization the stimulus allows.

An ambiguous shadow is read as a single object, not an improbable arrangement of separate ones, because the simpler reading is usually the correct one in a world built from coherent objects rather than coincidental alignments.

In behavioral ecology this logic appears in the principle of parsimony that governs how scientists interpret animal cognition, often called Morgan’s Canon. Formulated by the comparative psychologist C. Lloyd Morgan in 1894, it instructs researchers never to attribute an animal’s behavior to a higher mental faculty if it can be explained by a lower, simpler one.

A bird returning to a feeder need not be credited with abstract reasoning if simple associative learning suffices. Morgan’s Canon is Occam’s razor turned on the study of minds, and like the razor it is a default rather than a dogma: when the evidence genuinely demands a richer explanation, as it increasingly does for the cognition of corvids, great apes, and cephalopods, the canon yields.

Even nature’s own information processing, from the foraging paths of ants to the predictive coding of the mammalian cortex, appears to economize, favoring the leanest model that still works. The razor, in other words, may be less a rule we imposed on reasoning than a strategy that living systems discovered long before any philosopher named it.

Practical Uses in Everyday Reasoning

For day-to-day thinking, the razor’s value is as a brake on a specific human weakness: the appetite for elaborate stories. When something goes wrong, the mind reaches eagerly for hidden causes, malicious actors, and grand conspiracies, because such stories feel meaningful and assign blame. The razor asks a deflating question first.

Could this be incompetence rather than malice, coincidence rather than design, a single ordinary cause rather than a web of extraordinary ones? More often than not, it could.

This is the wisdom behind related folk maxims like Hanlon’s razor, “never attribute to malice that which is adequately explained by stupidity,” which is really the parsimony principle applied to human motives. The conspiratorial explanation requires many people to coordinate flawlessly and keep silent; the simple explanation requires only the ordinary frictions of error, accident, and self-interest that we know to exist.

The razor does not prove the simple story, but it correctly places the heavier burden of proof on the elaborate one.

The discipline cuts both ways, and that is its strength. Applied honestly, the razor restrains not only paranoia but also wishful thinking, since wishful explanations are often the ones that quietly multiply convenient assumptions. Used well, it is less a weapon for winning arguments than a habit of asking, before reaching for a complicated story, whether a plainer one would do.

The person who internalizes the razor does not become cynical or credulous; they become someone who demands that the weight of an explanation be matched by the weight of the evidence behind it, and who is willing, when that evidence arrives, to let even a beloved simple theory go. That last willingness is the part most people skip, and it is the part that separates a genuine reasoner from someone merely fond of the word “simple.”

References

Part of This Series

Frequently Asked Questions

What is Occam's razor in simple terms?

Occam’s razor is the principle that when two explanations account for the same evidence equally well, you should prefer the one that makes the fewest assumptions. It is also called the law of parsimony. If a shattered glass can be explained by the dog knocking it over or by an intruder who broke in to smash one glass and left, the razor favors the dog, because it requires fewer improbable claims. Importantly, the razor does not say the simplest explanation is always true. It is a tiebreaker that tells you which explanation to prefer and test first when the available evidence cannot yet distinguish between them.

Did William of Ockham actually say the famous quote?

Not exactly. The Franciscan friar William of Ockham, writing in the early 1300s, never penned the famous Latin line ‘entities must not be multiplied beyond necessity’ that bears his name. He did write closely related statements many times, consistently arguing against needless metaphysical assumptions in scholastic philosophy. The crisp Latin formulation was coined by later writers and attached to him centuries afterward. The underlying idea is even older: Aristotle and Ptolemy expressed versions of it long before Ockham. His real contribution was wielding the principle so famously and so often that his name became permanently fused to the blade.

Why is the simpler explanation usually more likely to be true?

There are two main reasons, and neither is purely aesthetic. First, every independent assumption an explanation requires is another claim that could be false, so the probability that all of several assumptions are true at once can only shrink as you add more. Complexity multiplies the chances of being wrong. Second, in prediction, a model with many adjustable parts can be bent to fit even the random noise in your data, a problem called overfitting, and such models generalize poorly. A simpler model captures only the genuine pattern and predicts new cases better. Modern statistics builds this trade-off directly into model selection.

When does Occam's razor fail or give the wrong answer?

The razor fails whenever reality is genuinely complicated and the simplest story is simply false. The geocentric universe felt parsimonious yet was wrong; continental drift was rejected for decades because a fixed Earth seemed simpler. In medicine, a counter-maxim called Hickam’s dictum warns that a patient can have as many diseases as they please, pushing back against the urge to force one elegant diagnosis. The lesson is that simplicity is a prior expectation, not a guarantee. The razor is a default that must always yield to evidence. When new data cannot be accommodated by the simple theory, parsimony itself demands you abandon it.

What is the difference between Occam's razor and Hanlon's razor?

Occam’s razor is the general principle of preferring explanations with fewer assumptions. Hanlon’s razor is a narrower, folksy application of that idea to human motives: ‘never attribute to malice that which is adequately explained by stupidity.’ A conspiratorial explanation usually requires many people coordinating flawlessly and staying silent, which demands a great many assumptions. The simpler explanation needs only the ordinary frictions of error, accident, incompetence, and self-interest that we already know exist. Hanlon’s razor does not prove the innocent story is true, but it correctly places the heavier burden of proof on the elaborate, malicious one, just as the parent principle would.

How does Occam's razor relate to Bayesian reasoning?

Bayesian probability theory gives the razor a rigorous mathematical home. A theory that can explain many possible outcomes must spread its probability thinly across all of them, so it assigns low probability to the specific result you actually observe. A simpler theory that predicts a narrow range of outcomes concentrates its probability, and when one of those outcomes happens, it earns a larger boost in credibility. This is called the Bayesian Occam’s razor, and it shows parsimony is not an arbitrary preference but a consequence of probability itself. Ray Solomonoff extended this through information theory, linking the best explanation to the most compressed one.

Does Occam's razor apply to how animals think?

Yes, in two distinct ways. Animal nervous systems themselves favor parsimony: visual systems resolve ambiguous images into the simplest stable interpretation, a tendency captured by the Gestalt principle of Pragnanz. Separately, scientists who study animal minds apply Morgan’s Canon, formulated by C. Lloyd Morgan in 1894, which says never to attribute behavior to a higher mental faculty when a simpler one suffices. A bird at a feeder need not be credited with abstract reasoning if associative learning explains it. Like the razor, Morgan’s Canon is a default, not a dogma, and it yields when evidence genuinely demands a richer explanation, as for corvids and great apes.