A Bayesian posterior can look rigorous precisely because the same modeling bias entered several times: when the hypotheses were chosen, when their priors were assigned, and when the likelihoods of the evidence were estimated. All of those choices happen before Bayes’ theorem performs the update.
Bayes then combines those inputs correctly, so the error emerges from the calculation with the appearance of mathematical discipline. “Before Bayes” does not mean “before the prior.” It means before the inference rule receives its model.
Robin Hanson’s Aliens Are Not Crazy Unlikely argues that some-UFOs-are-aliens deserves a prior of roughly 10⁻⁴ to 10⁻³ rather than the near-zero prior most intellectuals give it. The arithmetic is not the weak point. Hanson understands priors, selection effects, strategic behavior, and conditional reasoning. The distortion enters in how the model is built.
Precision without calibration
Hanson is right that extraterrestrial life should not receive a paranormal prior. Nothing in known physics makes life elsewhere absurd, but “life exists elsewhere” is far weaker than “advanced extraterrestrials are operating near Earth and explain some UAP.”
The prior he is defending comes from On UFOs-As-Aliens Priors, which decomposes the stronger claim into four conditional steps: that Earth was seeded by panspermia within its stellar nursery, that a sibling star produced a long-lived advanced civilization well before us, that this civilization suppressed its own expansion and travelled here to wait and suppress ours, and that it permits UFO-style encounters while it waits. He puts each at at least ten percent. Four such factors give 0.1 × 0.1 × 0.1 × 0.1 = 0.0001, which he offers as a floor rather than an estimate, and which is why his current range runs an order of magnitude above it.
There is nothing illegitimate about subjective priors. A Bayesian must assign probabilities even where frequencies are unavailable, and Hanson’s willingness to state numerical assumptions localizes disagreement: if he assigns ten percent to a step and someone else assigns one percent, the dispute becomes explicit instead of hiding inside words like “plausible” or “unlikely.” Framing those assumptions as lower bounds concedes something real about the uncertainty, and it blunts the usual complaint about false precision.
But a bound inherits everything that went into it. “At least ten percent” is still a judgment about a quantity nobody can estimate: the rate of interstellar panspermia, the fraction of seeded systems producing long-lived civilizations, the prevalence of expansion-suppression policies, the frequency of deliberately ambiguous contact. Calling the output a minimum fixes the direction of the error by assumption, which is a strong claim about quantities we have no way to check.
Point estimates and one-sided bounds hide the same thing: uncertainty in the estimates themselves. If one conditional probability plausibly spans several orders of magnitude, replacing that spread with 0.1 makes the product look far more stable than it is, whichever way the bound points. Sensitivity analysis answers this. Refusing to put numbers on anything does not. Explicit assumptions can be inspected, and inspection is not calibration.
The model generator
Before Bayesian updating begins, something has already decided which hypotheses exist. Call it the model generator: the process that determines which possibilities receive explicit representation, how finely they are partitioned, what causal structure they contain, what priors seem plausible, and what likelihoods are assigned to the evidence.
This creates model-generator bias. The same conceptual preference can shape hypothesis selection, granularity, priors, and likelihoods, so Bayesian updating may amplify the error instead of correcting it. Garbage in, garbage out understates the problem, because the inputs are not independently garbage. Correlated model error enters through several nominally separate terms and is then multiplied by formally correct inference. Occupying different terms in the equation does not make hypothesis space, priors, and likelihoods epistemically independent.
This connects to what I previously called tool bias. We formulate problems using the intellectual machinery we know how to operate, and Hanson operates the machinery of incentives, coordination, signaling, and equilibrium as well as anyone working.
Alien visitation becomes cognitively tractable once framed that way. Perhaps expansion is suppressed, concealment is strategic, ambiguity is intentional, and silence is an equilibrium. Hanson’s third and fourth conditions are exactly this: a civilization with the coordination capacity to police its own expansion and its siblings’, running a contact policy calibrated to show without telling. Each move is possible, but together they convert a weak-data astronomical problem into a strategic-agent problem, where his tools are richest. The resulting model acquires structure, motives, and explanatory depth, while its competitors remain a residual category.
Selecting the anomaly
UAP evidence is heavily post-selected. Over decades, pilots, radar systems, cameras, satellites, military sensors, and civilians have generated an enormous number of observations; most are ordinary, some are strange, and investigators concentrate on the strangest surviving cases.
That selection changes the likelihood calculation. Ordinary mechanisms get millions of opportunities to throw off an anomaly, and what we need to know is how often that many trials produce a case this bizarre. A one-in-a-million sensor failure is extraordinary in a single trial and expected a hundred times across a hundred million.
The strongest surviving anomaly is usually selected precisely because ordinary explanations fit it badly, so some of its evidential force comes from the search procedure. Evaluating a retrospectively selected case as though it had been specified in advance manufactures strength. Financial anomalies, medical miracles, and apparent prediction successes share the structure: large search spaces generate impressive tails.
This bears directly on Hanson’s legal analogy. He compares the prior for alien visitation with the roughly one-in-a-million prior that a particular person murdered a particular victim, and observes that ordinary testimony routinely overcomes odds that long. But a murder investigation begins from a prospectively fixed event: this person died, at roughly this time, under these circumstances. Evidence is then gathered about that event, and the suspect pool is built afterwards. “The dozen most solid UFO cases” are selected for their anomalousness out of eighty years of observation, and then used to estimate how anomalous UAP observations are. The analogy would hold if we named twelve encounters in advance and then asked what the evidence showed.
“Mundane” is not one hypothesis
The alien hypothesis is often represented as one coherent generative story while its competitor is called “mundane explanations.” Hanson does better than that. He partitions the field into error, hoax, secret Earth organizations, and aliens, and he is explicit that he is estimating only one of the eight numbers a full analysis needs. But three residual categories still make a coarse partition of something that is not three things. “Error” alone is a distribution over sensor artifacts, atmospheric effects, software failures, calibration errors, perceptual mistakes, bad metadata, radar propagation, and combinations of these; “secret Earth organizations” covers everything from a classified airframe to a mislabelled test range.
That creates an asymmetric contest. Failure to find one convincing mundane explanation for a case can feel like evidence for aliens, even though the quantity Bayes needs is the aggregate likelihood across the whole ordinary hypothesis space. A case can have no known specific explanation while still being unsurprising as the tail of a large, messy error-generating process.
“Unexplained” and “improbable under the null” must be kept separate. A case may remain unresolved because crucial information is missing, which makes the likelihood uncertain rather than low. Ignorance is uncertainty, not evidence for either side.
Conditional decomposition and hypothesis-space partition are different operations. Breaking one generative story into successive conditions is not the same as dividing the space into competing alternatives. Hanson does the former; the granularity problem concerns the latter.
We could treat “alien visitors” as one family while dividing mundane mechanisms into dozens, or reverse the partition and split alien visitation into thousands of incompatible origins, motives, technologies, and contact policies. The world does not supply a canonical partition. The model generator chooses one, and that choice affects priors, apparent simplicity, and posterior comparisons.
Likelihoods are where the disagreement lands
Hanson and I do not differ dramatically on the prior. His floor for alien visitation is 10⁻⁴, with a stated range reaching 10⁻³, and 10⁻⁴ is roughly where I would center a broad prior distribution.
The disagreement appears when the UAP evidence enters. Hanson thinks the strongest cases supply a large enough likelihood ratio to move a long way beyond the prior: combining priors and posteriors across his four categories, he judges hoax and aliens to be the most likely explanations for the hardest-to-explain cases. I do not, because the likelihoods themselves are distorted by post-selection, heterogeneous mundane alternatives, missing data, and poorly characterized error processes.
Likelihoods are not handed to us by Bayes’ theorem. They are constructed by the same model generator that constructs the hypothesis space and the priors, which is why the disagreement is still a case of bias before Bayes even though it surfaces at the evidence.
For UAP evidence E, we want to compare P(E | aliens) with P(E | ordinary mechanisms), and neither quantity is directly known. The alien likelihood requires assumptions about hypothetical alien behavior. The ordinary likelihood requires a model of sensors, human observers, environmental effects, institutional filtering, and post-selection.
A rich strategic model makes P(E | aliens) easy to imagine, while P(E | ordinary mechanisms) feels small because no single ordinary explanation reproduces the whole case. That feeling is the granularity problem arriving in the denominator.
An unmeasured likelihood should not silently become a tiny likelihood, and uncertainty should not automatically be resolved in favor of mundane explanations either. Unmodelled error processes should widen the posterior distribution instead of pushing it toward whichever story is easier to model.
Extraordinary evidence
Hanson’s post opens by asking whether Sagan had it wrong. Defended properly, “extraordinary claims require extraordinary evidence” is not a sociological rule at all. It follows from prior odds, including his own.
Take his floor of 10⁻⁴. The prior odds are then about 10,000 to 1 against. Reaching even odds requires a Bayes factor of about 10⁴; reaching 95 percent requires about 1.9 × 10⁵. There is nothing mystical here. Low prior odds require a correspondingly large likelihood ratio.
A recovered machine containing unmistakably nonhuman technology could supply such a ratio, and repeated measurements from independent calibrated instruments might accumulate it. Ambiguous testimony and retrospectively selected sensor anomalies have a harder time, because nobody has measured their likelihood ratios. Hanson is right that strangeness alone is no reason for a near-zero prior. The question is whether the evidence discriminates strongly enough between visitation and the alternatives.
So what should the posterior be?
Criticizing Hanson’s numbers without supplying my own would be cheap. I put my current probability that extraterrestrial technology is operating on or near Earth at roughly 10⁻⁴: about one chance in ten thousand. The fourth decimal place means nothing. Under other plausible model choices I could defend values from 10⁻⁶ to 10⁻², and that spread is a sensitivity range over model assumptions, not a probability distribution over my own credence. The number is a working credence, not an estimated physical frequency.
The prior should be broad because the biological, technological, spatial, and behavioral terms are almost entirely uncalibrated. So I resist multiplying several sharp point estimates, and I resist Hanson’s alternative of multiplying four lower bounds, because the multiplication is only as informative as the bounds themselves. If every conditional probability really is at least 0.1, then 10⁻⁴ really is a valid floor; the arithmetic is not at fault. What is unsupported is the claim that four quantities this badly characterized each deserve that floor: interstellar panspermia, civilizational longevity, expansion suppression, contact policy.
The UAP record provides some positive evidence. There are credible observers, occasional multi-sensor cases, and a residue harder to dismiss than an arbitrary collection of anecdotes. Some of that residue would be less likely if nothing unusual were happening at all.
Selection and absence offset it. The strongest cases were drawn from an enormous population of observations, their error distributions are poorly characterized, many lack the data needed for reconstruction, and decades of increasingly capable cameras, radar, satellites, and other sensors have produced no publicly verified observation that sharply discriminates extraterrestrial technology from the mundane aggregate.
Once post-selection, unresolved error distributions, and the persistent absence of decisive observations are included, I would assign the total UAP corpus a Bayes factor of order one. Some cases push upward. The long absence of reproducible, discriminating evidence pushes downward, unless the alien hypothesis is supplemented with a contact policy designed to prevent exactly that evidence from appearing. Hanson’s fourth condition does this work. The supplement is available, but it costs him something: once ambiguity is predicted by a behavioral assumption introduced to explain the absence of clarity, ambiguous encounters carry less evidential weight, not more.
That leaves the posterior close to the prior, around 10⁻⁴. Hanson does not stay near his. Ranking aliens among the likeliest explanations for the hardest cases puts his posterior three or four orders of magnitude above mine on identical evidence. We start from the same prior, we apply the same theorem correctly, and we end up that far apart, because the only quantity separating us is a likelihood ratio neither of us derived from anything. It was constructed.
This number could change quickly. A reproducible material sample, an independently observed artifact, or prospectively specified multi-sensor evidence showing capabilities strongly excluded by ordinary mechanisms could move 10⁻⁴ to high confidence in one update. Bayesian skepticism does not require slow belief revision; it requires evidence that discriminates.
The uncertainty around my own estimate is large, and pretending otherwise would reproduce the error I am objecting to.
Postscript
Bayesian updating is an error-correction mechanism inside a represented hypothesis space, and it cannot correct errors in the representation itself. If the correct hypothesis is omitted, Bayes cannot assign it probability; if one hypothesis is represented at favorable granularity while another is an incoherent residue, Bayes does not repair the partition; and if the same conceptual bias shapes priors and likelihoods, Bayes may multiply the distortion while leaving the arithmetic flawless.
That is why Bayesianism is not a complete epistemology. The model needs its own error-correction mechanisms: alternative hypothesis generation, adversarial modeling, sensitivity analysis, explicit treatment of selection effects, and uncertainty over priors and likelihoods. Hanson’s decomposition supplies some of this, because four numbered conditions can be attacked one at a time in a way that a verbal judgment cannot.
Bayes can update the uncertainty we represent, but it cannot represent the uncertainty we omitted. It can compare the hypotheses we supplied, but it cannot discover which alternatives our preferred conceptual tools prevented us from constructing. When the same bias shapes the hypothesis space, the priors, and the likelihoods, a mathematically correct posterior may be a more confident expression of the bias that came before Bayes.


