The Wrong Word Is Doing the Damage
“Hallucination” names an exception that does not exist.
“Hallucination” names an exception that does not exist. The machine is not lapsing. It is uncalibrated — and that difference decides what we ought to be building, what we ought to be buying, and what we ought to stop paying for.
Thesis. The word “hallucination” describes a language model as a normally reliable perceiver suffering intermittent episodes, and thereby sets the industry’s target at the elimination of those episodes. No such episodes occur. True and false outputs are produced by one process, in one voice, with no internal moment at which the system departs from its ordinary operation. The correct description of that condition is not intermittent error but absent calibration: the confidence a system attaches to a claim does not predict whether the claim is true. Once the description is corrected, the target moves. It stops being fewer false statements, which is unreachable, and becomes confidence that means something, which is both reachable and measurable. And the moment you accept the corrected target, a great deal of what is currently sold as progress is revealed as the transfer of unpriced verification work from the vendor to the user.
Abstract. This essay argues that the dominant metaphor for language-model error is not merely imprecise but actively harmful, because it licenses a research and procurement programme aimed at an unattainable endpoint while ignoring an attainable one. I set out what the metaphor smuggles in; why the generative mechanism is symmetric between truth and falsehood; why zero error is not a coherent goal for any system permitted to speak about uncertain matters; what calibration is as a testable property, and how it differs from accuracy; and the arithmetic by which a less accurate system that reports its own limits can be worth more than a more accurate one that does not. I then take the strongest objections in turn — that the metaphor is harmless shorthand, that confidence scores already exist, that retrieval solves the problem, and that calibration is a solved training problem — and show why each fails. I close with what the corrected frame demands of anyone building or buying these systems, and with the conditions under which my own argument would be wrong. Two figures. Both are schematic or stipulated; neither reports measurement, and both say so.
I. What the word smuggles in
Words do work while you are not watching them. “Hallucination” arrived early, spread fast, and is now embedded in product documentation, regulatory drafts, board papers and the vocabulary of people who have never opened a model. It sounds technical. It is a metaphor, and a load-bearing one, and the load it bears is an entire picture of what these systems are.
Consider what the word presupposes. A hallucination is a perceptual failure in an organism that ordinarily perceives. The concept requires a baseline of veridical contact with the world, from which the episode is a departure. It requires, in other words, rails — and an event in which the thing leaves them.
Note how much that gives away for free. It concedes that the system ordinarily knows. It frames error as exceptional rather than constitutive. It implies that between episodes, the output is trustworthy in a way it would not be if we had described the situation differently. And it suggests a natural remedy: find the episodes, count them, reduce them. Publish the count. Reduce it further next quarter.
The metaphor also flatters. Only a mind can hallucinate. To say the model hallucinates is to say, in the same breath and without argument, that the model is the sort of thing that ordinarily perceives — that there is someone home, having a bad day. This is a substantial philosophical claim delivered as a bit of engineering shorthand, and it has never once been defended by the people who use it most.
And it consoles. If the machine is having an episode, the failure is the machine’s. It is not the user’s failure for having trusted a confident sentence, nor the vendor’s for having shipped a system with no mechanism for distinguishing what it knows from what it has composed. Everyone gets to keep their position. The word does that work quietly, which is how words do their best work.
The alternative description is unglamorous and has no medical grandeur about it at all. It is this: the machine does the same thing every time, and sometimes what it produces happens to be true.
II. One mechanism, two outcomes
Nothing in a language model perceives. What it does is produce continuations: given a sequence, it produces a distribution over what comes next, and something is sampled from it. That is the operation. It is the operation when the output is a correct citation and it is the operation when the output is a citation to a case that does not exist. The temperature is the same. The architecture is the same. The weights are the same. No subroutine fires on the false one.
When a continuation corresponds to the world, we call it knowledge. When it does not, we call it hallucination. The naming happens outside the system, afterwards, by a human who checked. Inside the system, the two events are indistinguishable, because there is nothing inside the system whose job is to distinguish them.
So there is no moment at which the machine leaves the rails, for the sufficient reason that there are no rails. There is one process, producing true and false statements by the same mechanism, at the same temperature, in the same voice.
The name for that condition is not “occasionally hallucinating”. It is uncalibrated.
I want to be careful about what this claim is and is not, because it is the load-bearing claim of the essay and it deserves to be stated precisely enough to be attacked.
It is an architectural claim, not a mystical one. It says: in a standard autoregressive generative system, truth and falsehood are not separated at the point of generation, because the generative objective was to model the distribution of text and not to model the world. That is a claim about design, and it follows from the design.
It is not a claim that no such separation could exist. It could. A system with a genuinely separate verification path — one that takes a produced claim, checks it against a source or a database or a proof, and reports the outcome of that check as something distinct from the fluency of the sentence — would break the symmetry. That is precisely what one should build. The observation that current systems lack it is not a metaphysical limit; it is a description of an omission, and omissions can be repaired. Everything in this essay is an argument for repairing this one.
Nor is it a claim that models are useless, or that they are stochastic parrots, or any of the other slogans that tend to arrive at this point in the conversation and end the thinking. A system can be enormously useful and completely uncalibrated. Most of them are. That combination is exactly what makes the situation dangerous rather than merely disappointing: if they were useless we would not use them, and if they were calibrated we could rely on them. They are neither, and so we use them and cannot rely on them, and the gap between those two facts is filled, at present, by hope.
III. Why zero is not a target
Under the hallucination frame, the goal is fewer false statements, and the implied endpoint is zero. Every roadmap that speaks of “eliminating hallucinations” has that endpoint written into it, usually without noticing.
The endpoint is unreachable, and not for reasons of engineering difficulty. It is unreachable for reasons of logic.
Some questions have genuinely unknown answers. Not unknown-to-the-model: unknown. Nobody has established them. Ask about the outcome of pending litigation, or the correct interpretation of a statute on which appellate courts have split, or what a dataset that has not been collected would show. A system that never errs on such questions is a system that never speaks about them.
Some evidence is genuinely ambiguous. Two credible sources conflict. The document is silent on the point at issue. The best available inference is an inference, and inferences are wrong at some rate, and the rate is not zero for anyone — human, machine, or committee.
And some claims are true when written and false later. Everything the system has absorbed was true at some moment. Facts have half-lives, and they vary by orders of magnitude between an atomic mass and the holder of an office. A system that never states anything that will later be false is a system that states nothing time-dependent, which is to say almost nothing anyone wants to ask about.
So the zero-error system exists, and it is the silent one. Any system permitted to speak about uncertain matters — which is the only kind worth having — will make false statements at some rate. The interesting question was never whether. It was always: does the system tell you which ones?
That question has an answer, and the answer is a number, and the number can be measured. That is the entire attraction of the reframe. It replaces an aspiration with a measurement.
IV. Calibration: the property that replaces it
Under the calibration frame, the target is different and achievable: the confidence attached to a statement should predict whether the statement is true.
This is a testable property of a system, not an attitude it strikes. The test is simple to describe and laborious to run. Collect every claim to which the system attached confidence 0.7. Check them against ground truth. If roughly seventy per cent are true, the system is calibrated at 0.7. Repeat across the range. In linear notation, the condition is:
P(Y = 1 | q) ≈ q, for every confidence level q the system uses.
Plot it and you get a reliability diagram: stated confidence on one axis, observed proportion true on the other. Perfect calibration is the diagonal. Below the diagonal is overconfidence. Above it is underconfidence, which is rarer and, in commercial systems, close to extinct.
Figure 1. Three ways to be right, only one of which is usable. The dashed diagonal is perfect calibration. The red curve is an overconfident system: its ordering is honest — claims it rates higher really are more often true — but its numbers are inflated, so a stated 0.9 buys you rather less than 0.9. The blue curve is calibrated: the numbers mean what they say. The gold point is the flat reporter, which is 90% accurate and expresses that accuracy in a single confident register. It has no curve, because it has no variation in confidence. There is nothing to calibrate. Schematic: the shapes are illustrative, not measured.
Three consequences follow that are routinely missed, including by people who use the word “calibration” in their marketing.
Calibration is not accuracy, and the two can move in opposite directions. A system that answers every question with “0.5” and is right half the time is perfectly calibrated and perfectly useless. What is needed is calibration and discrimination — the confidence must be honest, and it must separate the true from the false. In the forecasting literature this second property goes under names like resolution or sharpness, and the standard formulation is sharpness subject to calibration. Either property alone is a party trick. Together they are the whole of what we mean by a trustworthy source.
Calibration is a property of a population of claims, not of any single claim. “How confident are you?” is meaningless asked about one sentence and meaningful asked across ten thousand. This has an uncomfortable implication that the field has not fully absorbed. Validating calibration requires scoring; scoring requires ground truth; and ground truth is scarcest in exactly the domains where the machine is most wanted — novel legal questions, frontier science, anything not yet settled. Where verification is cheap, calibration is measurable and matters least. Where it matters most, it is hardest to measure. This is not a reason to abandon the project. It is a reason to be extremely suspicious of any calibration claim that does not name the distribution it was measured on.
Calibration measured on one distribution guarantees nothing about another. Calibration on general knowledge is not calibration on case law. A model’s confidence was learned where feedback was dense; deployment happens where it is sparse; and the system has no way to notice it has crossed the boundary, because the out-of-domain question is made of the same words in the same grammar as the in-domain one. This is not a footnote. It is the central engineering difficulty, and it means calibration is not a property you install once and own. It is a property that degrades under shift and must be re-measured per domain, at recurring cost. That does not appear on a model card.
None of this apparatus is new. Meteorologists solved the measurement problem for rain long before anyone needed it for text: the squared-error scoring rule dates from 1950, its decomposition into components including reliability and resolution from 1973, and the formal comparison of forecasters from 1983. The references are at the foot of this essay, with a note on what I have and have not read. The relevant point here is that the machinery has been sitting on the shelf for seventy years, fully worked out, while an industry with more capital than any in history has been publishing accuracy figures.
V. The arithmetic of trust
Here is the practical consequence, and it is the part that should interest anyone who signs a purchase order.
Take two systems and a hundred claims.
System A is right sixty per cent of the time, and it tells you which sixty. In the limiting case — and I will come back to the fact that it is a limiting case — it flags exactly the forty claims it gets wrong. You check the flagged forty and proceed with the rest.
System B is right ninety per cent of the time and sounds identical throughout. Ten claims in the hundred are false and there is no signal indicating which. To reach certainty you must check all one hundred, because the ten are hiding among the ninety and nothing distinguishes them.
Figure 2. Higher accuracy, lower value. Left: what each system hands you out of 100 claims. Right: how many you must check by hand to reach certainty. Stipulated arithmetic, not measurement. System A is idealised — perfect discrimination is a limiting case and an upper bound on the benefit — and System B is the opposite limit. Real systems sit between. The ordering of the two right-hand columns is the point.
System B has higher accuracy and lower value. It saved you nothing. It moved the work downstream, and it disguised the move by sounding assured while doing it.
Now relax the idealisation, because System A as described is a bound rather than a product. Suppose System A’s flagging is imperfect: it flags fifty claims, of which thirty-five are the errors, and five errors escape unflagged. You now check fifty and carry five unlocated errors instead of ten. Half the verification burden of System B, half the residual risk. The advantage narrows but does not reverse, and it does not reverse until the flagging is so poor that it is nearly uninformative — at which point System A is simply System B with worse accuracy, and the comparison collapses back into the one the leaderboards already make.
That is the real design frontier, and it is worth stating in one sentence, because it is the sentence the industry has not said out loud: the value of a system is not its accuracy but its accuracy conditional on its own reported confidence, integrated over how much of the work it lets you skip.
Three observations follow.
First, the comparison is not exotic. It is what every professional does when deciding whether to rely on a colleague. A junior who says “I’ve checked these six and I’m unsure about the seventh” is more useful than one who returns seven with equal assurance and a better underlying hit rate, because the first has done part of your verification for you and the second has done none of it while appearing to.
Second, the accuracy figure alone cannot distinguish the two systems. It is not that the figure is imprecise; it is that the quantity it measures is not the quantity that determines value. Reporting accuracy without discrimination is like reporting a portfolio’s return without its variance, and we have known for seventy years why that is not permitted.
Third — and this is where the essay stops being about statistics — everyone selling accuracy numbers is selling the second system. Not through fraud, in most cases. Through the ordinary operation of measurement: the number that is easy to compute becomes the number that is reported, the number that is reported becomes the number that is optimised, and the number that is optimised becomes the definition of the product. The metric was chosen because it was tractable. It then quietly redefined what the industry thought it was building.
VI. Why the market prefers the worse system
Suppose I am right that calibration matters more than accuracy. Why has the market not noticed?
Because of what each property costs to demonstrate and what each does for the seller.
Accuracy is a single number, comparable across vendors, computable on a fixed benchmark, and it goes up. It fits in a headline and on a chart with an arrow. It permits ranking, and ranking permits competition, and competition on a public metric is the most efficient marketing apparatus ever devised, because your rivals do the promotion for you by contesting the leaderboard.
Calibration is a curve, not a number. It requires stating the distribution it was measured on, which invites the question of what happens off that distribution, which is a question with a bad answer. It is domain-specific and therefore not comparable across vendors in any clean way. Improving it can reduce the headline accuracy figure, because a system that abstains on the hard cases is a system that answers fewer questions. And its principal benefit accrues to the user’s verification budget — a cost the vendor does not carry and cannot see.
Add the training dynamics. A system penalised for saying “I do not know” learns to guess, and learns to guess fluently, because fluent guesses score better than honest silence with any evaluator grading on satisfaction. Optimising for user satisfaction and optimising for epistemic honesty are distinct objectives that agree most of the time and diverge exactly where the stakes are highest. Where they diverge, the deployed system does what it was scored on. It always does. There is no version of this where the incentive does not win.
So the structure of the market selects for confident systems, and the structure of the metric conceals the cost. That cost is real, and it is being paid, but it is paid in the private time of users checking things, which appears in no vendor’s accounts and no benchmark’s numbers. It is an externality in the technical sense: a cost of production borne by someone other than the producer, invisible to the price. And like every externality, it will continue to be produced in excess until someone measures it.
I should mark the epistemic status of this section honestly. That accuracy figures dominate public claims while calibration, where reported at all, sits in appendices, is an observation about public presentation. It is open to correction by anyone with a better survey of current practice than mine, and I would welcome the correction. The argument about incentives does not depend on it; it depends only on the relative cost of demonstrating the two properties, which is structural.
VII. Four objections, taken seriously
An argument that has not survived its best counterarguments is a slogan. Here are the four strongest I know.
“It’s just shorthand. Everyone knows what it means.”
This is the most reasonable objection and it is still wrong, for a reason that has nothing to do with pedantry.
Metaphors are not inert. They select the research programme. If error is an episode, you count episodes, and you build detectors for episodes, and you report the episode rate falling, and every one of those activities is coherent, fundable and misdirected. If error is a symptom of missing calibration, you build the missing measurement apparatus instead. Those are different capital allocations. The metaphor decided which one you made, before anyone reasoned about it.
The shorthand also does specific damage in front of non-technical audiences, which now includes courts, regulators and boards. A judge told that a system “sometimes hallucinates” understands: usually reliable, occasionally not, exercise care. That understanding is wrong in a way that will produce bad rules — rules aimed at reducing an error rate rather than at requiring systems to report their own uncertainty and to abstain where they cannot. The first kind of rule is unsatisfiable and will therefore be gamed. The second is satisfiable and auditable. We are, at present, drafting the first kind.
And notice that people who use the shorthand do not in fact know what it means, because they routinely draw the inference the shorthand licenses: that a confident output is more likely true than a hedged one. In an uncalibrated system that inference is unsupported. If everyone really knew what the word meant, nobody would ever say “but it seemed sure”.
“Confidence scores already exist. Look at the token probabilities.”
They exist and they are not the thing.
The distribution is over word forms, not propositions. “Paris” and “the capital of France” express one claim with entirely different token mass. Paraphrase the answer and the number moves while the belief does not.
Long claims decompose into many tokens, most of which are grammatically forced. In “the ruling was handed down in 1987”, nearly every token is determined by syntax and one carries the factual load. The joint probability of the sequence is dominated by grammar. Averaging it tells you the sentence is well formed, which you could already see.
Decoding parameters move token probabilities. Temperature, sampling strategy, prompt phrasing, the presence of an earlier example — all shift the numbers. None of them shift the truth of the claim.
And the mapping between tokens and claims is not one-to-one in either direction. One claim can be spread across a paragraph. One token can carry a whole claim, and the token is frequently “not”.
So a high token probability sitting on a false statement is not an anomaly requiring explanation. It is the expected behaviour of a system trained to model text rather than to model the world.
The same objection sometimes arrives in verbal form: the model says it is confident. But where does that phrase come from? From pre-training on text in which humans hedged, and from preference optimisation in which raters rewarded outputs that sounded a certain way. Neither process had access to whether the specific claim was true; both had access to whether the phrasing went down well. The predictable result is that verbal confidence tracks the register of the question rather than the evidence for the answer. Ask in a crisp technical voice and receive assured prose. Ask tentatively and receive hedges. Same claim, same evidence, different confidence language — because the confidence language is a costume selected to match yours, not a measurement taken before the sentence was composed.
What is needed instead is claim-level uncertainty: decompose the output into propositions, attach an estimate to each proposition, and score those estimates against outcomes. That is a different object from the softmax. It does not fall out of scale. It has to be built, deliberately, as its own thing.
“Retrieval fixes it. Ground the model in documents.”
Retrieval fixes one failure mode and introduces three.
It fixes the obvious one: information that was absent is now present.
It introduces, first, retrieval error — the wrong document, the outdated version, the superseded ruling, the second-best match returned confidently. Confidence in the answer now inherits confidence in the search, and almost nobody scores the search. The system reports how sure it is of the sentence and says nothing about how sure it is that it found the right page.
Second, misreading. A model can retrieve a correct, authoritative, current source and misstate what it says: reverse a holding, drop a qualifier, convert a conditional into a flat assertion, attribute the dissent to the majority. This is the dangerous mode, because the citation is real. The reader performs the check they know how to perform, confirms the source exists and is on point, and stops. The error survives verification because the verification was aimed at the wrong layer.
Third, authority laundering. A footnote confers the appearance of provenance on a sentence the model composed. Sentence and source become visually bound, and only reading the source separates them. The apparatus of scholarship is now available to anyone, at no cost, with no scholarship behind it. For law, medicine and peer review this is not a minor inconvenience; it is a solvent applied to a load-bearing structure, since all three run on the assumption that the cost of producing a citation is high enough to make citations informative.
Provenance done properly means: this specific proposition is supported by this specific passage, and here is the passage, adjacent, in full. Anything weaker is decoration — and decoration on a false claim is worse than none, because it consumes the reader’s suspicion.
“Calibration is a training problem and it is being solved.”
Partly true, and the part that is true does not rescue the position.
Confidence can be trained. There is active work on training models to abstain when uncertain and on optimising proper scoring rules so that a model’s reported probability is rewarded for being honest rather than for being high. I am not going to characterise the results of specific papers here, because I have not read them in full and this essay does not permit itself to describe work it has only skimmed. What can be said from the structure of the problem is enough.
Calibration is measured on a distribution and deployed off it. Three specific breaks follow, each sufficient alone.
Domain shift, as above: confidence learned where feedback was dense is meaningless where it was sparse, and the system cannot see the boundary.
Leading prompts: a confident premise in the question raises the confidence of the answer. Calibration measured on neutral questions predicts nothing about behaviour on loaded ones, and real users load their questions constantly, usually without intending to. “Explain why X causes Y” has already conceded the point that should have been tested.
Long chains: confidence in a conclusion ought to reflect confidence in each step that produced it. When a system reports high confidence in a ten-step argument while holding modest confidence in several steps, it is not propagating uncertainty; it is re-deriving confidence from the fluency of the final sentence. I want to be careful here, because there is a tempting overstatement available and I decline it: whether step errors compound multiplicatively depends on independence between steps, which is an untested assumption and not one I will smuggle in. The narrower claim is sufficient. The conclusion’s stated confidence is generated by the same next-token process as everything else in the paragraph, and is therefore not an aggregation of anything.
The honest summary is that calibration is trainable and not durable. It is a maintained property, like a certification, not an installed one, like a feature. Anyone claiming to have solved it owes you the distribution.
VIII. What the corrected frame demands
If the reframe is right, it generates requirements. Here they are, briefly, because each deserves its own essay and several will get one.
Uncertainty attached to propositions rather than tokens. The unit of truth is the claim. The unit of measurement must match.
Epistemic and aleatoric uncertainty reported separately. Epistemic uncertainty is reducible by going and finding out: the document exists and has not been read. Aleatoric uncertainty is not: the coin has not landed. “Sixty per cent” is an incomplete report, because sixty per cent because-the-world-has-not-decided and sixty per cent because-I-have-not-looked licence opposite actions. The first says decide now under risk. The second says stop and check, and the uncertainty disappears for the price of a search. Collapsing them into one number destroys the only information you needed.
Abstention that is selective, and measured as such. “I do not know” is scored as failure and is the opposite: it is the only output that locates the boundary. But it must be earned. The measure is not the refusal rate — a system that abstains constantly is as useless as one that never does. The measure is whether abstentions land on the questions the system would have got wrong. Plot accuracy on answered questions against coverage. If accuracy does not rise as coverage falls, the abstention is noise and the humility is theatre.
Abstention paired with action. Silence leaves the user where they started. “I do not know” plus “here is the fact that would settle it, here is who holds it, and here is how the answer changes either way” is a different product. That is the operational test of whether uncertainty is real: a system that can name the decisive missing evidence has a model of what it does not know. A system that only shrugs has learned a phrase that raters liked.
Assumptions enumerated rather than buried. Every point at which the system filled a gap rather than read a fact should be visible. Hidden assumptions are the serious failure precisely because they contaminate the conclusion while pretending not to exist. An assumption that announces itself can be tested, rejected, replaced, and the argument repaired. One that stays quiet cannot be, and is inherited downstream in silence until it has been built upon.
Verification proportional to consequence. Uniform checking is unaffordable, so it gets abandoned, and the abandonment is uniform too. The rule is decision-theoretic and unglamorous: weigh the cost of checking against the cost of being wrong. A date misremembered in conversation costs a moment. A citation in a filed document costs a career and, in some jurisdictions, a sanction. Same model, same confidence number, radically different obligation — and the system cannot help you allocate unless it knows what each claim is load-bearing for.
Notice what is absent from that list: being right more often. Capability and epistemic honesty are different variables. You can move either without moving the other, and essentially all of the industry’s effort, capital and publicity has gone into the first, because the first photographs well.
IX. What would show me wrong
An argument that cannot fail is not an argument, so here is what would defeat this one.
If a system were demonstrated whose reported confidence remained calibrated across domains it was not tuned on, under adversarial and leading prompts, and over multi-step reasoning with tool use — with the reliability curves published and the distributions named — then the substantive complaint dissolves. The metaphor would still be wrong, but wrong in the harmless way that dead metaphors are wrong, and I would file the objection under pedantry and stop writing about it.
If it turned out that discrimination in confidence is impossible in principle for systems of this architecture — that there is a proof, not a difficulty — then the reframe is cruel rather than useful, since it would name a target nobody can hit. I know of no such proof, and the existence of trainable abstention suggests the opposite, but I hold the position provisionally.
And if the verification burden I describe turned out to be empirically small — if users of confident-sounding systems in fact incur little checking cost because errors are rare and low-consequence in the domains where deployment actually happens — then the externality argument weakens considerably. That one is an empirical question I have not answered and cannot answer from an armchair. It is measurable. Someone should measure it. It would make a better paper than most of what is currently being written about hallucination rates.
X. Coda
The useful machine is not the one that knows the most. It is the one whose reported uncertainty means something in domains it was not tuned for — which knows when further evidence is required, which can name what that evidence would be, and which does not convert not-knowing into fluent invention.
We do not have that machine. We have something that produces the surface of one, at scale, cheaply, and with a vocabulary that describes its failures as episodes in an otherwise sound mind.
The first repair is not technical. It is to stop using the word that told us the problem was rare.
Classification
Reframing — argumentative. The claim that true and false outputs share a generating mechanism is architectural and follows from how autoregressive generation works; a system with a genuinely separate verification path would break the symmetry, which is the point of the essay rather than an objection to it. The 60/90 comparison is stipulated arithmetic, not measurement, and its idealised form is explicitly a bound. The claim about how the market presents accuracy versus calibration is an observation about public presentation and is open to correction. The compounding of errors across reasoning steps is deliberately not asserted, because it depends on an independence assumption I have not tested.
A note on sourcing
I have not read the full texts for this essay cover to cover, and I have therefore restricted myself to not characterising their arguments in detail. That distinction matters, and stating it is cheaper than pretending otherwise.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3. doi:10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2
Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12, 595–600. doi:10.1175/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2
DeGroot, M. H., & Fienberg, S. E. (1983). The comparison and evaluation of forecasters. Journal of the Royal Statistical Society, Series D (The Statistician), 32(1/2), 12–22.




So when I was on a beautiful magic mushroom trip a few years back was that a hallucination or another dimension? 😂