At a country fair in 1906, hundreds of people guessed the dressed weight of an ox. Statistician Francis Galton later reported that the middle estimate was close to the actual weight. His famous example shows a crowd doing something striking: making a better collective estimate than many of its members could make alone.¹
The mechanism is less mysterious than it sounds. One person guesses too high, another too low. If errors differ and do not all point the same way, combining the guesses can cancel some of them. You do not need every participant to be an expert. You do need a way to gather useful information without letting one mistake infect everybody else.
Independence is the hidden ingredient
Imagine asking ten people to estimate the number of sweets in a jar. If they write answers privately, you get ten partly independent observations. If the first person announces a confident guess, the other nine may drift towards it. The group now looks like ten votes but contains less than ten independent pieces of information.
Experiments have found that social influence can make estimates converge without making them more accurate.² Later research has also explored situations where sharing information can help.³ The effect depends on who knows what, how information flows and how the answers are combined. “Crowds are wise” is therefore a conditional claim, not a law.
Shared bias is another problem. If everyone uses the same faulty source, averages cannot wash it away. A group of people repeating one mistaken figure is not a crowd of independent observers. Nor will a poll of novices necessarily beat a specialist when the task requires rare expertise.
What kind of question suits a crowd?
Crowd aggregation works especially well for questions with a measurable answer and errors that can balance, such as estimating a quantity. It is harder to apply to questions of values or taste. Averaging opinions on what a neighbourhood ought to build does not produce a scientifically correct answer in the same sense as estimating a weight.
There are choices about the calculation, too. A median limits the influence of wild guesses; a mean uses all values but can be pulled by extremes. Galton used the median in his original report.¹ The method should fit the data, rather than being chosen after seeing which answer looks impressive.
A practical way to use collective judgement is simple: ask people for their first answers separately, gather a range of perspectives, then discuss why they differ. The discussion may reveal information, but the initial independent guesses should not be lost. That is where much of the crowd’s value lives.
Why aggregation works mathematically
Suppose each estimate equals the true value plus an error. If those errors are roughly independent and have no common direction, averaging makes some positive and negative errors offset. The typical size of the remaining random error falls as more independent estimates are added. But if everyone shares a bias—perhaps all misread the same misleading image—the average preserves that bias. A thousand repetitions of one bad assumption do not create wisdom.
The median offers a different kind of protection. It is the middle response after estimates are sorted, so a wildly high or low answer has little influence unless many people share it. That helps explain Galton’s interest in the middlemost guess.¹ Neither mean nor median is automatically best; the choice depends on the pattern of errors and the question being asked.
Discussion can help, but timing matters
It would be too simple to insist that group members must never speak. An expert may correct a misunderstanding, or someone may reveal a fact others lack. The danger is that early social influence can also make people converge before their independent information has been collected. Experiments on social influence show that the outcome depends on the network of discussion and what participants learn from one another.²,³
One useful method is to separate stages. First gather private estimates. Then share reasons and evidence. Finally, if appropriate, ask for revised estimates and compare the change. The first answers preserve diversity; the conversation can add information without erasing the record of what people thought independently.
This also clarifies why crowd wisdom should not be confused with majority rule. Voting on a policy requires weighing values and interests. Estimating a physical quantity has a right answer against which a collective estimate can be checked. The same phrase “the crowd knows” should not blur those very different tasks.
When a crowd should not decide
Aggregation cannot create information that nobody has. If every participant is guessing blindly, the average may be precise-looking but wrong. A specialist may also be necessary when the answer depends on technical evidence unavailable to the crowd.
Good group judgement therefore begins with the question: who has useful, distinct information? Diversity is helpful when it brings different observations or models, not merely different labels on otherwise identical guesses. The wisest procedure may combine expert analysis with independent public estimates rather than setting them in competition.
The smartest crowd is not necessarily the biggest or loudest. It is one whose members notice different things and whose errors have room to cancel.
