Wine critic scores: how much they move bottle prices
Updated
Wine critic scores do most of their price work above 86 points. Below that, the natural experiment that tested them could not tell their effect apart from zero.
Wine critic scores are a taster's quality judgement expressed as a number, almost always on the 50 to 100 scale Robert Parker popularised through The Wine Advocate, the newsletter he founded in 1978. Their effect on price is real, measurable and uneven. Hadj Ali, Lecocq and Visser exploited the one year Parker did not taste Bordeaux en primeur in spring, which left château owners setting prices with no grades to work from, and found that being graded at all was worth €2.80 a bottle on average across 233 châteaux. Split by grade, that average hides everything: the effect ran at about zero for a wine graded 84.5 and reached roughly €14 a bottle at 93.5. Jones and Storchmann report a single Parker point at 7% on average for auction prices across 21 prestigious Bordeaux, ranging from under 4% to more than 10%. So the question is never what a wine scored. It is what the market already expected it to score.
What is a wine critic score, and what does 90 points mean?
It means whatever the publication says it means, and the publications do not agree.
Parker's system starts every wine at 50 and ranks it,'s summary of his method, "on a scale from 50 to 100 points based upon the wine's color and appearance, aroma and bouquet, flavor and finish, and overall quality level or potential". Because nothing scores below 50, "51 rather than 100 different ratings are possible", and in practice the working range is narrower still. The Wine Advocate puts wines scoring 80 to 89 in a band it describes as "barely above average to very good", which is a wide door for a ten-point band. The newsletter takes no advertising and publishes in excess of 12,000 reviews a year.
Wine Spectator lists its own bands, and they do not line up with Parker's:
| Band | Wine Spectator's published definition |
|---|---|
| 95 to 100 | "Classic: a great wine" |
| 90 to 94 | "Outstanding: a wine of superior character and style" |
| 85 to 89 | "Very good: a wine with special qualities" |
| 80 to 84 | "Good: a solid, well-made wine" |
| 75 to 79 | "Mediocre: a drinkable wine that may have minor flaws" |
| 50 to 74 | "Not recommended" |
Others do not use 100 points at all. Jancis Robinson and Michael Broadbent score out of 20. Parker also handled unfinished wine differently from bottled wine: en primeur grades were published as intervals such as 88 to 90, and across the 2001 and 2002 vintages in that Bordeaux sample the interval was one or two points wide 91% of the time. Wine Spectator describes its own version of that practice as "rolling four-point spreads" for wines tasted from barrel.
Two scores that read the same are not the same measurement. Before you compare them, check the publication, the scale, and whether the wine was in barrel or in bottle.
How do critic scores affect wine prices?
They move price, and the size of the move depends on what else the researcher controlled for. Six studies, six answers, and the spread between them is the useful part:
| Study | What it measured | Sample | Finding |
|---|---|---|---|
| Jones & Storchmann (2001) | Auction prices against Parker points | 21 prestigious Bordeaux | One point worth 7% on average, "ranging from less than 4% to more than 10%" |
| Hadj Ali & Nauges (2004) | En primeur prices against Parker points | 132 châteaux, 16 vintages | One point worth 1.01% |
| Dubois & Nauges (2005) | The same, controlling for quality the producer knows and the buyer does not | 108 châteaux, 1994 to 1998 | One point worth 1.38% controlled, 3.95% uncontrolled |
| Hadj Ali, Lecocq & Visser (2007) | Price when graded minus price with no grade | 233 châteaux, 2002 and 2003 openings | €2.80 a bottle on average, zero below 86 points, about €14 at 93.5 |
| Masset, Weisskopf & Cossutta (2015) | Score surprise against en primeur prices | 12 critics, 2003 to 2012 vintages | For Parker and Jean-Marc Quarin, "a 10% surprise in their scores leads to a price increase of around 7%" |
| Ashton (2016) | Parker and Robinson ratings against futures prices | More than 1,700 red Bordeaux, 2004 to 2012 | Both significant, Parker larger than Robinson, the two combined larger than Parker alone |
Read the Dubois and Nauges row twice. Their point is that the château already knows how good the wine is, prices it accordingly, and the critic then scores the same wine. Ignore that and the score appears to be worth 3.95% a point. Account for it and the number falls to 1.38%. Most of what looks like critic influence is the producer and the critic reacting to the same bottle.
The 7% figures in the Jones and Storchmann row and the Masset row are not the same claim either. Jones and Storchmann measured a point on the auction market, where the wine is finished, bottled and traded. Masset and colleagues measured a surprise on the en primeur market, where the wine is not yet bottled and the buyer cannot taste it. Their effect was strongest for estates and appellations outside the 1855 classification, and for the best vintages.
Why is the effect near zero below 86 points?
Because in that sample the wines graded below 86 sold for what they would have fetched ungraded. The grade carried no information the price had not already absorbed.
Hadj Ali, Lecocq and Visser report the effect grade by grade rather than as one average, and the shape is the finding. The hypothesis that a grade was worth nothing was accepted for grades below 86 and rejected above it. From there the curve climbs steeply, reaching roughly €14 a bottle at 93.5 points, on wines whose average opening price in that sample was €19.01 for the 2001 vintage and €15.65 for the 2002.
The same convexity shows up across the sample. Top classified châteaux got €3.73 a bottle from being graded, against €2.18 for the rest. By appellation the spread was wider still: Pomerol led at €8.97 on 15 châteaux and Pauillac followed at €6.19 on 19, both significant at the 5% level, while Haut-Médoc sat at €0.42 and Pessac-Léognan at minus €0.04, neither of them significant. Pomerol topping that list is a detail worth holding on to. Hadj Ali, Lecocq and Visser note that it is one of the appellations Parker liked most, and it is the home of Pétrus. A Pauillac first growth such as Château Mouton Rothschild sits in the second band on the same table.
Which critic should you follow?
More than one. That is the closest thing to a settled answer in this literature.
Robert Ashton tested Parker and Jancis Robinson against more than 1,700 red Bordeaux across the 2004 to 2012 vintages. Ashton found both ratings had "a statistically and practically significant impact on prices after controlling for the effects of other known determinants of price". Parker's impact was the larger of the two. But the result that matters for a buyer is the next one: "combining the quality ratings of both experts has a significantly greater impact than Parker's ratings alone". Ashton found that the strength of the result also differed by region, because the amount of other quality information available differs across Bordeaux.
Masset, Weisskopf and Cossutta report something similar from the other direction, after widening the field to 12 critics. The critics reached "a relatively strong consensus on overall wine quality" and parted company on "wines that achieve a surprising level of quality given the vintage, the ranking, or the appellation from which they originate". European critics were "less transparent and in general more severe in their scoring than their American counterparts". Parker and Jean-Marc Quarin came out as the most influential.
Where the critics agree, you learn little the classification would not have told you. Where they disagree is where the price is unsettled.
Do community scores tell you anything the critics do not?
They tell you what a wine did over dozens of bottles rather than one, and they sit systematically lower than the professionals.
Gökçekus and Nottebaum took 120 wines at random from the 4,212 Bordeaux 2005 listings on CellarTracker, covering 2,928 bottles bought by the community members who scored them, an average of 35 notes a wine. Parker had reviewed 107 of the 120. On those, Gökçekus and Nottebaum found the community averaging 91.66 against Parker's 93.24, a gap of 1.58 points that held up statistically. Against Wine Spectator, on 104 wines, the gap was 1.51. Against Stephen Tanzer, on the 61 wines he had covered, the gap was 0.08 and could not be told apart from zero.
The gap is not flat, either. At the bottom the crowd marks wines up, and from there it falls progressively behind: "the community awards 91 to wines that RP scores at 90, 91.5 to wines with a 93 RP score, 92 to wines with 95 RP score, and 95 to wines with 98 RP score."
Then the finding that should interest anyone bidding. Correlating average price paid against each score, they got 0.76 for the median community score and 0.73 for Tanzer, against 0.64 for Wine Spectator and 0.53 for Parker. The crowd tracked what wine actually cost more closely than the most influential critic in the market did.
Community scores carry their own distortions, and the honest reading includes them. The authors point out that community notes are not blind, are written after the expert ratings are already public, and are written by someone who knows what the bottle cost. Cognitive dissonance does the rest: having already bought and opened an expensive bottle makes it harder to score that bottle low. The Wine Gourd found, in an analysis of 1,569,655 public CellarTracker scores, that among scores above 80, 57% sat in the 88 to 92 range, with 88, 89 and 90 over-represented and 93, 94 and 95 under-represented against expectation. It is a crowd with a comfort zone.
How reliable is a single score?
Less reliable than the number of digits implies. The only large replication study on wine judging is uncomfortable reading.
Robert Hodgson ran replicate samples at the California State Fair competition from 2005 to 2008, testing between 65 and 70 judges a year. Each panel of four received a flight of 30 wines with triplicate samples poured from the same bottle hidden inside it. The median spread across a judge's three identical glasses was about 4 points. Only around 10% of judges were, in his phrase, "consistently consistent" to a single medal range. Another 10% "on occasion, scored the same wine Bronze to Gold", a 12-point swing, or worse. Hodgson found that judges hit an identical score all three times about 18% of the time, and that this "usually occurred for wines that were rejected". The pooled standard deviation was 3.6 points, which puts the 95% confidence interval at 14 points, almost the entire range of medals available.
Two caveats, because they matter. Those were competition judges scoring on a converted 80 to 100 medal scale, not the specialist critics who cover fine wine, and Hodgson designed the flights to favour consistency, pouring each triplicate from one bottle on one flight. Parker's en primeur tastings, by contrast, were run in peer-group, single-blind conditions, with the producer's name withheld.
Even at the top, the accuracy case is not clean. Ashenfelter and Jones tested whether buyers pay for expert opinion because it is accurate, and found that public weather data improves the expert's predictions of subsequent prices. Their verdict: "expert opinions are not efficient, in the sense that they can be easily improved, and that these opinions must be demanded, at least in part, for some purpose other than their accuracy."
Are wine scores inflating?
The people inside the system say so. James Laube writes in a Wine Spectator column from May 2013 that "critics are giving more 100-point scores than ever before", noting that his own magazine had awarded two perfect scores to new releases in 30 years, and reviewed more than 17,000 new releases in 2012 without giving a single one, while "one well-established publication gave 100-point scores to more than 50 new releases in 2012".
Laube explains the drift as procedural rather than moral: "It's the non-blind method, especially with a vintner present, that I believe is at the root of today's escalating ratings." Knowing the label opens the door to confirmation bias. "Tasting blind forces the taster to be cautious and critical. In that context, perfection is elusive." Wine Spectator states that its finished wines are reviewed blind and that most barrel tastings are blind, with the exceptions flagged in the report.
There is a second, quieter source of drift: the institutions change. A majority stake in The Wine Advocate was sold to investors from Singapore in December 2012, Parker formally retired from the publication in 2019, and on 22 November 2019 the Michelin Guide became its sole owner. A 96 from 1996 and a 96 from today did not come off the same desk.
What should you do with a score before you bid?
Three things, in this order.
Price the surprise, not the score. A published number is already in the ask by the time you read it. Masset and colleagues measured the effect of the surprise for exactly this reason. If a wine has carried 96 points for a decade, that 96 is not information, it is furniture.
Check where you are on the curve. Below 86, in the one experiment that could test it, a grade bought nothing. Above 86 it climbed steeply, peaking at 93.5 in that sample. The same upgrade is a rounding error on one wine and a repricing on another. Track the whole market on the market index, then price the individual bottle on its own curve.
Never buy on one palate. Ashton's result was that two critics beat the better critic alone. The generalisation is not that critics are wrong, it is that any single reading of a subjective instrument carries the reader's noise as well as the wine's quality.
One number, and not one palate behind it
You get a quality reading that a single taster having a poor morning cannot move. Every wine on this site carries a score that is 40% published critic and 60% CellarTracker community, because the critic is the sharpest single signal in the market and the crowd is the half that, in that Bordeaux 2005 sample, tracked price paid at 0.76 where Parker managed 0.53. Each half corrects the other's known failure: the taster who cannot replicate his own score, and the crowd's habit of piling up at 88 to 90.
Put the wines you care about on a watch list and we will tell you when the 40/60 score and the price stop agreeing. A re-tasted score that lifts before the market notices, a headline number doing all the pricing while the cellar quietly disagrees, a vintage where the critics split and the ask has not settled: those are the moments worth an alert, and they are invisible if you are reading one number on one page.
Start with our methodology if you want to check how the score is built before you trust it, the Bordeaux producer atlas if you want to see which appellations the critic effect has historically landed hardest on, or best Bordeaux vintages if the question is which year to buy rather than which bottle.
