Skip to content

How to judge wine quality: a repeatable framework

Updated

Quality gets judged in three passes, and the order is not optional: rule out faults, score four markers, then place the wine in time. Price and label come last, because both change what tasters report. So here is how to judge wine quality the way the two grids used in professional training do it. Pass one is condition. The WSET Level 3 tasting card asks whether the nose is "clean" or "unclean (faulty?)" before it asks anything else, and the Court of Master Sommeliers lists "Minor Presence of Fault(s)" at the head of its nose section. Pass two is the four markers the grids converge on: intensity, complexity, balance and length. The Court rates complexity "Simple" to "High" and finish "Short" to "Long"; WSET rates finish "short" to "long" and flavour intensity "light" to "pronounced". Pass three is time: WSET closes on a readiness line running from "too young" to "too old". Only then do you look at the price, the label or anyone else's score.

What are the four markers of wine quality?

Intensity, complexity, balance and length. Both grids score them in nearly the same words, which is the closest thing the trade has to an agreed definition of quality.

Marker Court of Master Sommeliers, Americas (2024) WSET Level 3 tasting card The question underneath
Intensity Aromatic Intensity: "Delicate, Medium Minus, Medium, Medium Plus, Powerful" Nose intensity: "light – medium(-) – medium – medium(+) – pronounced" Does the wine project, or do you have to go looking for it?
Complexity Complexity: "Simple, Medium Minus, Medium, Medium Plus, High" Aroma and flavour characteristics, plus development How many distinct things are in there, and do they come from more than the fruit?
Balance Balance: "Yes or No: What main characteristic(s)/element(s) dominate the wine?" Other observations: texture, balance Does one element stick out?
Length Length of Finish: "Short, Medium Minus, Medium, Medium Plus, Long" Finish: "short – medium(-) – medium – medium(+) – long" How long after swallowing does the flavour hold its shape?

The WSET card has no line called complexity. It arrives at the same judgement by asking for aroma and flavour characteristics and then a development verdict of "youthful – developing – fully developed – tired/past its best". Asking a taster to list what they smell is a complexity test run as an inventory.

Balance is the marker people fudge, so both grids force it. WSET lists the prompts in its lexicon: acid is "sour, refreshing or flabby, heavy?", red tannin is "well-integrated, soft or harsh, bitter?", alcohol is "delicate, light or hot burning", fruit is "hollow, thin, neutral or juicy, fruit-driven?" and the overall call is "elegant, harmonious or shapeless, clumsy". The Court compresses all of that into one line. If you can name the element that dominates, the wine is not balanced.

Which faults disqualify a bottle, and at what concentration?

A fault is a dose question rather than a yes or no question, and the doses are published. The Australian Wine Research Institute lists a sensory threshold for each of the common ones, and they sit orders of magnitude apart.

Fault Compound AWRI sensory threshold What it smells of
Cork taint TCA "the aroma threshold of TCA in a Pinot Noir wine as 1.4 ng/L" "distinct musty, mouldy aroma"
Brettanomyces 4-ethylphenol "368 µg/L in a neutral red wine" "'Band aid®', 'medicinal' or 'pharmaceutical' character"
Brettanomyces 4-ethylguaiacol "158 µg/L in Australian wine styles" "'clove', 'spicy' or 'smoky' aroma"
Volatile acidity Acetic acid "as low as 0.1 – 0.125 g/L" Vinegar
Volatile acidity Ethyl acetate 12 mg/L Nail polish remover
Oxidation Acetaldehyde "ranges from 100-125 mg/L" "over-ripe bruised apples", "sherry" and "nut-like" characters
Reduction Hydrogen sulfide "about 1 – 2 µg/L" "rotten egg gas"
Reduction Methyl mercaptan "0.02 – 2.0 µg/L" "rotten eggs" and "cabbage"

Cork taint and oxidation sit at opposite ends of that range. TCA registers at 1.4 nanograms per litre while acetaldehyde needs 100 milligrams, so the two thresholds sit roughly seventy million times apart. That is why one corked bottle in a case is obvious to everyone at the table within seconds, while mild oxidation gets argued about for twenty minutes.

The dose point also decides what counts as a fault at all. The Court lists "Struck Match" and "Barnyard" among its faults, both of which read as style at trace levels and as spoilage above them, and the line is headed "Minor Presence of Fault(s)" rather than "Faults". Musty is the one worth treating as absolute: see what mould on a cork does and does not tell you. And if you are buying rather than opening, the odds of a fault are partly readable from the outside, which is what fill level and ullage are for.

Why do two tasters score the same wine differently?

Because a single score is a point estimate with a wide interval around it, and that has been measured. Robert Hodgson tested judges at the California State Fair wine competition from 2005 to 2008 by hiding triplicate pours from the same bottle inside flights of 30 wines, with between 65 and 70 judges tested each year.

Hodgson reports his results in the Journal of Wine Economics in 2008: "About 10 percent of the judges were able to replicate their score within a single medal group. Another 10 percent, on occasion, scored the same wine Bronze to Gold." The median pooled standard deviation across all judges was 3.6 points, and "the corresponding 95 percent confidence interval is 14 points, which includes almost the entire range of medals possible". Judges hit an identical score about 18 percent of the time, usually on wines they rejected, which is Hodgson's basis for saying that "judges tend to be more consistent in what they don't like than what they do". Consistency did not persist year to year either: across the 26 judges who worked both 2005 and 2006, the correlation was -0.01.

The rule that falls out of a 14-point interval is simple. Treat a two or three point gap between critics as noise, and treat agreement across many tasters as the actual signal.

Does the label change what you taste?

Yes, and the effect is large enough to invalidate any judgement made with the bottle in view. Having dyed a white wine red with an odourless dye, Morrot, Brochet and Dubourdieu report in Brain and Language that it "was olfactory described as a red wine by a panel of 54 tasters. Hence, because of the visual information, the tasters discounted the olfactory information."

Price does the same work. Plassmann and colleagues, writing in PNAS in 2008, scanned 20 subjects who tasted the same wines at two stated prices, one pair at $5 and $45, another at $10 and $90. Plassmann reports that "increasing the price of a wine increases subjective reports of flavor pleasantness as well as blood-oxygen-level-dependent activity in medial orbitofrontal cortex". Take the price away and the relationship inverts: across 6,175 double-blind tasting observations, from 506 participants and 523 wines, Goldstein and co-authors found "the correlation between price and overall rating is small and negative". For the 12 percent of tasters who had some wine training, they report "indications of a non-negative relationship", while noting that the expert price coefficient is not statistically strong.

Which is the case for the professional protocol rather than against it. Wine Spectator states that all reviews are done 'blind', so the taster does not know the producer or the price, and for international competitions the OIV states the same point in stronger terms: "Absolute anonymity shall be a fundamental principle of competitions."

What does a 95-point score actually mean?

It depends whose 95 it is, and the published scales do not line up.

Band Wine Spectator CellarTracker community scale OIV competition medals
95-100 "Classic: a great wine" 98-100 "A+"; 94-97 "A" Grand gold from 92 points
90-94 "Outstanding: a wine of superior character and style" 90-93 "A-"
85-89 "Very good: a wine with special qualities" 86-89 "B+" Gold from 85 points
80-84 "Good: a solid, well-made wine" 80-85 "B" Silver from 82, bronze from 80
75-79 "Mediocre: a drinkable wine that may have minor flaws" 70-79 "C"
50-74 "Not recommended" 50-69 "D", avoid

An OIV gold medal starts at 85 points, two bands below what Wine Spectator calls classic. It is also a rank rather than an absolute: the OIV specifies a cap on awards so that "the sum of all the medals awarded to the samples must not exceed 30 % of the total of samples presented at the competition", with juries of seven tasters, five at minimum.

What the number is meant to contain also varies. Wine Spectator says its taster "judges the wine's structure, flavors and typicity (how well it reflects its grape variety, region and vintage)" and describes the result as "an attempt to quantify a judgment that includes both subjective and objective components". Typicity is doing quiet work there: a wine can be balanced, complex and long and still score poorly for tasting nothing like its appellation. That is one reason the same vintage reads differently across critics, and why we publish per-vintage detail on pages like Bordeaux 2016 rather than a single house number.

What does the crowd know that a critic doesn't?

When it drank well, and how often the bottle disappointed. Those are sample-size questions, and no individual critic can answer them.

The published comparison is Kopsacheilis and colleagues in the Journal of Wine Economics in 2024, matching amateur platform ratings for Bordeaux against professional reviews. Kopsacheilis reports: "Vivino ratings correlate substantially with those of professional critics, but these correlations are smaller than those among professional critics." Correlations ran from 0.50 against the Wine Advocate and 0.48 against Jeff Leve down to 0.16 against Decanter, averaging 0.40. The authors attribute the gap to scope rather than competence: "Whereas amateurs focus on immediate pleasure, professionals gauge the wine's potential once it has matured."

That is the argument for using both, and for weighting them differently. Critics taste young wine and forecast. Community tasters, on CellarTracker's published 50-to-100 scale where 94-97 is "A" and 86-89 is "B+", record what a bottle did on a date, at an age, after real storage. Our score is 40 percent critic and 60 percent CellarTracker for that reason, and both components stay visible so you can see when they disagree. The full weighting is in our 40/60 critic and CellarTracker methodology.

How do you judge quality on a wine you cannot taste?

You judge the evidence instead of the wine, in the same three passes. Condition first, then consensus, then timing.

Condition is the fill level, the storage history and the label state, which is where an auction lot either earns or loses its estimate. Consensus is the two score components together plus the number of notes behind them, because Hodgson's 14-point interval applies to a single taster and narrows as tasters are added. Timing is the drink window: the WSET card ends on whether a wine is "too young", can drink now, but has potential for ageing, drink now: not suitable for ageing or further ageing or "too old", and a bottle bought at the wrong end of that line is a quality problem you paid for.

Run it on a wine with deep coverage, Pétrus or Romanée-Conti, and the consensus is dense enough to trust. Run it on something with nine notes and the honest answer is that you are guessing. Drink now shows which of the wines we cover are in their window today, and the market index shows what the market thinks of the same wines with quality stripped out.

See both scores before you pay for one

The disagreement between a critic and a room full of people who actually opened the bottle is the most useful thing on a wine page, and most sites average it away.

Every wine here carries the critic score and the CellarTracker community score side by side, with the note count under each, so you can tell a 95 backed by four hundred notes from a 95 backed by four. Compare them before you bid. When the crowd sits well below the critic on a mature vintage, you are usually looking at a wine that was scored on promise and never delivered, and that is a discount you should be taking rather than paying.

Wines we track under this

Reference cheat sheets

Reference Cheat Sheets

1855, Premier vs Grand Cru, Cru Bourgeois, and the château map, on two pages.