Methodology

How We Score

Most travel sites tell you the best time to visit a place. That question has no answer, because it depends entirely on what you plan to do when you get there.

Split in July scores 100 for swimming and close to nothing for hiking. Rio's beaches peak in January; its walking trails peak in July. A single number per destination cannot be right about both, so we do not publish one.

Every destination is scored once per activity, for every month of the year. 139 destinations, 9 activities, 415 graded windows. Here is how each number is built and, more usefully, where it goes wrong.

Two halves, doing different jobs

A score has a human half and a computed half, and they answer different questions.

A person decides which months are prime. That judgement takes in things no weather model can see: when a place is alive rather than merely warm, when the ferries run, when the vines have leaves on them, when a town is celebrating something.

A climate model decides the order within each band. Ten years of daily weather — temperature, rainfall, wind, cloud, snow — reduced to monthly normals, then scored against what each activity actually needs.

The human sets the tier. The model sets the ranking inside it.

80–100Prime — the months the place is known for
55–79Good — a real trade, usually crowds or cost
30–54Possible — it works, with caveats you should read
0–29Not the season — shut, snowbound, flooded or unpleasant

So an April at 95 and a December at 82 are both prime, and the gap between them is the model's opinion about the weather. A month at 34 is never promoted by good conditions, and a month at 96 is never demoted by a cold snap in the data.

One destination, two answers

JFMAMJJASOND
Rio de Janeiro — swimming
JFMAMJJASOND
Rio de Janeiro — hiking

January feels like 36°C and rains eighteen days a month. Wonderful for the beach; punishing on Corcovado. July is 26°C and rains seven. The two answers are six months apart, and both are correct.

How to read the crowd bars

JFMAMJJASOND
Yellowstone — counted visitors, National Park Service
JFMAMJJASOND
Joshua Tree — counted visitors, National Park Service

Taller is busier. Both of these are people counted through a gate, published monthly by the National Park Service — not an estimate.

July is Yellowstone’s busiest month of the year and Joshua Tree’s emptiest. The height of a bar is never a statement about July. It is a statement about that place’s own year.

Which is why the scale never crosses between destinations. Each row is drawn 0 to 100 against its own quietest and busiest month, so Joshua Tree’s 100 and Yellowstone’s 100 are not the same number of people — and nothing on this site says they are.

The third thing, which is not a score

Two more facts about a month refuse to fit into a number, and we stopped trying to make them.

Closure — is the place open

A mountain refuge with the shutters up, a ferry that stops in October, a lift that does not turn, a park road under snow until late May. This is not bad weather; it is a closed door, and no amount of sunshine compensates for it. So closure is the one thing that does multiply the score, from 1 (everything running) down toward 0 (shut). 98 of the 139 destinations carry one for at least one month.

It took us two attempts to get this right. The first version quietly meant “high season is over” — a Greek island in November was marked down because the beach clubs had closed, when the island itself was perfectly open and rather nice. That is a crowd fact wearing a closure costume, and it was pushing real scores around. Closure now means one thing only: can you get in and is anything running.

Hazards — what could go wrong, which never touches the score

Hurricane season. Monsoon. Wildfire smoke. Avalanche risk. Polar winter. 116 warnings across 95 destinations, in 14 kinds. None of them move a single point.

That is deliberate, and it is the decision we are least likely to reverse. Folding risk into the number would turn September in the Caribbean into a mediocre month — a 62, say — which tells you nothing useful. September in the Caribbean is not mediocre. It is usually excellent and occasionally catastrophic, and those are different claims that average into a lie. So the score says what the weather is normally like, and the warning sits beside it saying what the tail risk is, and you decide. You are the one who knows whether you can move your flights.

The two tones, as they appear on a card

Hurricane season

Peak risk August to October; storms can close airports for days and most travel insurance excludes named storms once one is forecast. Aug · Sep · Oct

Wet season

Afternoon downpours most days; mornings usually clear, and the landscape is at its greenest. May · Jun · Oct · Nov

Severe means it can end your trip. Notable means it will shape your days. Both carry their label in words, never colour alone.

Every hazard names its months, because “hurricane season” spanning June to November is not a useful thing to be told when you are looking at June, which is nearly always fine.

How we found this out. For a long stretch the site computed both of these and then applied neither — the numbers were built, stored, and silently dropped before they reached a score. Hazard was also a single unlabelled number doing eight different jobs at once, so a wildfire and a monsoon and a polar night were the same value with no way to tell them apart or say which months they applied to.

Nothing in the output looked obviously wrong, which is exactly why it went unnoticed. It surfaced only when we checked the two halves against each other destination by destination.

Reading the filter

Pick an activity and a month. Results are grouped by tier rather than cut off at a number, because the prime group is the shortlist — typically four to twelve places, occasionally none.

Selecting two activities scores the months where both work, using a harmonic mean so one weak half drags the month down instead of averaging out. A beach month that is unwalkable in the heat does not come back as "fine".

The crowd slider is yours, not ours. At “not at all” you get the conventional answer, ranked on conditions. Move it and quietness is mixed in, which reorders the list — a beach at 88 in July and 85 in September flips to September well before the slider reaches halfway. Destinations with no crowd data are moved into their own group rather than ranked alongside, because an unmeasured place is not a quiet one.

When nothing is prime, we say so and tell you when it would be. Roughly one filter in five returns no prime result — nowhere skis well in July in the northern hemisphere, and no aurora is visible anywhere in June. An empty answer is usually the honest one.

Checking the work

The two halves are built separately, which makes them a partial test of each other. That is worth saying carefully, because it is the claim on this page most easily overstated: the windows were written by the same hands that built the site, not graded by an outside panel, so when a window is corrected the agreement figure moves with it. A high figure over a small number of cases is the weakest line in the table below, not the strongest. If a person marks September prime for hiking in Tuscany and a climate model built from ten years of weather lands on the same month without being told, that agreement means something.

We measure it two ways: how often the model's best month falls inside a human-graded prime window, and how closely the twelve-month curves track each other.

ActivityCasesAgreementCorrelation
Swimming5190%0.95
Skiing2383%0.95
Lakes & Coastline8178%0.91
Wine & Vineyards not counted2095%0.89
Northern Lights11100%0.88
Hiking10274%0.86
City & Culture5480%0.84
All, excluding wine32280%0.88

Wine is excluded from that total on purpose. Its graded windows were rewritten to describe the growing season — which is the same thing the model measures — so the two halves are no longer independent there and their agreement is arithmetic rather than evidence. Leaving it in would flatter the figure by 1 in a thousand — less than the table above shows — which is not worth the dishonesty.

The model was built and tuned against the first 36 destinations. It has since been applied to 103 more, and the correlation did not fall. A model quietly fitted to its own examples does not behave that way.

The honest caveat on that claim. This is not a blind holdout. The human grades for those 103 destinations were written by the same person who built the model — without consulting its output, but not in ignorance of how it works. A genuinely independent test would need a second grader who had never seen it.

And the bands were chosen against these grades. Where a scoring band had a defensible range — what counts as comfortable, how much rain is too much — the value was picked by testing options against the hand grades. That is the model being fitted to the answer, in the opposite direction. We hold the line where we can: a band that scored better was rejected once as noise, and the last change to the coastline band was validated by re-picking it on 74 destinations and scoring the 75th. But it is selection on the target, and it belongs in this box.

We think the agreement is real. We would not stake the site on it being as real as the number looks.

What we get wrong

Every number on this site comes from a model, and models are wrong in specific, knowable ways. These are ours.

Weather data cannot see into valleys

The climate record works on a grid roughly 25km across. In ordinary terrain that is fine. In the Khumbu it averages Everest into the valley floor and reports 245mm of rain in January, where the real figure is nearer 25. We dropped that destination rather than publish it. Mountains with severe relief are where to distrust us first.

Crowds are measured for two-thirds of the list, and guessed for none of it

85 of the 139 destinations carry a measured monthly crowd curve. The other 54 show nothing at all — not a default, not an estimate. Where we could not verify a figure we would rather leave the space empty, because a crowd number nobody has checked is worse than no crowd number.

A rain day is a blunt instrument

Any day with a millimetre of rain counts as wet. That treats a twenty-minute tropical downpour the same as all-day drizzle, and it is the weakest input we have.

We have tried to fix it and not yet succeeded. Total rainfall divided by wet days separates the two cases in principle — convective storms run 12–20mm per wet day, frontal drizzle 3–6 — and discounting the showery months looked like a clear improvement until we noticed that most of our cached records had no rainfall totals at all, the arithmetic was producing NaN, and the nonsense correlated with the hand grades better than the real model did. On the records that do carry rainfall, the effect is indistinguishable from noise. The idea may still be right; the evidence for it was not.

The coastline model does not know the water is the point

Rio's beaches are a clean failure. The model ranks its coastline months almost exactly backwards — it likes July, and July is winter. The reason is that the comfort term scores anything above 34°C as unbearable, and Rio's summer runs 34–36°C, while rain days counts summer rain as a pure negative when summer rain is simply what Rio's beach season comes with. Both terms point the same wrong way.

We assumed for some time that this was a missing sea-temperature input. It is not: we tested thirty-two ways of adding water temperature, and the set median moved by 0.00. Naming the wrong cause confidently for weeks is its own kind of error, and the fix is still open. The damage is contained — the hand grades set the tier, so Rio's beach months are still graded prime and its winter months still are not; what suffers is the ordering within the prime band.

Wine will not score better than it does

0.66 is a ceiling, not a defect. The model says conditions are good across the whole growing season, which is true. A person says harvest is special, which is also true and is not a fact about weather. We leave the gap rather than fake it.

Bugs we found by checking, and what they cost

Each was found by the two halves disagreeing. That is what the checking is for, and it is why we publish the disagreements rather than tune them away.

How the crowd figures are built

Two governments publish exact monthly visitor counts for free — the US National Park Service, and Eurostat for EU regions. They cover a fraction of this list. Everywhere else we use Wikipedia pageviews as a proxy for travel interest, divided by Wikipedia's own total monthly traffic to remove the fact that people read more encyclopedias in winter.

The proxy is not trusted by default. It has to pass three checks, and where it fails, the destination is published with no crowd data.

27 of them are checked directly against official government statistics, at a median correlation of 0.894. Where an official series disagrees with the proxy by more than two months we discard the comparison rather than report it — a correlation that needs the calendar shifted half a year is not a prediction, however good the number looks. And where the official counts are clean and simply say the proxy is wrong, the official counts win: the three checks above can only look inside the proxy, and it is indefensible to hold a measured answer and publish a guess instead.

The thing we did not expect

The proxy works when the Wikipedia article is a place people sleep, and fails when it is a subject people study. “Serengeti National Park” is read by schoolchildren and documentary viewers in numbers that swamp its actual visitors. “Burgundy” is read about wine. Swapping those articles for the town you actually stay in — Beaune for Burgundy, Moshi for Kilimanjaro, Logroño for Rioja — fixed 19 destinations across two rounds and broke none.

Two caveats we owe you. The plausibility check uses our own hand-graded months as the filter, so it is not fully independent of the rest of the site — the correlations against government statistics are, and they are the part to trust most.

And the crowd scale is within a destination, never between them. 100 is that place's own busiest month and 0 its quietest. It does not say Santorini is busier than Tahoe, and it cannot.

Where there is no crowd data at all

Safari is the honest failure. Serengeti, Masai Mara, the Okavango and Kruger have no usable monthly crowd signal from any source we tried, including their base towns. That is most of our Africa coverage. It costs less than it looks — safari camps have fixed bed capacity and sell out months ahead, so the real question there is when the animals are there, which the wildlife scores already answer — but it is a gap, and we would rather name it.

Where the numbers come from

Daily weather 2015–2024 from the Open-Meteo historical archive (ERA5 reanalysis), sea-surface temperature from its marine archive, 2018–2024. Daylight is calculated from latitude rather than downloaded. Geomagnetic activity for the aurora model follows the Russell–McPherron effect, which is why our northern-lights scores peak at the equinoxes rather than in midwinter — a disagreement with most published advice that we are confident about.

Crowd figures from US National Park Service monthly visitation, Eurostat monthly nights spent at tourist accommodation, and Wikimedia pageview statistics, 2016–2025 excluding 2020 and 2021.

Everything else — which months are prime, what is shut, what to avoid, what the walking is like — is written by hand, destination by destination.