Methodology
How the Which AI Score works
Every number on this site comes from somewhere, and every score can be taken apart. This page explains exactly how the Which AI Score is calculated, what we are confident about, and what we are not.
The six pillars
Every plan is scored 0–100 on six independent dimensions. They are deliberately non-overlapping: “how good are the models” and “how much of them do you get” are different questions, and most comparisons collapse them into one.
- Intelligence
- Quality of the models this plan actually gives you access to.
- Capacity
- How much of it you get — usage limits, where the provider states them.
- Features
- Breadth of included capabilities, weighted for this category.
- Value
- Composite capability relative to what the plan costs each month.
- Reliability
- Incident history from the provider's own status page.
- Sentiment
- Whether people paying for it are currently happy with it.
Weights differ by category
A ranking category is not just a re-sort. It carries three vectors: how much each pillar counts, which benchmarks are relevant, and which features matter. That is why Coding genuinely measures something different from Overall rather than reshuffling the same number.
| Pillar | Overall | Coding | Best value | API |
|---|---|---|---|---|
| Intelligence | 25% | 26% | 12% | 30% |
| Capacity | 18% | 20% | 15% | 12% |
| Features | 14% | 12% | 10% | 16% |
| Value | 12% | 8% | 33% | 17% |
| Reliability | 6% | 9% | 5% | 10% |
| Sentiment | 25% | 25% | 25% | 15% |
Where Intelligence comes from
Model quality comes from Epoch AI’s benchmarking data, ingested nightly: the Epoch Capabilities Index, which combines dozens of benchmarks into one score per model, and for coding DeepSWE, FrontierCode, CursorBench and WebDev Arena. One independent source, one methodology, every lab — rather than launch-day figures each vendor chose to publish. Where a model was run at several effort settings we use its best. Data © Epoch AI, used under CC BY 4.0.
A plan is scored on the best model it officially includes. Models that cost extra on top of the plan do not count, and a model the provider caps tightly — GPT-6 Astra on ChatGPT Plus, available in Codex at roughly 5–45 messages per five hours — counts at a discount. New models take Epoch one to three weeks to measure; until then a plan is scored on its models that have been measured.
Listed, not ranked
If we have no independent measurement for any model a plan includes — usually because the provider does not say which model you get — we do not rank it. It is listed below the ranking with its price and source. A position based on a guess would look exactly as authoritative as a real one, which is the problem.
Missing data lowers confidence, not the score
If a plan-level figure is missing — a provider that states no usage limits, say — that pillar’s weight is redistributed across the pillars we can measure, and the confidence rating drops. Scoring an unknown as zero is the most common failure in ranking sites.
Sentiment and reliability are handled differently, because they are measured per provider and we cannot measure them for everyone. Redistributing around the gap would turn absence into an advantage: a provider whose outages and complaints we can see would be marked down, while one we cannot see would skip the pillar. So each of these counts only when every provider in the category is measured. Until then it is shown on provider pages but struck through in the weighting and left out for everyone.
The same rule applies to individual features. A feature counts only if every ranked plan’s provider says whether it is included; if one is silent on, say, memory, memory is left out for all of them rather than quietly dropped for the provider that said least. Feature support comes from each provider’s own plans page or help centre.
Scores are absolute, not relative
Each pillar maps onto a fixed range rather than being normalized against the current field. A plan’s score therefore moves only when something about that plan moves. If scores were relative, a competitor shipping a new model would silently change everyone else’s history, and month-to-month comparison would be meaningless.
Capacity: shown, not yet scored
Providers do not publish usage limits in comparable units, and a plan gives you several models with different context windows, so any single “plan context window” would be a number we picked. We therefore score Capacity only from multiples a provider states itself: Claude Max 20x is “20 times Pro’s per-session usage”, Google AI Pro is “4x higher usage limits than Free”. Each plan’s stated multiple of its own provider’s free tier is placed on a log scale, from 1x (0) to 200x (100).
That scale only compares fairly within one provider: each multiple is relative to that provider’s own free tier, and free tiers differ. And Perplexity and xAI state no multiples at all. Scoring Capacity for some plans and not others would reward the providers that say least, so Capacity is left out of every score until every plan can be measured. Each plan’s stated usage is still shown next to it. In practice this means a $100–200 tier ranks below its $20 sibling: what you pay extra for is usage we cannot yet compare across providers.
Where plan data comes from
Every price, usage statement and “includes” line is taken from the provider’s own pricing page or help centre, with the link shown next to it. Some providers block automated reading of those pages; for them we record the figures from the page’s indexed text and mark the price unverified until a person has checked it directly. A model that costs extra on top of a plan — Claude’s Fable on usage credits, for instance — does not count towards that plan’s Intelligence score.
How Value is calculated
Value is capability per pound, not cheapness. We take the Intelligence, Capacity and Features composite and divide it by the monthly price, with two stated adjustments: a $5 friction constant (no plan is infinitely good value however cheap) and a 0.75 exponent (the step from $0 to $20 matters more to a buyer than the step from $180 to $200). Free tiers are excluded from the Value ranking, because the question that category answers is which paid plan is worth it.
How API plans are priced
API pricing is per token, so “value” needs a fixed basket. We price a reference workload of 10M input and 2M output tokens per month against each provider’s live token rates. Your workload will differ; the figure exists to make providers comparable, not to predict your bill.
How sentiment works
Sentiment is worth 25% of the score on consumer categories — the second-heaviest pillar, and deliberately so. Whether people paying for a subscription think it is worth it right now is the thing benchmark-led comparisons cannot see, and it moves fastest when a provider tightens limits or ships a regression. The question we ask everywhere is simply: is it worth paying for? — worth it, mixed, or not worth it.
Four channels answer that question:
- App Store reviews — star ratings from the last 30 days of reviews of each provider’s official iOS app, US and UK stores (4–5 stars: worth it, 3: mixed, 1–2: not). Lifetime ratings are useless for comparison — every one sits near 4.8 — so only recent reviews count. They cover the app rather than the subscription alone, and include free users.
- Reader votes on this site.
- Hacker News and Bluesky discussion about paying for each provider, classified by an inspectable word list rather than a model call — a factor worth a quarter of the score should not be a black box.
Two rules keep the pool fair. No channel wins on volume: each contributes its proportions but counts as at most 100 opinions, so a thousand app reviews cannot drown reader votes. A channel counts only if it covers every provider: Hacker News and Bluesky reach our minimum only for the most-discussed providers and run more negative than app reviews, so counting them for some providers and not others would mark those providers down for being talked about. They are shown on provider pages until they cover everyone.
Each category listens to its own audience. App Store reviewers are mostly people using the phone app; they are not the people choosing a coding subscription, who work in a terminal or an editor. So the Coding ranking uses only developer channels — reader votes, Hacker News and Bluesky — and until one of those covers every provider, Coding runs without sentiment. Overall and Best value use every channel.
Sentiment keeps its stated share. When another pillar cannot be scored yet, its weight goes to the other measured pillars — but not to sentiment, which stays at 25% rather than quietly growing into a third of the score.
Safeguards are load-bearing at this weight. The pool is scored with a Wilson lower bound, which pulls small samples toward the middle, and below 30 opinions we display nothing. Voting is limited per provider per day, each public channel needs a minimum number of recent opinions before it counts, and a weak product with adoring users still cannot overtake a strong one outright — there is a test asserting exactly that.
We rate-limit voting with a daily-rotating hash of IP and user-agent. This raises the cost of casual ballot-stuffing. It is not identity enforcement, and we will not pretend otherwise.
We store counts and links, never review or post text. No sentiment is ever seeded or simulated: fabricated community opinion is the one thing this product must never ship.
Local prices, not conversions
Providers set regional prices deliberately rather than tracking exchange rates. Google charges €22.99 and £18.99 where it charges $19.99 — converting the dollar figure would show €17.17, which is not imprecise but simply wrong. So where a provider publishes a real local price we show that price and mark it official.
Where one does not exist we convert at ECB reference rates and label the result an estimate with a ≈. OpenAI, for instance, publishes no local price list at all: it bills in USD and converts at checkout before adding local VAT, so an estimate is the honest representation. A converted figure is never presented as an official price.
How reliability works — and its blind spot
We count incidents from providers’ own status pages over a trailing 90 days, weighting a major incident as three minor ones, on a curve that falls steeply and then flattens.
There is an asymmetry here worth stating plainly: only OpenAI and Anthropic publish a machine-readable incident history. xAI’s status API is behind Cloudflare and Google publishes an HTML dashboard, so those providers have no reliability score and the pillar is redistributed for them. A provider that documents every incident in detail can therefore look worse than one that documents nothing. We use a saturating curve and a small weight to limit the damage, but we cannot fully solve it, and you should read the pillar with that in mind. Severity labels are also each provider’s own judgement, not a common standard.
We do not publish uptime percentages. Status pages report incidents, not uptime, and deriving a precise-looking “99.92%” from an incident list would be invented precision.
Sponsorship and affiliate links
Which AI takes money from sponsors and affiliate links. Neither can affect a ranking. The scoring engine's input type contains no commercial fields at all — there is no property on it for an affiliate rate or a sponsorship tier — so paid placement has no path into a score even by accident. Sponsored units always display the product's true, unmodified rank.
Manual corrections
Automated and curated data is sometimes wrong. Editors can override a field, but never by editing the original record: the ingested value is stored permanently and the correction is a separate layer applied on read, with a required reason. The resolution order is raw → normalized → override → effective value.
History starts when we started
Score history begins on the day our daily ranking run first went live, 31 August 2026. We do not backfill earlier history or simulate votes: a trend chart only shows what we actually recorded.
What this score cannot tell you
No ranking is fully objective. Choosing that coding benchmarks matter more than image generation for a developer is an editorial judgement, and a different weighting would produce a different order. We publish the weights instead of pretending they are neutral.
Benchmarks also measure what is easy to measure. None of them capture whether a model is pleasant to work with, how it handles your particular domain, or whether its refusals will annoy you. A score of 93.8 is not a promise; it is a summary of the things we can actually observe.