How it works
How the numbers are made
What is collected, what is counted, how opinions are scored, and what the figures can and cannot tell.
- 1Collect22 subreddits, read over RSS
- 2Matchcode finds every model an item names
- 3Scoreone opinion per model, 1–5
- 4Index7 days pooled, 0–100, with its ±
How to read the board
Index
n and ±
Rank
Colour
Mentions
Change
Models and series
A sample, not a census
Collection
22 subreddits, two RSS feeds each: new posts and the latest comments (the quietest subreddits share one feed per kind), one request at a time, at most one a minute (one every three minutes for a day after Reddit asks to slow down). Each feed is fetched as often as its pace needs, from every 10 minutes to every 6 hours (posts up to 12). A feed returns at most 100 items; when all 100 are new, the next 100 are fetched too, so the busiest comment feeds are covered almost in full. What remains is a sample, not a census. Nothing before a subreddit's first fetch counts; RSS has no archive. Comments by bot and moderator accounts (AutoModerator, mod teams, reminder bots) are dropped on arrival.
Mentions
Counted in code on every item: an alias of a series or a model in the item's own text (a post's title and body, a comment's body), with coding-tool names (Claude Code, Codex, Cursor, Pragma, Claude Cowork) removed first; "Codex" right after a version number ("GPT-5.3 Codex") is a model and stays. Series marked mentions only are counted but not yet scored: they join the judges' question with the next question set. A family series (claude, gpt, gemini) counts only when no specific series of the same lab is named. Series totals count distinct items, so an item naming two of its models counts once for the series.
Models and series
The models are found in code. Which models exist comes from OpenRouter's public model list and, for models it no longer carries, Epoch AI's, both read daily, and for open-weight models neither has, the labs' own Hugging Face pages, where a model counts once 5 items in 7 days name it; batch, free and dated snapshots count as one model, and a Pro/Flash/Max variant is its own model only once it has 5 mentions. A series the registry lacks (the name before a listed model's number, "Llama" in "Llama 3.3") is added by itself once 10 items in 7 days name one of its models, and is scored from then on. A mention goes to the most specific model named ("Opus 5.5" is never "Opus 5"). Each model an item names gets its own opinion; a comment in a thread whose title names a model of a series the comment names no model of is asked about the title's model. An item that names only the series ("Opus") is a mention of the series and gets no opinion: which model was meant is not guessed. A generation named after its models ("GPT-6", "MiMo V2.6") with no entry of its own is not a model here. Version-like names that are not in the list ("Sonnet 5.5") are counted separately, on Unverified.
Sentiment
Every item that names a model (in its text or thread title) is scored while that fits the model budget of $40 a month. Should the volume outgrow it, a fixed-seed random sample is scored instead: each series gets its own share, recomputed daily from the last 7 days; the series with the fewest items stay in full, and every other series gets the same expected number of scored items. An item takes the largest share among the series it names, and its opinion is weighted by 1 / that share (1 when everything is scored), so a heavy series' index is not tilted toward the items that also name a thin series.
Code lists every model an item names; Jev (TypeSafe, a pinned model) answers for each one whether the author judges it. Those it passes go to DeepSeek V4.1 Flash one model at a time, which decides whether the author states their own opinion of that model and scores it 1–5, so a comparison gives an opinion for each model it judges. A model that is only the yardstick for another ("better than Sol") gets no opinion from it, and neither does an image, video or audio generator sharing a model's name (Grok Imagine). Complaints about the app, limits or plans do not count.
The index maps the mean of those scores linearly to 0–100, weighted as above. Day = the item's publish time, UTC. The daily chart also draws the 7-day pooled index up to each day (the headline's estimator over that day and the six before it), with no value below 10 opinions. A dashed vertical line marks the model's release.
Benchmarks
Each model page lists outside measurements beside Reddit's index, read once a day: Arena's blind head-to-head votes (the style-controlled text arena and Code Arena, from the LMArena leaderboard dataset, CC BY 4.0), the Artificial Analysis indices as OpenRouter passes them on, Epoch AI's Capabilities Index (CC BY), and OpenRouter's prices and limits. A source that ran a model at several reasoning efforts shows each; the one with the most votes stands for the model. None of them feeds the index.
Limits
- The numbers come from the feeds' latest items; every one that names a model is scored while that fits the budget.
- Not every counted opinion is one. On 2026-09-28, 160 opinions the pipeline counted were drawn at random from the kept text and judged by two independent reviewers (Claude models that did not see the pipeline's answer; a third settled disagreements). The reviewers could not tell for 3, and of the other 157, 142 were the author's own opinion of that model: 90 % (95 % interval 85–94 %).
- Post and comment text is kept 90 days (rumor posts longer) and never shown here; thread titles are shown while kept.
Contact
Questions, corrections or a model we miss: [email protected].
Right now
Last Reddit fetch just nowScoring spend $6.10 of $40.00 this month54 waiting to be scored (oldest 34m)Scoring every item that names a model