modelsentiment

How it works

How the numbers are made

What is collected, what is counted, how opinions are scored, and what the figures can and cannot tell.

  1. 1Collect22 subreddits, read over RSS
  2. 2Matchcode finds every model an item names
  3. 3Scoreone opinion per model, 1–5
  4. 4Index7 days pooled, 0–100, with its ±

How to read the board

Index

The mean opinion of a random sample of the items that name a model, among those that hold an opinion, weighted by 1 / the item's sample share, on a 0–100 scale: 0 = all opinions very negative, 50 = neutral, 100 = all very positive. The headline figure pools the last 7 UTC days (fewer while collection is younger, as its label says) and all subreddits, so a shift in the subreddit mix moves it; the model page shows the same figure per subreddit. The daily values are detail. A model's index holds only the opinions that name it (see Models and series).

n and ±

n is the effective number of opinions behind an index, (Σw)² / Σw², which is below the count when the weights differ; ± a 95% interval, cluster-robust by thread: opinions in one thread tend to agree, so each thread counts as one unit of the variance, and the interval widens when there are few threads. The ± covers the sample and that clustering only, not the scorer: when the question set changes, the whole history is rescored under the new one, and the switch of the second stage from Claude Haiku to DeepSeek V4.1 Flash (2026-09-27) moved series by up to 6.5 points. Fewer than 10 opinions: no index, only the n. Fewer than 30, or from fewer than 20 threads: a hollow dot, thin. Beside the ±, in words: positive or negative when the whole interval is above or below 50, otherwise no clear lean.

Rank

The board ranks models by index, highest first; a series is never ranked. Two models share a rank only with the same index as shown. A rank is not a verdict: two models a point or two apart are often within each other's 95% interval, which each row's card shows.

Colour

Sentiment runs cold to warm: grey around 50, shading towards blue below and orange above. Mentions bars are grey; lab logos keep their own colours. The status line at the foot of this page uses colour for operational state (amber, red).

Mentions

Posts and comments that name the model, counted in code on every fetched item. Bot and moderator comments are dropped before counting. On the board's strip, a larger dot is a model more talked about.

Change

The 7-day index minus the previous 7 days' index, shown only when both windows are full (not thin) and collection covers all of the previous 7 days. Its ± is √(±₁² + ±₂²) of the two windows; a change no larger than its ± is marked within noise.

Models and series

A model is one entry of the model list ("Claude Opus 5.5"); a series ("Claude Opus") groups models for navigation and matching and has no index, change or rank of its own. An opinion counts only for a model the author named, in the text or, for a comment, in the thread title; an opinion that names only the series ("Opus is lazy") counts for no model, and no model is guessed from the date. A generation name that is not itself a listed model ("GPT-6", "MiMo V2.6") does not count either. The family series ("Claude", "GPT / ChatGPT", "Gemini") are brand mentions: items naming the brand with no series; a model named with no series ("GPT-4o") belongs to them. Gemini has two series, Gemini Pro and Gemini Flash, as Claude has Opus and Sonnet.

A sample, not a census

Every number carries its n, and the source is the feeds' latest items only.

Collection

22 subreddits, two RSS feeds each: new posts and the latest comments (the quietest subreddits share one feed per kind), one request at a time, at most one a minute (one every three minutes for a day after Reddit asks to slow down). Each feed is fetched as often as its pace needs, from every 10 minutes to every 6 hours (posts up to 12). A feed returns at most 100 items; when all 100 are new, the next 100 are fetched too, so the busiest comment feeds are covered almost in full. What remains is a sample, not a census. Nothing before a subreddit's first fetch counts; RSS has no archive. Comments by bot and moderator accounts (AutoModerator, mod teams, reminder bots) are dropped on arrival.

Mentions

Counted in code on every item: an alias of a series or a model in the item's own text (a post's title and body, a comment's body), with coding-tool names (Claude Code, Codex, Cursor, Pragma, Claude Cowork) removed first; "Codex" right after a version number ("GPT-5.3 Codex") is a model and stays. Series marked mentions only are counted but not yet scored: they join the judges' question with the next question set. A family series (claude, gpt, gemini) counts only when no specific series of the same lab is named. Series totals count distinct items, so an item naming two of its models counts once for the series.

Models and series

The models are found in code. Which models exist comes from OpenRouter's public model list and, for models it no longer carries, Epoch AI's, both read daily, and for open-weight models neither has, the labs' own Hugging Face pages, where a model counts once 5 items in 7 days name it; batch, free and dated snapshots count as one model, and a Pro/Flash/Max variant is its own model only once it has 5 mentions. A series the registry lacks (the name before a listed model's number, "Llama" in "Llama 3.3") is added by itself once 10 items in 7 days name one of its models, and is scored from then on. A mention goes to the most specific model named ("Opus 5.5" is never "Opus 5"). Each model an item names gets its own opinion; a comment in a thread whose title names a model of a series the comment names no model of is asked about the title's model. An item that names only the series ("Opus") is a mention of the series and gets no opinion: which model was meant is not guessed. A generation named after its models ("GPT-6", "MiMo V2.6") with no entry of its own is not a model here. Version-like names that are not in the list ("Sonnet 5.5") are counted separately, on Unverified.

Sentiment

Every item that names a model (in its text or thread title) is scored while that fits the model budget of $40 a month. Should the volume outgrow it, a fixed-seed random sample is scored instead: each series gets its own share, recomputed daily from the last 7 days; the series with the fewest items stay in full, and every other series gets the same expected number of scored items. An item takes the largest share among the series it names, and its opinion is weighted by 1 / that share (1 when everything is scored), so a heavy series' index is not tilted toward the items that also name a thin series.

Code lists every model an item names; Jev (TypeSafe, a pinned model) answers for each one whether the author judges it. Those it passes go to DeepSeek V4.1 Flash one model at a time, which decides whether the author states their own opinion of that model and scores it 1–5, so a comparison gives an opinion for each model it judges. A model that is only the yardstick for another ("better than Sol") gets no opinion from it, and neither does an image, video or audio generator sharing a model's name (Grok Imagine). Complaints about the app, limits or plans do not count.

The index maps the mean of those scores linearly to 0–100, weighted as above. Day = the item's publish time, UTC. The daily chart also draws the 7-day pooled index up to each day (the headline's estimator over that day and the six before it), with no value below 10 opinions. A dashed vertical line marks the model's release.

Benchmarks

Each model page lists outside measurements beside Reddit's index, read once a day: Arena's blind head-to-head votes (the style-controlled text arena and Code Arena, from the LMArena leaderboard dataset, CC BY 4.0), the Artificial Analysis indices as OpenRouter passes them on, Epoch AI's Capabilities Index (CC BY), and OpenRouter's prices and limits. A source that ran a model at several reasoning efforts shows each; the one with the most votes stands for the model. None of them feeds the index.

Limits

  • The numbers come from the feeds' latest items; every one that names a model is scored while that fits the budget.
  • Not every counted opinion is one. On 2026-09-28, 160 opinions the pipeline counted were drawn at random from the kept text and judged by two independent reviewers (Claude models that did not see the pipeline's answer; a third settled disagreements). The reviewers could not tell for 3, and of the other 157, 142 were the author's own opinion of that model: 90 % (95 % interval 85–94 %).
  • Post and comment text is kept 90 days (rumor posts longer) and never shown here; thread titles are shown while kept.

Contact

Questions, corrections or a model we miss: [email protected].

Right now

Last Reddit fetch just nowScoring spend $6.10 of $40.00 this month54 waiting to be scored (oldest 34m)Scoring every item that names a model