Changelog
Notable updates to Sophon - new features, fixes, and improvements, newest first.
August 25, 2026
- AddedHead-to-head pages for the models people actually compare: /compare/claude-opus-4-8-vs-gpt-5.5, /compare/gpt-5-vs-gpt-5-mini and 433 more. Comparing two models used to live behind a temporary address that could not be linked or found, so the single most common question in the field - "is A better than B?" - had no page to land on. Each pair now has its own, with the benchmarks both models report, price per million tokens, context window, release dates, and a short plain-language answer on who leads and where. The most-compared pairs are listed on /compare.
- FixedA model with no published price no longer reads "$0.00" in a comparison. 195 of 454 price records store a zero to mean "not reported", which made a paid model look free; those cells now show a dash.
- ImprovedSearch engines are now pointed at the pages that carry Sophon's own work rather than at everything. Roughly 130,000 of the 205,000 pages are a name attached to a single paper, or an abstract copied from arXiv with nothing added; those stay on the site and stay linked, but are no longer offered up for indexing. What is offered is the catalog itself plus the 42,000 people and papers with a real profile, an employer, a contribution, citations, code, or a link into the benchmark graph.
August 19, 2026
- FixedModel names no longer carry a stray "(batch)" tag. OpenRouter recently began listing a discounted batch endpoint alongside each model, and that listing's name was overwriting the real one, so the flagship read "Claude Opus 5 (batch)" on its own page, in every leaderboard row and in search results. 49 models were affected, among them GPT-5.5, Gemini 2.5 Pro and Claude Sonnet 5. The same listing was also setting their cached-input price to the batch discount while input and output stayed at standard rates; all three now read as the standard rate.
- FixedBenchmark scores again come from a model's strongest setting when the benchmark is not Artificial Analysis. The rule that picks between a model's reasoning settings only ever applied to AA's numbers, so on a source that publishes several settings the most recently written one won at random: Claude Opus 4.6 showed 68.8 on Thematic Generalization where it actually scores 80.6, which is first place. 76 scores were affected.
- FixedLMArena ratings are current again. They had not refreshed since 28 May, so the site showed three-month-old human-preference Elo and was missing every model released since. The boards now refresh weekly on their own schedule, and a missed run shows up in the ingest health view instead of ageing unnoticed.
- FixedA model no longer risks appearing several times on an LMArena board, once per reasoning setting. LMArena lists "claude-opus-5", "claude-opus-5-high" and "claude-opus-5-max" as separate entries; each now folds into the one model, represented by its strongest setting that has the votes to back it.
- AddedTwo benchmarks from Lech Mazur's independent suite: Creative Story-Writing (each model writes a story that must use 10 assigned elements, graded by a panel of LLM judges) and Thematic Generalization. 71 model scores, with each reasoning setting kept separate. Claude Opus 5 at extra-high effort leads the writing board with a 96% expected win rate against the field.
- AddedSix new LMArena leaderboards, ranked by human votes rather than an LLM judge: Creative Writing, Writing and Literature, Instruction Following, Hard Prompts, Coding and Math. Claude Opus 4.6 leads creative writing at 1506 Elo, Claude Fable 5 leads writing and literature. The site carried only LMArena's blended overall rating until now, so there was no way to ask which model writes best.
- AddedCreative-writing and emotional-intelligence scores are now in the catalog, from EQ-Bench: 214 scores across its Longform Creative Writing board (a 0-100 rubric on an 8-chapter story), Creative Writing v3 and EQ-Bench 4 (both Elo). Claude Opus 5 leads all three. Nothing in the catalog covered writing quality or emotional intelligence before, and the boards refresh with the nightly sync.
July 26, 2026
- ImprovedPaper pages now live at readable addresses named after the paper - /papers/attention-is-all-you-need instead of /papers/pwc-19007 - across all 61,000 papers. Old links keep working via redirect.
- ImprovedSearch-result titles now say what a page actually offers: benchmark pages read "GPQA Diamond benchmark: model scores & leaderboard" instead of just the benchmark's name, leaderboard pages name the current top models in the preview text, and a person's page leads with who they are ("Bob McGrew - Former Chief Research Officer of OpenAI") rather than a generic "researcher" label.
July 25, 2026
- FixedA model no longer shows up once per reasoning-effort setting. Artificial Analysis benchmarks each setting separately, so Claude Opus 5's launch filled the feed with five near-identical cards - low, medium, high, xhigh and max - and the same duplication ran through the models list, search and compare. Each family is now a single entry, and new launches stay that way. 26 duplicate entries were folded back into 11 models, including Claude Opus 5, Claude Sonnet 5, and GPT-5.6 Sol, Terra and Luna, which each carried six.
- AddedModel pages now show how the model scores at each reasoning effort, on 90 models. Claude Opus 5 ranges from 60.7 at max effort down to 50.6 at low, on identical pricing; GPT-5.6 Luna spans 51.2 to 26.6. Previously you could only see that by comparing what looked like separate models.
July 24, 2026
- AddedFirst-time visitors now get a short three-step intro: what the catalog covers, how to search all of it with ⌘K, and how to save a set of feed filters as a view and have new matches emailed to you. Each step shows the actual screen it describes. It replaces the "Start here" band that used to sit on the home page.
July 21, 2026
- FixedBenchmark scores on model pages now come from the model's strongest setting. Vendors ship many models at several reasoning efforts (GPT-5 at high, medium, low and minimal; GLM-5 with reasoning on or off), and every setting's scores were overwriting the others, so a page could show the weakest setting's numbers under the model's plain name - and which setting won could change from one night to the next. 81 models were affected. GPT-5's GPQA Diamond reads 85.4% instead of 80.8%, and its HLE 26.5% instead of 18.4%.
- FixedSearch no longer returns dead results. About 500 entries still pointed at models, tools, leaderboards, organizations, and evals that had been merged into another entry or removed, so they either led nowhere or pushed a real result off the list. Search now drops such entries as soon as the entry they point to goes away.
- FixedThe nightly refresh finishes again. It had been running over its time limit and getting cancelled partway on 4 of the last 10 nights, so fresh benchmark scores, release dates, and newly added papers and people could show up a day or more late, or wait for the next night that happened to finish. The slowest steps are now minutes faster, and the parts that record scores and update search no longer get dropped when an earlier step runs long.
June 26, 2026
- AddedModels that are announced but not generally available - like OpenAI's new GPT-5.6 Sol, Terra, and Luna, or the rolled-back Claude Fable 5 - now carry a "Not available" badge across the catalog, and their score bars render with a checkerboard fill on the charts, the same way Artificial Analysis flags them. Numbers that are still provisional (independent evaluation forthcoming) show a diagonal hatch.
June 10, 2026
- FixedPasting a Papers link into an AI assistant (or any reader that doesn't run JavaScript, like crawlers) now shows the actual paper list with the link's filters applied - week/month/year window, domain, search. Previously such tools got an empty page shell with no papers at all.
June 9, 2026
- FixedNew models, scores, and papers land in the feed every night again. The nightly data sync had been timing out and getting cut off partway for a week (Jun 2-8), so several days of updates surfaced late or not at all.
- AddedClaude Fable 5 and Claude Mythos 5 (released today) are in the catalog with their full launch benchmark table - 13 evals including seven new ones (FrontierCode, ExploitBench, BioMysteryBench, and more), pricing, and specs. Each benchmark shows Fable 5 against Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, with scores marked by who reported them.
June 8, 2026
- AddedAI models now have a "Mentions on X" tab on their page, showing what people are saying about the model on Twitter/X - filtered for spam, ranked by signal, and tagged by what kind of mention it is (comparison, result, news, opinion). Sort by signal, most recent, or most-liked.
- FixedAn eval's README tab now shows the actual README again, instead of a block of raw configuration text (title, tasks, tags, and so on).
- ImprovedA model page now loads its benchmark-score charts in batches as you scroll, instead of rendering all of them at once, so pages for models measured on dozens of benchmarks no longer stall.
- ImprovedThe Papers "Trending" tab now defaults to the last month instead of all time, so it reflects what's actually trending now; switch to this week or this year anytime.
June 6, 2026
- AddedSearch results in the API and `sophon` CLI can now be sorted: pass `sort=title` or `sort=recent` (prefix `-` to reverse) to `/api/v1/search`, or `sophon search … --sort title`. The default stays relevance.
June 3, 2026
- ImprovedOpening a paper's first-page preview now expands it into the reader with a quick animation, and collapses back when you close it, instead of opening abruptly.
June 2, 2026
- AddedBranded social preview cards for every page - the image shown when a Sophon link is shared on Slack, X, Discord, or LinkedIn. Previously only paper pages had one; the homepage card shows live catalog counts.
- AddedA public API and a `sophon` command-line tool for querying the catalog (evals, models, tools, papers) from code, agents, or the terminal - find them on the new API & CLI page, linked in the sidebar.
- FixedCollapsing the sidebar no longer briefly flashes its floating version before it closes.
- FixedThe feed's New tab no longer shows older papers as if they were just published; papers are ordered by their real publication date again.
- FixedWhen a paper's preview image fails to load, the feed and paper page now render the paper's real first page instead of a broken image or blank placeholder.
- FixedPaper abstracts no longer show raw LaTeX code (like `\textsc{}` or `\href{}{}`); the markup is cleaned up so the text and links read normally.
- ImprovedThe Capabilities page is now grouped into categories (reasoning, language, safety, agents, coding, and more) you can filter by, instead of one long flat list.
- ImprovedOn list pages, moving between the filter menus now glides the dropdown across to the next filter (or the Add filter button) in place, instead of closing one and opening the next.
- ImprovedHovering a benchmark bar on a feed model card now explains what the benchmark measures, instead of repeating its name and score.
- ImprovedIn a paper's PDF, hovering a highlight you saved now shows its note inline; click it to edit.
- ImprovedThe OSWorld, DABstep, Arena-Hard, and AlpacaEval leaderboards now refresh automatically every day, instead of only when updated by hand.
June 1, 2026
- AddedFresh arXiv papers are discovered and added automatically every day.
- AddedNew benchmarks from AfterQuery (agentic IDE coding, quant trading, web-app generation, finance) and Vals AI (tax, finance, legal, and medical domain evals), with model scores.
- FixedThe feed could freeze on stale content and stop showing new items; it now updates as soon as new evals, models, tools, or papers land.
- FixedModel cards no longer show "$0/M" or "0 tok/s" before a model's price and speed have been measured.
- ImprovedThe feed's paper stream now shows only notable papers.
- ImprovedPapers in the feed now display their own generated preview cards instead of external thumbnails.
May 30, 2026
- AddedSubscribe to any feed view by email digest (weekly or monthly) or RSS.
- AddedRename saved feed views inline.
- FixedLeaderboard pages no longer scroll sideways on mobile.
- FixedLeaderboard pages no longer freeze on mobile when a board tracks many charts.
- FixedPrivate notes no longer zoom the page on iOS or flash a phantom scrollbar.
- ImprovedModel list shows price, prompt caching, speed, latency, and context as optional inline tags you toggle per field.
- ImprovedFeed model rows explain each spec (context, price, speed, license) on hover.
- ImprovedEval pages group similar evals into their own tab.
- ImprovedLarge leaderboards load faster, revealing more models as you scroll.
- ImprovedLeaderboards show when their scores were last refreshed, next to how often they update.
May 29, 2026
- AddedBento landing homepage with per-feature previews.
- AddedFilterable feed header with saved views.
- FixedCmd-K search shows the most relevant results first.
- FixedFeed "New" orders by real release date.
- FixedHorizontal tab strips no longer scroll vertically.
- ImprovedModel titles are cleaner and more consistent.
- ImprovedOrganization pages rank lab models by intelligence and release date.