<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>trackers-and-leaderboards on tomrochette.com</title>
    <link>https://tomrochette.com/tags/trackers-and-leaderboards/</link>
    <description>Recent content in trackers-and-leaderboards on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Thu, 24 Sep 2026 11:16:33 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/trackers-and-leaderboards/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>AI Release Tracker</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>ai</category><category>releases</category><category>open-data</category>
      <description>&lt;p&gt;AI Release Tracker is a free, ad-free timeline of major frontier AI model releases since ChatGPT&amp;rsquo;s launch on November 30, 2022, with launch-day benchmark scores, pricing, and machine-readable exports for each entry.&#xA;Facts below verified as of 2026-09-24; the live site blocks automated fetchers (HTTP 429), so the current-content evidence comes from the 2026-08-26 archived copy.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It answers the one question every other site in this category layers something on top of: what shipped, when, with which launch-day numbers.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A timeline site whose FAQ describes &amp;ldquo;a free, continuously updated timeline of every major AI model release from the world&amp;rsquo;s leading AI labs&amp;rdquo;, tracking 231 models from 10 companies as of the archived copy (OpenAI, Anthropic, Google, Meta, xAI, DeepSeek, Mistral, Moonshot AI, Z.ai, Qwen).&#xA;Each entry records the release date, benchmark scores as published at launch (GPQA Diamond, SWE-Bench Verified, MMMU, Terminal-Bench, and others), parameter counts, context windows, license tier, and first-party API pricing where the lab publishes one.&#xA;Scores are framed as historical records, not live standings, and third-party-leaderboard figures are marked with their source.&#xA;The site also publishes &amp;ldquo;Expected&amp;rdquo; next-release dates per lab, which the author described as an average of the intervals between previous releases.&#xA;Surfaces: the timeline, per-benchmark ranking pages, model compare pages, analytics, pricing comparison, free email notifications on new releases, and flat-file exports: &lt;code&gt;/models.json&lt;/code&gt; and a &lt;code&gt;llms-full.txt&lt;/code&gt; corpus described as &amp;ldquo;plain text for LLM ingestion&amp;rdquo;, free to use with an attribution link.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active but nearly invisible: the dataset was last updated 2026-08-26 (newest tracked release Qwen3.8-Flash-Next), and Wayback captures through June, July, and August 2026 show the corpus growing month over month.&#xA;It was Show HN&amp;rsquo;d on 2025-12-02 by user curlii and drew 2 points and 6 comments; that thread is still the site&amp;rsquo;s entire public footprint, and no GitHub repository exists.&#xA;The footer reads &amp;ldquo;© 2026 To sider ApS · CVR DK42753491&amp;rdquo;, a Danish private limited company, so it is a side-project-grade operation with corporate paperwork.&#xA;&lt;strong&gt;Treat it as alive, useful, and one person away from dormant.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The flat files (&lt;code&gt;/models.json&lt;/code&gt;, &lt;code&gt;llms-full.txt&lt;/code&gt;) are the category&amp;rsquo;s friendliest data source for agents: one fetch, structured, free, attribution-licensed.&lt;/li&gt;&#xA;&lt;li&gt;Launch-day scores are quotable in a way leaderboard standings are not, because a record of what the lab claimed on release day does not silently move.&lt;/li&gt;&#xA;&lt;li&gt;No ads, no account wall, and a scope (releases only) it does not pretend to extend past.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Single-operator bus factor: one person, one company registration, no repository, no visible team.&lt;/li&gt;&#xA;&lt;li&gt;Verification is thin: the community footprint is one small Show HN thread, so data-entry errors would surface slowly.&lt;/li&gt;&#xA;&lt;li&gt;Naming quirks (xAI labeled &amp;ldquo;SpaceXAI&amp;rdquo;) and via-attributed benchmark numbers mean the corpus needs spot checks before you cite it.&lt;/li&gt;&#xA;&lt;li&gt;The live site 429s automated fetchers, so agents should use the flat files or the archive rather than scraping the UI.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to use, no ads.&#xA;Free tier includes email notifications on new releases; a paid &amp;ldquo;Pro instant alerts&amp;rdquo; tier sends email the moment a model is added, with no price published.&#xA;No paid price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;: current composite rankings with an API; AI Release Tracker is the historical log LLM Stats&amp;rsquo;s score churns over.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: live quality, price, and speed; choose it for what is true today, the tracker for what was claimed at launch.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;: trend datasets and research; the tracker is the raw release stream Epoch-style analysis consumes.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for engineers and agents who want a dated, quotable &amp;ldquo;what shipped when&amp;rdquo; record with a machine-readable corpus, not for deciding what to deploy this week.&lt;/strong&gt;&#xA;My disagreeable claim: this is the most valuable site in the category for agents and the least visited by humans, because &amp;ldquo;what changed since last time&amp;rdquo; is the first question of every refresh run, and this is the only member that answers it as a flat file.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt; - the live rankings layer that consumes the same release stream&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - independent quality, price, and speed measurement of what the tracker logs&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the long-run trend layer built from release-adjacent data&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/keeping-up-with-ai/&#34; &gt;Keeping Up With AI Is a Losing Strategy&lt;/a&gt; - why a cheap change-detector beats trying to read everything&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&lt;/a&gt; - archived homepage: FAQ (231 models, 10 companies, free/pro alerts, no ads), footer &amp;ldquo;To sider ApS · CVR DK42753491&amp;rdquo; (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/llms-full.txt&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/llms-full.txt&lt;/a&gt; - dataset scope, last-updated date, benchmark glossary, attribution terms (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/models.json&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/models.json&lt;/a&gt; - structured export: 231 models with release dates, parameters, context windows, benchmarks (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=46119002&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=46119002&lt;/a&gt; - the Show HN thread and the author&amp;rsquo;s description of purpose and notifications (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/items/46119002&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/items/46119002&lt;/a&gt; - full thread JSON backing the footprint claim (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=aireleasetracker&amp;amp;tags=story&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=aireleasetracker&amp;tags=story&lt;/a&gt; - the single-story footprint evidence (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/cdx/search/cdx?url=aireleasetracker.com&amp;amp;matchType=domain&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/cdx/search/cdx?url=aireleasetracker.com&amp;matchType=domain&lt;/a&gt; - snapshot cadence and subpage inventory through 2026-08-26 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Artificial Analysis</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>ai</category><category>benchmarks</category><category>evaluation</category>
      <description>&lt;p&gt;Artificial Analysis is an independent benchmarking company whose site measures AI at four layers, agents, models, cloud inference providers, and chips, and publishes the results as leaderboards, price and speed comparisons, and a public changelog.&#xA;Facts below verified as of 2026-09-24.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It is the closest thing the field has to a consumer reports for model inference: one place where quality, cost, and speed are measured the same way across 671 models.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A website and data business covering models (proprietary and open weights), coding agents, inference providers, and accelerator hardware.&#xA;The flagship Artificial Analysis Intelligence Index (v4.3.2 as of 2026-09-24) incorporates ten evaluations with published weights (agents 30%, coding 20%, scientific reasoning 20%, general 30%), and a separate Coding Agent Index (v1.5) combines DeepSWE, Terminal-Bench, and SWE-Atlas-QnA.&#xA;The homepage compares 671 models on price per token, output speed, and latency, and its Endpoint Accuracy Index re-runs the same evals against 16 third-party providers of one model to measure how much accuracy each endpoint loses to quantization or configuration.&#xA;Products around the data include Optima (build-your-own benchmarks), MicroEvals, a Model Recommender, and a Data Playground.&#xA;Scale claims from the about page: 500+ models benchmarked, 100+ inference providers, 1,000+ endpoints, 1T+ evaluation tokens.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Very active and heavily cited: the changelog had entries dated 23 September 2026, one day before verification, and the Intelligence Index itself iterates weekly (v4.2 on September 5, v4.3 two days later, v4.3.2 by September 24).&#xA;Its numbers are market-moving enough that HN threads are titled by its rankings (&amp;ldquo;GLM-5.2 is the new leading open weights model on Artificial Analysis&amp;rdquo;, 916 points in June 2026), and providers market against its measurements (Baseten&amp;rsquo;s &amp;ldquo;fastest Kimi K2.5&amp;rdquo; post).&#xA;Founded by Micah Hill-Smith (CEO, ex-McKinsey) and George Cameron (CPO); started as a side project in 2023, launched January 2024, went viral after a Swyx retweet, and raised a seed from Nat Friedman and Daniel Gross&amp;rsquo;s AI Grant with angels including Andrew Ng, Adam D&amp;rsquo;Angelo, Clem Delangue, Guillermo Rauch, and swyx.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Breadth no rival matches: quality, price, speed, providers, and now chips, in one methodology.&lt;/li&gt;&#xA;&lt;li&gt;The methodology hub publishes evaluation lists, weights, and scoring detail, and the Endpoint Accuracy Index measures an error source (endpoint drift) nobody else prices in.&lt;/li&gt;&#xA;&lt;li&gt;A dated public changelog makes its own update cadence auditable.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Its own methodology concedes prompts are partly written in-house and grading partly uses LLM judges, so the indexes measure the measurer as much as the models.&lt;/li&gt;&#xA;&lt;li&gt;HN criticism is persistent: &amp;ldquo;the benchmarks are bunk&amp;rdquo;, a &amp;ldquo;lost all trustworthiness with the astra blunder&amp;rdquo; comment on the v4.2 thread, and a pay-to-play analogy for the benchmark-firm business model; a counterpoint in the same thread calls it &amp;ldquo;the best option currently available&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;It sells private custom benchmarking and an enterprise insights subscription to AI companies, the same companies it ranks; the founders state &amp;ldquo;no one pays to be on the public leaderboard&amp;rdquo; and describe a mystery-shopper policy, which you can accept or audit.&lt;/li&gt;&#xA;&lt;li&gt;Index version churn means any citation needs the version number and date or it goes stale within weeks.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;The public site is free.&#xA;Revenue comes from an enterprise benchmarking-insights subscription and private custom benchmarking; a paid data platform exists (its terms PDF is published) with no public prices.&#xA;No public price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;: blind human votes instead of controlled evals; choose the arena for preference, Artificial Analysis for price and speed.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;: trends and open datasets rather than leaderboard-style comparisons.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;: revealed preference (actual spend) rather than constructed measurement.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended as the default first stop when a decision needs a specific model&amp;rsquo;s quality against its price and speed across providers, not for trend or adoption questions.&lt;/strong&gt;&#xA;My disagreeable claim: the Endpoint Accuracy Index is the most underrated page on the site, because quantization and endpoint defaults are the silent error source teams accept when they pick the cheap provider, and this is the only ranking in the category that makes that tradeoff visible.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt; - the crowd-preference counterpart to controlled benchmarking&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - what developers actually spend, as a check on what benchmarks imply&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the trend-and-dataset layer this site&amp;rsquo;s snapshots feed&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - the per-token economics these benchmarks feed&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/&lt;/a&gt; - homepage: 671 models, Intelligence Index v4.3.2, Coding Agent Index v1.5, Endpoint Accuracy Index across 16 providers (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/about&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/about&lt;/a&gt; - founders, backers, scale claims, four-layer scope (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/methodology&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/methodology&lt;/a&gt; - methodology hub: scope, blended-price definition, benchmark inventory (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/methodology/intelligence-benchmarking&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/methodology/intelligence-benchmarking&lt;/a&gt; - the ten evaluations, weights, and scoring detail (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/changelog&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/changelog&lt;/a&gt; - update cadence, entries dated 23 September 2026 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.latent.space/p/artificialanalysis&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.latent.space&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.latent.space/p/artificialanalysis&lt;/a&gt; - founding story, seed round, revenue model, mystery-shopper policy (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=48567759&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=48567759&lt;/a&gt; - 916-point HN thread showing citation scale (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=49586403&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=49586403&lt;/a&gt; - the &amp;ldquo;astra blunder&amp;rdquo; criticism on the v4.2 thread (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=45706969&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=45706969&lt;/a&gt; - the &amp;ldquo;benchmarks are bunk&amp;rdquo; skeptical take (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysiscdn.com/legal/ProDataPlatformTerms.pdf&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysiscdn.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysiscdn.com/legal/ProDataPlatformTerms.pdf&lt;/a&gt; - existence of the paid data platform terms (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Epoch AI</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>ai</category><category>datasets</category><category>trends</category>
      <description>&lt;p&gt;Epoch AI is a 501(c)(3) research nonprofit that maintains open datasets on AI models, compute, data centers, chips, and companies, runs its own benchmarks, and publishes research, a newsletter, and a podcast, all under a transparency policy that itemizes its funders and paid consultations.&#xA;Facts below verified as of 2026-09-24.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Where the other sites in this category tell you what is winning this week, Epoch tells you how fast the whole field is moving, and lets you download the series behind the claim.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A research organization, roughly 50 people, whose public products are data explorers (AI Models with 3,200+ entries from 1950 to today, AI Data Centers, Chip owners, Companies, Polling), benchmarks (the Epoch Capabilities Index, FrontierMath, MirrorCode, EBR-bench), papers and reports, short Data Insights, the Gradient Updates newsletter, and the Epoch After Hours podcast.&#xA;Data is free to use with attribution under Creative Commons (CC BY), and a Python client library and data repositories live on its GitHub org.&#xA;Its numbers are the citation of record for AI trend claims, used by Our World in Data, government AI-safety reports, and the financial press.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Actively maintained at near-daily granularity: the data page was stamped &amp;ldquo;Updated Sep. 24, 2026&amp;rdquo;, the day of verification, and its GitHub repositories show pushes the same day.&#xA;HN traction is substantial and recurring: the FrontierMath launch drew 185 points in 2024, &amp;ldquo;FrontierMath was funded by OpenAI&amp;rdquo; drew 483 points in January 2025, and &amp;ldquo;Epoch confirms GPT5.4 Pro solved a frontier math open problem&amp;rdquo; drew 480 points in March 2026.&#xA;Founded by Jaime Sevilla and collaborators; funding is donations (Coefficient Giving grants of $8.5M in 2025, $4.13M in 2024, and more, plus the Survival and Flourishing Fund, Jaan Tallinn, and Schmidt Sciences), with paid consultations disclosed on the transparency page.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The most checkable actor in the category: funders, grants, and paid consultations (OpenAI, Google DeepMind, the EU AI Office) are named on one page.&lt;/li&gt;&#xA;&lt;li&gt;CC BY data plus a Python client means their series can go straight into your own analysis instead of a screenshot.&lt;/li&gt;&#xA;&lt;li&gt;Coverage spans hardware, data centers, and companies, not just model leaderboards, which is where the multi-year questions live.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The FrontierMath funding controversy is the load test it had to pass: TechCrunch reported in January 2025 that Epoch &amp;ldquo;didn&amp;rsquo;t disclose that it had received funding from OpenAI&amp;rdquo; while developing a benchmark OpenAI&amp;rsquo;s models were scored on.&lt;/li&gt;&#xA;&lt;li&gt;Its response was structural (full disclosure plus a 2026 FrontierMath data audit), but the structural tension remains: the labs it benchmarks are also its consulting clients and neighbors in the funding ecosystem.&lt;/li&gt;&#xA;&lt;li&gt;Trend series are backward-looking; nothing here tells you which provider to buy this week.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free: a donation-funded nonprofit whose data is CC BY licensed.&#xA;It sells consultations and custom research to organizations, with no public price list.&#xA;No reader-facing price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: live model-versus-model comparisons; Epoch answers the multi-year questions those snapshots sit inside.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;: the raw release stream; Epoch&amp;rsquo;s model dataset is the curated, parameterized version of the same history.&lt;/li&gt;&#xA;&lt;li&gt;Stanford HAI&amp;rsquo;s AI Index: a single annual curated report; Epoch is continuous and downloadable.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for anyone who needs to cite how fast AI capability, compute, or cost is moving, with the data behind the number, not for weekly model selection.&lt;/strong&gt;&#xA;My disagreeable claim: the FrontierMath episode made Epoch more trustworthy, not less, because it is the only member of this category that responded to a conflict-of-interest scandal by publishing an itemized funder and client list, and none of its commercial peers has matched that.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the present-tense measurement layer above Epoch&amp;rsquo;s trend lines&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt; - the raw release stream behind the model dataset&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt; - preference-based ranking, the opposite timescale of an Epoch series&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/&lt;/a&gt; - homepage: tagline, 3,200+ model dataset, product map, update stamps (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/about&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/about&lt;/a&gt; - &amp;ldquo;data-first research nonprofit&amp;rdquo; self-description and mission (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/about/transparency&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/about/transparency&lt;/a&gt; - 501(c)(3) status, itemized funders, disclosed OpenAI and DeepMind consultations, the 2026 FrontierMath data audit (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/data&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/data&lt;/a&gt; - data explorers, CC BY licensing, &amp;ldquo;Updated Sep. 24, 2026&amp;rdquo; stamp (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/team&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/team&lt;/a&gt; - team size and leadership including Jaime Sevilla (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/frontiermath&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/frontiermath&lt;/a&gt; - the FrontierMath benchmark surface (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/01/19/ai-benchmarking-organization-criticized-for-waiting-to-disclose-funding-from-openai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/01/19/ai-benchmarking-organization-criticized-for-waiting-to-disclose-funding-from-openai/&lt;/a&gt; - the funding-disclosure controversy (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/epoch-research&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/epoch-research&lt;/a&gt; - active data and benchmark repositories, pushes on 2026-09-24 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epochai.substack.com/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epochai.substack.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epochai.substack.com/&lt;/a&gt; - the Gradient Updates newsletter (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>LLM Stats</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>llm</category><category>benchmarks</category><category>api</category>
      <description>&lt;p&gt;LLM Stats is a model-comparison platform that ranks 398 canonical models on a composite &amp;ldquo;LLM Stats Score&amp;rdquo;, publishes task-level leaderboards, pricing, and comparison pages, and exposes the whole dataset through a REST API and an MCP server aimed at agents.&#xA;Facts below verified as of 2026-09-24.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It is the only member of this category that treats an agent, not a human, as the primary consumer of the leaderboard.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A website and data service: a composite leaderboard (&amp;ldquo;Rankings for 300+ Top AI Models by Intelligence, Speed &amp;amp; Price&amp;rdquo;), category boards for reasoning, coding, agents, math, long context, vision, and more, model compare and pricing pages, an AI news feed, a weekly newsletter, and arena products hosted under its sibling brand huggle.ai.&#xA;The methodology (score v3.1, updated 2026-09-02) is more careful than the aggregator genre demands: scores normalize within each benchmark, missing results are treated as missing rather than as failures, models with less evidence carry wider uncertainty, and lab-reported numbers are labeled separately from independently verified ones.&#xA;The developer surface lists endpoints for models, benchmarks, scores, rankings, and updates, plus &amp;ldquo;11 MCP tools&amp;rdquo; with one-click OAuth for Cursor, Claude Code, and Windsurf.&#xA;It is operated by ZeroEval Inc., which describes itself as building &amp;ldquo;the independent measurement layer for AI&amp;rdquo; and names LLM Stats as its first product; the founder is Jonathan Chavez, who Show HN&amp;rsquo;d earlier versions in November 2024 and January 2025.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Actively maintained: the methodology page was modified 2026-09-02, the docs repository was pushed 2026-09-08, and the news page is titled by the current month.&#xA;Traction is modest but growing: two small Show HNs (3 and 7 points), 20 HN comments referencing the site, and a small GitHub org, with self-reported reach (people at OpenAI, Anthropic, Google, Meta, &amp;ldquo;400,000+ more&amp;rdquo;) that cannot be verified.&#xA;The corporate backing is self-displayed: a &amp;ldquo;Backed by&amp;rdquo; strip on zeroeval.com lists Y Combinator, Hugging Face, Harvard Medical, Google, and Datadog.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The agent surface is the category&amp;rsquo;s best: a free API tier plus MCP tools means your next session can query leaderboard data directly.&lt;/li&gt;&#xA;&lt;li&gt;The uncertainty-aware, missing-is-missing scoring policy is more statistically defensible than the genre&amp;rsquo;s usual single-number bravado.&lt;/li&gt;&#xA;&lt;li&gt;Breadth of coverage (398 models, 50+ benchmarks claimed) with a compare tool and per-task boards for quick narrowing.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The site is dressed for search engines: dozens of &amp;ldquo;best AI for X&amp;rdquo; landing pages (lawyers, architects, YouTube thumbnails) whose purpose is search traffic rather than measurement.&lt;/li&gt;&#xA;&lt;li&gt;Methodology details are partly gated: &amp;ldquo;We share additional methodology details with researchers, customers, and partners when appropriate&amp;rdquo;, and pages sit behind a Cloudflare Turnstile for non-interactive clients.&lt;/li&gt;&#xA;&lt;li&gt;Early HN feedback flagged coverage gaps (Cerebras missing, &amp;ldquo;makes me wonder if other evaluations are missing&amp;rdquo;) and the author&amp;rsquo;s own admission that &amp;ldquo;some labs cherry pick the benchmarks they want to report&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;Paid evaluation services (custom benchmarking, data labeling) put revenue on the same side of the table as the companies being scored.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to browse, with a free Community API tier (250 data responses/day, 50 req/min, rolling 6-month history).&#xA;Builder: $99/month with 5,000 responses/day, 300 req/min, 12-month history, bulk snapshots, and an incremental-updates API.&#xA;Commercial: contract pricing with redistribution licensing and signed webhooks.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Price history&#xA;    &lt;div id=&#34;price-history&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#price-history&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Date&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Plan&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Change&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Source&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-09-24&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Community&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Baseline: free tier at 250 data responses/day, 50 req/min, 6-month history&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;a href=&#34;https://llm-stats.com/developer&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/developer&lt;/a&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-09-24&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Builder&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Baseline: $99/month at 5,000 responses/day, 300 req/min, 12-month history&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;a href=&#34;https://llm-stats.com/developer&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/developer&lt;/a&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: first-party controlled evals with published weights; LLM Stats aggregates public evidence and labels it.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;: the dated release log; LLM Stats is the current-rankings view of the same population.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;: human preference rather than benchmark aggregation.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for wiring leaderboard data into agents and scripts on the free tier, and for a quick conservative read across many models; not as the deciding source for a deployment commitment.&lt;/strong&gt;&#xA;My disagreeable claim: the MCP server is the most consequential feature introduced in this category in years, because the reader of these leaderboards is becoming an agent that cannot click through a chart-dense SPA, and LLM Stats is the only member that noticed.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the first-party-measurement alternative to aggregation&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt; - the release log whose entries become LLM Stats rows&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - usage data as the counterweight to benchmark aggregates&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - what a composite score is and is not useful for when picking a model&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/&lt;/a&gt; - homepage: 398 canonical models, composite score, task boards, newsletter (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/methodology/llm-stats-score&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/methodology/llm-stats-score&lt;/a&gt; - score v3.1 construction, evidence policy, limitations, 2026-09-02 modification date (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/developer&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/developer&lt;/a&gt; - API and MCP endpoints, plan tiers and quotas, &amp;ldquo;updated within hours&amp;rdquo; claim (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/about-us&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/about-us&lt;/a&gt; - founder Jonathan Chavez and the zeroeval relationship (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/ai-news&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/ai-news&lt;/a&gt; - the news feed and self-reported reach claim (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://zeroeval.com&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=zeroeval.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://zeroeval.com&lt;/a&gt; - ZeroEval Inc. description, product list, self-displayed &amp;ldquo;Backed by&amp;rdquo; strip (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/orgs/zeroeval&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/orgs/zeroeval&lt;/a&gt; - org metadata and docs-repo activity (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=42590841&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=42590841&lt;/a&gt; - the January 2025 Show HN and authorship evidence (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=42231372&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=42231372&lt;/a&gt; - the November 2024 launch thread with the Cerebras coverage-gap criticism (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>LMArena</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>evaluation</category><category>leaderboards</category><category>human-feedback</category>
      <description>&lt;p&gt;LMArena ranks AI models by blind human preference votes: two anonymous models answer the same prompt, you pick the winner, and Bradley-Terry-style statistics turn millions of those picks into leaderboards spanning text, image, video, vision, search, web development, and agents.&#xA;Facts below verified as of 2026-09-24; lmarena.ai is a client-rendered app, so vote and model counts beyond the founding paper&amp;rsquo;s figures could not be read from its HTML.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;It is the field&amp;rsquo;s mood ring: the most cited signal of which model people prefer, and the easiest leaderboard in existence to game.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A website (lmarena.ai), a family of leaderboards (Agent Overall, Text, WebDev, Image, Video, Vision, Document, Search), a WebDev arena (web.lmarena.ai), a blog, and open methodology repositories, run by Arena Intelligence Inc., the company that grew out of the UC Berkeley and LMSYS Chatbot Arena project.&#xA;The founding paper (arXiv:2403.04132, March 2024) describes the pairwise crowdsourcing method and 240K+ votes at the time; the leaderboard methodology source is published as the &lt;code&gt;arena-rank&lt;/code&gt; repository, pushed August 2026.&#xA;The company raised $100M at a $600M valuation in May 2025, led by Andreessen Horowitz and UC Investments, and labs including OpenAI, Google, and Anthropic partner with it to put flagship models in front of voters.&#xA;Recent product motion includes AutoEval scores added to the leaderboards (to complement slowly collected human votes), agent leaderboard categories with task costs (August 2026), and a HarnessTax research post (September 2026).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;The most-cited leaderboard in the field: a critical paper describing Chatbot Arena as &amp;ldquo;the go-to leaderboard for ranking the most capable AI systems&amp;rdquo; is itself the best evidence of that status.&#xA;The blog posts within days of verification, the arenas run continuously, and HN threads routinely open with its rankings as the premise.&#xA;The company is well capitalized and has converted the academic project into a venture-scale business, which is also the source of its hardest questions.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Preference at scale is a measurement no benchmark suite replaces: it captures whatever makes people pick one answer over another, including style, format, and thoroughness.&lt;/li&gt;&#xA;&lt;li&gt;The methodology code and the voting procedure are public, so the statistics are checkable even when the data pipelines are not.&lt;/li&gt;&#xA;&lt;li&gt;The arena expansion (WebDev, agents, image, video) follows usage: it measures the surfaces people actually use models on.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The gaming record is documented: Meta&amp;rsquo;s Llama 4 Maverick episode (an &amp;ldquo;experimental chat version&amp;rdquo; tested on the arena that differed from the shipped model) forced a policy update, and LMArena&amp;rsquo;s own statement conceded &amp;ldquo;Meta&amp;rsquo;s interpretation of our policy did not match what we expect from model providers&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;The Leaderboard Illusion paper (arXiv:2504.20879) documents &amp;ldquo;undisclosed private testing practices&amp;rdquo; that &amp;ldquo;benefit a handful of providers&amp;rdquo;, counting 27 private Meta variants tested before the Llama 4 release and sampling-rate asymmetries favoring closed models.&lt;/li&gt;&#xA;&lt;li&gt;The sharpest criticism (Surge AI&amp;rsquo;s &amp;ldquo;LMArena is a cancer on AI&amp;rdquo;, 246 points on HN in January 2026) argues the format &amp;ldquo;rewards superficiality over accuracy&amp;rdquo; because &amp;ldquo;the easiest way to climb the leaderboard isn&amp;rsquo;t to be smarter; it&amp;rsquo;s to hack human attention span&amp;rdquo;.&lt;/li&gt;&#xA;&lt;li&gt;A preference rank is not a capability claim: verbosity and sycophancy win votes that lose tasks, so the number is routinely over-read by headlines.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to use and to vote.&#xA;No public pricing page exists; the company is venture funded, and its &amp;ldquo;Try Arena&amp;rdquo; product surfaces are free at the time of verification.&#xA;No reader-facing price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: controlled first-party evals with published weights; choose it when you need price and speed, the arena when you need preference.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;: revealed preference (spend) versus stated preference (votes).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;: benchmark aggregation, which at least labels what it cannot verify.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended as the fastest read on which models feel better to people, and as a required filter on any headline of the form &amp;ldquo;X tops the arena&amp;rdquo;; not as the number you commit money against.&lt;/strong&gt;&#xA;My disagreeable claim: the leaderboard is the least valuable thing the arenas produce, because the vote stream is quietly one of the largest human-feedback datasets ever assembled, and whoever holds it holds a training asset, not just a ranking.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the controlled-eval counterpart to crowd preference&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - what people pay for, against what they vote for&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the research-nonprofit model of measurement this company left behind&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/you-cannot-out-review-a-machine-by-hand/&#34; &gt;You Cannot Out-Review a Machine by Hand&lt;/a&gt; - human judgment as a bottleneck, applied to review instead of ranking&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://lmarena.ai/&lt;/a&gt; - homepage meta: blind comparison, vote-driven leaderboards across text, image, and code (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/&lt;/a&gt; - Arena Intelligence Inc. identity, leaderboard families, 2026 post dates (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://blog.lmarena.ai/how-it-works/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=blog.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://blog.lmarena.ai/how-it-works/&lt;/a&gt; - the vote flow and identity reveal procedure (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2403.04132&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2403.04132&lt;/a&gt; - the founding paper: method and 240K+ votes (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2504.20879&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2504.20879&lt;/a&gt; - The Leaderboard Illusion: private testing, 27 Meta variants, sampling asymmetries (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&lt;/a&gt; - $100M seed, $600M valuation, investors, Berkeley origin (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.theverge.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming&lt;/a&gt; - the Maverick gaming episode and LMArena&amp;rsquo;s policy response (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=surgehq.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://surgehq.ai/blog/lmarena-is-a-plague-on-ai&lt;/a&gt; - the strongest critical essay on the format (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/lmarena&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/lmarena&lt;/a&gt; - methodology repositories including arena-rank, pushed 2026-08-04 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.lmarena.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.lmarena.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.lmarena.ai/&lt;/a&gt; - the WebDev arena surface (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>OpenRouter Rankings</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>trackers-and-leaderboards</category><category>usage</category><category>market-share</category><category>open-data</category>
      <description>&lt;p&gt;OpenRouter Rankings ranks models by the tokens actually processed through the OpenRouter gateway, published with CC BY 4.0 licensing, a public Data API, and views by task, cost per session, market share, and app, and it is the only ranking in its category built on revealed spend rather than votes or benchmarks.&#xA;Facts below verified as of 2026-09-24.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Every other leaderboard in this category measures opinion; this one measures invoices.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A rankings surface inside openrouter.ai, the LLM gateway founded in early 2023 that claims 500T+ monthly tokens, 10M+ users, 80+ providers, and 500+ models.&#xA;The page states plainly what it does not measure: &amp;ldquo;They do not rank models by accuracy, reasoning ability, or benchmark performance&amp;rdquo;, and it excludes private requests while bucketing usage into daily UTC aggregates per model variant.&#xA;Views include top models (today, week, month), top models by task share of spend, cost per session for coding agents, market share by model author, fastest models, and top apps (Hermes Agent, Claude Code, Kilo Code, and Cline led in tokens as of 2026-09-24).&#xA;Data access is unusually open: rankings are CC BY 4.0 with a required citation format, and the Data API serves &lt;code&gt;rankings-daily&lt;/code&gt; (top 50 public models per day, history back to 2025-01-01, 30 requests/minute, live-updating) to any inference key.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Fresh and commercially central: the page showed &amp;ldquo;Usage data through Sep 23, 2026&amp;rdquo; on the day of verification, a one-day lag, and updates flow live as traffic arrives.&#xA;The underlying gateway is a top-of-category business: $113M Series B led by CapitalG in May 2026 at a reported $1.3B valuation, then an announced acquisition by Stripe on August 19, 2026 at a reported price above $7B, with the company claiming 10T+ tokens per day.&#xA;Its numbers are routinely quoted as market signal in press and on HN, and OpenRouter itself turned the dataset into research (a 100T-token empirical study that drew 207 points on HN).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Revealed preference: tokens and spend through a paid gateway are harder to game with a demo day prompt than a vote or a benchmark run.&lt;/li&gt;&#xA;&lt;li&gt;Openness is a policy, not a favor: CC BY 4.0, a documented citation format, and a free JSON API with a year-plus of history.&lt;/li&gt;&#xA;&lt;li&gt;The by-task and cost-per-session views answer questions (what does a coding-agent session actually cost) that no benchmark board even frames.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The sample is one gateway, and the page says so: &amp;ldquo;They do not include usage on a model provider&amp;rsquo;s own API or across the whole market&amp;rdquo;, and &amp;ldquo;token volume is not a count of requests, users, or spend&amp;rdquo; since verbosity and tokenization differ.&lt;/li&gt;&#xA;&lt;li&gt;Free-tier promotions distort the ranks: Max Woolf&amp;rsquo;s analysis of the Tencent Hy3 model topping the leaderboard documents a free period on a single provider driving the rank, and the Kilo Code free-Grok episode did the same in 2025.&lt;/li&gt;&#xA;&lt;li&gt;Independence is now a clock: the rankings belong to a gateway being absorbed by Stripe, so treat &amp;ldquo;market share&amp;rdquo; as &amp;ldquo;OpenRouter&amp;rsquo;s customers&amp;rsquo; share&amp;rdquo; and re-verify who publishes the numbers after the deal closes.&lt;/li&gt;&#xA;&lt;li&gt;Figures restate as snapshots refresh, so any number you cite needs its as-of date recorded alongside it.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to view and to query, CC BY 4.0 licensed with attribution.&#xA;OpenRouter monetizes the gateway (inference and enterprise controls), not the rankings.&#xA;No reader-facing price is stated, so no price history applies.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;: constructed measurement of quality, price, and speed; the rankings measure adoption, and the two disagree for good reasons.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;: stated preference versus revealed preference.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;: benchmark aggregates with an agent-facing API; the rankings&amp;rsquo; API is the usage-side mirror.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for answering &amp;ldquo;what are developers actually paying for this week&amp;rdquo; with numbers you can legally republish, not for any claim about which model is more capable.&lt;/strong&gt;&#xA;My disagreeable claim: this is the most misused ranking in the category, cited as market truth from a sample that excludes every provider&amp;rsquo;s first-party API traffic, and the Stripe acquisition makes reading its provenance, not its numbers, the skill that matters going forward.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - the quality-side measurement these usage numbers need as a counterweight&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt; - votes versus invoices as two theories of preference&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt; - the release stream that explains rank jumps the usage data cannot&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - the cost-per-task thinking the session-cost view extends&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/rankings&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/rankings&lt;/a&gt; - methodology, caveats, windows, top apps, CC BY 4.0 note, data-through date (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/docs/cookbook/administration/data-api&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/docs/cookbook/administration/data-api&lt;/a&gt; - Data API endpoints, limits, citation format, dataset history from 2025-01-01 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/about&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/about&lt;/a&gt; - company scale claims and founding date (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/blog/announcements/series-b/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/blog/announcements/series-b/&lt;/a&gt; - $113M Series B and investors, May 2026 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&lt;/a&gt; - the Stripe acquisition announcement, August 2026 (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openrouter-for-7b/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openrouter-for-7b/&lt;/a&gt; - the reported price and valuation (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://minimaxir.com/2026/05/openrouter-hy3/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=minimaxir.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://minimaxir.com/2026/05/openrouter-hy3/&lt;/a&gt; - the free-tier distortion analysis behind the Hy3 rank (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=48317294&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=48317294&lt;/a&gt; - the HN thread questioning free-driven ranks (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=48330499&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=48330499&lt;/a&gt; - the &amp;ldquo;market signal&amp;rdquo; counterpoint (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=46154022&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=46154022&lt;/a&gt; - the 100T-token study showing the dataset&amp;rsquo;s research use (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Trackers and leaderboards</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/</guid>
      <category>agents</category><category>trackers-and-leaderboards</category>
      <description>&lt;p&gt;Sites that watch the AI field itself and publish a continuously refreshed number from that watching: release timelines, independent benchmarks, open datasets, composite rankings, preference arenas, and usage statistics.&#xA;The category&amp;rsquo;s dividing line is what the number measures, since each member answers a different question about the same field.&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt; - a free, ad-free timeline of frontier releases since ChatGPT, with launch-day scores and flat-file exports for agents.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt; - independent quality, price, and speed measurement across 671 models, 100+ providers, and the chips underneath.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt; - the research nonprofit&amp;rsquo;s open datasets and trend series on models, compute, data centers, and AI&amp;rsquo;s trajectory.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt; - a conservative composite leaderboard over 398 models with the category&amp;rsquo;s only MCP server for agents.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt; - blind human-vote arenas ranking models by preference, the field&amp;rsquo;s most cited and most gamed signal.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt; - real token spend through the OpenRouter gateway, CC BY licensed, the category&amp;rsquo;s only revealed-preference number.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;Its members are compared on shared rows in the &lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/&#34; &gt;Trackers and Leaderboards Feature Matrix&lt;/a&gt;.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Added AI Release Tracker.&lt;/li&gt;&#xA;&lt;li&gt;2026-09-24 - Added Artificial Analysis.&lt;/li&gt;&#xA;&lt;li&gt;2026-09-24 - Added Epoch AI.&lt;/li&gt;&#xA;&lt;li&gt;2026-09-24 - Added LLM Stats.&lt;/li&gt;&#xA;&lt;li&gt;2026-09-24 - Added LMArena.&lt;/li&gt;&#xA;&lt;li&gt;2026-09-24 - Added OpenRouter Rankings.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Trackers and Leaderboards Feature Matrix</title>
      <link>https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/</link>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/trackers-and-leaderboards/trackers-and-leaderboards-feature-matrix/</guid>
      <category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>comparison</category><category>trackers-and-leaderboards</category><category>benchmarks</category><category>leaderboards</category><category>open-data</category>
      <description>&lt;p&gt;This matrix compares the six members of the Trackers and leaderboards category: sites whose product is a continuously refreshed number about the AI field itself.&#xA;The members split cleanly on what their number measures: what shipped (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;), what the operator measured (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;), how the field moves (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;), what public evidence aggregates to (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;), what people prefer (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;), and what people pay for (&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;).&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;No row of this matrix crowns a winner, because the rows are different questions; the failure mode is citing a site for a question it does not answer, usually LMArena ranks quoted as capability or OpenRouter tokens quoted as market share.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;Legend: ✓ supported, ✗ not supported, ~ partial or conditional, ? not verified.&#xA;Each column links to the full research note; every cell traces to a source cited there or in the references.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;The matrix&#xA;    &lt;div id=&#34;the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/ai-release-tracker/&#34; &gt;AI Release Tracker&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/artificial-analysis/&#34; &gt;Artificial Analysis&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/epoch-ai/&#34; &gt;Epoch AI&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/llm-stats/&#34; &gt;LLM Stats&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/lmarena/&#34; &gt;LMArena&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;OpenRouter Rankings&lt;/a&gt;&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Kind&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;release timeline plus flat-file corpus&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;independent benchmarking site and data business&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;research nonprofit with open datasets&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;composite aggregator with agent-facing API&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;blind preference arena&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;gateway usage rankings&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;The number measures&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;launch-day facts: what shipped, when, with which claimed scores&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;the operator&amp;rsquo;s own controlled evals, prices, and speed runs&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;long-run trends: compute, cost, capability over time&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;public benchmark evidence, normalized with uncertainty&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;blind human preference votes&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;tokens processed through one gateway&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Run by&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;To sider ApS, a Danish side project (one visible operator)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;venture-backed independent company (AI Grant seed)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;501(c)(3) nonprofit, itemized donors, about 50 people&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;ZeroEval Inc. (self-displayed YC backing)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Arena Intelligence Inc. ($100M seed at $600M)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;OpenRouter, a gateway acquired-by-Stripe (announced 2026-08-19)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Coverage&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;231 releases, 10 labs (as of 2026-08-26)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;671 models, 100+ providers, 1,000+ endpoints, chips&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;3,200+ models since 1950, data centers, chips, companies&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;398 canonical models, 50+ benchmarks claimed&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;text, image, video, vision, search, webdev, agent arenas&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;500+ models via the gateway, 80+ providers&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Update cadence&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;on each release; dataset last updated 2026-08-26&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;daily changelog; index versions iterate weekly&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;near-daily data updates, page-stamped&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;continuously; &amp;ldquo;within hours of release&amp;rdquo; claimed&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;continuous votes; product posts within days&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;daily UTC buckets, about one day of lag&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Methodology published&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ FAQ and data notes, no formal methodology&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology hub with evaluation lists and weights&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ transparency page, papers, and data documentation&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ score construction published, details gated&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology repo (arena-rank) plus founding paper&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ on-page caveats plus Data API docs&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Data access&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ /models.json and llms-full.txt, free with attribution&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ web charts free, data platform paid, no public API documented&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ CC BY datasets and a Python client&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ REST plus 11 MCP tools, free tier&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ arenas open, methodology open, raw votes not re-offered&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ CC BY 4.0 JSON via Data API, history to 2025-01-01&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Reader pricing&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, no ads; Pro instant alerts, price unpublished&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free site; enterprise and data-platform tiers, prices unpublished&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, donation funded&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free tier; Builder $99/month; Commercial contract&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;free, attribution required&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Independence caveat&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;one operator, one company registration, no second pair of eyes&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;sells private benchmarking to the labs it ranks (disclosed policy)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;consulted for OpenAI and DeepMind, disclosed and audited&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;paid eval services and a consumer sibling in the same company&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;lab partnerships and a documented private-testing history&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;owned by the gateway it measures, being absorbed by Stripe&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Verification hooks&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ full corpus re-downloadable and diffable&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ changelog and methodology public, raw runs private&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ data, notebooks, and code public&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;~ API re-queryable, scoring internals partly gated&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ methodology code public, vote data not re-offered&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;✓ daily snapshots re-fetchable under CC BY&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Reading the matrix&#xA;    &lt;div id=&#34;reading-the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#reading-the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;The row that sorts the category is &amp;ldquo;the number measures&amp;rdquo;, and it doubles as a citation guide: release questions go to the tracker, present-tense quality to Artificial Analysis, trend claims to Epoch, quick composites to LLM Stats, preference to LMArena, spend to OpenRouter.&lt;/strong&gt;&#xA;The two most-quoted members are also the two most misquoted: LMArena ranks get read as capability, and gateway tokens get read as market share, when the matrix&amp;rsquo;s own rows show both measure something narrower.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Verification strength tracks institutional type, not popularity.&lt;/strong&gt;&#xA;The nonprofit (Epoch) publishes its funders, its consultations, and its data; the gateway (OpenRouter) publishes its raw daily numbers under CC BY; the side project (AI Release Tracker) lets you diff its whole corpus; while the three venture-scale companies publish methodology and products but keep raw runs, votes, or scoring internals private.&#xA;&lt;strong&gt;If you need to check the work rather than read the work, the checkable columns are the nonprofit&amp;rsquo;s, the gateway&amp;rsquo;s, and the side project&amp;rsquo;s, which is the opposite of what citation frequency would predict.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;Every column carries a conflict row, and the conflicts are structural rather than scandals: the benchmark firm sells benchmarking to the benchmarked, the arena partners with the labs it ranks, the aggregator sells evaluation services, the gateway&amp;rsquo;s data is its own marketing, and the nonprofit consults for the labs.&lt;/strong&gt;&#xA;Epoch&amp;rsquo;s FrontierMath episode and LMArena&amp;rsquo;s Maverick episode are the two documented failures, and both produced their institution&amp;rsquo;s strongest disclosure artifacts, which is the pattern worth watching on every refresh of this page.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Choosing from the matrix&#xA;    &lt;div id=&#34;choosing-from-the-matrix&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#choosing-from-the-matrix&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Needing a dated record of what shipped and what the lab claimed that day: AI Release Tracker, via its flat files.&lt;/li&gt;&#xA;&lt;li&gt;Picking a model or provider this week and needing quality against price and speed: Artificial Analysis, with the Endpoint Accuracy Index for the cheap-provider question.&lt;/li&gt;&#xA;&lt;li&gt;Citing how fast the field moves, with downloadable data: Epoch AI.&lt;/li&gt;&#xA;&lt;li&gt;Wiring leaderboard data into an agent or a script: LLM Stats&amp;rsquo;s free API and MCP tools, accepting the gated internals.&lt;/li&gt;&#xA;&lt;li&gt;Sensing which model people prefer right now: LMArena, read as a mood ring and never as a capability claim.&lt;/li&gt;&#xA;&lt;li&gt;Asking what developers actually spend on: OpenRouter Rankings, quoted with its as-of date and its single-gateway caveat.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-24 - Created with six columns (AI Release Tracker, Artificial Analysis, Epoch AI, LLM Stats, LMArena, OpenRouter Rankings) when the category was seeded at the owner&amp;rsquo;s request.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-selection-for-coding-tasks/&#34; &gt;Model Selection for Coding Tasks&lt;/a&gt; - the decision layer these measurements feed, and the leaderboard skepticism it argues for&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/people-and-publications/people-and-publications-feature-matrix/&#34; &gt;People and Publications Feature Matrix&lt;/a&gt; - the voices interpreting these numbers, compared on their own matrix&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/keeping-up-with-ai/&#34; &gt;Keeping Up With AI Is a Losing Strategy&lt;/a&gt; - the filtering argument for keeping this category small and pull-driven&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=web.archive.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://web.archive.org/web/20260826134434/https://aireleasetracker.com/&lt;/a&gt; - AI Release Tracker column: scope, footprint, alerts, operator (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/&lt;/a&gt; - Artificial Analysis column: model count, indexes, provider coverage (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://artificialanalysis.ai/methodology/intelligence-benchmarking&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=artificialanalysis.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://artificialanalysis.ai/methodology/intelligence-benchmarking&lt;/a&gt; - Artificial Analysis column: published evaluation weights (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://epoch.ai/about/transparency&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=epoch.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://epoch.ai/about/transparency&lt;/a&gt; - Epoch AI column: nonprofit status, funders, consultations (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://llm-stats.com/developer&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=llm-stats.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://llm-stats.com/developer&lt;/a&gt; - LLM Stats column: API tiers, MCP tools, quotas (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=techcrunch.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://techcrunch.com/2025/05/21/lm-arena-the-organization-behind-popular-ai-leaderboards-lands-100m/&lt;/a&gt; - LMArena column: funding and company identity (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2504.20879&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2504.20879&lt;/a&gt; - LMArena column: the private-testing findings (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/rankings&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/rankings&lt;/a&gt; - OpenRouter Rankings column: methodology, caveats, licensing (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=openrouter.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://openrouter.ai/blog/announcements/openrouter-is-joining-stripe/&lt;/a&gt; - OpenRouter Rankings column: the ownership question (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://minimaxir.com/2026/05/openrouter-hy3/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=minimaxir.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://minimaxir.com/2026/05/openrouter-hy3/&lt;/a&gt; - OpenRouter Rankings column: the free-tier distortion evidence (fetched 200, 2026-09-24)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
