<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>inference on tomrochette.com</title>
    <link>https://tomrochette.com/tags/inference/</link>
    <description>Recent content in inference on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Sat, 26 Sep 2026 03:01:49 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/inference/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Cerebras Code</title>
      <link>https://tomrochette.com/agents/model-access/cerebras-code/</link>
      <pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/model-access/cerebras-code/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>model-access</category><category>inference</category><category>coding</category>
      <description>&lt;p&gt;Cerebras Code is a subscription from chipmaker Cerebras that sells fast inference on one open coding model at $50/month (Pro) and $200/month (Max).&#xA;&lt;strong&gt;The pitch is speed, and the record shows the speed claim and the quota fine print are the two things to verify before paying.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A hosted coding-inference plan on Cerebras wafer-scale hardware, consumed by pointing any OpenAI-compatible editor or agent (Cline, OpenCode, Crush, Cursor) at a Cerebras API key.&#xA;It launched August 1, 2025 with Qwen3-Coder-480B advertised at up to 2,000 tokens per second and a 131k context window.&#xA;As of 2026-09-26 the product page promotes GLM 4.7 at &amp;ldquo;1,000 tokens+ per second&amp;rdquo;, so the headline model has already been swapped once.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Launched August 1, 2025 and drew 449 points and 172 comments on Hacker News the same day.&#xA;Launch windows sold out repeatedly, and as of 2026-09-26 both Pro and Max are marked &amp;ldquo;sold out&amp;rdquo; on cerebras.ai/code, with a limited free trial still open.&#xA;The model changed from Qwen3-Coder to GLM 4.7 between launch and now, which shows the plan follows whichever open model is fastest rather than committing to one family.&#xA;I could not verify funding or subscriber counts.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Speed is genuinely differentiated: even its harshest reviewer calls Cerebras the fastest provider of its model, bar none.&lt;/li&gt;&#xA;&lt;li&gt;Flat monthly pricing with published daily allowances (24M tokens Pro, 120M Max) instead of per-token anxiety.&lt;/li&gt;&#xA;&lt;li&gt;Bring-your-own-editor stance with OpenAI-compatible endpoints, no proprietary IDE lock-in.&lt;/li&gt;&#xA;&lt;li&gt;By October 2025 the original tokens-per-minute caps had been raised in response to criticism.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The marketing number did not survive contact: InfoWorld measured well under 500 tokens/second and often under 100, against the &amp;ldquo;up to 2,000&amp;rdquo; claim.&lt;/li&gt;&#xA;&lt;li&gt;Undocumented throttles drove the experience: 300k TPM on Pro and 400k on Max produced 429 errors mid-session, and an early buyer reported a 7.5M-token daily cap hidden behind an advertised 1,000-request limit.&lt;/li&gt;&#xA;&lt;li&gt;131k context is about half the model&amp;rsquo;s native window and demands careful context management.&lt;/li&gt;&#xA;&lt;li&gt;At launch there was no prompt caching, which made agent loops expensive at the $2/1M API rate, and the usage console had defects (a Max purchase provisioning as Pro).&lt;/li&gt;&#xA;&lt;li&gt;Cerebras declined InfoWorld&amp;rsquo;s request for comment on these issues.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Pro costs $50/month with up to 24M tokens/day, and Max costs $200/month with up to 120M tokens/day.&#xA;Both plans were marked sold out as of 2026-09-26; a free tier with limited tokens remains for connection testing.&#xA;The underlying API price at launch was $2 per 1M input and $2 per 1M output on Qwen3-Coder.&#xA;The current per-token table on cerebras.ai/pricing renders client-side and I could not extract it, so treat current API rates as unverified.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Price history&#xA;    &lt;div id=&#34;price-history&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#price-history&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Date&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Plan&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Change&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Source&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2025-08-01&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Pro / Max&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Launched at $50/month (24M tokens/day) and $200/month (120M tokens/day)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Cerebras blog and HN thread&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2025-09-15&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Pro / Max&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;TPM caps (300k/400k) documented as the binding limit; prompt caching promised&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;InfoWorld review&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2025-10-28&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Pro / Max&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Caps reported improved; Qwen3 deprecated in favor of GLM-4.6 from November&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;InfoWorld follow-up&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-09-26&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Pro / Max&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Prices unchanged at $50/$200 but both marked sold out; model now GLM 4.7&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;cerebras.ai/code&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/chutes/&#34; &gt;- Chutes&lt;/a&gt; is the cheap multi-model pay-as-you-go option; choose Chutes for price and variety, Cerebras when seconds per response matter.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/glm-coding-plan/&#34; &gt;- GLM Coding Plan&lt;/a&gt; undercuts Cerebras on quota per dollar and now serves the same GLM family; choose it for volume, Cerebras for raw tokens per second.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Recommended for developers whose bottleneck is iteration latency and who code in bursts that fit the daily token allowance.&#xA;Not for heavy all-day agent runs, where the TPM throttles and 131k context cut the advertised advantage.&#xA;Claim to disagree with: at the documented throttles, Max at $200 was worse value than four Pro accounts, because 4x300k TPM beats 400k TPM.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-26 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/chutes/&#34; &gt;- Chutes&lt;/a&gt; - the low-price multi-model alternative.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/glm-coding-plan/&#34; &gt;- GLM Coding Plan&lt;/a&gt; - the quota-heavy alternative now running the same model family.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/cline/&#34; &gt;- Cline&lt;/a&gt; - the editor integration Cerebras documents first.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/opencode/&#34; &gt;- OpenCode&lt;/a&gt; - a terminal harness that consumes Cerebras keys.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-provider-feature-matrix/&#34; &gt;- Model provider feature matrix&lt;/a&gt; - where Cerebras Code sits among access providers.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.cerebras.ai/blog/introducing-cerebras-code&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.cerebras.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.cerebras.ai/blog/introducing-cerebras-code&lt;/a&gt; - launch post: $50/$200 tiers, 24M/120M tokens/day, 2,000 tok/s claim, 131k context (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.cerebras.ai/blog/qwen3-coder-480b-is-live-on-cerebras&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.cerebras.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.cerebras.ai/blog/qwen3-coder-480b-is-live-on-cerebras&lt;/a&gt; - Qwen3-Coder launch, $2/1M API rate, &amp;ldquo;20x higher coding speed&amp;rdquo; claim (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.cerebras.ai/code&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.cerebras.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.cerebras.ai/code&lt;/a&gt; - current product page: GLM 4.7 at 1,000+ tok/s, Pro and Max both marked &amp;ldquo;sold out&amp;rdquo; (fetched, HTTP 200, as of 2026-09-26).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://news.ycombinator.com/item?id=44762959&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=news.ycombinator.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://news.ycombinator.com/item?id=44762959&lt;/a&gt; - launch-day reception: 449 points, 172 comments, top comment flags missing caching (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.infoworld.com/article/4055909/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.infoworld.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.infoworld.com/article/4055909/&lt;/a&gt; - critical review, Sep 15, 2025: disputed tok/s claims, TPM caps, 131k context, billing mixups, no vendor comment (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.infoworld.com/article/4075825/how-to-vibe-code-for-free-or-almost-free.html&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.infoworld.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.infoworld.com/article/4075825/how-to-vibe-code-for-free-or-almost-free.html&lt;/a&gt; - follow-up: caps improved, Qwen3 deprecated for GLM-4.6, &amp;ldquo;fastest bar none&amp;rdquo; (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.reddit.com/r/LocalLLaMA/comments/1mfeazc/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.reddit.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.reddit.com/r/LocalLLaMA/comments/1mfeazc/&lt;/a&gt; - critical post, Aug 2, 2025: advertised 1,000 requests/day actually a 7.5M-token cap, &amp;ldquo;request&amp;rdquo; defined at ~8k tokens (retrieved, HTTP 200, via the Arctic Shift archive API).&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>Chutes</title>
      <link>https://tomrochette.com/agents/model-access/chutes/</link>
      <pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/model-access/chutes/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>model-access</category><category>inference</category><category>pay-as-you-go</category><category>bittensor</category>
      <description>&lt;p&gt;Chutes is a model-access platform built on Bittensor (subnet 64) that sells per-token inference on open-weight models, optional monthly plans, and private GPU deployments.&#xA;&lt;strong&gt;Its per-token prices were the lowest I verified in this category, and price is the main reason to pick it.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Chutes serves open-weight models (Qwen, DeepSeek, Kimi, GLM, Mistral, Gemma) from hosted instances and charges by the million tokens, with no subscription required.&#xA;Featured models run on confidential TEE compute, and billing accepts USD or TAO from a Bittensor wallet.&#xA;A CLI can also deploy a private chute on self-serve GPU capacity billed per second, with a one-time deployment fee scaled to the GPU choice.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active and shipping changes as of September 2026, with a live pricing page, status page, and regular community announcements.&#xA;2026 was a repair year for the business model: on February 27 the team retired the free 200-requests-per-day Early Access perk, capped subscriptions at 5x pay-as-you-go value, and cut frontier models from the $3 Base tier after its tables showed top users extracting up to 324x their subscription price.&#xA;A March 20 update reports tokens served down 45% over six weeks while revenue per million tokens rose 37.7% since February 1, after ending roughly 10B tokens/day of free-quota serving to about 11,000 users.&#xA;I could not verify funding or traffic figures.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Lowest per-token list prices I found this run, for example DeepSeek-V3.2 at $1.00/$1.00 per 1M in/out and Mistral-Nemo at $0.0245/$0.0978.&lt;/li&gt;&#xA;&lt;li&gt;Many model families behind one OpenAI-compatible key, so switching models costs nothing.&lt;/li&gt;&#xA;&lt;li&gt;Pay-as-you-go works with no plan attached, keeping commitment near zero.&lt;/li&gt;&#xA;&lt;li&gt;TEE enclaves on featured models, a privacy posture most cheap providers lack.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Terms changed three times in early 2026 (free tier removed, 5x cap, Base tier model cuts), so treat any quota as provisional.&lt;/li&gt;&#xA;&lt;li&gt;InfoWorld&amp;rsquo;s reviewer found performance underwhelming, called the privacy policy ambiguous, and needed a VPN workaround to sign up.&lt;/li&gt;&#xA;&lt;li&gt;The March 2026 update cites a global GPU shortage limiting inventory, so capacity depends on miner participation.&lt;/li&gt;&#xA;&lt;li&gt;The 6%/10% plan discounts apply only beyond the bundled daily quota, so heavy users are pushed back to full PAYG rates.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;PAYG is the default as of 2026-09-26: GLM-5.2 $1.25 in / $3.95 out per 1M, Kimi-K3 $3.00/$15.00, Qwen3-235B-A22B-Thinking $0.2989/$1.1957, DeepSeek-V3.2 $1.00/$1.00.&#xA;Plus costs $10/month for a bundled daily quota plus 6% off PAYG beyond it, and Pro costs $20/month for a larger quota plus 10% off.&#xA;Since February 27, 2026, every subscription is capped at 5x the equivalent PAYG value, with overflow billing at standard PAYG.&#xA;Private chutes run on an RTX Pro 6000 at $1.80/hour plus a $5.40 one-time deployment fee (3x the hourly rate).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Price history&#xA;    &lt;div id=&#34;price-history&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#price-history&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Date&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Plan&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Change&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Source&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2025-07-31&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Subscriptions&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Pricing tiers announced, &amp;ldquo;coming next Monday&amp;rdquo; (live early August 2025)&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Chutes news post, July 31, 2025&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-02&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Base / Plus / Pro&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Baseline: $3 / $10 / $20 per month tiers existed&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Feb 27, 2026 community announcement usage tables&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-02-27&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Early Access&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Free 200-requests/day perk cut off from TEE models, fully retired March 15, 2026&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;same announcement&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-02-27&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;All subscriptions&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Value capped at 5x PAYG equivalent; GLM-5, Kimi K2.5, Qwen 3.5, MiniMax M2.5 removed from Base&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;same announcement&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;2026-09-26&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Base&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;No longer listed on the pricing page, which shows only Plus $10, Pro $20, and Enterprise&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;chutes.ai/pricing&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/cerebras-code/&#34; &gt;- Cerebras Code&lt;/a&gt; sells raw speed on one coding model at $50/$200; choose it when latency dominates, Chutes when price dominates.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/glm-coding-plan/&#34; &gt;- GLM Coding Plan&lt;/a&gt; sells a large fixed quota on one model family; choose it if you will burn the quota, Chutes for variety without commitment.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Recommended for hobbyists and cost-driven developers who want many open models cheaply and can tolerate churn in terms and capacity.&#xA;Not for teams needing SLAs, contractual stability, or closed frontier models.&#xA;Claim to disagree with: after the 5x cap, Plus and Pro make little sense for coders, because PAYG with a small prepaid balance delivers the same upside without a recurring bill.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-09-26 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/cerebras-code/&#34; &gt;- Cerebras Code&lt;/a&gt; - the speed-first subscription alternative.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-access/glm-coding-plan/&#34; &gt;- GLM Coding Plan&lt;/a&gt; - the fixed-quota rival that wins on coding throughput per dollar.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/trackers-and-leaderboards/openrouter-rankings/&#34; &gt;- OpenRouter rankings&lt;/a&gt; - where provider model traffic becomes visible.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/model-provider-feature-matrix/&#34; &gt;- Model provider feature matrix&lt;/a&gt; - how access providers compare feature by feature.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://chutes.ai/pricing&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=chutes.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://chutes.ai/pricing&lt;/a&gt; - current PAYG model rates, Plus $10 / Pro $20 with 6%/10% PAYG discounts, private GPU pricing (fetched, HTTP 200, as of 2026-09-26).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://chutes.ai/news/community-announcement-february&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=chutes.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://chutes.ai/news/community-announcement-february&lt;/a&gt; - February 27, 2026 changes: Early Access retirement, 5x subscription cap, Base tier model removals, abuse tables (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://chutes.ai/news/from-volume-to-value-building-a-sustainable-ai-inference-platform-2&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=chutes.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://chutes.ai/news/from-volume-to-value-building-a-sustainable-ai-inference-platform-2&lt;/a&gt; - March 20, 2026 economics: tokens down 45%, revenue per token up 37.7%, free-tier costs (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://chutes.ai/news/coming-soon&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=chutes.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://chutes.ai/news/coming-soon&lt;/a&gt; - July 31, 2025 post announcing the first pricing tiers (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://chutes.ai/docs/cli/deploy&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=chutes.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://chutes.ai/docs/cli/deploy&lt;/a&gt; - private chute deployment mechanics and deployment fee structure (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.infoworld.com/article/4075825/how-to-vibe-code-for-free-or-almost-free.html&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.infoworld.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.infoworld.com/article/4075825/how-to-vibe-code-for-free-or-almost-free.html&lt;/a&gt; - critical third-party review: underwhelming performance, ambiguous privacy policy, signup bugs; confirms the $3 entry tier in October 2025 (fetched, HTTP 200).&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
