<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>training-free on tomrochette.com</title>
    <link>https://tomrochette.com/tags/training-free/</link>
    <description>Recent content in training-free on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Wed, 07 Oct 2026 06:38:49 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/training-free/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>AnyJev</title>
      <link>https://tomrochette.com/agents/hybrid-execution/anyjev/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/hybrid-execution/anyjev/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>hybrid-execution</category><category>decision-models</category><category>training-free</category>
      <description>&lt;p&gt;AnyJev is Nokia Applied Research&amp;rsquo;s Apache-2.0 method that turns any instruction-tuned LLM into a Jev-style decision model: it reads a typed choice as probabilities straight from the logits, then corrects the two biases that readout suffers, no gradient steps and no generated text.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;AnyJev is the training-free entry in the decision-model wave, and its arXiv-backed claim inverts the family&amp;rsquo;s economics: instead of paying for a fine-tune or a specialist checkpoint, you pay a few extra prefills per decision, spent only where a stopping rule says the calibration needs them.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A pip-installable Python method plus a paper (arXiv 2610.00831, submitted 2026-09-30, authors from Nokia, Tencent Hunyuan, and Hong Kong Polytechnic University) from nokia-applied-research (1,092 stars, pushed 2026-10-07, as of 2026-10-07).&#xA;The raw readout restricts the next-token distribution at the answer position to the option tokens; AnyJev then corrects its two documented defects by dividing out a label prior estimated from unlabelled inputs and averaging log-probabilities over the K cyclic rotations of the option list, which lowers the order-flip rate from 0.33 to 0.14 and 0.18 on two 20-option tasks and raises accuracy on 11 of 11 models tested.&#xA;A stopping rule selected on unlabelled splits cuts the rotation count: 10.6 of 18 rotations at a verified 0.008 disagreement bound on two of four cells, or 7.3 when selected on one split, which served 2.2 times as many decisions per second on vLLM.&#xA;The same team publishes Tacit, a self-distilled 1.7B-to-9B model line on Hugging Face that compresses the method to one forward pass per decision, with an adaptive mode routing a capped share of low-confidence decisions back to the base model&amp;rsquo;s reasoning.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active and freshly credible: created 2026-09-21, 1,092 stars, release v0.3.0 published 2026-10-07, as of 2026-10-07.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=nokia-applied-research/AnyJev&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=nokia-applied-research/AnyJev&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=nokia-applied-research/AnyJev&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;This note reverses this section&amp;rsquo;s 2026-09-25 rejection of the same project, which then had 506 stars and no independent coverage; the promised re-review trigger, a second implementation, now exists in kind.&#xA;The third-party local-jev-bench benchmark (2026-09-30, updated 2026-10-05) runs AnyJev against Kev, Winnow, Clef, Von, Jeff, Laya, and CLM on Apple Silicon and reports training-free AnyJev level with the trained engines on the six public transfer sources; community builds such as mesa-anyjev and a Tetris-playing AnyJev judge extend it.&#xA;The HN footprint remains thin (mentions inside decision-model threads, no thread of its own), which this note records as the standing signal.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;No training: any instruction-tuned model you already serve becomes a decision engine, which removes the fine-tune cost and the frozen-weights staleness the rest of the wave accepts.&lt;/li&gt;&#xA;&lt;li&gt;The corrections are principled and measured, with the rotation-averaging and prior-division effects reported across 11 of 11 models and the early-exit bound verified rather than asserted.&lt;/li&gt;&#xA;&lt;li&gt;The calibration pipeline (thresholds selected and bounded on unlabelled splits) addresses the abstention question most replicas leave open.&lt;/li&gt;&#xA;&lt;li&gt;Apache-2.0 code, a peer-readable technical report, and PyPI packaging make it the wave&amp;rsquo;s most reproducible-looking method entry.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Rotation reads cost prefills: the fair cost comparison is against trained single-pass models, and the stopping-rule numbers are per-cell, not a universal 7.3.&lt;/li&gt;&#xA;&lt;li&gt;Independent benchmark coverage exists (local-jev-bench) but is one hobbyist-scale suite; no lab replication yet, as of 2026-10-07.&lt;/li&gt;&#xA;&lt;li&gt;The project is three weeks old, so API stability and the Tacit line&amp;rsquo;s licenses are early-grade.&lt;/li&gt;&#xA;&lt;li&gt;The vendor is a telecom research lab, and the wave has already seen research artifacts go quiet after launch; the community builds are promising but 0-star so far.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open source under Apache-2.0; the Tacit checkpoints are published on Hugging Face, so pricing does not apply.&#xA;Costs are your own model serving and the extra prefills the rotations or early-exit read.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt;: the closed contract that defined the category, single-pass and priced per million tokens; AnyJev is the open, bring-your-own-model alternative with prefill-based costs.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt;: the strongest open trained family; choose AnyJev to avoid training entirely, Kev when a single-pass fine-tune beats multi-prefill reads on your latency budget.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/semif/&#34; &gt;SemIf&lt;/a&gt;: the other logit-readout entrant over frozen models, without AnyJev&amp;rsquo;s bias corrections or stopping rule.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for teams already serving an instruction-tuned model who need typed, calibrated decisions and cannot justify training or a specialist checkpoint, starting with the early-exit configuration and validating calibration on your own label splits.&lt;/strong&gt;&#xA;Not for single-digit-millisecond budgets, where the trained single-pass models keep the edge, and not for production trust before an independent lab replicates the paper&amp;rsquo;s bounds.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-07 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/jev/&#34; &gt;Jev&lt;/a&gt; - the closed contract AnyJev replicates without training&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/kev/&#34; &gt;Kev&lt;/a&gt; - the trained open alternative with the strongest eval discipline&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/semif/&#34; &gt;SemIf&lt;/a&gt; - the frozen-model logit readout without the bias corrections&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/ollaya/&#34; &gt;Ollaya&lt;/a&gt; - the runtime that could serve a Tacit checkpoint beside the trained families&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/hybrid-execution/hybrid-execution-feature-matrix/&#34; &gt;Hybrid Execution Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/nokia-applied-research/AnyJev&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/nokia-applied-research/AnyJev&lt;/a&gt; - repository, Apache-2.0 license, the readout and correction design, and the Tacit line (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/nokia-applied-research/AnyJev&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/nokia-applied-research/AnyJev&lt;/a&gt; - stars, forks, created date, and push date for the as-of status (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2610.00831&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2610.00831&lt;/a&gt; - the technical report: rotation-averaging and prior-division results, the early-exit bounds, and the vLLM serving numbers (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/nokia-applied-research/AnyJev/releases&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/nokia-applied-research/AnyJev/releases&lt;/a&gt; - the v0.0.2 through v0.3.0 releases, v0.3.0 on 2026-10-07 (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/tak-bro/local-jev-bench/HEAD/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/tak-bro/local-jev-bench/HEAD/README.md&lt;/a&gt; - the third-party Apple Silicon benchmark running AnyJev against Kev, Winnow, Clef, Von, Jeff, Laya, and CLM (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/collections/morriszjm/tacit-6ac41d0b50af9e5417c5c234&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/collections/morriszjm/tacit-6ac41d0b50af9e5417c5c234&lt;/a&gt; - the Tacit collection: five checkpoints (1.7B to 9B) self-distilled on Qwen bases, one forward pass per decision (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/idss-mesa/mesa-anyjev&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/idss-mesa/mesa-anyjev&lt;/a&gt; - a second implementation of the method, calibrated ontology and schema decisions for the MESA stack, part of the re-review evidence (fetched 200, 2026-10-07)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
