<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>paper-generation on tomrochette.com</title>
    <link>https://tomrochette.com/tags/paper-generation/</link>
    <description>Recent content in paper-generation on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Sun, 11 Oct 2026 02:01:18 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/paper-generation/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>AutoResearchClaw</title>
      <link>https://tomrochette.com/agents/automated-research/autoresearchclaw/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/automated-research/autoresearchclaw/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>automated-research</category><category>agentic-llm</category><category>open-source</category><category>paper-generation</category>
      <description>&lt;p&gt;AutoResearchClaw is the MIT-licensed 23-stage pipeline from the aiming-lab organization that turns a one-line research idea into a compile-ready paper, with sandbox experiments, multi-agent debate, a four-layer citation-verification layer, and seven human-in-the-loop modes.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;AutoResearchClaw is the idea-to-paper pipeline that treats failure and fabrication as first-class problems: its executor heals through Pivot/Refine loops, its citations pass arXiv, CrossRef, DataCite, and LLM checks before delivery, and its 54.7 percent win over AI Scientist v2 is measured on its own ARC-Bench, a self-run benchmark the README has since widened from 25 to 55 topics.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A Python pipeline (created 2026-03-15) run from the &lt;code&gt;researchclaw&lt;/code&gt; CLI, standalone, through an OpenClaw bridge to Discord, Telegram, Lark, or WeChat, or on any ACP-compatible agent backend (Claude Code, Codex CLI, Copilot CLI, Gemini CLI, Kimi CLI).&#xA;The arXiv paper (2605.20025, by Jiaqi Liu, Shi Qiu, Mairui Li, Bingzhou Li, Haonian Ji, Siwei Han, Xinyu Ye, and Peng Xia) presents five mechanisms: structured multi-agent debate, a self-healing executor with a Pivot/Refine decision loop, verifiable result reporting against fabricated numbers and hallucinated citations, human-in-the-loop collaboration across seven intervention modes, and cross-run evolution that converts past failures into future safeguards.&#xA;Deliverables per run: a paper draft, conference-ready LaTeX (NeurIPS, ICML, ICLR templates), a BibTeX file with references pulled from OpenAlex, Semantic Scholar, and arXiv, a verification report, sandbox experiment code and metrics, charts, multi-agent reviews, and evolution lessons.&#xA;MetaClaw adds cross-run learning (pipeline failures become structured lessons injected into later runs), and the skills library loads 20 preloaded skills plus community contributions.&#xA;The companion ARC-Bench dataset ships on Hugging Face under the AIMING-Lab-UNC account, widened at v0.5.0 (May 2026) to 55 topics across machine learning, high-energy physics, quantum, biology, and statistics.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Dormant since August: 14,618 stars, 1,706 forks, created 2026-03-15, last push 2026-08-19, latest release v0.5.0 on 2026-05-20, as of 2026-10-10.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;&lt;strong&gt;The repository has gone quiet while attention held: no commits in eight weeks and no release in five months as of 2026-10-10, a Hacker News footprint of two stories at two and one points with zero comments, and the eight showcase papers are self-published, so the 14.6k stars currently measure attention the repository is no longer feeding.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Citation integrity is a design point, not an afterthought: the four-layer check (arXiv, CrossRef, DataCite, LLM) kills fabricated references before delivery, the failure class the no-kernel columns are most exposed to.&lt;/li&gt;&#xA;&lt;li&gt;Failure handling is architectural: Pivot/Refine turns failed experiments into information, and MetaClaw accumulates them into reusable skills.&lt;/li&gt;&#xA;&lt;li&gt;The seven-mode human-oversight ladder (full-auto through step-by-step) is the widest in this category, an answer to the trust problem rather than a denial of it.&lt;/li&gt;&#xA;&lt;li&gt;ARC-Bench gives the loop a rubric-scored eval set spanning five domains instead of a single ML sandbox.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The 54.7 percent improvement over AI Scientist v2 is self-reported on the project&amp;rsquo;s own benchmark, with no independent replication surfaced in this run&amp;rsquo;s searches.&lt;/li&gt;&#xA;&lt;li&gt;Development is quiet: last push 2026-08-19 and last release 2026-05-20, which is a poor sign for a 23-stage pipeline whose stages depend on external APIs and agent backends.&lt;/li&gt;&#xA;&lt;li&gt;The two zero-comment Hacker News stories say there is no practitioner debate to check the claims against.&lt;/li&gt;&#xA;&lt;li&gt;A 23-stage pipeline is heavy to adopt and heavier to fork: the runtime surface (LLM providers, Docker, LaTeX, five agent backends) is wide.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and MIT; there is no paid tier, so pricing does not apply.&#xA;Costs are your own LLM API keys and compute (a Docker executor with GPU, MPS, or CPU auto-detection).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt;: the other idea-to-paper column; Agon is a Claude Code plugin with adversarial critics and a failure taxonomy, AutoResearchClaw a standalone pipeline whose added layer is citation verification and a human-oversight ladder.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/aris/&#34; &gt;ARIS&lt;/a&gt;: the markdown-skill loop ported across harnesses with cross-model review gates; AutoResearchClaw is a fixed 23-stage pipeline you run rather than a method you install.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt;: the cited-report baseline with no experiments; AutoResearchClaw runs sandbox experiments and ships LaTeX.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for studying how an autonomous pipeline can make citation fabrication and silent failure first-class problems, and for teams that want idea-to-paper runs with explicit human-oversight modes.&lt;/strong&gt;&#xA;Not for anyone who needs an actively maintained tool (eight weeks without a commit) or independently verified benchmark claims.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt; - the producer-critic idea-to-paper column and the Agon paper&amp;rsquo;s comparator&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/aris/&#34; &gt;ARIS&lt;/a&gt; - the skills-based cross-model loop in the same family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt; - the cited-report baseline AutoResearchClaw extends to experiments&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/automated-research-feature-matrix/&#34; &gt;Automated Research Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/aiming-lab/AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/aiming-lab/AutoResearchClaw&lt;/a&gt; - repository and README: 23-stage pipeline, deliverables, OpenClaw and ACP-backend support, release news (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/aiming-lab/AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/aiming-lab/AutoResearchClaw&lt;/a&gt; - stars, forks, created and pushed dates, MIT license, and topics for the as-of status (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/aiming-lab/AutoResearchClaw/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/aiming-lab/AutoResearchClaw/main/README.md&lt;/a&gt; - the README source: stage outputs, citation-verification layers, HITL system, ARC-Bench v0.5.0 widening (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2605.20025&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2605.20025&lt;/a&gt; - the paper: five mechanisms, the 54.7 percent AI Scientist v2 comparison on ARC-Bench, and the seven intervention modes (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/aiming-lab/AutoResearchClaw/releases?per_page=5&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/aiming-lab/AutoResearchClaw/releases?per_page=5&lt;/a&gt; - the release train through v0.5.0 (2026-05-20) (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=AutoResearchClaw&lt;/a&gt; - the two zero-comment Hacker News stories behind the discussion-footprint claim (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench&lt;/a&gt; - the ARC-Bench dataset page under the AIMING-Lab-UNC account (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
