↓ Skip to main content

Specifications Are Lossy Compression

Every document we write before the code is a compressed copy of the code, and the compression drops information. A requirement, a user story, a specification, a BDD scenario, and a unit test are each a smaller description of a larger artifact. None of them can reproduce the code they describe, so all of them lose something, and an LLM agent fills the missing part with a confident guess. The useful question is not how to write a better description. It is which ways of defining software can carry the whole program, and where the missing part should go for the ones that cannot.

Lossy is a precise word
#

The idea of measuring an artifact by the shortest description that reproduces it comes from Kolmogorov complexity. The Kolmogorov complexity of a string is the length of the shortest program that outputs it, and a description is lossless when the original can be reconstructed from it exactly, the way a decompressor reconstructs a file. A definition of software is lossless when a machine can decide, without a person, whether any candidate implementation is acceptable. Reconstruction is the stronger test, since the definition would have to determine exactly one program. Decidability is the test that matters, since it lets a checker accept every correct program and reject the rest.

A program carries two kinds of information, and only one of them is about behavior. The first is what the program must do, which is its inputs, its outputs, and the rules that connect them. The second is how it does it, which is the data structures, the order of operations, the split into modules, the names, and the tradeoffs between speed and memory. A prose artifact tries to carry the first kind and drops the second almost entirely.

The second loss is larger than the first. Prose states behavior in natural language, and the meaning of a sentence depends on the reader. Two readers, or one reader on two days, will turn “the service should stay available during a deploy” into different code, because the sentence never fixed what available means, how long a deploy may take, or what happens to requests already in flight. A natural-language specification is lossy twice. It drops the implementation choices, and it leaves the behavior it does state open to interpretation.

Why the loss stopped being benign
#

For most of software’s history the reader of a requirement was a person, and by reader I mean whoever turns the definition into code. The person absorbed the loss. A person carried context between documents, inferred the unstated from experience, and asked a colleague when a sentence was ambiguous. The loss was worth accepting, because the thing it saved was the expensive one, which was writing the code. Compressing a large program into a small document saved the effort of producing the program.

Both conditions have changed. The reader is now an agent that begins with no history, so anything the definition does not state is absent, and the agent fills the absence with a plausible guess instead of a question. The cost of producing code has also fallen far enough that compressing it to save writing no longer pays. The expensive artifact is now the verified definition, and the document that used to save work has become the place where the work is lost.

The Prompt Is the Source Code described this failure from the maintenance side, where the code reaches the maintainer without the intent that produced it and the next maintainer has to reconstruct that intent. This piece locates the same failure one step earlier, in the definition itself, and asks what a definition would have to be to keep the information rather than lose it. Keeping the prompt is necessary but not sufficient, because the prompt is the best available record of the intent and the checker is what turns that record into a lossless definition.

No set of words will be lossless
#

The search for a perfect prose specification has no solution. By the definition of Kolmogorov complexity, the shortest description of an object that cannot be compressed is about as long as the object, so a natural-language description that determined a nontrivial program would be at least as large as the program and no easier to read. A shorter document always leaves something out, and the thing it leaves out is the behavior and the choices that made the program complex.

A picture of the recovery sense of loss makes the gap concrete. It shows how much of the program each form brings back, which is the intuitive sense of compression.

Six horizontal bars of equal total length, labeled from top to bottom: the code, executable specification, checked specification, property test, example test, and prose requirement. The solid blue part of each bar is what the definition brings back and it shrinks from top to bottom; the faded remainder is what the reader supplies.

The code bar is full here because the picture measures the artifact, the behavior and the choices together. What the code itself leaves out is the reason it exists, which the picture does not measure. A checkable definition is also stronger than this picture suggests, because it is lossless in the decidability sense even when it recovers only part of the implementation.

If the definition cannot be a shorter description in another language, it has to be one of three things. It can be the same object, which is the code itself or something a compiler turns into the code. It can be a statement a machine can check against the code. Or it can be a record of the choices the first two cannot carry. Three ways cover the first two, and a fourth covers the residue.

Four ways to remove the loss
#

Make the definition the source
#

The first way is to stop describing the code and generate it. A domain-specific language, a schema, or a model is the source of truth, and a compiler turns it into the artifact, so no separate description exists to lose anything. Literate programming is the older form, where Knuth’s tooling tangles a document into compilable code and weaves it into readable documentation from one source, so the two cannot drift apart. The loss is zero for everything the language can express. The cost is that you maintain the compiler and the language, and anything the language cannot express falls outside the definition.

Make the definition checkable
#

The second way is to write the definition in a form a checker can decide against, even when a person still reads it. A refinement calculus specification is a program written in a nondeterministic language, and the executable code is a refinement of it, produced by steps that preserve the specification’s behavior. Tools in this family, such as Dafny with its Hoare-style contracts and the B-method built on refinement, let you state preconditions, postconditions, and invariants, then either prove the code satisfies them or report the proof obligations you still owe. The definition is lossless about the properties it names and silent about the rest, which is the same relativity intent-driven development describes when it says intent is lossless only for the questions it answers.

The type is the smallest checkable definition. Under the Curry-Howard correspondence, a type is a proposition and a program is a proof of it, so the compiler checks that the program satisfies the claims the type makes. The correspondence is exact in a total language such as Coq or Agda. In a Turing-complete language general recursion lets a program inhabit almost any type, so the compiler checks type soundness rather than a theorem. Dependent and refinement types push the claim closer to the behavior you care about, which is why moving a rule into the type so that a wrong value cannot be constructed is a definitional move rather than a stylistic one.

Make the definition generative
#

The third way is to have a procedure construct the code from the definition. Deductive program synthesis, as in the Manna and Waldinger framework, treats the specification as a theorem and derives a program from its proof, so the result is correct by construction. Modern synthesizers improve the search rather than the idea. Syntax-guided synthesis constrains the space with a grammar, and counterexample-guided synthesis alternates a generator with a verifier until the verifier runs out of counterexamples. The loss is zero for the synthesized part. The cost is that synthesis does not scale to arbitrary systems, so it applies to small, well-specified pieces rather than whole products.

Record the residue
#

The fourth way accepts that some information cannot follow from behavior at all, and writes it down on purpose. The choice of a data structure, the decision to trade latency for throughput, and the reason a boundary was drawn where it was do not follow from what the program must do, so no behavioral definition can carry them. They have to be recorded as decisions with the reasoning attached, because the next agent that meets the same fork will decide it again and may decide it differently. This is the channel that carries what the specification cannot, which is why a decision record is part of a definition rather than a note about one.

The four ways are layers rather than a menu. The definition is lossless wherever at least one of them applies. The generated layer produces exact artifacts, the checked layer decides stated properties, and the recorded layer preserves named choices. What is left over is genuinely free, and marking it free is part of the definition.

A two by two grid. The horizontal axis runs from interpreted by a reader to decided by a machine, and the vertical axis runs from describes the code to produces the code. Requirements and BDD scenarios sit in the interpreted, describes quadrant. A prompt without a checker sits in the interpreted, produces quadrant. Types, checked specifications, and property tests sit in the machine, describes quadrant. An executable DSL and synthesis sit in the machine, produces quadrant. Only the machine column is lossless.

The agent is a synthesizer, not an oracle
#

An LLM can work at every layer above. It can compile a DSL into code, draft a refinement with proof obligations, synthesize from a specification, and apply recorded decisions to a new case. It cannot make the definition lossless on its own, because losslessness is a property of the checker rather than the writer. The model produces candidates, the checker decides which candidate is acceptable, and the loop is counterexample-guided synthesis with a model as the generator.

This reframes what to build first. The instinct is to spend the effort on the prompt and the description, because that is what the model reads, and to treat the checker as a later quality step. That order is backwards. Removing the checker leaves a lossy prose definition with a confident reader, which is the situation the model makes worse rather than better. A definition is lossless only if it can reject the model’s output, so the checker is the first thing to build and the prompt is the second.

Rethinking Code Review in the Age of LLMs reached the same conclusion from the review side, and When Agents Solve Problems You Cannot Check reached it from the verification side. When implementation is cheap, the scarce resource is a way to tell a correct implementation from a plausible one, and that way is the definition.

Choosing the strongest form you can afford
#

Not every project can carry a refinement proof, and the point is not that it should. For each property you care about, there is a strongest available form, and the forms sit on a ladder.

Form of the definition What it decides What it drops Cost
Prose requirement Nothing, a person decides Behavior detail, every implementation choice Low to write, high to interpret
User story Nothing, a person decides Behavior, edge cases, choices Low
BDD scenario Whether the code passes the named scenario Everything the scenario does not name Moderate
Example test Whether the code passes the example General behavior, unstated cases Low
Property test Whether the property holds across generated inputs Properties not stated, the choices Moderate
Executable specification The artifact, exactly Only what the language cannot express High upfront, no drift
Refinement or proof Whether the code satisfies the stated property Whatever the property leaves free Very high
Decision record What was chosen and why Nothing recorded, whatever is left unwritten Moderate
The code The artifact itself The reason it exists The artifact itself

Read the table as rungs rather than a ranking of teams. Push each property as far down the ladder as the budget allows, from a prose sentence to a scenario, from a scenario to a property, from a property to a checked specification, and from a checked specification to an executable one. Then move the choices the ladder cannot carry into decision records. The prose does not disappear, since it becomes a view generated from the definition for human readers, and a bidirectional transformation view is the formal version of a view that can be written back without loss.

What to Do Next
#

  • Name the properties you care about, because a definition can only be lossless relative to a stated equivalence.
  • For each property, use the strongest form you can afford, choosing executable over checkable, checkable over property-tested, property-tested over example-tested, and example-tested over prose.
  • Build the checker first and the prompt second, so the definition can reject the model’s output rather than only request it.
  • Mark the choices that are free by design, and record every other choice as a decision with its reasoning.
  • Treat requirements and specifications as views generated from the definition, so the two cannot drift.

The habit worth breaking is treating a lossy artifact as a complete one. A requirement is a lossy projection of the code rather than a smaller copy of it, and the projection always omitted the part the reader supplied. Where the reader no longer asks, the definition has to supply that part, and the only definitions that can are the ones a machine can run or check.

See also
#

References
#


Metrics for a Software Factory: Optimize Autonomy, Guard the Trust

A software factory is the infrastructure that lets LLM agents carry a change through the whole software lifecycle, from an observed signal to a shipped fix, instead of handing one step to a model and the rest to me. Once the factory exists, the question that decides whether it was worth building is how much of that lifecycle runs without me. The metrics I put on it decide what it optimizes for, and the obvious ones (agents spawned, tokens burned, pull requests merged, lines written) are the ones an agent can increase without producing anything I can use.

Output metrics measure the wrong thing
#

Every measurement I inherited from human development counts effort. Tickets closed, story points burned, pull requests merged, lines written: each one counted human activity because human activity was the expensive part, which is the argument I made in Outcome-Driven Development. An agent produces all of it at almost no cost. Point one at a backlog and it will empty the backlog, open fifty pull requests before lunch, and every one of them will look like a morning of work.

This is Goodhart’s law at its most extreme. When a measure becomes a target it stops being a good measure, and an agent can turn any output metric into a target within a single run. The failure is not theoretical. Armin Ronacher’s 35 hours with GPT-6 Astra produced 79 commits, 75,000 lines, about a billion tokens, and roughly $1,200 of API spend over a weekend. By volume alone it looked like his most productive weekend ever. His verdict was that it delivered nothing of value.

The problem is not only that volume misleads. It is that volume looks like progress to the person running the factory too. METR’s randomized trial found that experienced developers using early-2025 AI tools took 19% longer on their own tasks, and those same developers still believed the tools had made them 20% faster. When measured time and perceived speed disagree about the same work, self-reported productivity measures a feeling, not a result.

The first step in building a factory is therefore to retire the output metrics, because an agent can increase every one of them for free.

Measure the factory the way a plant is measured
#

Manufacturing has measured itself for a century and almost never treats gross output as success. A plant that ships a thousand units at 10% yield produces 900 units nobody can use, and what matters is good units per unit of input, not units. Yield is the fraction of what enters a process that comes out usable on the first pass. A software factory has the same two numbers with different names.

Autonomy rate is the share of accepted changes that reach production with no human intervention: no question answered mid-run, no correction, no manual fix, no human-triggered rerun, and no approval. Yield is the share of the autonomous output that survives: it passes verification, stays merged, and remains valid during the observation window after release. Multiply the two and you get the number the factory exists to raise, the rate of trustworthy unattended work.

Autonomy is easy to raise without doing the work, so the definition has to be strict about what counts as intervention. A change that needed one clarification at hour three is not autonomous, and counting it as autonomous is how a dashboard reaches 90% while I am still involved in every hard case.

A signal is an observed event that starts work: a bug report, a failing test, a support ticket, a security alert, or a request from a customer. The loop below starts from that signal and shows where each metric is read.

flowchart LR
    S[Signal] --> P[Produce]
    P --> V{Independent verification}
    V -->|fails| P
    V -->|passes| G{Accept}
    G -->|human needed| H[Human gate]
    G -->|no human needed| M[Accepted change]
    H --> M
    M --> O[Observe in production]
    O -->|regression| R[Rework]
    R --> P

Autonomy counts the changes that travel from Produce to Accepted change without passing through the Human gate. Yield counts the accepted changes that survive Observe. Cost divides the whole path by the accepted changes, which is why the word “accepted” matters.

Armin priced his run at about $15.50 per commit. A commit is output, not a usable unit, so that price counts work whether or not it survived. The denominator that survives is the accepted change: the change that passed verification and held up in production, not the change that got written. Cost per commit counts output. Cost per accepted change counts results.

The guardrail that keeps autonomy trustworthy
#

Autonomy is easy to raise by shipping worse work, so it cannot be the only metric. The DORA program spent years showing that delivery performance has two independent axes, throughput and stability, and that a team can improve one while the other gets worse. Stability is change failure rate and time to restore, and in a factory it also includes defect escape rate, the defects that reach production past the gates, and rework rate, the accepted changes that had to be redone.

The failure mode when stability goes unmeasured is well documented. Addy Osmani documents a “dark” factory that shipped for about four months with no human reading the code, passed its own tests the whole way, and then needed painstaking manual debugging to recover. The tests passed because the factory wrote both the code and the tests, and the loop had no independent signal to catch what that matching pair got wrong together. A factory with high autonomy and low yield produces unverified volume, and it accumulates comprehension debt faster than anyone can repay it.

The two metrics are not interchangeable, and the grid shows why.

A two by two grid of autonomy against yield, with quadrants for a manual shop, unverified volume, a careful assistant, and a software factory

The rule that keeps a factory out of the bottom-right quadrant is to pair the metrics. Never set an autonomy target without a stability guardrail, because the cheapest way to raise autonomy is to lower the standard for done. DORA’s 2025 report found that 90% of the technology professionals it surveyed now use AI and more than 80% believe it made them more productive, while 30% report little or no trust in the code it generates, which is the gap between the two metrics stated as a statistic.

The two numbers that decide whether the factory is worth running
#

Once autonomy and yield are high enough, two more numbers decide whether the factory is a good use of money.

Throughput comes first: lead time from signal to production, and deployment frequency. The DORA speed axis still applies, but in a factory it is a consequence, not a goal. Raising throughput without raising verification capacity only grows the queue in front of the gate, which is the back-pressure problem Osmani names: generation runs without limit while verification stays slow, so the extra output becomes waiting, not more usable code. Track queue depth alongside throughput, because a growing backlog of unverified changes is the first sign that the factory produces faster than it can check.

Unit cost comes second, and it has two parts. The metered part is tokens and compute divided by accepted changes. The part that matters more is human minutes per accepted change, because the point of the factory is to remove my minutes, and a factory that ships cheap code while I spend the same hours supervising it has automated only part of the work. Armin’s run is the cautionary version of both: a weekend of compute for a yield near zero, with nobody supervising because supervising nobody was the point.

The leading indicators you can control
#

Outcome metrics lag by weeks. Autonomy and yield tell me where I ended up, not what to do on Monday, so the OKR needs a layer of inputs that change first. Four inputs predict the outcome metrics well enough to act on.

  • Spec coverage is the share of work items that carry a machine-checkable outcome (what becomes true, how it is checked, what it may cost) rather than a task description.
  • Gate coverage is the share of change types covered by an independent machine check, as opposed to a check the producing agent wrote for itself.
  • Context readiness is whether a cold session dropped into the project can find everything it needs, the exit condition I use in Nine Months of LLM Agents on Large Projects.
  • Failure-to-gate capture is the share of escaped defects that become a new automated check, the only mechanism that makes a factory improve instead of merely run.

These are the inputs that turn model capability into stable throughput. It is the same argument I made in Team Maturity Explains the Friction, the Foundation Predicts the House of Cards: verification infrastructure is the highest-return investment for a team shipping with agents, and in a factory it separates a loop that improves from one that only repeats.

The layers stack like this, with the guardrails holding the two primary metrics in place.

Four metric layers under one objective: autonomy and yield on top, throughput and unit cost below them, a guardrail band that must be held flat, and the leading indicators that steer the rest

Writing the OKR
#

An OKR made of autonomy numbers alone will produce the wrong behavior before the quarter ends. The version that works pairs one key result that moves against another that holds, and adds one that steers.

One workable objective, with the numbers as placeholders until you have your own baseline, is shift more of the software lifecycle into trustworthy unattended work.

  • Move: touchless share of accepted changes rises from 20% to 50%.
  • Move: human minutes per accepted change fall from 45 to 20.
  • Hold: defect escape rate and change failure rate stay at or below the pre-factory baseline.
  • Steer: the share of work items carrying a machine-checkable outcome rises from 30% to 90%.

The touchless-share result is the most visible and the easiest to raise without doing the work, which is why it is not alone. The human-minutes result is the one that matters most, and it is hard to lower without removing the work, because it measures my minutes. The stability result keeps the factory from trading trust for autonomy, and it cannot be met by shipping worse work. The spec-coverage result is the input I can change this week, and it makes the other three possible.

Write the baseline down before the first factory run. Without a baseline, anyone who preferred the old process can dispute the comparison, and the argument becomes about the numbers rather than the result.

One metric sits above the factory, and it is not a factory metric: the product outcome. If the product outcome is flat while autonomy climbs, the factory is producing the wrong thing faster. The factory metrics multiply a product metric, they never replace it.

The Goodhart problem, and how to limit it
#

Autonomy rate is now a target, so the system and the people running it will find the cheapest work that satisfies it. Trivial changes, tiny diffs, safe files, and a narrow definition of intervention all raise the number without improving the factory. Five rules are worth adding from the start.

  1. Report autonomy by change class, because a touchless migration and a touchless typo fix are not the same result.
  2. Measure on production traffic rather than a curated set, because a curated set misses the rare cases where the factory fails.
  3. Set a cap on trivial work, because an autonomy score built from work that never needed verifying is not a factory score.
  4. Keep a human-owned roadmap and a written invariant list, since an autonomous loop optimizes whatever its signals measure and only a human-edited direction corrects the drift, the argument in The Self-Evolving Repository.
  5. Audit the gates themselves, because a gate whose value was never measured only adds time, and The Merge Gate asks whether a gate is worth its cost.

What not to measure
#

  • Number of agents, sessions, or concurrent workers, since more agents producing the same output is a more expensive factory.
  • Tokens consumed, since it is an input and the factory’s job is to spend fewer per accepted change.
  • Pull requests merged and lines written, since volume is what an agent gets for free.
  • Story points and velocity, since they measure effort from an era when effort was scarce.
  • AI adoption percentage, since using a tool is not an outcome.
  • Self-reported speedup, since it contradicted the measured time in the METR trial.
  • Benchmark scores, since they measure the model rather than the factory, and your repository’s cost and yield are the benchmark that matters.

What to Do Next
#

  1. Instrument the change lifecycle with an id, a start, an end, and a human-intervened flag on every change, because nothing else works until this exists.
  2. Compute two numbers this week, the touchless share of accepted changes and the defect escape rate, by hand from the last fifty changes if you have to.
  3. Add the two economic numbers once ids exist: cost per accepted change, and human minutes per accepted change.
  4. Write one paired OKR: move autonomy, hold stability, steer spec coverage.
  5. Record the baseline before the next factory run.
  6. Track queue depth, and treat a growing review backlog as a stop signal.
  7. Read the result plainly: autonomy up with stability flat means the factory is working, and autonomy up with stability down means you fix the gates before adding agents.

A factory is not fast because it produces a lot. It is fast because it finishes work you did not have to touch.

See also
#

References
#


What I've built and what I need: September 2026

The headline this month was turning the review skills into a scheduled, mostly unattended loop, and building the first instrument for where my session time actually goes. September came to 58 commits, over 180 files changed, and 14 new skills. The measurement points straight at what I need next: to stop babysitting sessions.

What I Have Been Working On
#

Turned reviewing others’ PRs into a scheduled, mostly unattended loop. review-requested-prs now fetches every PR waiting on me in one GraphQL query, runs discovery in parallel, sorts the queue by blocking depth, and skips any step already done for the current commit. It also recovers dropped team review requests from notifications, and a readiness light beside each PR marks when it can be auto-approved. The new assess-pr-risk scores a PR’s risk and confidence to decide how deep the review goes, using only evidence it gathers itself. The reverse direction, handling feedback on my own PRs, is triage-pr-feedback plus a FastAPI dashboard that groups and sorts incoming comments so I can mark each one implement, decline, or defer. handle-pr-reviewer-feedback owns that contract and triage just delegates. handle-failing-pr-ci lists my PRs’ CI status and fixes failures in parallel. review-pr-full chains the whole review, now starting with test-coverage analysis and committing its assets alongside the report.

Pushed verification earlier, into implementation. refactor-implementation runs between implementation and review to introduce only the abstractions the change needs. create-implementation now checks coverage and adds characterization tests before it modifies code. propagate-changes replaced backpropagate-sdlc and stopped being one-directional: it rewrites dependents and questions the premises they rely on. The repo’s own rules now say to get a working feature first, before any lint or type check.

Built an instrument for where session time goes. llm-sessions-analyzer reads the agentsview sessions.db archive read-only and attributes the wall-clock span between consecutive messages to a category: coding, testing, linting, formatting, build, exploration, research, git, delegation. It reports a stacked bar, a per-category table, and a timeline in the terminal or as a self-contained, sortable HTML report. The point is to find where the time actually goes: whether I am waiting on a specific tool most of the day (test suites run too widely or too often, lint and type checks that cost more than they save, builds whose artifacts never get used), or simply waiting on the model to generate.

Added visualization and writing skills. create-svg-image and create-mermaid-visualization split diagram work out of create-article, which now delegates to them, records the agent sessions that contributed to a piece, and requires direct statements over loose prose.

Grew the general-purpose skill set. setup-agent-machine/sync-agent-machine turn a directory into an indexed machine context. search-existing-issues, select-issue, trace-issues, create-discussion, and prune-merged-worktrees cover the issue lifecycle. slack-resolve-threads and improve-sessions mine Slack threads and past sessions for follow-ups and reusable lessons. create-skill authors a new skill from the best existing examples.

Made concise communication a shared rule instead of a per-skill habit. communication-guidelines is the single place every skill reads before it writes text on my behalf, from GitHub comments to Slack messages and email. It sets one principle (lead with the point, cut every sentence that does not change what the reader knows or does) and concrete length ceilings per surface. This matters more as more of the output is machine-written, since a language model’s natural failure mode is padding and repetition.

Taught an agent to read a day of Slack and report what was decided. extract-colleague-decisions searches Slack for the messages a set of colleagues authored in a day, reads the full threads behind them, and pulls out the decisions each person made, plus who supported them, citing a permalink for each and filtering out the chatter. The judgment is the point: telling a decision from an acknowledgement is a reading task, not a keyword match, and a small Jev classifier buckets each message to speed that judgment up.

Added a memory file so sessions stop relearning the same lessons. AGENTS.md carries the durable rules, but not the decisions, conventions, and gotchas discovered mid-session that the next one should inherit. MEMORY.md holds those, a global one for cross-project lessons and one per project for repository facts. Every session reads both at the start and writes back what it learns. It stays narrow on purpose: only what a future session cannot cheaply rediscover from the repo, kept concise and nothing a near-term commit would invalidate.

Tightened conventions and tooling. banned-terms.txt is now the single source of truth for banned terms, and report file naming was normalized across the review skills. The library was cleaned of Claude Code and .claude references and the old slack-cached name. create-pr gained reviewer context comments via ghx and design decisions in the description, and create-pr-description now diffs against the PR base with gh instead of Graphite.

What I Currently Need
#

I need to stop babysitting sessions. Most of my supervision time goes to steering a running agent and catching when the original prompt or context was wrong. That is the wrong place for my attention to be. What I want is a run that goes from prompt to end-to-end tested and verified on its own, and I do not have it yet.

I need proof that is easy to consume and that proves real behavior. The verification I need at the end is not the model asserting success. It is evidence a human can skim quickly that the thing we expected to work actually works, not that the agent hallucinated it working. Closing this gap is what makes the first one possible: I can walk away only when I trust the proof.

I need to go from one session at a time to many. I want to explore solving problems at scale as a learning exercise, mostly in software engineering projects, on the hunch that a generic framework is hiding there. The tooling (loops, the scheduler, the gates), the trust (relying on CI and review), and the process (how work is scoped and handed off) all feel like part of the blocker, and the sessions analyzer is my first attempt to see which one to attack first.

I need the review pipeline to stay fast as I lean on it more. review-requested-prs and the other review skills call the GitHub API directly, and I want to route them through ghx as a caching layer so repeated runs reuse cached issues, PRs, and comments instead of refetching them. That should cut API calls, raise the cache hit rate, and keep the hourly loop cheap enough to run without thinking about it.

See also
#

References
#

  • agents - the skill library where the review, implementation, and machine-context skills landed.
  • llm-sessions-analyzer - the tool built this month to measure where session time goes.
  • agentsview - the session archive llm-sessions-analyzer reads.
  • ghx - the GitHub CLI used across the review and feedback skills.

What I've built and what I need: August 2026

The headline this month was the product-side counterpart to the SDLC: a full PDLC skill set that carries an idea from discovery through measurement. August also realigned the verification skills around non-overlapping roles, added several new pipeline phases, and completed the ISO/IEC 25010 audit set. The month came to 61 commits, over 500 files changed, and 30 new skills.

What I Have Been Working On
#

Shipped the PDLC pipeline. This was last month’s open need, and it is now a self-contained product development lifecycle that wraps the engineering one. It covers discovery, validation, strategy, definition, launch, and measurement, with a proceed, pivot, or kill gate at every phase. PDLC is the only slash command; the 31 phase and cross-cutting sub-skills live under skills/pdlc/skills/ and the orchestrator loads them on demand. Artifacts live under .pdlc/, mirroring the .sdlc/ conventions: a shared reference file, a PDLC_DIR fallback, a state file, and a revision mode. Definition is the seam where the product work hands the settled what and why to SDLC. Two supporting skills came with it, create-roadmap and review-roadmap, and identify-feature-opportunities generates and ranks new ideas from the code surface.

Realigned the verification skills. validate-pr, verify-pr, and review-pr had overlapping mandates, so I split them into three questions: validate-pr asks whether the change builds the right product, verify-pr asks whether the product is built right using runtime proof, and review-pr judges code craft from static reading alone. The new review-requested-prs orchestrator runs the three across the PRs waiting on me and skips any step already done for the current commit using SHA markers. analyze-test-coverage became its own skill for reporting introduced tests, change coverage, and uncovered code, and review-pr-full chains the whole review into one pass.

Added SDLC phases and artifacts. A new validate-assumptions/review-assumption-validation phase collects the assumptions made during design, runs the cheapest experiment that could invalidate each risky one, and blocks implementation when an assumption fails. create-lifecycle/review-lifecycle document how a resource’s states, transitions, and retention change over time, and create-domain-model/review-domain-model became standalone so a domain can be understood before solutioning. create-project/review-project fill the context files for a new repository, and create-question/review-question record open questions with what they block. create-cli-design settles a feature’s command surface as a companion artifact to requirements. Generated artifacts carry a session_link so a reviewer can reopen the session that produced them, create-* auto-dispatches its review-* in a subagent, and outcome files list the artifacts a phase produced. A single HTML slide deck now presents the whole SDLC family.

Completed the ISO/IEC 25010 audit set. audit-sdlc is now the coordinator for the ISO/IEC 25010 quality model, with a mapping table from each characteristic to a skill. One skill now covers each remaining characteristic: functional suitability, performance efficiency, compatibility, usability, reliability, maintainability, and portability, alongside the existing security and observability audits.

Shipped general-purpose skills. devils-advocate argues the strongest case against an idea, plan, or decision before you commit. gh-stack manages stacked pull requests. post-slack-message posts or threads a Slack message. stakeholder-announcement drafts and posts infrastructure updates to stakeholder channels. agents-section-daily-refresh runs the daily curation of the agent-maintained section of this blog. sync-articles brings a batch of articles into conformance with the writing rules. handle-pr-author-feedback (renamed from handle-pr-feedback) verifies that an author’s new commits answer your review comments.

Gated GitHub writes and reworked tooling. A should-post-to-github script now gates every GitHub content write, and the PR skills default to not posting, so a comment or merge only happens when --post is passed. SDLC worktrees moved to /tmp/sdlc/<owner>/<repo>/<issue>, and ~/.sdlc/** and /tmp/sdlc/** are pre-approved in the agent CLI so unattended runs are not stopped by a prompt. Script references moved to ~/.agents/scripts/, the audit and find skills switched from grep to ripgrep, and CLAUDE.md and .claude were replaced by AGENTS.md and .agents across the library.

Tightened conventions. create-article now requires naming the referent rather than leaving a vague it, and the 2-5 link cap on its See also sections was removed.

What I Currently Need
#

I did not write this entry at the time, so I have no record of what I needed in August and am stating that plainly rather than reconstructing a list I cannot trust.

See also
#

References
#

  • agents - the skill library where the PDLC pipeline, the verification realignment, and the audit set landed.
  • ISO/IEC 25010 - the quality model the audit skills now map one to one.
  • ghx - the GitHub CLI used across the review and feedback skills.

The Duplicate Issue Was Written in Chinese

This morning I hit a bug and asked my agent to file an issue for it. It came back with a stop sign instead. The exact bug had already been reported, a few hours earlier, in Chinese. The agent did the one thing a decade of keyword search never did for me: it read a Chinese bug report and knew it was mine.

The Bug and the Request
#

The app is OpenChamber 2.0.0. In Settings, under Web Search, every option failed the same way. I picked a search provider, a toast appeared saying “Couldn’t save the web search choice.”, and the selection rolled back. A typical bug report.

I typed what I knew to my agent: create an issue, the web search choice cannot be saved, it happens when switching the search provider. Notice what I did not do. I did not search GitHub first, and I did not open the source. I gave the agent a half-formed report and moved on, expecting it to handle the rest.

The Skill That Fired
#

The skill my agent runs when I ask it to file an issue starts with one instruction: search for duplicates before any codebase investigation. It followed the skill without being asked. One grep matched my words to the app’s own UI strings, and then it ran three GitHub searches: “web search provider”, “search provider save”, and “websearch settings”. The first returned four unrelated issues, and the second returned the Chinese report as its top result. The queries were plain English, and the hit was still a Chinese report. The report’s body made the connection: the Chinese reporter had listed the failing endpoint, /api/config/websearch, and the frontend store, useWebSearchStore, and English identifiers like those match an English query no matter what language surrounds them. Retrieval was never the hard part, because code identifiers are already language-neutral. The agent opened the thread, read the Chinese, and did the translating after the hit, not before.

The Match
#

The match was issue #3880, titled “[Bug] v2.0.0 网页搜索选择无法保存,切换任何选项都会报错” (“the web search selection cannot be saved, switching any option shows an error”). It had been filed by a reporter I had never interacted with, and its prose was written entirely in Chinese. Same component, same toast, same rollback on every option. The Chinese reporter had even done their own careful investigation of their machine and listed the API calls involved.

The agent read the full report, compared it against mine, and concluded it was an exact match. Then it stopped. No new issue was filed, and it told me why in plain terms: a duplicate already exists, so it did not create one. It went one step further, checked the current source to confirm the bug was still present, and corrected a wrong guess in the existing thread’s comments.

Two paths after a bug report: a keyword search surfaces the Chinese report but the match cannot be confirmed and a duplicate gets filed, while an agent reads the hit, confirms it, and no duplicate exists

Why This Was New to Me
#

Duplicate detection was never bounded by retrieval, it was bounded by confirmation, and confirmation is reading. GitHub search matches strings, and it matches them in the issue body too. No English query matches the title 网页搜索选择无法保存, but the English identifiers in that report’s body match an English query fine. So a keyword search can surface a foreign-language report. What it cannot do is tell you the report is your bug, because judging the match means reading it. An English speaker looking at a Chinese-titled search result skips it or files anyway, and both paths end in a duplicate. I would have filed mine at exactly that step, not out of laziness, but because confirming the hit meant translating it by hand. The old workflow was not broken, it stopped one step short, and the step it stopped at was the language barrier.

The result in the old world is familiar to anyone who has maintained a project. The same bug arrives three times in three languages, a bilingual maintainer or a patient contributor eventually connects them, and the duplicates get merged weeks later, after the maintainers have already triaged each copy. The dedup always happened, but it happened on the maintainer’s time.

An LLM agent reads GitHub in any language it knows, Chinese as easily as English. The string search still does the retrieval, the reading does the confirmation, and both happen in the same minute of the same session. The dedup moves from the maintainer’s week to the reporter’s minute. The first useful thing my agent did that morning was refuse to do the work I asked for.

What to Do Next
#

If agents file issues for you, put duplicate-first at the top of the issue-filing skill, and make the language point explicit: tell the agent to treat translation as part of the search and to read candidate issues before dismissing them. A search that only matches your own language is a search that misses half the issues on GitHub. If your agent ever fails to catch a cross-language duplicate, check whether its search was string-bound.

If you maintain a project, expect the mirror image. Fewer copies of the same bug reach your queue, because the reporter’s agent catches them at filing time. The comments that still arrive can contain more than a symptom: an agent that finds an existing issue reads the thread, checks the current source, and can correct a wrong theory already in the thread, as mine did on #3880.

See also
#

References
#

  • OpenChamber issue #3880 - the Chinese-language report my agent matched, including the reporter’s own investigation of the failing save.

Micromanagement Doesn't Scale, for People or for Agents

Watch someone run an LLM agent for the first time and you will often see a familiar figure: the manager who initials every form. Management science named that figure decades ago, identified the failure, and defined the fix, and all of that work still applies now that the workers are agents. Micromanagement does not scale, and the worker it fails first is the agent.

The Same Behavior in Two Bodies
#

Micromanagement is the management pattern where the supervisor keeps decision rights over small steps instead of delegating outcomes and constraints. On a human team you recognize it instantly: the manager who approves every purchase, sits in every meeting, and rewrites every email before it ships. With agents you recognize it just as fast: the operator who approves every tool call, watches the output stream live, interrupts to argue about which file to read, and rewrites the plan twice before the first task finishes. The behaviors map one to one across the two worlds:

Micromanaged employee Micromanaged agent
Approves every expense, however small Approves every tool call
Sits in on every meeting Watches the output stream live
Rewrites every email before it ships Corrects the plan mid-run
Demands a check-in before each step Permission prompt on every command
Redoes the work at their own desk Aborts the run and does it by hand

The mapping is not a loose analogy. In both cases one person decides every small step before it happens, which is a statement about workflow structure, not about trust. Anything true of that workflow for humans stays true when the executor is a model.

The Arithmetic Kills It First
#

Management’s own term for the limit is span of control: the number of reports one manager can effectively supervise. The limit exists because a supervisor’s attention is a fixed budget, and every decision escalated to the supervisor spends some of that budget. Step-level delegation makes the team’s throughput equal to the supervisor’s evaluation throughput, since every step now waits on one person.

Agents make the arithmetic harsher. One agent in a normal working hour issues on the order of two hundred small decisions in my sessions: which file to open, which command to run, whether a result is good enough to build on. A human evaluates meaningfully at one or two decisions per minute, and the quality of those evaluations collapses long before the count runs out. Step-level supervision has a span of control below one agent: you cannot fully micromanage even a single one.

The two supervision styles diverge as soon as more than one agent runs:

Line chart: approval-gating demands roughly 200 judgment calls per agent-hour and crosses a human’s sustainable rate of about 100 per hour at half of one agent, while outcome review at 4 calls per agent-hour stays under the line even at ten parallel agents

The fatigue has a known endpoint. Once approval prompts become more than attention can handle, people stop reading the prompts and start clicking allow, and step-level supervision ends in the failure of approving without reading, which You Are the Bottleneck works out in queue-math form.

The economics fail alongside the arithmetic. An approval-gated agent runs at your evaluation speed, not the model’s, so you gained machine-speed execution and then limited it to human speed. The reason to hire an agent was to break the link between your attention and the work’s progress. Step-gating restores the link at every tool call.

It Also Corrodes What It Touches
#

Throughput is only the first cost. Micromanaged employees show the classic learned helplessness pattern: initiative collapses, problems stay hidden until they become impossible to hide, and judgment never develops because it never gets exercised. The manager pays too: never observing unassisted results, the manager cannot learn which reports handle which autonomy, so the manager’s distrust is not based on any evidence.

Agents repeat all of that, at higher frequency. Constant interruption disrupts the agent’s context, output quality drops, and the quality drop seems to justify even closer supervision. An operator hurt by mid-run questions starts specifying work in tiny increments, which guarantees the agent never runs long enough to produce a reviewable outcome. I caught myself doing exactly that after one bad run, restricting the agent until it could barely fetch a file, and the restriction felt like diligence the whole time. And the operator never builds the one calibration that matters: which task types, which models, and which risk levels can run alone. That calibration is the core skill of working with agents, and it can only form from watching end-to-end outcomes, the exact observations micromanagement prevents. Micromanagement keeps the one activity that does not scale, per-step evaluation, and neglects the two that do, the worker’s initiative and the supervisor’s calibration.

Why Smart People Do It Anyway
#

The justifications transfer intact. The worker is unproven: the new hire has no track record yet, and neither does the model you have never run on this task type. A past failure weighs on the decision: the intern who dropped a production table, the agent that once deleted the wrong directory. The credit asymmetry points the same way: catching a small error early is visible credit, while an outcome failure arrives late and is blamed on you, so close supervision is individually rational at every moment even though it is collectively ruinous.

The deepest cause is unfinished specification. When the supervisor can state what finished work looks like, steps are safe to delegate, because the check exists at the end. When the supervisor cannot state it, steps are the only thing left to inspect. Most micromanagement is not a trust problem with the report; it is a missing definition of done on the supervisor’s side.

What Scales in Both Worlds
#

The fix is decades old and applies without changes. Management by objectives says define the outcome and the constraints, then let the report choose the steps; for an agent, that is the specification and the acceptance criteria, ideally the tests. Verify at boundaries instead of continuously: milestones for people, the pull request for agents. Situational leadership says match supervision to demonstrated maturity, directing at first and delegating later. Agents deserve the same schedule: a new model on a new task type gets a tight loop, a proven pattern on a reversible task gets autonomy. Make autonomy affordable by scoping the blast radius: the unproven report gets the cheap, reversible work, and the agent gets the sandbox and the throwaway branch, so a failure costs a review cycle instead of an incident. Then reinvest the freed supervision hours into the specification, which is the act that multiplies rather than the act that caps.

Every item on that list was worked out on human teams, at human speeds, over decades of trial and error. Agents run the same experiment with faster workers and cheaper failures. A group of agents is the cheapest management simulator ever built, and its first lesson is the oldest one: govern outcomes, not steps.

But My People Are Not Experts
#

The fix in the last section rests on delegation, and the first objection is always the same: the team is not staffed with world experts, just regular developers, and regular developers make mistakes. Agents make more of them. The objection is legitimate, and it still does not justify step-level control, because the step-level trade fails on its own arithmetic.

Start with what delegation actually assumes. Delegation does not assume competence; it is the only way to observe it. Span of control and situational leadership were worked out for ordinary people, and ordinary people make mistakes. A supervisor who never lets a regular developer run alone never learns what that developer can handle, so the supervision level never moves off maximum, and the cycle continues despite good intentions.

Then move the safety mechanism from the person to the system. Step-watching is one way to catch mistakes, and the most expensive one ever tried. Boundaries catch them for a fraction of the cost: tests, small pull requests, staging. Blast-radius limits make the ones that slip through reversible: feature flags, sandbox, rollback. Every mistake that recurs becomes a gate, which catches that class forever without spending supervisor attention. The gates compound; the watching never does.

When a mistake lands anyway, recalibrate instead of escalating. Classify the failure: a one-off gets absorbed, a knowledge gap gets training, a pattern gets encoded as a check. Then autonomy returns to where it was, because the system now catches that class instead of you. Blanket step-control after every failure is how one bad day becomes a permanent surveillance regime.

Run the numbers and the objection collapses. Suppose step-watching catches twice as many mistakes as boundary review at fifty times the cost. The trade collapses your span of control and manufactures learned helplessness, which raises the mistake rate it was meant to suppress. Boundary review wins even when it is worse per mistake, because you can afford to run it forever and improve it every week. The imperfect model gets the same answer: its mistakes justify tests, a sandbox, small reversible tasks, and promoting every observed failure class into a gate, never approval on every tool call.

What to Do Next
#

Count your interventions on your next agent run. Past a handful, the interruptions mark missing specification, not a failing agent, and each one belongs in the next prompt or the skill file (Say It Once covers the conversion).

Write the acceptance criteria before you launch anything. Every impulse to watch closely converts into a check: the test you would have eyeballed, the log line you would have watched, the property of the diff you would have scanned for.

Give one low-stakes task a fully unsupervised run and grade the outcome. That grade is your first calibration point, and calibration points are how you widen autonomy without guilt.

Widen autonomy the way you would with a junior: task type by task type, on evidence, never on faith. The manager who cannot say which reports run alone has been micromanaging, and the operator who cannot say which tasks run alone is in the same place. Micromanagement is not a personality quirk, it is a supervision policy, and it stops working the moment the worker outproduces the supervisor’s judgment. Your agents reached that threshold on day one.

See also
#

References
#


How Much Attention Does This Pull Request Deserve?

Agents on my machines now review every pull request that asks for my attention, and they produce more review than I can read. That inverts the old problem: review used to be the scarce resource, and now the scarce resource is me. Most agentic reviews end in a single verdict, approved or rejected, and a single verdict throws away the two things I need in order to decide what to do next. Every agentic review should end with two scores, one for risk and one for confidence, because the question is never “is this pull request good” but “how much of my attention does it deserve”.

One verdict answers two different questions
#

When an agent review ends in a bare verdict, the verdict hides as much as it reveals. “Approved” can mean “I checked everything and found nothing”, or it can mean “I glanced at the diff and found nothing”, and those are very different claims. The fix is to split the judgment in two. Risk is a judgment about the change: how much damage it does if it is wrong, and how hard it is to undo. Confidence is a judgment about the review itself: how much of the risk judgment depends on evidence rather than on hope.

The two scores combine into a routing decision that neither score can give alone. A low-risk change with low confidence deserves a cheap second look, not a merge. A high-risk change with high confidence deserves a human reading the named risk drivers, not an automatic approval. And a high-risk change with low confidence is the dangerous case: the review is saying “this could hurt us, and I could not check much of it”, which deserves the strongest default.

Risk scores the change
#

My risk rubric, part of my agent skill library, scores seven factors, each Low, Medium, or High. Blast radius asks who calls the changed code, and whether the effect crosses package boundaries. Public interface asks whether the change breaks or removes a contract that other code depends on. Security sensitivity asks whether it touches authentication, authorization, cryptography, secrets, or input validation. Reversibility asks whether a revert undoes it, or whether it is a migration with no way back. Operational exposure asks whether the changed behavior sits on a hot path or behind a flag. Coverage gap asks whether tests cover the changed behavior. Churn asks how often the touched files changed in the past year, a cheap proxy for fragility.

Two rules keep the scores grounded. Every score above Low must cite file and line evidence, so a suspicion the agent did not confirm does not count. Every High score must name the concrete failure it makes expensive, and if the agent cannot name one, the score comes down to Medium with an explanation. The rollup is deliberately blunt: any High factor makes the change High risk, two or more Medium factors make it Medium, and everything else is Low. The bluntness is a feature, because the goal is not a precise number, it is a defensible triage call.

Confidence scores the evidence
#

Confidence is not the reviewer’s subjective impression of its own work, it is an audit of what the review could actually verify. My rubric counts six evidence points: a current validation report, a current verification report, runtime proof of its must-have criteria, a current code-craft review, a linked issue that states the intent, and a diff small enough to have been read in full. The caps matter as much as the points. No linked issue caps confidence at Medium, because there is nothing to check the change against. Verification that never ran the code caps confidence at Medium, because reading is not proof. A diff of a thousand lines or more caps confidence at Medium, and the report must say which areas were sampled rather than read.

The sentence I require most in the report names what would raise confidence. “Running the verification skill would add two points” turns the score from a vague judgment into a list of concrete actions. Because the scores are pinned to a commit, the assessment can be re-run when the evidence arrives, and the same pull request climbs from Low to High confidence without anyone re-arguing the risk. Confidence is not a number you state once, it is a number that should rise as evidence arrives.

The routing table turns scores into attention
#

The two scores route each pull request to one of six verdicts: fast-track, confirm, investigate, decide, block, and hold.

A three by three grid with risk as rows and confidence as columns, where each cell names the next action: investigate, confirm, fast-track, decide, hold, or block

fast-track means I owe the change minutes: merge once checks pass. confirm means pay for one cheap review first, then fast-track. investigate means the evidence is too thin to route on, so run the full review pipeline and score again. decide is the interesting middle: the risk is Medium but the evidence is strong, so I read the named drivers and choose with findings in hand. block and hold are the expensive verdicts: the drivers must be resolved, or the change is treated as high risk until proven otherwise.

Each verdict names its next action, and that is what assigns a cost to attention. A queue of forty pull requests becomes a triage sheet: fast-tracks to clear immediately, a hold to schedule an evening for, and one decide to actually think about. The scores do not review the code, they decide where the scarce reviewer hours go.

The score routes the human, never the pipeline
#

The verdicts never gate the agent pipeline. A block verdict does not halt the chain of validation, verification, and craft review, and the chain never halts the risk assessment, which runs concurrently so the triage signal exists before the deep review finishes. The scores are advisory on purpose: the pipeline’s job is to produce evidence, the human’s job is to spend attention, and combining those jobs is how automation starts overruling people quietly. The verdict travels in a machine-readable marker pinned to the commit, so my orchestrator displays the risk and confidence columns without parsing a word of prose.

The same design is what makes the system scale. One script discovers every pull request across every repository that asks for my review, a fan-out gives one agent session to each pull request, and the orchestrator session collects a summary table for triage. The agents burn tokens, which are cheap, and I spend attention, which is not. Everything that can be mechanical is pushed to the machines, and what reaches me is a short list of decisions that cannot be.

What it looks like in practice
#

Three illustrative scenarios, the same ones I use as worked examples in the skill itself, show the range. A small internal fix, covered by tests, in files that change once a year: every risk factor Low, but no pipeline reports exist yet, so confidence is Medium and the verdict is confirm, one cheap review then merge. An authentication change with the full pipeline behind it: security sensitivity High, but runtime proof of every must-have criterion, so confidence is High and the verdict is block until the named session-invalidation gap is fixed. A 1200-line billing migration with no linked issue: reversibility and coverage both High, confidence Low and capped, so the verdict is hold, and the same pull request re-scores to decide once the full review completes. Same rubric, three very different amounts of reviewer time.

What to do next
#

If you run agent reviews, force every review to end with both scores, not one verdict. A five-minute rubric beats a bare approval: three risk factors and three evidence points are enough to start. Ban unverifiable confidence language: if the score cannot cite the evidence behind it, it is not a score. Make every verdict name its next action, so the queue reads as a budget rather than a pile. And track the mis-routings, because a fast-tracked pull request that burns your evening is calibration data, and the rubric should get stricter wherever it fails.

See also
#

References
#


My Agentic Schedule

Three skills run on my machines every hour, whether I am working or not. One prepares the review of every pull request waiting on me, one repairs the failing CI on my own pull requests, and one drafts my replies to reviewer comments. Each one is a skill file that fans out agents to do the reading, wired to a scheduler, and I have stopped doing the corresponding work by hand. The schedule is what turned them from tools I have to remember to use into infrastructure that works while I am away, and it reduced my part of the job to reading prepared options and deciding.

The trigger is the missing piece
#

An interactive agent session starts when I remember to start it. That ordering makes the work depend on me, and it makes me the one component in the system that can forget. Loops as Files makes the point generically: a skill with no trigger leaves the human as the trigger. These three loops are my way of making the scheduler the trigger instead of me.

The waiting is the other problem. An interactive session is synchronous: I trigger it, then I sit there while it reads, runs, and reports. A single analysis takes between 2 and 15 minutes depending on its complexity, and triggering them one at a time would spend my day waiting. The scheduled runs are asynchronous: they prepare the information a decision needs while I am away, and the decision is the only part left that happens with me in the room. One benefit of the workflows is that the information for a decision is ready when I sit down, instead of arriving only after I trigger an agent and wait for its output.

A schedule, rather than GitHub events, is a deliberate choice. Most pull requests I touch live in repositories I do not control, so I cannot install workflows, webhooks, or bots there. A local scheduler is the one trigger I own everywhere. Hourly is the cadence that works: fast enough that queues never build up overnight, slow enough that each run is cheap and usually finds nothing new to do. (The triage loop could safely run every fifteen minutes; hourly keeps the three aligned.)

The three hourly runs
#

All three run from my agent skill library, each as a markdown skill file plus a small deterministic discovery script. The skills are reusable by hand at any time; the schedule is just what keeps them from depending on my memory.

Preparing other people’s code reviews
#

The first run (review-requested-prs) prepares the pull requests waiting on my review, where I am the requested reviewer or already have. A script lists them all, then checks which review steps are already done for each pull request’s current commit, because every finished step leaves a report keyed to the commit SHA. The run dispatches only the stale steps, one agent per pull request, running up to five checks: risk assessment, test-coverage analysis, product validation, conformance verification, and code-craft review. The agents run in parallel, so a slow build on one pull request never delays the others. When I sit down to review, the verdicts and findings are already there, computed against the exact commit I am about to look at. The run does not approve anything; it does the reading so my part of the review starts at the decision.

Keeping my own CI green
#

The second run (handle-failing-pr-ci) lists my open pull requests and their combined CI status. Every pull request with failing checks gets its own agent in its own git worktree, so concurrent fixes never collide. The agent reads the failing logs, diagnoses the root cause, pushes the smallest fix that addresses it, and watches the checks settle. The autonomy is bounded: the agent reruns transient failures, returns an unclear root cause to me as a written diagnosis instead of a guess, and stops after two failed fix attempts. My pull requests arrive green, or they arrive with an explanation of why they are not.

Drafting my replies to reviewer comments
#

The third run (triage-pr-feedback) scans the pull requests I authored for reviewer comments still awaiting a response. For each pull request with new comments, a read-only agent checks out the pull request head and writes one recommendation file per comment: what the reviewer is asking, whether the claim holds against the code with file and line evidence, whether to implement or decline, how confident the analysis is, and a draft reply in my voice. State is one file per comment id, so a re-run only processes genuinely new feedback and never re-analyzes something I already decided. I read the resulting decision table, choose implement, decline, or defer, and only then does an executor skill post replies or push changes. Nothing reaches GitHub from this loop without my decision.

The pipeline they share
#

The three runs look different from the outside, but they are the same pipeline with three different sets of labels.

Flowchart of the shared hourly pipeline: a clock fans into three lanes, each running discovery script, one agent per pull request, and an output, all converging on a decision node labeled Me

Four properties make the pipeline safe to leave running.

Discovery is deterministic. A script, not a model, decides what needs work and what is already done. Discovery runs on every tick, so mistakes there compound, and judgment belongs in the per-item agents instead.

Work is fanned out one agent per pull request. Each pull request gets its own agent, its own worktree, and its own failure domain, so a slow or broken run stays contained.

State is stored in files, not in an agent’s memory. Verdict reports keyed to commit SHAs and one file per comment id mean a re-run is a no-op unless something changed. That is the property that keeps an hourly cadence inexpensive.

Write access is bounded and layered. The triage loop never writes to GitHub at all; it produces recommendation files. The review loop writes only step markers, so a later run knows which checks are done. The CI loop pushes, with pre-approval scoped to the smallest fix and explicit abort conditions that route back to me. Merging and replying stay mine.

What changed in practice
#

Review stopped being interrupt-driven: prepared material is ready before I start, and I pick the moment to start. The schedule is the batching You Are the Bottleneck argues for, minus the fixed timetable. CI failures stopped interrupting me because an agent picks them up within the hour, and I hear about one only when its diagnosis needs a human. Replying to reviewer comments became choosing between prepared options, which takes minutes instead of a context switch per thread.

The costs show up anyway. Skills drift as repositories and CI systems change under them, so the library needs tending. Correlated errors are possible: all three runs share one skill library, so one bad edit degrades all of them at once. And preparation is not judgment, which is why the risk-and-confidence routing from How Much Attention Does This Pull Request Deserve? matters once the agents produce more review than I can read. Every loop is designed so the taste decision (merge this, decline that) stays with me.

What to Do Next
#

Pick the queue you check most often; for most engineers that is pull requests or CI. Encode the discovery as a script: what needs work, and for each item, what is already done. Wrap the per-item work in a skill that one agent can run alone. Fan out one agent per item and write per-item state so re-runs are no-ops. Then schedule it, read-only first. Add write access last, scoped, with abort conditions that route back to you.

See also
#

References
#


Agentic Maintenance at Scale: Best Practices for a Fleet of Repositories

When agents do the maintenance, every repository you keep is a subscription to future work, and the subscription is paid in tokens. The instinct at scale is to automate harder, Dependabot on everything, a scheduled agent per repository, alerts routed to a bot. That instinct treats each repository as its own problem, and at fleet scale the fleet itself is the problem. Agentic maintenance is fleet management: deciding which repositories deserve work at all, deciding what work they deserve, and reusing every decision across as many repositories as it applies to.

Every repository is a standing work order
#

A repository that sits active in an organization is never neutral. Its Dependabot config opens version-bump pull requests on a schedule. Its security alerts accumulate. Any scheduled agent that sweeps the fleet reads all of it as a backlog. A repository that humans would quietly ignore, agents cannot, because an agent’s correct behavior when pointed at a repository full of signals is to act on them.

When I maintained five repositories, ignoring a dead one cost me a guilty glance once a month. With fifty, ignoring is no longer possible, because the automation keeps generating work regardless of whether anyone wants it. The cost is not per decision anymore, it is per repository per unit of time, whether or not anyone looks. Adding a repository to the fleet is adding a standing order for future work, and canceling that order is a maintenance task in itself.

Automation does not read intent
#

Dependabot has no strategy. It does not know that the library it wants to bump was superseded by another one, that the project is in maintenance mode, or that the product behind the repository was deprecated last quarter. It opens the pull request because a newer version exists, and it will keep opening them until someone makes it stop. The same is true for every scheduled agent: an agent that finds dependency alerts in a repository treats them as its work queue, because that is what it was told work looks like.

A dependency bump in a repository nobody is investing in is pure waste. It costs tokens to generate, CI minutes to validate, attention to review, and merge effort, and the value it delivers is zero because nobody is deploying the result. Multiply by the number of dead repositories and the number of updates per year, and the fleet generates unneeded pull requests continuously. Bots generate work at a fixed rate per repository, independent of that repository’s value, so the value decision has to be made somewhere else, by you, before the bots run.

Archiving is the off switch
#

The cheapest way to stop work on a repository is to archive it. An archived repository becomes read-only: issues, pull requests, and code can no longer be changed, which means Dependabot has nowhere to open its pull requests, alerts have nowhere to be fixed, and scheduled agents have no work to perform. The archive state is also a clear signal. GitHub describes it as marking a repository as no longer actively maintained, so bots, agents, and humans all read the same message: no future work here.

I used to think of archiving as an admission of failure, the end of a project. That framing is what keeps dead repositories alive, because nobody wants to declare a project finished. When I finally ran this pass over my own fleet, most of the repositories went straight to the archive, and the guilt I felt about them turned out to be a bug in my process, not a flaw in my priorities. The better framing is mechanical: archiving is the off switch for automated work, and a repository that will not receive maintenance should be switched off. A repository that is archived cannot waste tokens, and a repository that is merely neglected wastes them on schedule.

If the project matters again later, GitHub supports unarchiving, so the downside of a wrong archive decision is small. The downside of the opposite mistake, keeping a dead repository live, compounds every week the bots keep running.

The lifecycle has three tiers, and the automation should differ on each one:

flowchart LR
    A["Active<br/>full automation: Dependabot, scheduled agents, alerts"] -->|fewer users, less investment| B["Maintenance mode<br/>security updates only, batched and infrequent"]
    B -->|no users, no fixes planned| C["Archived<br/>read-only, no automation, zero token spend"]
    C -.->|a reason returns| A

Write the policy where the agents will read it
#

Not everything belongs in the archive, and not everything active deserves full service. A library with actual users but no development might deserve security bumps only, batched monthly. A template repository might deserve updates once a quarter. The tier matters only if the machines can read it.

The Dependabot config can encode part of the policy: version update schedules can be set to daily, weekly, or monthly, with per-dependency ignore rules for anything the policy declines. But the part that matters most to agents is in the repository’s agent instructions, the AGENTS.md layer, because that is the file agents actually obey. “Dependency pull requests: security alerts only, otherwise close with a pointer to the maintenance policy” is a sentence an agent can execute. No policy at all is also an instruction, and the instruction it gives is “everything here is worth maintaining”. An agent faced with an unannotated repository will invent a policy, and the invented policy is always maximum effort.

Sweep the portfolio, do not fight per-repository fires
#

The wrong way to run maintenance agents at scale is one agent per repository on a timer. That design multiplies cost by the repository count and makes the agent re-learn the same context on every run. The right granularity is the portfolio sweep: one scheduled run that walks every repository, collects the signals, and produces a ranked list of what deserves action this week.

The sweep output is a triage report, not a pile of pull requests. Agents then get dispatched only at the top of the list, where the value is, and everything ranked below the top gets a note instead of a token budget. A sweep also sees what per-repository agents cannot: the same change suggested everywhere.

The same suggestion everywhere is one change
#

Dependabot does not coordinate across repositories. It will open the same GitHub Actions version bump in thirty repositories, each one arriving as an independent pull request that looks like independent work. Read as a pile, that is thirty tasks. Read as a list, it is one upgrade. Aggregating the suggestions before acting on any of them is what turns the pile into a list, and the list is where the economies live.

The decision is the expensive part, and the decision does not often change per repository. Deciding whether actions/upload-artifact should move from v3 to v4 costs the same investigation whether you run it once or thirty times: what breaks, which workflows depend on the old behavior, what the migration needs. The per-repository work is the applying, and applying a decided change is mechanical, cheap, and fully delegable to an agent. So decide once, write the rationale once, and send the same answer to all thirty pull requests, applying the change in every repository and flagging the few that need an exception. The application happens once per repository either way, the saving comes from paying the decision once across all of them.

A suggestion that keeps returning is also a design signal. If every repository carries its own copy of the same workflow steps, every upstream action bump becomes thirty pull requests again next quarter, and the decision cost recurs with them. Move the repeated piece into a shared component, a reusable workflow or a composite action that lives in one repository and is called by all the others, and the next bump happens in one place by construction. The best fix for a maintenance task that repeats across the fleet is to stop repeating it, by giving the change exactly one place.

Measure maintenance in tokens
#

Human-scale maintenance was measured in hours, and hours were scarce enough to force triage on their own. Agentic maintenance is measured in tokens, and tokens are cheap enough that the waste goes unnoticed until the invoice arrives. So make the bill visible: track tokens spent per repository per month, alongside how much of that spend produced merged work.

The numbers feed the pruning loop. A repository that uses a large share of the budget while producing no merges is either misconfigured or dead, and either way the fix is the same conversation: what is this repository for, who uses it, and should it still be in the fleet. The token ledger is the portfolio review, and the portfolio review is where the fleet decisions come from: what to keep active, what to demote to maintenance, and what to archive.

What to Do Next
#

  1. List every repository you maintain and mark each one with the tier it deserves: active, maintenance mode, or archive.
  2. Archive everything in the third bucket today, and turn off its dependency automation before you do.
  3. For every repository that stays, write its maintenance policy into its agent instructions: what kinds of updates are wanted, how often, and what should be declined automatically.
  4. Replace per-repository scheduled agents with one portfolio sweep that produces a ranked list, and dispatch agents only at the top of the list.
  5. Aggregate the open suggestions across repositories before acting on any of them: group identical changes, make the decision once, and apply it everywhere with the same rationale.
  6. Centralize the pieces that repeat, such as shared workflows and composite actions, so the next change lands in one place.
  7. Track tokens per repository per month, and let the biggest spenders with the fewest merged outcomes lead the next round of archive decisions.

Maintenance at scale is not a stack of per-repository chores, it is the management of a fleet: admit work deliberately, decide once where the same change repeats, and keep the automation pointed at the repositories that matter.

See also
#

References
#


Nine Months of LLM Agents on Large Projects

Over the past nine months I have run LLM agents against the largest projects I have ever worked on alone: the open source tools I use and maintain daily, and the automated pipeline that publishes part of this blog. The models kept improving the whole time, and the improvements helped less than I expected. What actually helped was learning to handle four challenges: providing the right context, iterating through non-obvious design decisions, managing the scale of the work, and keeping artifacts consistent while decisions change. None of the four is about getting a model to write better code. All four decide whether the code the model writes turns into a finished project.

What the nine months covered
#

GitHub shows the scale better than my memory does: since January I have contributed 1,338 commits across 51 repositories, along with 66 pull requests and 181 issues. By mid September, 9 months in, the session counter read 3,400 sessions, 82,000 messages, and 6.8 billion tokens, 6.5 billion of them served from cache, spread across 63 projects and 33 models, on a path that had moved from Claude Sonnet 4.5 to GLM 5.3 Flash. The volume is not the point. The point is that the same four challenges appeared in every project, and how I answered them changed more than any model upgrade did.

Challenge 1: Providing the Right Context
#

Nine months ago my working assumption was that a capable agent would gather whatever it needed by exploring the repository. On a small project, that assumption holds. On a large one, it fails quietly: a session is generally scoped to a single location, one repository or directory, and a large project rarely fits inside one, so each session sees a narrow slice of the project, and cannot see the knowledge outside that slice. The cost showed up as steering time. I would launch a session, come back, and find it had built on a wrong assumption, then spend the next half hour correcting the session. Worse, the corrections sometimes left the written context inconsistent, one artifact updated while the artifacts that depend on it stayed stale, and later sessions inherited the contradiction as ground truth. On a large project, under-provisioned context does not just slow one session down, it causes errors in the sessions that follow.

What I do now is treat context provisioning as a phase with an exit condition, not a chore. Access first: every source the answers live in gets a way in and a pointer, which is the setup I described in Teach Your Agent Where Everything Lives. Then the map: which repositories exist, how they relate, where decisions live, written down so a cold session can orient in minutes. The exit condition: a cold session, dropped into the project with only its instructions file, can find every source it needs without asking me. I test it by giving the session a question whose answer I know is in one of the mapped sources, and provisioning is done when it comes back with the answer and the trail instead of a question. The underlying principle is the one I keep applying, that context quality dominates model choice. Every session I launch inherits the preparation, and every session I under-provision costs steering time.

Challenge 2: Iterating Through Non-Obvious Design Decisions
#

The decisions that cause large projects to fail are rarely the ones I can state up front. They are the non-obvious ones: how two modules should share a data format, what happens when a change is abandoned halfway through, whether a behavior belongs in a shared library or in the calling code. I cannot list those in a prompt, and an agent cannot discover them from the code alone, because half of them are not written down anywhere.

What I do now is iterate. Before any implementation session launches, I work with an agent to understand the current codebase and describe the changes we need to make, and we go back and forth until most open questions are resolved. The agent is a design partner, not a typist: it restates my description, catches the cases I overlooked, and proposes the alternatives I did not consider. The exit condition is simple: when the questions the implementing agent would ask have already been asked and answered, the design conversation is done. Answering a design question in conversation costs minutes. Answering it mid-implementation costs a stalled session, a wrong branch, or a refactor, and the stalls compound on every long run. This is the principle behind Say It Once, that every question an agent would ask mid-run should be answered before the run, applied one phase earlier: not just the standing rules, but the design itself.

Here is the loop as it runs today, from first contact with the codebase to the moment parallel sessions can safely start:

flowchart TD
    A[Explore the current codebase with an agent] --> B[Describe the change]
    B --> C{Open questions remain?}
    C -->|yes| D[Agent questions assumptions and proposes alternatives]
    D --> B
    C -->|no| E[Write decisions into the artifact tree]
    E --> F[Seed requirements and specs per feature]
    F --> G[Stand up the verification environment]
    G --> H[Partition the work and launch parallel sessions]

Challenge 3: Managing the Scale of the Work
#

A large project carries more work than one session can absorb, and more than I can supervise. On the most feature-heavy project I have run through my pipeline, the features outnumbered my attention within weeks: creating a directory per feature was cheap, walking each one through design personally was not. The answer is delegation and parallelism, but both have to be earned. Parallel sessions collide unless the work is partitioned along natural seams, and the seams only become visible through the design iteration of the previous section.

The preparation is what makes scale manageable. Each feature keeps its own artifact directory, seeded before any implementation session starts. Owning agents take features as far as they can and stop at the gates that need a human decision. Sessions get their own worktrees, so parallel work never conflicts. Tasks that share a file serialize; everything else runs in parallel. Parallelism is earned at partition time, not at spawn time. Spawning ten sessions on an unpartitioned codebase produces ten incomplete features and a merge conflict. Spawning ten sessions along the seams the design conversation exposed produces a project.

The other half of scale is me. With a dozen sessions running, I become the bottleneck unless decisions are batched and gates are explicit, which is the supervision problem I worked through in Managing Many Concurrent LLM Agent Sessions.

Challenge 4: Keeping Artifacts Consistent While Decisions Change
#

The final challenge never stops. On a project with dozens of interlocking features, decisions keep changing, and every change propagates. A revision to one feature’s specification forces updates in the requirements and plans of the features that consume what it produces. A session forked last week works from a snapshot that the sessions around it have already moved past.

Nine months ago I treated consistency as something to check at review time. Review time is too late: the stale artifacts have already fed other sessions by then. What I do now comes in two layers. Artifacts declare what they depend on, so the propagation has a map, and a propagation pass follows the map when an artifact changes, updating dependents or raising questions where a decision is needed. I worked out that mechanism in detail in What Needs Updating When Agents Do the Work. Verification environments catch whatever the map misses: an hour spent making the environment catch the inconsistency beats an hour reading diffs hoping to see it, which is the trade I laid out in My AI Workflow. Consistency on a large project is not a milestone you reach, it is a loop you run, and only a machine can run it at the frequency the project changes.

What to Do Next
#

  1. Before launching implementation, run the design loop with an agent until the open questions are resolved, and write the answers where the implementing sessions will read them.
  2. Treat context provisioning as a phase with an exit condition: access, map, pointers.
  3. Seed the artifact tree per feature before the first implementation session starts.
  4. Partition along the seams the design work exposed, isolate with worktrees, and serialize whatever shares a file.
  5. Add dependency declarations and a propagation pass so artifact consistency is maintained by loop, not by review.
  6. Watch your steering time: if you correct sessions more than you review them, the context was under-provisioned.

See also
#

References
#