Six months into running a working shadow team, every CEO I have set up hits the same problem. The CLAUDE.md is now 1,000 lines. The morning briefing has grown to five pages. The Researcher's digest started arriving with seven stories instead of three because three felt restrictive in week four and nobody trimmed it back when it should have been trimmed. The Connector's content calendar has three competing themes running because each one was important when it was added. The outputs are still arriving on time. The outputs are also drifting. Something feels off but you cannot name what.
The System Tuner is the agent that names it. Once a week, it reviews the run logs from every other agent, reads what each one produced, compares the outputs against the goal stated in the agent definition, and writes a short report on which agents are still earning their place. Once a month, it audits your CLAUDE.md and the scoped instruction files for size, recency, and structure, and recommends what to prune, refactor, or retire. The system stays sharp because one of the agents is paid to watch the others.
This is the per-agent breakdown for the System Tuner in the Agent OS Webinar series. Week nine of the 90-day install order. It comes last because it needs the other eight to review.
What the System Tuner reads (Inputs)
Five inputs. The System Tuner is the most meta agent in the shadow team — its subject is the other agents.
- Run logs from every other agent. When the Researcher runs at 6 AM, the run produces a log entry. Same for the Speechwriter, the Personal CFO, the Executive Assistant, the Chief of Staff, the Connector, the Compass, and the Health Team. The System Tuner reads every log.
- Sample outputs from each agent. Last week's digests, drafts, briefings, flags. The agent reads enough to know whether the outputs are still landing on the goal.
- The CLAUDE.md file and any scoped instruction files. Size, structure, recency-since-last-edit on each section. The System Tuner runs a monthly audit on this against the practical instruction-budget that frontier models can hold without dropping things.
- The agent definition files. Each agent's stated goal, inputs, skills, scheduled tasks, and outputs. The System Tuner compares the definitions against what is happening in practice, and flags drift.
- Your feedback. Whatever you have written in your daily notes about the agents — what landed, what missed, what you ignored. The System Tuner reads your reactions to the system as another signal.
The inputs accumulate quietly. The other agents produce their work; the System Tuner reads it. The owner does nothing extra.
What it does with that material (Skills)
Three skills. Like the Compass, the System Tuner is narrowly scoped by design.
- Agent performance review. For each agent, compare the last week of outputs against the stated goal. Are the outputs still useful? Are they being read? Are they drifting toward vagueness or sliding away from the goal? The agent grades each one in plain language: earning its keep, drifting, candidate for removal.
- CLAUDE.md bloat detection. A growing CLAUDE.md is a quiet failure mode. Once it crosses about 200 lines, the model starts dropping instructions silently. The System Tuner reads the file monthly, names sections untouched in 90 days, flags duplications, and proposes a prune list.
- Prune, refactor, retire. For each agent the review surfaces as drifting or no longer useful, the System Tuner proposes one of three actions: prune the scope, refactor the definition, or retire the agent. The recommendation includes a one-line rationale and a one-line proposed change.
The System Tuner does not edit anything autonomously. The agent does not rewrite the Researcher's definition. It does not edit your CLAUDE.md. It does not retire the Personal CFO. It writes recommendations into a report; you make the edit.
The schedule (Scheduled Tasks)
Three cadences.
- Sunday 9:00 PM weekly. System review. One report, in a System folder, covering all eight other agents. Each one gets a paragraph: what ran, what produced, whether the output landed on the goal.
- First Sunday of the month, monthly. CLAUDE.md audit. Size check, section-age check, duplication scan, prune list. One page, in the System folder.
- On demand. When you suspect a specific agent is drifting, you can ask the System Tuner for a focused diagnostic. The agent reads two weeks of that agent's outputs and writes a short note on what it sees.
The Sunday 9 PM time is the last slot in the planning cluster. The Connector runs at 7 PM. The Health Team runs at 8 PM. The System Tuner reads what they produced and writes the system review at 9 PM. By the time you sit down with your weekly planning on Sunday evening or Monday morning, the three meta-reports are already waiting.
What lands on your desk (Outputs)
Three outputs, one folder.
- The weekly system report. One page, with one paragraph per agent. What is earning its keep, what is drifting, what needs your attention. Action items at the top.
- The monthly CLAUDE.md audit. Size of the file, sections untouched in the last 90 days, proposed prunes, proposed refactors. One page. You make the edits.
- The on-demand diagnostic. When you ask for a deep read on a specific agent, a short note on what is happening with it. Two weeks of outputs, the drift signal if any, a proposed fix.
The reports live in a System folder. Most weeks you skim the weekly report in two minutes and confirm nothing needs your attention. Occasionally you find a real prune-or-refactor decision waiting for you, and you make it.
The goal: the stack stays sharp, not bloated
The System Tuner exists for one outcome. The shadow team gets sharper over time, instead of bloating. The agents earn their place every month. The CLAUDE.md stays small enough that the model can hold the whole thing. Outputs stay crisp because someone is watching for vagueness.
The second-order effect is structural. Without a System Tuner, every shadow team I have seen drifts the same way over six to nine months. The CLAUDE.md doubles in size because every small fix gets a paragraph added. The Researcher's filter quietly loosens because the GM wanted more stories one week and forgot to tighten the prompt back. The Connector's content calendar accumulates three running themes because each one was important when added and nobody removed the prior ones. None of these are dramatic failures. They are quiet drift, and the system as a whole stops being the system the owner built. With a System Tuner running, the drift gets named in the weekly report. The fixes are small. The stack stays close to the version that worked.
A worked example at Not Really Wines
At Not Really Wines, our fictional demo winery at NunnCurtis Labs, Maren Holloway has been running the full shadow team for six months. The Sunday system report on the last Sunday of October opens with three flags:
System review — last Sunday in October.
CLAUDE.md is at 1,047 lines. Up from 612 at the start of August. The growth is in two sections: the wholesale-strategy section now has 280 lines covering the Oregon push, and the Society-content section has 195 lines that were added by Priya in September. Both sections were useful at the time they were added; both are now bigger than they need to be to do their job today. Recommendation: prune both sections back to roughly 80 lines each. Move the historical detail to dedicated reference files in
3_Resources/that the agents can pull in when asked.The Researcher's digest is producing 6 stories most mornings. The definition says 3 to 4. The drift started in week 18 when you asked for "more on retail signals" once and the agent quietly relaxed the cap. Recommendation: edit the agent definition to enforce the 3-story cap and add the retail-signal angle as a Tuesday-and-Friday rotation rather than a daily addition.
The Compass's drift letter from last month was 1.5 pages. The definition says one page. It has been creeping for three months. The October letter was not less useful for being longer, but the value of the Compass is partly in its discipline. Recommendation: tighten the page cap. If the letter wants to be longer, the Compass is overreaching its lane.
Maren reads the report on Monday morning during her weekly planning block. She spends 30 minutes that morning making the three changes the System Tuner proposed. By the following Sunday, the CLAUDE.md is at 540 lines, the Researcher is back to 3 stories, and the October Compass letter is appropriately tight. The system feels sharper that week than it has for two months.
She did not catch any of those three problems on her own, because she was inside the system. The System Tuner caught them because that is its job.
How the System Tuner demonstrates Step-Back
The System Tuner is the cleanest example in the shadow team of Step-Back applied to the system itself. Step-Back is the move of zooming out and asking what is still earning its place. The Compass does it for your calendar. The System Tuner does it for the agents.
The pattern is the same. You cannot Step-Back on a system you are inside of. You can read a weekly report from an agent that is outside the system, looking at it. The artifact is what makes the move possible at all. Without the weekly report, you would notice the drift in six months, when something obvious finally broke. With the weekly report, you notice it the Sunday after the drift starts.
Why this is the last agent in the install order
The System Tuner comes ninth for the obvious reason: it needs the other eight to review. Build the System Tuner first and the agent has nothing to read. Build it last and it inherits eight weeks of agents to audit, eight CLAUDE.md sections to review for bloat, and eight goal documents to test the practice against.
Two more reasons:
- The System Tuner is the agent that teaches you the discipline of the system. Owners who run the System Tuner from week nine onward report that within three months they start running their CLAUDE.md, their agent definitions, and their goal docs with a tighter pen, because they know the weekly report will read them. The agent's existence improves the leader's discipline upstream of its actual reviews.
- By week nine, you have seen enough of the other agents to know which ones you want to keep, which ones you tuned away from the default, and which ones you would consider retiring. The first System Tuner report is mostly a confirmation of choices you already made. The second one starts catching drift you would have missed without it. The agent's value compounds the longer it runs, same as the rest of the team.
Build this one in week nine. By the end of week ten, the shadow team is running and one of its agents is keeping the others honest. That is the version of the system that lasts.
The webinar recording, the rest of the Agent OS series, and the agent files for download live at nunncurtislabs.com/webinars.