Manager OS Build Log: Building Memory That Can Survive the Chat
Updated: 3 hours ago
ARCHIVED REFERENCE
This build log is retained as historical reference. Routine Manager OS updates now live on the canonical MOS page rather than in the main blog feed.
Manager OS is an experiment in making long-running AI work behave less like a forgetful chat and more like a recoverable operating system.
The private system tracks projects, decisions, sources, next actions, tool failures, and recovery state. This public build log records only the architecture and lessons. Personal, health, financial, legal/compliance, family, customer, employer, credential, account, and security-sensitive details are deliberately removed.
BUILD LOG — SEPTEMBER 2026
Canonical memory layer
The system now has a compact home/index, project-specific canonical files, a private changelog, and operating rules. A fresh assistant starts at the home file, then loads only the relevant project context instead of vacuuming up the entire attic.
Clear work states
Work is tracked as ACTIVE, NEXT ACTION, WAITING ON, DONE, or ABANDONED / DEFERRED. This prevents completed work from quietly climbing out of the grave and demanding another review.
Confidence and evidence
Important claims can be marked CONFIRMED, PROBABLE, ESTIMATED, or UNKNOWN. Evidence moves through OBSERVED, CANDIDATE, VERIFIED, CANONICAL, SUPERSEDED, and EXPIRED states, so an old fact does not stay immortal just because an AI once wrote it down.
Attention budget
The default is no more than three major active fronts at once. Small maintenance work can continue, but new multi-step projects have to displace something else. Side quests now need a parking permit.
Stop-check
Before repeating a review, the system asks whether looking at the same evidence again could materially change the answer, plan, filing, risk, or decision. If not, stop.
Failure budget
A tool or approach gets the original attempt plus one targeted retry. If it still fails, the system records the exact checkpoint and dependency instead of escalating into a giant panic-rebuild.
Action receipts
Consequential state-changing actions can carry a stable action ID, pre-state check, authority basis, result, verification, and replay rule. The goal is simple: do not accidentally do the same important thing twice because a spinner looked suspicious.
Recovery testing
A tabletop recovery test asks a brutal question: if the current chat vanished, could a fresh assistant recover the working system from the canonical files and backups without conversational memory? If not, patch the smallest missing locator or state line.
Shadow review across models
I am also testing a multi-model advisory layer where additional AI systems can critique or stress-test work without automatically gaining authority to rewrite canonical state. Advice can be distributed. Authority stays narrow.
Portable operating rules
Human-readable, AI-readable, and multi-agent guardrails are being separated from private project data so the useful operating ideas can travel without dragging a person's entire digital life behind them.
PUBLIC REDACTION WALL
This public log is not a copy of the private changelog. It is rewritten through a redaction gate. It will not publish personal health information, family details, financial or tax specifics, legal/compliance specifics, private identity links, customer or employer details, account identifiers, credentials, lock/key data, site-security details, or other material that would make a stranger unusually well-informed about my life or infrastructure.
NEXT
Keep measuring whether Manager OS reduces work instead of generating bureaucracy. Keep recovery boring. Keep permissions narrow. Keep side quests on a leash. And keep the public notes useful enough that someone can steal the good ideas without stealing my life.
Update: The council gets a job
A useful Manager OS experiment this week was deliberately simple: give another AI the same bounded problem, keep one source of truth, and see whether disagreement improves the result instead of multiplying noise.
It did. The independent review caught two things worth keeping visible: a temporary safety workaround is not the same as a finished repair, and an incomplete audit must stay labelled incomplete. Neither observation required creating a second memory system or handing another model authority.
What changed
The primary Manager remains responsible for canonical state, evidence reconciliation, and final decisions.
A second model is useful as an on-demand red team when there is a real assumption to challenge.
Additional tools have to earn a distinct job. Repeating the same opinion in a different voice is not collaboration.
Manual handoffs are acceptable when they are occasional and bounded. Building more automation is not automatically an improvement.
Consensus is not evidence. Several AIs repeating the same unsupported claim is still one unsupported claim wearing several hats.
Current rule
Use extra models only when they can materially change the decision, confidence, risk, or implementation plan. Keep the packet small, preserve the source of truth, review the answer before acting, and stop when the extra perspective becomes theatre.
The practical result is less glamorous than a swarm of autonomous agents, which is probably why I trust it more: one manager, specialist reviewers, explicit boundaries, and receipts when something actually changes.
PUBLIC-SAFE SYNC RULE — SEPTEMBER 19, 2026
Usable, general-purpose Manager OS improvements now flow into the public build log or public MOS resources by default after a privacy and security scrub. The public material is a rewritten derivative, never a raw mirror of the private operating system.
Architecture, recovery patterns, evidence controls, workflow design, multi-AI coordination lessons, anti-loop rules, and reusable setup improvements are good public candidates. Personal case history, health or family information, finances and tax/legal details, employer or customer specifics, account identifiers, credentials, security-sensitive data, and identity links stay private. Tiny changes may be batched so the log stays useful instead of noisy.
Update: Stop when the dependency is real
A useful upgrade this week was learning to distinguish an AI-solvable problem from a provider-only dependency. Once the available evidence proves that the next fact lives behind a specific portal, account, device, or human decision, repeating broader searches is no longer diligence. It is just expensive pacing.
What changed
Broad searches now get a stop line once the remaining source is named precisely.
Visual assets are screened in bounded batches, with sensitive identifiers, topology, credentials, and context rejected before anything reaches a public queue.
Cross-model review gets one bounded job: challenge assumptions or find a missed risk, then stop if another round will not materially change the result.
Material canonical updates get a pre-change backup and readback verification instead of trusting a successful-looking spinner.
Current rule
An AI should not keep searching merely because another search is technically possible. When the remaining dependency is real and named, record it, protect the evidence already gathered, and move the machine to work it can actually finish.
2026-09-20 — Portable bootstrap refresh
The portable Manager OS boot now starts smaller, adds an explicit share/export privacy gate, extends regression coverage from E01–E30 to E01–E40 with generalized multi-agent controls, and adds graceful tool/agent degradation plus cost guards.
The public package remains a scrubbed derivative of the private system: architecture, templates, recovery patterns, and reusable operating rules are shareable; live IDs, personal history, client or employer details, financial/legal/health data, secrets, and identity-linking breadcrumbs stay private.
2026-09-20 — Grey Goo joins the public boot
The public Manager OS package is now aligned with the current multi-AI operating rules rather than carrying only a simplified summary.
The public bootstrap now includes the privacy/export gate, minimum-viable setup path, E01–E40 regression coverage, cost guards, and graceful degradation when an AI or tool is unavailable.
Grey Goo is published as its own model-agnostic interop standard with Task Envelopes, GREEN / AMBER / RED privacy classes, Claims Ledgers, SOLO / RED_TEAM / PARALLEL / VERIFY / STOP routing, Agent Cards, SHADOW-to-canon promotion gates, capacity circuit breakers, and write-collision controls.
The public references remain scrubbed of real account names, IDs, private capacity history, employer/client specifics, financial/legal/health material, credentials, and other identity or security-sensitive details.
The practical rule remains boring on purpose: one source of truth, bounded specialist reviews, explicit evidence labels, no infinite robot debates, and no adding paid infrastructure just because another AI exists.
2026-09-20 — Tool-first routing, change control, and live publishing
The operating model now prefers the authoritative connected tool or source before reconstructing state from memory. Versioned change control tracks reusable public-safe operating files and maintenance notes, while publication stays a separate deliberate step.
What changed
Tool-first retrieval became the default when an authoritative connected system is available.
Reusable MOS and website-maintenance changes gained explicit version history and rollback instead of relying on one mutable copy.
Public publishing remains permission-bounded and separate from internal preparation, so a saved change is not automatically treated as a public change.
Multi-AI assistance remains advisory by default: extra seats can review, research, or prepare bounded work without becoming canonical authorities.
Capacity, failure, and cost handling now preserve checkpoints and route around unavailable tools rather than spawning uncontrolled retries or duplicate loops.
The public website log continues to receive only generalized, privacy-scrubbed architecture changes.
Current rule
Use the best authoritative tool first, keep one source of truth, version reusable changes, publish deliberately, and leave enough evidence that the next session can tell what changed without excavating the chat history.


Comments