Managed Agents Now Answer to the Harness
At four in the afternoon, a developer's Claude Code session opens a work session on the Billing API to add proration to plan changes. Twenty minutes later, while that session is still working, a scheduled Archyl run on the backend-fixer profile picks up a ticket about invoice rounding, which also lives in the Billing API.
Until this week, here is what the scheduled run did. It opened its own work session, as managed runs already did. The preflight gate returned warn, with the reason 1 target element(s) are being worked on by other active sessions — coordinate before changing them. That verdict appeared in the run's event feed. Then the run started anyway, because nothing acted on the verdict, and nobody was watching a scheduled run to coordinate with.
I'd rather say it plainly: for hosted runs, the wrapping was mostly cosmetic. The gate was displayed and ignored. The files the agent changed were never tied back to its session, so memory landed on whichever elements the task text happened to mention. And the session closed with a one-line summary, which gave the next agent very little to learn from.
That changed this week. The harness now steers managed runs, and hosted runs and local coding agents coordinate on the same model.
Running an agent is becoming the easy part
Three of the biggest names in the field now offer a hosted place for agents to run. Anthropic's docs describe Claude Managed Agents as "a pre-built, configurable agent harness that runs in managed infrastructure", with cloud or self-hosted sandboxes, persistent sessions and scheduled runs on a cron. OpenAI's Agents API, announced in public beta on 10 September, runs the agent loop on OpenAI's infrastructure and handles "orchestration, long-running sessions, and context management". Cursor Projects, announced the same day, puts a coordinator agent on its own cloud machine that delegates to subagents and can run on a schedule or follow your pull requests.
These are good products, and they point the same way: sandbox, tool loop, sessions and schedules are becoming something you pick from a list.
What a runtime can't know on its own is your system. Which container the invoice code belongs to. Who owns it. The contract its consumers depend on. The decision that made a boundary deliberate. And the part that matters when nobody is watching: which other agent, including ones this runtime didn't start, is working on which part of it right now.
That knowledge can't live in a runtime, because your agents don't all run in one. Some run on a laptop, some in CI, some in a hosted worker. So the way we think about it is your runtime, our architecture: whatever runs the agent, the model it answers to is the same.
Archyl's own managed agent runs, launched in April, are one runtime among those. This week they started behaving like one. You can also watch and steer a run live.
A run can now be refused, and told why
Coordination is opt-in. Each agent profile has a new setting under Coordination, Respect other agents' work, off by default. With it on, the run opens its work session exclusively, and the gate verdict decides what happens next.
If another agent holds a lease on the same elements, the run is refused, whether that agent is Claude Code on someone's laptop or another managed run. The run ends before anything is cloned, the gate event reads Run refused, and the reason names who is working there and on what. For the afternoon above:
session denied by preflight gate: "Billing API" is being worked on by
claude-code/vincent (add proration to plan changes)
The format is the real one; the task is an example. Whoever reads the run in the morning knows exactly who to talk to.
If the gate only warns, for instance because an error-level guardrail applies to the task, the run neither starts nor fails. It waits in Awaiting approval without taking one of your organization's concurrent run seats. The run page lists the reasons, with Approve and start and Cancel run.
Approving doesn't rubber-stamp the old verdict. It evaluates the gate again, so if a local agent took a lease on those elements while the run was waiting, the approved run is still refused. A run nobody approves within 24 hours is cancelled.
Why a waiting run holds nothing
While a run waits, its session is dropped. The session that raised the warning is cancelled, its leases go with it, and approval opens a new one. Keeping the leases would sound more careful and be worse: a run nobody ever approves would keep every other agent off those elements for a day, telling each of them that something was working there when nothing was.
This follows a line we drew in August. Leases are advisory, not locks. Exclusivity is something a profile asks for, not something the platform imposes, and a run waiting for a human claims nothing.
The Guard moved into the worker
The session covers intent. The Guard covers what gets written. Local agents have had it as a Claude Code hook since August, and managed runs now have the same check inside the worker.
When the project has a linked repository, every write_file and edit_file the agent makes is checked against the project's conformance rules before the change lands. A critical violation refuses the write, and the agent reads which rule and why:
blocked by Archyl Guard: 1 critical architecture violation(s) in
backend/internal/adapter/http/handlers/invoice.go. Change the content so it
conforms, then write again:
- [critical] No direct database access from HTTP handlers — move the query behind a service
The agent moves the query and writes again. A high-severity violation lets the write land with a warning. If the check itself fails or times out, the write goes through: the Guard never blocks an agent on its own errors, because a governance check that breaks the work it governs gets switched off.
Each check also carries the session ID, which attributes the file to the run's session while the agent works. That matters two sections down.
The agent knows its session, but doesn't own it
The run's prompt now names its session and tells the agent the platform opened it and will close it, "so there are no session tools to call." The agent can't open a second session or finish this one early. The lifecycle belongs to the platform, the only party that knows when a run has ended.
What the agent does own is its account of the work. Right before finishing, it calls report_outcome once, with an honest summary, the decisions worth an ADR, follow-ups for work left undone, and the briefing memories it relied on. The tool's description tells it to keep situational findings out of decisions, because decisions are replayed to future sessions and a debugging note shouldn't bind anybody.
The outcome lands where the work landed
When the run ends, Archyl first ties the files the run changed to the session. This fixes the quietest of the three gaps. The elements a session leases are inferred from the task text, and task text is a guess. The changed files are what happened. Memory now goes to the elements real work touched, not to everything the ticket mentioned.
Then it closes the session, stores the summary as memory on those elements, and credits the memories the agent says it used, which is the signal memory ranking learns from.
Decisions are recorded only when the run succeeded. They become project memory and open a draft Architecture Change Request, so a person reviews how the model should catch up. A failed run keeps its summary and follow-ups, which are what the next attempt needs, but records no decisions. Work that didn't complete has nothing to stand by.
The run page shows all of this in a Work session outcome card: summary, decisions, follow-ups, the elements touched, the memories cited, and a link to the Change Request. An element another session holds is flagged Held by another work session, so a run that edited code inside someone else's lease doesn't pass unnoticed.
Two kinds of agent, one model
This is the piece many agents, one architecture was reaching for: coordination between agents that never share a process, in both directions.
A managed run opens its session as managed-agent/<profile name>. Reverse the scene from the top of this post: the scheduled run is in the Billing API when a developer starts Claude Code on it. Their briefing now carries the run as a conflict:
- **Conflicts** (someone else is already working here):
- Billing API — held by managed-agent/backend-fixer: fix rounding on prorated invoices
The local agent is told to coordinate before changing that element, the Fleet console shows both sessions, and whichever finishes leaves memory the other reads next time.
Also shipped this week
Briefly, because the rest stands on it. Profiles now carry built-in skills and tool allowlists, so a read-only reviewer can be limited to read_file, list_* and get_*. Worker heartbeats detect dead runs and reap them, each run's API key is revoked when the run ends, and a failed run still publishes its work as a draft pull request. Runs always execute in Archyl's own worker, on your organization's AI provider when bring-your-own is enabled.
What it doesn't do
Coordination only sees agents that use the harness. A lease exists because an agent opened a session. An agent editing the repository without one is invisible to the gate, and a run that respects other agents' work will start right next to it.
The worker doesn't run your tests. It has workspace file tools (read, write, edit, list, grep) and no shell. It can't build the project or run the suite, so a clean outcome card means the writes passed the Guard, not that the code compiles. CI still does that job.
An agent's decisions are unreviewed. A decision in the outcome card is the agent's claim until someone reviews the Change Request. That review is the point, not a formality.
The Guard needs something to check against. Without a linked repository there is no write check, and an empty conformance rule set passes every write.
Where to start
Pick the one profile that runs on a schedule against code people also change by hand, and turn on Respect other agents' work under Coordination. Leave the others as they are.
If it refuses a run, the reason names the agent and the task that overlapped with it. Until now, that overlap would have gone through without anyone hearing about it.
Managed agent runs, work sessions, the Guard and memory are part of archyl. Every setting named above is in the managed agent runs docs, and the local side is in the Harness guide. Related reading: many agents, one architecture, work sessions for coding agents, memory for AI coding agents, and the launch of managed agent runs.