AI Coding Agent Harness for Reviewed Engineering Work

Run one or several AI coding agents inside a Mrrlin harness: repository context, isolated git worktrees, checkpoints, reviews, and deploy control.

Chat tools vs Mrrlin

The execution layer beats another blank chat window.

Chat tools make you re-send context and manage the work. Mrrlin keeps project memory and spends tokens deliberately.

Chat-based AI tools
Mrrlin
Token cost
Full context re-sent with every prompt
Progressive context compression — the fewest tokens per task
Project memory
Forgets between sessions — you re-explain and re-instruct
Memory is captured continuously — context stays attached to the work

Why it matters

Coding agents need walls, a shared spec and proof to hand over.

Mrrlin helps coding agents work inside a reviewed harness: tasks carry repository context, worktrees isolate changes, checks create evidence, and deploy-sensitive steps wait for approval.

01Claude Code starts editing the checkout you are already using, and your half-finished migration and its changes end up in the same working tree. Untangling whose edit is whose costs more than the task saved.
02You split a signup-page rework across three agents and paste the spec into three sessions. By the second revision, the copy agent and the form agent disagree about which fields the form even has.
03The agent's final message says every test passed. Nothing was saved, so the reviewer reruns the suite from scratch, and the claim added no information at all.
04A pull request arrives with a large diff and no statement of intent. Reviewing it means reverse-engineering the task from the code, then guessing which changes were deliberate and which were side effects.

Workflow

From a repository task to a pull request a reviewer can trust.

01

Attach the repository context

The task carries the spec, acceptance checks and the repository facts that matter for it. Claude Code reads the repo and docs before the plan is final, and open questions about expected behavior go to the inbox instead of being settled by assumption halfway through the change.

02

One worktree per agent

Every run gets an isolated git worktree and its own branch, with the base commit recorded. Three agents working toward one goal therefore edit three separate trees, and your checkout stays exactly as you left it while they work.

03

Turn checks into artifacts

Build output, test results, command output and a diff summary are recorded on the run as they happen. The run's self-review is kept as well, but the verdict belongs to a second model from another vendor, which reads that evidence rather than the agent's account of it.

04

Hand off at the pull request

When review passes, the run moves to pull request status on your repository. The task's deploy policy decides whether it stops there, goes to a preview, or heads toward production, and production deploys run through your own GitHub Actions workflows only after you approve.

Use cases

Where it fits.

Repository implementation

Ask for the landing page hero and signup form to be rebuilt against the new spec. The AI coding agent harness hands the agent that spec, the component conventions recorded in the wiki and the acceptance checks, then records its checks and diff summary before anyone opens the branch.

Parallel agents in separate worktrees

One agent rewrites the hero copy, Claude Code rebuilds the form component, and a third verifies routes and redirects. Each works in its own worktree against the same spec in project memory, so running them in parallel does not put three agents in one directory.

Build and test checkpoints

Midway through the form rebuild, a checkpoint records a failing validation test. The reviewer returns the run with that output attached, and the resumed run carries only the delta since the last attempt instead of replaying the whole conversation from the start.

Cross-agent handoff to a pull request

The route-check task is blocked by the implementation task, so it starts only once the form change exists. It reads that run's diff summary and handoff notes, adds its own verification, and the reviewer sees every agent's evidence on one trail before the pull request opens.

Comparison

What changes when coding agents stop sharing your working tree.

Usual approach
Mrrlin
Where edits land
Directly in the checkout you are using, mixed with your own uncommitted changes and whatever the last session left behind.
In an isolated worktree per run, on its own branch, with the base commit recorded so the change can be compared cleanly.
Several agents, one goal
Separate terminal tabs, each holding its own pasted copy of the spec, reconciled by you whenever their outputs disagree.
Tasks that read one spec from project memory, each in its own worktree, with blocked-by links deciding who hands off to whom.
Test claims
The agent says the suite passed, and you either believe it or run everything again yourself before merging.
Test results and command output are artifacts on the run, read by a reviewer from a different provider than the one that wrote the code.

FAQ

Before you start.

Can one AI coding agent harness run several agents on the same goal at once?

Yes. Mrrlin splits the goal into tasks, gives each running agent its own worktree and branch, and has all of them read one spec from project memory. Blocked-by dependencies order the handoffs, so a verification agent waits for the implementation it checks, and the Autopilot policy caps how many runs execute concurrently.

Does the harness replace Claude Code or Codex?

No, it runs them. Claude Code and Codex CLI start on your own computer under your existing subscriptions, and they can read and update their tasks through the Mrrlin MCP server. What the AI coding agent harness adds is what the agents do not keep on their own: scope, isolation, evidence, review and approval.

What happens when a coding agent's change fails review?

It goes back to the agent with the reviewer's notes instead of landing in your inbox as finished. The task stays open, the failed attempt keeps its artifacts, and a resumed run sends only what changed since the last try. If Autopilot sees the same task fail repeatedly, it stops instead of retrying indefinitely.

Tell us the outcome you want AI to execute.

Share the workflow you want to automate. We’ll map the first Mrrlin run — plan, agent routing, review loops, and approval checkpoints.

No credit card · No migration · One goal