Your agents: governed on every call. Assured for life.
Anyone can ship an agent. Staying in control of it is the hard part — and that is what Brutor Agent Control is built for. You decide what each agent may do; every call it makes is enforced against that decision; and the agent is watched for as long as it runs — still behaving the way you approved, still showing up for work, still worth its cost per completed task. And when something does need a person, it lands in one queue, gets a decision with a name on it, and that trail is in the report.
Day one is easy. These questions decide day 201.
Will the agent still do its job in six months?
Baselines learned from its real runs; any change reported with the likely cause.Watch · baselines & drift →
Can we answer “what did that agent actually do?”
One signed run per task — every action, one outcome, one cost.Run · the signed ledger →
How do agents leave pilot — without betting the company?
A written contract, a test gate, and approvals exactly where you want a human.Define & Promote →
Two teams, one agent — and Finance kept to itself?
Build it once, share it governed — per-team rules, limits and visibility.The agent registry →
What stops a runaway loop or a burned budget?
Budgets, rate and action ceilings enforced before the bill exists.Run · limits →
Who notices the agent that didn’t come back?
Brutor does — every agent has a learned rhythm; silence raises the alert, in Slack.Watch · liveness →
Who actually looked at it — and what did they decide?
Every finding lands in one queue. Closing one takes a disposition and a note, kept as evidence.Respond · the assurance inbox →
One lifecycle answers all of them — and it’s a loop, not a line.
Follow an agent from your IDE to retirement. You build; Brutor governs everything after the handover — and keeps governing, every day it runs.
Seven stages, one boundary, one gate, one closed loop. Respond changes enforcement before the next request lands — which is what makes this a loop rather than a launch checklist.
Develop with whatever you prefer
LangGraph, CrewAI, AutoGen, the Claude Agent SDK, or plain code — your agent, your framework, your process. When it’s ready, one line of code points it at Brutor: swap the base URL, and the agent is on the control plane. The brutor CLI does it for your local tools in about a minute.
Write down what the agent may do
Every surface it touches, in one contract: which models it may call, which tools it may use, which agents it may talk to, which skills it may run — plus what it may spend and where its data may go. Versioned and hashed, so “what was it allowed to do in March?” has an exact answer.
Two entries in the agent registry: a passport and a listing
Defining an agent gives it an Identity — its passport, issued through your IdP and carrying its grants — and an Agent Card, its listing: what it advertises for others to discover, signed, over the open A2A standard. One card can serve many clients, each under its own rules and audit trail.
Tools connect once, at team level
Plug an MCP server in once, at team level, and every agent below inherits exactly the access you set — the same way you publish skills and clear the agents it may call over A2A. Nobody wires the same tool up twice.
Production is earned, not clicked
Before the agent goes live, three things have to be true: its contract resolves, recorded test tasks replay successfully, and a named person owns it — and for a high-risk system, a fourth. Pass, and it’s active — with the approved contract stamped on everything it does from then on.
High-risk systems clear a higher bar — automatically
Declare a system’s EU AI Act risk tier and your role — provider, deployer, or both — and the gate tightens itself. A high-risk system will not transition without a named approver on every step and an unexpired impact assessment attached as evidence. Nothing to remember at audit time: it could not have gone live without it.
That same declaration is what the obligations register reads. Which articles apply to which system is derived, not ticked — from the tier, the role and the models each system is bound to — and every obligation carries its date, so “what changes in December, and for which systems?” is a question with an answer rather than a project.
And nothing is unchecked in between
The replay suite covers day one: recorded tasks, run against the agent before it ships. Baselines need about fifty real runs before they mean anything. The two cover different windows on purpose — replay proves it works now, assurance proves it still works later.
Every call decided in flight
Each call is checked against the contract as it happens: allowed, denied, or held for approval. Content is guarded in both directions, budgets are counted, and it all lands in one signed run — one task, one outcome, and the cost per completed task — not just cost per call. Enforcing guardrails on live traffic is what a good gateway does; Brutor does it on every call, and then keeps going.
Secrets and recipes never reach the agent
Skills run in Brutor’s sandbox — the agent gets the result, the recipe stays protected. The OAuth broker attaches real tokens at the gateway, so no credential ever rides in agent code. And agent-to-agent chains run only as deep as policy allows.
It learns first, then it judges
Assurance earns its verdicts before it gives them. Until an agent has about fifty runs to learn from, Watch is building that agent’s own baseline from its real traffic — action counts, tool mix, cost per completed task and how it grows, rhythm — and while it does, the verdict reads learning. Never a green light it hasn’t earned. A busy agent gets there in a day; a weekly batch job takes longer, and Brutor says so rather than guessing. Once the baseline is the agent’s own, movement comes with its likely cause and silence comes with an alert.
Verdicts — and every agent starts at learning
Drift arrives with a cause
“Something changed” is an alert nobody can act on. Brutor names the likely reason — a new model version, a changed tool definition, different inputs — so the first question in the incident channel is already answered.
Baselines catch what changed. Checks catch what you asked about.
A baseline notices movement you didn’t predict. A continuous check answers a question you did: “flag any run that used a payment tool without an approval”, “alert me if more than 6% of claims runs end in an error over 24 hours.” Deterministic checks compile to a bounded expression over the run ledger and are evaluated the moment a run closes — no model involved, fully reproducible. Where a question genuinely needs to read content, a governed model judges it with its sampling rate declared and its token budget capped, and every result carries the fingerprint that makes it repeatable.
Plain English is how you write a check, never how it runs. What executes is a versioned, hash-pinned spec you approved — and before it is allowed to raise anything, you can backtest it against your recorded runs: “this would have flagged 187 of the last 4,737.”
Measured action, automatically
When something moves, the response fits the problem: tighten one permission, reroute to a safer model, drop the agent to approval-required — or pause it. Previewed first, recorded always, and in effect before the next request lands. That last part is what closes the loop: enforcement changes, and Run picks the change up immediately.
One queue — and a name against every decision
Everything worth a human’s attention lands in a single Assurance Inbox: a guardrail block, a drift finding, a check breach, a budget hard-stop, an agent that went quiet. Recurring signals deduplicate — a storm is one row with a counter, not five hundred alerts. Closing an item takes a disposition and a note — true positive, false positive, expected, or fixed — and those are kept as an append-only trail that becomes the report’s human-review evidence.
A critical item nobody has looked at past the review SLA is itself a health signal: review debt drags the score down, so an ignored queue cannot quietly read green.
Told, not just recorded
Findings route to Slack or a signed webhook as they are raised, filtered by severity, category or system, and coalesced per finding so a recurring problem is one message per window rather than a pager storm. Every send is recorded — and a channel that stops delivering becomes a finding of its own, because a dead alert channel is the one failure that hides all the others. A weekly digest summarises the same numbers as the report, never a second calculation.
Everything the loop produces — runs, drift events, check results, responses, and the dispositions someone signed — rolls up into an assurance report you can hand to an auditor. It is also why agents elsewhere get demoted or switched off: when the only available response is all-or-nothing, all-off wins.
And when its work is done — a clean exit
Deprecated, then retired: access wound down, record intact. The agent’s whole history stays answerable, even after it’s gone.
New white paper: AI Systems Assurance
The full method behind the Watch and Respond stages you just read — why AI systems fail differently from traditional software, and how the four guarantees (bounded, alive, non-drifting, doing its job) plus the human-review evidence that sits alongside them answer the two questions that stall adoption: who signs their name under this, and how do we know it still works six months in.
A human in the loop, at four points in the lifecycle
Not a switch you flip in an emergency — a thread that runs the whole way through. Mark an action approval-required when you define it. Run pauses it and waits for sign-off rather than guessing. Respond can dial a whole agent down to approvals-only when its behaviour moves. And when something needs judgement rather than a decision in flight, it waits in the Assurance Inbox until a person dispositions it — with their name on the record. The same mechanism at all four points, so “where does a person decide?” has one answer instead of four.
What it looks like on one real system.
The Claims Co-Pilot: five governed actions, the rule that applied on each, and the telemetry landing in Mission Control.
Clear lines, on purpose.
Nobody should need a meeting to work out whose job something is. Here is the split, stage by stage: what Brutor decides, and what stays yours.
| Stage | Brutor handles | You decide |
|---|---|---|
| Define | The contract — grants, limits, residency — versioned, hashed, dry-runnable. | What the agent is for, its code and prompts, which tools to trust. |
| Promote | The gate: contract checks, replay suite, evidence for high-risk, autonomy level. | Whether the risk is acceptable, and when to ship. |
| Run | Allow / deny / approve on every call; budgets; the signed record. | What the agent does between calls, and the quality of its answers. |
| Watch | Baselines, drift causes, liveness, continuous checks, the health score. | Which questions are worth asking — and whether the outcomes serve the business. |
| Respond | Tighten, reroute, de-escalate, pause — previewed and audited; the finding queue and its alerts. | The response policy, the disposition on every finding, and fixing or redeploying the agent. |
Five things that make the lifecycle real.
Permissions you can point to
Every agent’s rules are written down, versioned, and stamped on everything it does.
Production is earned
An agent goes live when its contract resolves, its replays pass, and a named person owns it.
Every action on the record
One task becomes one signed run — every action, one cost, nothing to reconstruct later.
You hear it from Brutor first
Change arrives with its likely cause, your own standing checks run on every task, and silence arrives as an alert.
Measured — and signed for
Tighten, add an approval, or pause — matched to the problem. Then a person closes the finding, on the record.
The whole story, on video.
Deep dives into every capability on this page — recorded on the live platform, a few minutes per episode.
Agent Governance with Brutor
Five episodes: AI Systems, identity & authorization, the agent network (A2A), and observability.
Watch the series →AI Systems Assurance with Brutor
Seven episodes: the run ledger, liveness, drift with a cause, contracts & the gate, replay, and the health report.
Watch the series →Run one agent through the whole loop.
Download the trial, point an agent at the Gateway, and watch a contract, a run, a baseline, a check and a response appear against traffic you recognize.
