The AI control plane: define what your AI may do, and prove it still does.

An AI control plane is the management and governance layer above your AI traffic. Your AI Systems, policies, access, limits and contracts are defined here; the AI Gateway enforces them on every call; and every recorded run comes back here to be monitored and assured against what was approved.

This page is for the engineer evaluating one: what a control plane is, how it relates to the gateway, what sets Brutor’s apart, and how it goes into production. All of it ships together. Nothing is priced per module.

The Brutor AI Gateway and the Brutor AI Control Plane as two layers. At the top, the AI Gateway is the data plane: AI agents, workflows and applications, and chat clients such as the Brutor User Portal and Goose used by end users, send every call through it. Each request passes identity, policy, guardrails and limits before it is routed to models, MCP servers, A2A agents or skills; each response passes output guardrails and cost metering and is recorded. The gateway reads its configuration and never writes it. In the middle, PostgreSQL, state you own, holds configuration, runs and evidence. At the bottom, the AI Control Plane, used by admins through the Admin Console, writes that configuration (AI Systems and agents, policies and contracts, access and permissions, limits and budgets) and reads the recorded runs back for the asset registry, usage and cost, assurance and drift, and reports and evidence: a compliant system continues, a deviation is blocked, alerted or investigated. Non-routable vendor chat apps (ChatGPT, Claude Desktop, Microsoft Copilot) are observed only, through imported usage, and shadow AI is found by discovery agents.
what a control plane is

The layer that decides what your AI should do.

The idea comes from Kubernetes. Its data plane runs the workloads and carries the traffic; its control plane decides what should run, where, and under which configuration and policies. An AI control plane does the same for AI: it holds what each AI System is allowed to do, keeps the configuration that makes that true, watches what actually happened, and governs the whole estate, including AI that no gateway will ever see.

01Define

What each AI System is

An owner, a purpose and a contract: the models, tools, skills and agents it may use, what it may reach and what it may spend.

02Manage

The configuration behind it

Policies, access, guardrails, limits and budgets, managed in one place, resolved down your organization tree, and exportable as policy-as-code.

03Monitor

What actually happened

Every run the gateway recorded and its cost per completed task, plus the usage of the AI you buy and the AI nobody registered.

04Assure

Whether it still holds

Each system judged against what was approved: drift, silence and cost changes flagged with their likely cause, and evidence an auditor can read.

It governs the whole estate, not only the traffic it can route

Not all AI can pass through a gateway, so the control plane’s registry carries one of three states for every asset. Nobody mistakes one number for the whole estate.

Governed

Enforced in the request path

Your own AI Systems route through the gateway. Everything applies: guardrails both ways, policy, budgets, identity, and a record of every decision.

Observed

Imported and costed

The AI you buy runs in the vendor’s own app, so nothing can stand in the path of those calls. Usage is imported instead: inventory, cost in the same ledger, budget alerts.

Discovered

Found, and awaiting onboarding

Shadow AI, surfaced by adapters on the scanners you already run. It lands in the registry with an owner and a purpose, and a pre-filled path to Governed.


how it relates to the gateway

The gateway enforces. The control plane decides what to enforce.

An AI gateway and an AI control plane are related, but they are not the same thing. The gateway is the runtime traffic layer, in the path of every call. The control plane is the management and governance layer above it: it defines what the gateway enforces, and judges what the gateway recorded. Brutor ships both, as separate planes.

CapabilityBrutor AI Control PlaneBrutor AI Gateway
Primary roleManagement and governance layer: decides what is enforcedRuntime traffic layer: the enforcement point
In the path of AI callsNo, never in the request pathYes, every call
Model, MCP, agent and skill callsConfigures what may be reached, and by whomRoutes and proxies them
Authentication and accessDefines users, keys, groups and grantsEnforces on every call
Guardrails and policiesDefines and versions themEnforces, including as responses stream
Rate and spend limitsSets them down your organization treeEnforces before the call is made
ObservabilityBroader: runs, cost per task, and the AI you buy or never registeredRecords every call as it happens
AI System inventoryThe asset registry: Governed, Observed, DiscoveredNot its job
ContractsGenerates, hash-pins and promotes themEnforces per-run ceilings and stamps each run
Drift and behavior assuranceBaselines, drift, liveness and the verdictSupplies the evidence
Compliance evidenceReports, framework mappings and exportsWrites the audit record
Agent lifecycleDefine, promote, run, watch, respond, retireNot its job

Define, enforce, observe, assure

One loop, split across the two planes. The last step is the one a gateway on its own cannot take.

01020304
Define

Control plane. Contracts, policies, access and limits, set in the Admin Console or applied as policy-as-code.

Enforce

Gateway. Every call checked in the path before it leaves. The gateway reads the configuration and never writes it.

Observe

Gateway, then control plane. Every call recorded as it happens and grouped into runs, in a database you own.

Assure

Control plane. Each AI System judged against what was approved. Compliant: continue. Deviation: block, alert or investigate.

A note on names

The market has not settled its words yet. Some products sold as AI gateways carry a good deal of control-plane function; others call a dashboard a control plane, and a gateway can grow into one. Two questions cut through it: is the thing that enforces separate from the thing that decides, and does anything prove afterwards that what was decided actually held?

what sets brutor’s apart

Most control planes govern the call. Brutor proves the system still works.

Request-level governance answers “was this call allowed?” For an agent, that is the easy half. A model chooses its own next step, so an agent can take twelve steps where it used to take four, go quiet, or drift after a model update, while every call returns 200 and every dashboard stays green. Assurance is the layer that notices.

The unit is a run, not a request

01

Every call of one task, across models, tools, skills and delegated agents, is grouped into a single run with an honest terminal state: completed, degraded, errored or abandoned.

Normal is learned, per system

02

Baselines come from each AI System’s own runs: step counts, tool mix, cost per completed task, outcomes. While a baseline is still forming, the verdict reads learning, never a passing grade it hasn’t earned.

Health is the worst signal, never the average

03

Liveness, behavior, reliability, cost, conformance and oversight combine worst-of into one verdict. Too little evidence reads unknown. Absence of evidence is never green.

shaped to your organization

Your org chart is the policy model.

Everything in the control plane is scoped to a node in one tree that mirrors your organization. Models, tools, skills, keys and limits all hang off it, so a grant or a cap means the same thing wherever you look.

One tree, the way you actually operate

Model your real structure, not the one a template assumes. Organizations hold departments, departments hold teams, and each AI System sits where the people who own it do.

  • Organization. The ceiling: the budgets, guardrails and policies that finance, security and compliance set for everyone below.
  • Department. A business unit, such as Corporate Functions or Customer Operations, working within its share of that ceiling.
  • Team. The people who build and run AI day to day, such as Finance, HR or Legal and Compliance.
  • AI System. An agent, assistant or workflow: the unit that gets a contract, a run ledger and an assurance verdict.

Grant resources once, where they belong

Models, MCP servers, agents, skills and knowledge bases can be granted at any level of the tree, and everything below inherits them. Give the whole organization its approved models once, a department the MCP servers its work needs, and a team only the few tools that are truly its own. What a group can use is what it is granted plus what it inherits, so nobody lists every resource for every group, and a parent can still withhold an inheritance where a group should not have it.

A new team picks up the resources and governance above it the moment it exists. Reorganize, and every grant, limit and policy recalculates down the new tree, with nothing re-templated by hand.

Policies and limits travel down the same tree, but by a different rule. The diagram below shows how.

Brutor Admin Console, the Resource Groups tree. Beside Default (Custom) and ACME Corp (Organization), Nordica Insurance Group (Organization) is expanded: Corporate Functions (Department) holds the Finance team with the Invoice Processing Workflow AI System, the HR team with the CV Screening Agent and HR Policy Assistant, and the Legal and Compliance team with the Contract Review Assistant; Customer Operations (Department) holds the Claims Triage Agent, Complaint Summarization Pipeline and Customer Service Chatbot AI Systems. Collapsed below: a Discovered / Ungoverned team and the Engineering, Sales and Marketing, and Underwriting and Risk departments.

Two rules, and only two

Resources and limits travel down the same tree, and they compose in opposite directions. That asymmetry is what makes delegation safe: a team lead can tighten anything without a ticket, and nobody can grant themselves more than their parent allows.

A resource group tree runs Organization Nordica at one thousand dollars a day, Department Claims at four hundred, Team Claims Triage at five hundred, and the Triage Agent with no cap of its own. Resources compose additively, gated so a parent can withhold. Limits compose restrictively: the effective daily cap is four hundred dollars, set at the department. The team asking for five hundred does not raise the four hundred it sits under. The same most-restrictive rule covers tenant-wide limits set directly on a model, MCP server or skill.
Resources compose additively, and each inheritance is gated. Limits compose restrictively: the most restrictive value on the path is the one enforced.
Opt out of tooling, never out of governance

A sub-group that needs different tooling (an R&D sandbox, a coding environment) switches off resource inheritance and curates its own catalog of models, MCP servers, skills and knowledge bases. Every limit and policy above it still applies.


observe the estate

Observe your organization’s AI estate.

The tree that shapes governance is also how you watch it. Every department, team and AI System reports its spend, runs and verdict in one place; anything that needs a person lands in one inbox; and continuous checks you define test every recorded run against the rules your organization cares about.

Mission Control: the estate at a glance

Each group in the tree is a card: spend, runs and tokens for the window you choose, blended cost per completed task with its trend, the workload class that says how its cost grows, and a pinned health and assurance verdict. Sort by what needs attention and drill down from the organization to a department to a single AI System. A group without enough evidence says so, and discovered AI stays in the picture with its governance gaps counted.

Brutor Admin Console, Mission Control: the AI estate by resource group for Nordica Insurance Group, all time, sorted by what needs attention. Six cards, one per group: Engineering (4 systems, $146.49, 8,555 runs, 58.4M tokens, $0.0172 per task), Underwriting and Risk (3 systems, $50.35, 1,340 runs, $0.0376 per task), Sales and Marketing (3 systems, $22.63, 1,630 runs, $0.0139 per task), Corporate Functions (4 systems, $13.49, 1,085 runs, $0.0124 per task), Discovered / Ungoverned (3 systems, $0.00, 0 runs, insufficient evidence, 12 governance gaps) and Customer Operations (3 systems, $231.93, 12,278 runs, learning, $0.0189 per task). Each card carries a liveness chip, an assurance verdict such as Assured with exceptions, its workload class, a spend sparkline, and Drill down, Usage details and Metrics links.
Mission Control on the Nordica estate: each department carries its spend, runs, cost per task, workload class and verdict, and the discovered group shows its 12 governance gaps.

One inbox for everything that needs a person

Liveness, drift, run failures and breached checks from every AI System arrive as findings in the Assurance Inbox. Each is graded warning or critical, repeats fold into one finding with a count, and it stays open until someone acknowledges it. Nobody has to stare at a dashboard to notice that an agent has gone quiet.

Brutor Admin Console, Assurance Inbox: 8 critical, 27 open, 2 acknowledged, filtered to open and acknowledged, warning and critical, all categories, 18 shown of 34. Findings in time order, each with a severity, a category and a count: check breaches for customer-facing guardrail blocks (102) and the claims triage error rate, 8 of 96 over 24 hours (82); liveness findings for the Claims Triage Agent stalling and for the Credit/Risk Scoring Agent, Complaint Summarization Pipeline, SDR Outreach Agent, Invoice Processing Workflow, Policy Document Extraction and Release Notes Generator going silent; a run that finished errored; and behavioral drift detected on the terminal state mix (13).
The Assurance Inbox: silent and stalling systems, breached checks, errored runs and drift, graded by severity and counted rather than repeated.

Continuous checks, for the rules only you know

Checks are versioned specs that an operator approves before they run, evaluated over the run ledger. Tier A checks are deterministic expressions, evaluated in the gateway core as each run finishes; backtest one against the last seven days before you turn it on.

Brutor Admin Console, Continuous checks: 6 enabled of 7 defined, versioned, operator-approved specs evaluated over the run ledger, with deterministic checks run at finalization in the gateway core and semantic checks judged in governed batches. Three Tier A boolean checks, each approved by admin and enabled: claims triage error rate, customer-facing guardrail blocks, and costly run with no approval step. The last is expanded to its expression, run.total_cost_usd > 0.10 && run.approval_request_count == 0, with Backtest 7 days and Delete buttons and its recent results, one run per line, each evaluating to false.
A Tier A check: a costly run with no approval step, written as one expression over the run, with its backtest and a result for every run.

Tier B checks are semantic: a governed judge model reads a declared sample of runs against a question written in plain language, and every judgment records the tokens it spent. A breached check of either tier becomes a finding in the inbox.

Brutor Admin Console, Continuous checks with seven checks listed: five Tier A boolean checks (claims triage error rate, customer-facing guardrail blocks, costly run with no approval step, underwriting argument-policy denials, high-risk system approved models only) and two Tier B checks. The Tier B check for personal identifiers leaked by extraction is expanded: a plain-language instruction telling the judge to answer true if an insurance policy-document extraction contains a personal identity number, full residential address or bank account number, sampling 25 percent of runs, declared and disclosed, judged by gpt-5.2, and recent results per run with the tokens each judgment spent, two of them true. A second Tier B check, chatbot settlement amount promised, is disabled.
A Tier B check: a judge asked whether an extraction leaked personal identifiers, on a disclosed 25 percent sample, with the tokens behind each verdict.

assure the ai system

Assure each AI System runs according to its contract.

From the organization down to the leaf of the tree: the AI System, the agent, assistant or workflow that does the work. Here the question is no longer what the estate costs, but whether this one system is still doing what was approved. The control plane answers it in four steps: define the contract, record every run, learn the system’s normal, and assure it against both.

Define: a contract, generated rather than written

Before a system goes live, a contract is minted from the controls that actually resolve onto it: the models, MCP servers and agent identities it may use, the guardrails and policies it inherits, what each of its agents may do per action (allowed, held for approval, or denied because it was never granted), and its limits and governance, merged restrictively down the group chain. Each version is frozen, pinned by hash, approved by a named person, and lists what it tightened. What was signed stays readable after the live configuration moves on.

Brutor Admin Console, an AI System contract, version 5, hash 66e8396bdf74, minted and approved by admin, promoted 01/09/2026, active, with download and print buttons. What this version covers, frozen at mint time and pinned by hash: resources it may use (models OpenAI GPT-5.5 and OpenAI GPT-5.2 Auto-Routing, MCP servers Claims Core and DocStore, agent identity claims-triage-worker); policies in force (inherited guardrails Nordica Baseline PII and Secrets, and Customer Ops Injection and Toxicity); what its agents may do, per action, for claims-triage-worker: allow an A2A call to fraud-screening-agent with a maximum delegation depth, allow LLM calls with a rate, allow the MCP tools claims_get_document, claims_lookup, doc_search and fraud_flags_get, and require approval for claims_update_queue; limits and governance merged restrictively (an LLM budget of 40 dollars a day and 2,000 a month with a warning at 80 percent, 5,000,000 tokens and 20,000 requests a day, temperature clamped at 0.3, output clamped at 2,048 tokens); autonomy autonomous, inherited through 3 groups; grants nothing new; and four tightenings added in this version for context window and temperature.
Contract v5: what the system may use, the guardrails it inherits, what its agent may do per action with one tool held for approval, its merged limits, and what this version tightened.

Run: every task becomes a run

In production, every call the system makes through the gateway, whether a model call, an MCP tool call or a delegation to another agent, is grouped by task into a run with an honest outcome: completed, errored or abandoned. The ledger records who started it, how it closed, what it cost, and whether its chain of calls was signed by the gateway or only asserted by the client. Open a run to see the models it used and every action in order.

Brutor Admin Console, the Runs tab of an AI System. A table of runs with started and finished times, actor, outcome, turns, actions, tools, cost, closed by and chain: two abandoned runs closed idle with intact chains, an errored run closed explicitly with a client-asserted chain, another abandoned run, and a completed run of 7 actions and 1 tool costing $0.0583, closed explicitly with a client-asserted chain. The completed run is expanded: gpt-5.2 with 4 calls and 4,200 tokens for $0.0046, gpt-5.5 with 2 calls and 6,347 tokens for $0.0537, no errors and nothing blocked, and its seven actions in order, llm calls to gpt-5.2 and gpt-5.5 and an mcp call to doc_get, each with its duration.
The run ledger: each run with its actor, outcome, cost, how it closed and whether its chain was signed, and inside one run the models it used and its actions in order.

Learn: what normal looks like for this system

From those runs, the control plane learns each system’s own baseline, signal by signal, instead of relying on thresholds someone guessed. Liveness is the one expectation you set, and Brutor can suggest a cadence from the runs it has already seen. It is the only signal that fires when traffic stops, even though nothing errored.

Outcomes, tool errors, trajectory length and cost per completed run are observed with their history, so a change reads as a step, a slow drift or one bad afternoon. Money spent on runs that produced nothing is counted on its own.

Trajectory length, observed: median 6 actions, p95 9, maximum 10, exhausted 0 percent (0 runs), a chart of average actions per run holding near 6 over time, and a histogram of runs by action count (45 with 1 to 2, 1,014 with 3 to 5, 1,469 with 6 to 10, none above), with a note that 79 runs were closed by idle timeout, usually a client that stopped calling.
Trajectory length: how many actions a run takes, and how that is spread.
Cost per completed run, observed: $0.0189 across 2,335 completed runs, median $0.0162, p95 $0.0446, and $3.4995 spent on errored, abandoned or exhausted runs that produced nothing, with a line chart of cost per completed run over time moving between roughly one and two and a half cents.
Cost per completed run, and the spend that produced nothing.

Assure: a verdict in words, with its limits stated

Everything above rolls up into an assurance report for the system: a status, what we know, what prevents full assurance, the recommended action, and why it matters. Health is the worst of six signals, never the average, and a system scored as an assistant discloses which signals are relaxed for its kind.

The contract is checked item by item, and an evidence maturity ladder shows the current rung, what blocks the next one and what to do about it. Capture a snapshot, or export the report as JSON or PDF for whoever asked.

The agent lifecycle, end to end

The Assurance tab of an AI System, Contract Review Assistant, with Capture snapshot, Export JSON, Export PDF and Refresh buttons. Assurance status Assured, evidence maturity Observable. What we know: it operated within its approved operating envelope during the reporting period. What prevents full assurance: 48 percent of runs lack cryptographically signed execution chains, and no independent evidence is attached. Recommended action: remain under monitored operation until the outstanding requirements are completed. Five evidence dimensions: operational evidence passed, behavioral stability failed while baselines are still learning, contract assurance passed, execution evidence warned, independent evidence failed. Health still learning at 94 percent, with each signal scored and a note that scoring as an assistant relaxes 2 of 6 signals. Contract compliance met for tokens per run, cost ceiling for one run and model calls per run. Evidence maturity at level 0, Observable: L2 Controlled reached, L1 Baseline established blocked until baselines finish learning, and L3 Evidence-backed and L4 Independently assured not yet.

how it goes into production

From download to production.

Self-hosted, inside your perimeter, on the same images from the first evaluation to the production cluster. No rewrite, no migration project, no new SDK.

0102030405
Day one: evaluate

Download the trial and start it with Docker Compose on a machine you control. It is the stack we ship, not a sandbox, and the benchmark tooling is included so you can measure governance-on latency on your own hardware.

Day one: connect

Change one base URL. The gateway speaks the OpenAI-compatible API and the native Anthropic Messages API, so existing SDKs, agent frameworks and coding tools keep working.

First week: shape

Model your organization as resource groups, sign people in through your identity provider, and set budgets, guardrails and policies. Teams onboard, and the Portal gives them a better tool than the one it replaces.

Week two: assure

Baselines learned from real runs; drift and liveness verdicts begin. Until they are ready, the verdict reads learning.

Day thirty

An inventory, one cost ledger, evidence accumulating and agents under assurance. And if you walk away, that is a config change too.

Where it runs

On-premises or in your private cloud (AWS, Azure or GCP): your account, your VPC. There is no productized public SaaS; we can build and operate a managed deployment for you on request.

Kubernetes

A Helm chart for EKS, AKS, GKE or OpenShift. Pods run as non-root with every capability dropped and read-only root file systems on the Rust services. The gateway core autoscales; bring your own PostgreSQL, Redis and Qdrant.

AWS reference architecture

Terraform for ECS Fargate behind an Application Load Balancer, with Aurora PostgreSQL, ElastiCache Redis and Qdrant in private subnets, secrets in Secrets Manager and TLS at the load balancer.

On-premises and air-gapped

Your hardware, your perimeter. Air-gapped installs are possible under the right conditions; tell us about yours.

Operating it

Scale the hot path

Add gateway instances without ceremony: configuration invalidates across them through PostgreSQL, and rate-limit and quota counters are shared in Redis. The control plane only takes admin traffic.

Your identity provider

Single sign-on through SAML 2.0 or OIDC (Okta, Microsoft Entra ID, Ping, Google Workspace and others) with group-based role mapping. Agent identities anchor in your IdP too. Brutor decides; it doesn’t store identities.

Your observability

OpenTelemetry traces and metrics over OTLP to the stack you already run: your dashboards, not another silo.

Policy-as-code

The whole governance posture as YAML with a Git history: reviewable, promotable, revertible.

Upgrades you control

Pin a version and roll forward on a frequent release train. The documentation is version-locked to the platform, so it always matches what you run.

Your evidence stays yours

The asset registry and the audit evidence export in standard formats, on demand, so what you accumulate here does not become ours.


questions engineers ask

The short answers.

What is the difference between an AI gateway and an AI control plane?

The gateway is the enforcement point: it sits in the path of every call and applies policy, guardrails and limits. The control plane defines, manages, monitors and governs what the gateway should enforce, and judges what it recorded. Brutor ships both, as separate planes: the Brutor AI Gateway underneath, the Brutor AI Control Plane above it.

Does it add latency?

The full pipeline (authentication, access checks, routing, guardrails, quotas and cache) measured statistically indistinguishable from calling the model directly at tested concurrency. The benchmark and its method are published, and the trial includes the tooling to rerun it on your hardware.

Can Brutor govern ChatGPT, Claude Desktop or Microsoft Copilot?

No. A vendor app talks to the vendor service, so nothing can stand in the path of those calls. Brutor imports their usage and cost from the vendor enterprise APIs and marks them Observed rather than Governed.

Do I have to rewrite my agent?

No. Brutor is a base URL and a scoped key. Existing OpenAI and Anthropic clients work unchanged, and MCP and A2A are the real protocols rather than a Brutor dialect.

What happens when a guardrail provider or policy judge is down?

The call is refused. Guardrails and semantic policies fail closed by default, after retry, circuit-breaker and local-fallback stages, with an expiring break-glass that is itself recorded.

Is there a hosted SaaS?

Not a productized one. Brutor is self-hosted, on-premises or in your private cloud. We can build and operate a managed deployment for you on request.

/take control/

Put every model, tool and agent
on one governed path.

Stay connected. Sign up for updates from Brutor AI.