What each AI System is
An owner, a purpose and a contract: the models, tools, skills and agents it may use, what it may reach and what it may spend.
An AI control plane is the management and governance layer above your AI traffic. Your AI Systems, policies, access, limits and contracts are defined here; the AI Gateway enforces them on every call; and every recorded run comes back here to be monitored and assured against what was approved.
This page is for the engineer evaluating one: what a control plane is, how it relates to the gateway, what sets Brutor’s apart, and how it goes into production. All of it ships together. Nothing is priced per module.
The idea comes from Kubernetes. Its data plane runs the workloads and carries the traffic; its control plane decides what should run, where, and under which configuration and policies. An AI control plane does the same for AI: it holds what each AI System is allowed to do, keeps the configuration that makes that true, watches what actually happened, and governs the whole estate, including AI that no gateway will ever see.
An owner, a purpose and a contract: the models, tools, skills and agents it may use, what it may reach and what it may spend.
Policies, access, guardrails, limits and budgets, managed in one place, resolved down your organization tree, and exportable as policy-as-code.
Every run the gateway recorded and its cost per completed task, plus the usage of the AI you buy and the AI nobody registered.
Each system judged against what was approved: drift, silence and cost changes flagged with their likely cause, and evidence an auditor can read.
Not all AI can pass through a gateway, so the control plane’s registry carries one of three states for every asset. Nobody mistakes one number for the whole estate.
Your own AI Systems route through the gateway. Everything applies: guardrails both ways, policy, budgets, identity, and a record of every decision.
The AI you buy runs in the vendor’s own app, so nothing can stand in the path of those calls. Usage is imported instead: inventory, cost in the same ledger, budget alerts.
Shadow AI, surfaced by adapters on the scanners you already run. It lands in the registry with an owner and a purpose, and a pre-filled path to Governed.
An AI gateway and an AI control plane are related, but they are not the same thing. The gateway is the runtime traffic layer, in the path of every call. The control plane is the management and governance layer above it: it defines what the gateway enforces, and judges what the gateway recorded. Brutor ships both, as separate planes.
| Capability | Brutor AI Control Plane | Brutor AI Gateway |
|---|---|---|
| Primary role | Management and governance layer: decides what is enforced | Runtime traffic layer: the enforcement point |
| In the path of AI calls | No, never in the request path | Yes, every call |
| Model, MCP, agent and skill calls | Configures what may be reached, and by whom | Routes and proxies them |
| Authentication and access | Defines users, keys, groups and grants | Enforces on every call |
| Guardrails and policies | Defines and versions them | Enforces, including as responses stream |
| Rate and spend limits | Sets them down your organization tree | Enforces before the call is made |
| Observability | Broader: runs, cost per task, and the AI you buy or never registered | Records every call as it happens |
| AI System inventory | The asset registry: Governed, Observed, Discovered | Not its job |
| Contracts | Generates, hash-pins and promotes them | Enforces per-run ceilings and stamps each run |
| Drift and behavior assurance | Baselines, drift, liveness and the verdict | Supplies the evidence |
| Compliance evidence | Reports, framework mappings and exports | Writes the audit record |
| Agent lifecycle | Define, promote, run, watch, respond, retire | Not its job |
One loop, split across the two planes. The last step is the one a gateway on its own cannot take.
Control plane. Contracts, policies, access and limits, set in the Admin Console or applied as policy-as-code.
Gateway. Every call checked in the path before it leaves. The gateway reads the configuration and never writes it.
Gateway, then control plane. Every call recorded as it happens and grouped into runs, in a database you own.
Control plane. Each AI System judged against what was approved. Compliant: continue. Deviation: block, alert or investigate.
The market has not settled its words yet. Some products sold as AI gateways carry a good deal of control-plane function; others call a dashboard a control plane, and a gateway can grow into one. Two questions cut through it: is the thing that enforces separate from the thing that decides, and does anything prove afterwards that what was decided actually held?
Request-level governance answers “was this call allowed?” For an agent, that is the easy half. A model chooses its own next step, so an agent can take twelve steps where it used to take four, go quiet, or drift after a model update, while every call returns 200 and every dashboard stays green. Assurance is the layer that notices.
Every call of one task, across models, tools, skills and delegated agents, is grouped into a single run with an honest terminal state: completed, degraded, errored or abandoned.
Baselines come from each AI System’s own runs: step counts, tool mix, cost per completed task, outcomes. While a baseline is still forming, the verdict reads learning, never a passing grade it hasn’t earned.
Liveness, behavior, reliability, cost, conformance and oversight combine worst-of into one verdict. Too little evidence reads unknown. Absence of evidence is never green.
Everything in the control plane is scoped to a node in one tree that mirrors your organization. Models, tools, skills, keys and limits all hang off it, so a grant or a cap means the same thing wherever you look.
Model your real structure, not the one a template assumes. Organizations hold departments, departments hold teams, and each AI System sits where the people who own it do.
Models, MCP servers, agents, skills and knowledge bases can be granted at any level of the tree, and everything below inherits them. Give the whole organization its approved models once, a department the MCP servers its work needs, and a team only the few tools that are truly its own. What a group can use is what it is granted plus what it inherits, so nobody lists every resource for every group, and a parent can still withhold an inheritance where a group should not have it.
A new team picks up the resources and governance above it the moment it exists. Reorganize, and every grant, limit and policy recalculates down the new tree, with nothing re-templated by hand.
Policies and limits travel down the same tree, but by a different rule. The diagram below shows how.

Resources and limits travel down the same tree, and they compose in opposite directions. That asymmetry is what makes delegation safe: a team lead can tighten anything without a ticket, and nobody can grant themselves more than their parent allows.
A sub-group that needs different tooling (an R&D sandbox, a coding environment) switches off resource inheritance and curates its own catalog of models, MCP servers, skills and knowledge bases. Every limit and policy above it still applies.
The tree that shapes governance is also how you watch it. Every department, team and AI System reports its spend, runs and verdict in one place; anything that needs a person lands in one inbox; and continuous checks you define test every recorded run against the rules your organization cares about.
Each group in the tree is a card: spend, runs and tokens for the window you choose, blended cost per completed task with its trend, the workload class that says how its cost grows, and a pinned health and assurance verdict. Sort by what needs attention and drill down from the organization to a department to a single AI System. A group without enough evidence says so, and discovered AI stays in the picture with its governance gaps counted.
Liveness, drift, run failures and breached checks from every AI System arrive as findings in the Assurance Inbox. Each is graded warning or critical, repeats fold into one finding with a count, and it stays open until someone acknowledges it. Nobody has to stare at a dashboard to notice that an agent has gone quiet.
Checks are versioned specs that an operator approves before they run, evaluated over the run ledger. Tier A checks are deterministic expressions, evaluated in the gateway core as each run finishes; backtest one against the last seven days before you turn it on.
Tier B checks are semantic: a governed judge model reads a declared sample of runs against a question written in plain language, and every judgment records the tokens it spent. A breached check of either tier becomes a finding in the inbox.
From the organization down to the leaf of the tree: the AI System, the agent, assistant or workflow that does the work. Here the question is no longer what the estate costs, but whether this one system is still doing what was approved. The control plane answers it in four steps: define the contract, record every run, learn the system’s normal, and assure it against both.
Before a system goes live, a contract is minted from the controls that actually resolve onto it: the models, MCP servers and agent identities it may use, the guardrails and policies it inherits, what each of its agents may do per action (allowed, held for approval, or denied because it was never granted), and its limits and governance, merged restrictively down the group chain. Each version is frozen, pinned by hash, approved by a named person, and lists what it tightened. What was signed stays readable after the live configuration moves on.
In production, every call the system makes through the gateway, whether a model call, an MCP tool call or a delegation to another agent, is grouped by task into a run with an honest outcome: completed, errored or abandoned. The ledger records who started it, how it closed, what it cost, and whether its chain of calls was signed by the gateway or only asserted by the client. Open a run to see the models it used and every action in order.
From those runs, the control plane learns each system’s own baseline, signal by signal, instead of relying on thresholds someone guessed. Liveness is the one expectation you set, and Brutor can suggest a cadence from the runs it has already seen. It is the only signal that fires when traffic stops, even though nothing errored.
Outcomes, tool errors, trajectory length and cost per completed run are observed with their history, so a change reads as a step, a slow drift or one bad afternoon. Money spent on runs that produced nothing is counted on its own.
Everything above rolls up into an assurance report for the system: a status, what we know, what prevents full assurance, the recommended action, and why it matters. Health is the worst of six signals, never the average, and a system scored as an assistant discloses which signals are relaxed for its kind.
The contract is checked item by item, and an evidence maturity ladder shows the current rung, what blocks the next one and what to do about it. Capture a snapshot, or export the report as JSON or PDF for whoever asked.

Self-hosted, inside your perimeter, on the same images from the first evaluation to the production cluster. No rewrite, no migration project, no new SDK.
Download the trial and start it with Docker Compose on a machine you control. It is the stack we ship, not a sandbox, and the benchmark tooling is included so you can measure governance-on latency on your own hardware.
Change one base URL. The gateway speaks the OpenAI-compatible API and the native Anthropic Messages API, so existing SDKs, agent frameworks and coding tools keep working.
Model your organization as resource groups, sign people in through your identity provider, and set budgets, guardrails and policies. Teams onboard, and the Portal gives them a better tool than the one it replaces.
Baselines learned from real runs; drift and liveness verdicts begin. Until they are ready, the verdict reads learning.
An inventory, one cost ledger, evidence accumulating and agents under assurance. And if you walk away, that is a config change too.
On-premises or in your private cloud (AWS, Azure or GCP): your account, your VPC. There is no productized public SaaS; we can build and operate a managed deployment for you on request.
A Helm chart for EKS, AKS, GKE or OpenShift. Pods run as non-root with every capability dropped and read-only root file systems on the Rust services. The gateway core autoscales; bring your own PostgreSQL, Redis and Qdrant.
Terraform for ECS Fargate behind an Application Load Balancer, with Aurora PostgreSQL, ElastiCache Redis and Qdrant in private subnets, secrets in Secrets Manager and TLS at the load balancer.
Your hardware, your perimeter. Air-gapped installs are possible under the right conditions; tell us about yours.
Add gateway instances without ceremony: configuration invalidates across them through PostgreSQL, and rate-limit and quota counters are shared in Redis. The control plane only takes admin traffic.
Single sign-on through SAML 2.0 or OIDC (Okta, Microsoft Entra ID, Ping, Google Workspace and others) with group-based role mapping. Agent identities anchor in your IdP too. Brutor decides; it doesn’t store identities.
OpenTelemetry traces and metrics over OTLP to the stack you already run: your dashboards, not another silo.
The whole governance posture as YAML with a Git history: reviewable, promotable, revertible.
Pin a version and roll forward on a frequent release train. The documentation is version-locked to the platform, so it always matches what you run.
The asset registry and the audit evidence export in standard formats, on demand, so what you accumulate here does not become ours.
The gateway is the enforcement point: it sits in the path of every call and applies policy, guardrails and limits. The control plane defines, manages, monitors and governs what the gateway should enforce, and judges what it recorded. Brutor ships both, as separate planes: the Brutor AI Gateway underneath, the Brutor AI Control Plane above it.
The full pipeline (authentication, access checks, routing, guardrails, quotas and cache) measured statistically indistinguishable from calling the model directly at tested concurrency. The benchmark and its method are published, and the trial includes the tooling to rerun it on your hardware.
No. A vendor app talks to the vendor service, so nothing can stand in the path of those calls. Brutor imports their usage and cost from the vendor enterprise APIs and marks them Observed rather than Governed.
No. Brutor is a base URL and a scoped key. Existing OpenAI and Anthropic clients work unchanged, and MCP and A2A are the real protocols rather than a Brutor dialect.
The call is refused. Guardrails and semantic policies fail closed by default, after retry, circuit-breaker and local-fallback stages, with an expiring break-glass that is itself recorded.
Not a productized one. Brutor is self-hosted, on-premises or in your private cloud. We can build and operate a managed deployment for you on request.