CarbonylR&D

Platform · Agent infrastructure

Agent runtime & orchestration

The infrastructure for deploying, controlling and auditing systems of many agents. Every agent is treated as an untrusted process: what it can see, touch and run is scoped and enforced, every action is observable as it happens, and every finding can be traced to the agent that produced it and the evidence it cited.

DisciplineAgent infrastructure
IssuePublic extract
AccessBy invitation

The premise

Every agent is an untrusted process.

An agent driven by a language model is capable, fast and unpredictable. Giving it tools is giving it authority. The runtime starts from that fact: it owns orchestration, security and the lifecycle of every agent, and grants each one exactly the authority its task needs, and no more.

Untrusted by default

No agent inherits permission. Every tool, every file and every outbound request is granted explicitly, checked on every use, and refused loudly when it is not.

Boundaries, not imports

Platforms plug into the runtime; they never embed it. The boundary is a network interface and a standard tool protocol, never shared memory and never imported code.

Observable now

Everything an agent does is visible while it happens, not reconstructed afterwards from logs, and recorded where it cannot be quietly altered.

Architecture

One compiled runtime, with authority granted in tiers.

platformdomain tools platformdomain tools the boundary orchestrationpermissions & scopetamper-evident auditlive telemetry the runtime self-hostedmodels externalmodels, per agent
FIG. 1Platforms expose what they know as tools across a network boundary. The runtime owns orchestration, permissions, audit and telemetry, and routes each agent's inference to self-hosted or external models according to that agent's configuration.
A-1

A single compiled binary

The whole runtime is one binary in a memory-safe systems language. Nothing interpreted sits in the orchestration path, and nothing pauses agent execution to collect garbage.

A-2

Read-only

Inspection, analysis and search. An agent at this tier can learn anything it is allowed to see and change nothing.

A-3

Workspace write

Creating and modifying files inside a sandboxed working directory, with no ability to run processes.

A-4

Full access, never inherited

Running commands, spawning sub-agents and changing the system: granted explicitly to a named agent, and never passed on to the agents it creates.

A-5

Confinement before any I/O

Every path is resolved to its canonical form before anything is read or written, so links and relative tricks that would escape the sandbox are refused before they touch the disk.

A-6

Platforms own the domain

The runtime knows how to orchestrate, enforce and record. It does not know what a trade, a tax form or a satellite is. The platform does, and exposes it as tools.

Swarms

Agents that debate, converge and report, under a contract.

FIG. 2A hierarchical swarm inside the runtime: three groups of agents in different roles, debating in pairs, converging on an arbiter and passing their artifacts to the next group. The arc on each agent marks its authority. Every event is appended to the linked audit chain below.
W-1

Scoped tasks

Short-lived, bounded analysis with a single objective and a fixed number of rounds, ending in a report.

W-2

Continuous service

Long-running swarms that keep watch, investigate and maintain awareness of a system over time.

W-3

Hierarchical

Groups that run in phases, research, then review, then implementation, routing artifacts from one group to the next, with architect and critic pairs debating until they converge.

W-4

Full authority

Broad autonomy for the tasks that need it, still bounded by each agent's allow and deny lists and its file scope.

W-5

Agents as specifications

Each agent is defined by its domain of expertise, its style of argument, its role as worker, critic, arbiter or coordinator, its appetite for risk, its priority and its exact permissions.

W-6

The manifest is the contract

A declarative manifest fixes the roster, the debate, the thresholds for convergence, the iteration and timeout policies and what may be shared between groups. The runtime enforces it as written.

Zero trust

Authenticated, scoped and recorded, without exception.

Nothing fails silently. An agent that reaches for a tool outside its list, or a file outside its scope, receives an explicit denial, and the attempt is recorded alongside everything else it did.

Z-1

Signed requests

Every interface and every message between agents is authenticated with a keyed signature.

Z-2

Allow and deny lists

Each agent has its own list of permitted and forbidden tools. Where they disagree, the denial wins, always.

Z-3

File scope

Readable paths, writable paths and forbidden patterns are set per agent. An agent with no scope can read nothing at all.

Z-4

Tamper-evident ledger

Every action is written to an audit ledger that is linked cryptographically, so any later alteration can be detected and is.

Z-5

Session isolation

An agent writes findings only to its own session and reads only from sessions explicitly shared with it.

Z-6

Outbound requests contained

Requests leaving the runtime cannot reach private networks, loopback or link-local addresses, or the metadata services of cloud machines.

Telemetry

Every session, streamed as it happens.

FIG. 3A live session: one lane per agent with its turns, tool calls (ticks), findings (dots) and refused actions (crosses), debate phases shaded behind, consensus checks across every lane, and cumulative usage below. Drawn as an illustration.
T-1

Structured events, live

Turns, tool calls, findings, security events, debate phases, consensus checks, injected prompts, refused file access and usage, streamed to any subscriber the moment they occur.

T-2

Complete metadata

Each event carries its timestamps and duration, usage, the arguments and results of tool calls, the files touched and the permission decision that allowed or refused them.

T-3

Steerable while running

Instructions can be injected into a running session, and the effect is visible in the same stream that shows everything else.

T-4

Findings with their evidence

Every finding is traceable to the agent that produced it, the tools it called and the evidence it cited.

Code intelligence

Deterministic analysis that agents reason from.

No model inference sits in the analysis path. The results are exact and reproducible, and they are the ground truth that agents reason about, not the other way around.

K-1

Ground truth per file

Exact signatures, complexity, decorators, routes, imports and calls, read from the syntax tree, with a content hash that changes only when the file does.

K-2

Cycles in the dependency graph

Circular dependencies found across the whole import graph of a project.

K-3

What matters most

A centrality ranking over the project graph that identifies the parts of a codebase everything else depends on.

K-4

Project accounting

Files, lines and languages across the project, with an incremental hash of the whole, across several languages.

Inference

Sovereign by default, explicit when not.

N-1

Self-hosted models

Inference routed through a gateway to models served on our own graphics hardware, with no mandatory dependency on any external provider.

N-2

External, when chosen

When an external model is used, the choice is explicit, made per agent, and recorded with everything else.

N-3

Mixed within a swarm

One swarm can keep its security-critical agents on self-hosted models and give general-purpose agents an external one.

N-4

Configuration, not a global switch

Which model an agent uses is part of that agent's specification, and so part of the contract the manifest enforces.

Command center

Design, launch and watch an orchestration.

O-1

Visual architect

A canvas for composing groups of agents, their debate topology and the routing of artifacts between them.

O-2

Session manager

Every active swarm, with its live event stream.

O-3

Codebase explorer

Navigation through the code the agents are analyzing, built on the same syntax trees they use.

O-4

Security ledger

The audit trail, with verification that it has not been altered.

O-5

Platform registry

The platforms and tools available to agents, registered and managed in one place.

O-6

Limits

Resource governance and rate limits for every session.

Integration

Domain knowledge through tools, orchestration through the interface.

A platform connects in three steps: it exposes its capabilities as tools through a standard protocol, writes a manifest for the swarm it needs, and launches it over the network, subscribing to its event stream and retrieving its findings. Neither side imports the other.

I-1

A swarm that audits a trading system

One agent reads backtest results and one reads live fills. A third computes the divergences, two debate whether each one is a genuine defect or expected variance, and a sixth writes a structured report, with every call scoped, every file access recorded and every finding tied to its evidence.

I-2

A swarm that keeps watch

A continuous swarm monitors the health of infrastructure, correlates anomalies across services and escalates what it finds, unable by construction to read a credential or run a destructive command, with every action auditable.

Connections

The same ideas, elsewhere in our work.

X-1

Trading & research platform

Measured parity between backtest and live is exactly the kind of question a scoped swarm is built to audit: two sources of evidence, a computed divergence, and a debate over which differences are defects.

X-2

Accounting & tax system

A tamper-evident ledger of agent actions and a double-entry ledger of money rest on the same principle: a record that is only ever appended to, and that shows it if anyone tries otherwise.

X-3

Geospatial intelligence platform

Every agent action carries its source and its time, as every position on the globe does. Nothing is trusted that cannot say where it came from.

X-4

Security assurance platform

The runtime is held to the same standard it enforces: assessed the way an adversary would assess it, with every finding backed by evidence.

Engage

By invitation.

Every engagement begins with a conversation and is scoped before work begins.

Access

Teams that need many agents working under audit, by invitation.

Subject: Access

Custom builds

Orchestrations, manifests and tool integrations for your own platform.

Subject: Build

Partnerships

Platforms and compute partners.

Subject: Partnership

Investment

Briefings and a demonstration, under confidentiality.

Subject: Funding
contact@carbonyl.org