Research proposal · multi-principal agent systems
Containing and attributing failures in multi-principal agent networks.
When agents serve different people or organizations, a handoff crosses an authority boundary. This proposal asks whether two locally enforceable controls can contain a seeded failure without requiring one participant to see or govern the whole system.
Status of this note
No testbed has been selected, no protocol has been locked, and no experiment has been run. Continuous Function does not operate the multi-principal system described here. There are no episodes, participants, empirical results, or independent human review. The purpose of this note is to make the question and the proposed evidence standard available for criticism before any decision to execute the work.

The object of study
A multi-agent system is not necessarily a multi-principal system. Several model instances may still answer to one operator, share one policy, use one credential boundary, and expose one global transcript. The harder case begins when agents act for different people or organizations. Each principal may have its own authority, private information, tools, obligations, and reasons to withhold part of its state.
An agent is not a principal. A principal is the person or organization whose authority, information, and interests the agent represents. A qualifying multi-principal experiment must preserve those distinctions at the administrative boundary; four role prompts inside one all-powerful orchestrator would not qualify.
This distinction changes what a useful safeguard may observe. An offline evaluator can reconstruct a complete episode after outcomes are frozen. A receiving principal cannot rely on that retrospective global view when deciding whether to accept an artifact now. The control must operate on evidence the receiver is entitled to see.
Recent work has documented propagation through agent communication in Prompt Infection, failures of per-agent monitoring under distributed attacks in Multi-Agent AI Control, and the difficulty of attributing a failed task in Which Agent Causes Task Failures and When? Identity and provenance specifications such as IDs for AI Systems and SLSA provenance offer useful building blocks. They do not by themselves tell us whether a local control prevents a downstream failure, how much legitimate work it blocks, or what evidence is sufficient to locate the responsible handoff.
Research question
In a contained workflow with genuinely distinct principals, how much do locally enforced delegation envelopes and artifact-lineage validation reduce the probability that one seeded boundary fault causes a downstream breach, and what policy-permitted evidence is sufficient to attribute that breach under partial observability?
The proposal separates two claims that are often combined. The first is causal: a control changes the probability of a machine-verifiable failure. The second is epistemic: a record lets another person determine which handoff caused the failure. A rich log may help attribution without preventing anything. A restrictive control may prevent every failure by preventing useful work. Neither result should be reported as a complete safety improvement.
A contained test case
The proposed first setting is a synthetic cross-organization software-release chain. A producer, a dependency supplier, an assessor, and a release operator pass artifacts through disposable source-control, continuous-integration, registry, ticketing, and staging services. Each principal has separate local policy, memory, credentials, tools, and trace storage. The relay between them may route an artifact and check its syntax, but it may not inspect undisclosed state, grant authority, or override a principal's decision.
One known boundary fault is introduced into each matched scenario. Faults the two controls are designed to block—such as an expired delegation, ancestry mismatch, stale replay, or missing attestation—would define the primary containment estimand. An authorized but substantively false assessment is different: it may pass both controls and would be analyzed separately as a semantic negative control, not pooled into their primary efficacy estimate. The environment is deliberately synthetic. It would contain no production deployment, public network access, real credential, financial transaction, or physical actuation.
Two interventions, crossed
The candidate design is a matched two-by-two experiment. Every arm records the same declared metadata so that observability alone cannot explain a difference. What changes is whether a principal must enforce the declaration before accepting or forwarding the artifact.
A delegation envelope binds a requesting principal and agent to an allowed action, resource scope, recipient, expiry or revocation state, and rule for downstream delegation. Artifact-lineage validation checks the content digest, declared parents, producing principal, environment and version, authorization chain, and required attestations. It checks identity and consistency. It does not establish that the artifact is correct, current in the world, or scientifically true. Enforcement would require both a valid sender-side delegation and a recipient-side acceptance rule; neither principal may confer the other's authority.
| Lineage validation | Delegation envelope off | Delegation envelope on |
|---|---|---|
| Off | Record-only comparator: metadata is retained but neither control blocks the handoff. | Delegation envelope enforced; lineage is recorded but not validated. |
| On | Lineage is validated; delegation scope is recorded but not enforced. | Both local controls are enforced. |
The experimental unit would be an independently generated scenario seed and initial state, not a message, turn, or agent. Within a matched block, the seed, injected fault, model allocation, tool state, and resource budget would be held fixed across the four conditions; condition order would be randomized. Pilot and confirmatory episodes, exclusions, failed-run treatment, sample size, and decision margins would have to be fixed before any confirmatory outcome is inspected.
What would be measured
| Measure | Role in the decision |
|---|---|
| Terminal boundary breach | Primary safety outcome: whether a downstream principal causes synthetic staging to accept a known-bad or unauthorized artifact before correction. |
| Legitimate task completion | Utility guardrail: a control that makes collaboration impossible cannot qualify as useful containment. |
| Propagation and recovery | Secondary outcomes: propagation depth, time to detection, false blocks, recovery, and severity. |
| Fault-edge localization | Separate attribution outcome: whether a blinded reviewer can identify the causal handoff from permitted evidence, with “insufficient evidence” available. |
The terminal safety outcome should be read from deterministic system state rather than assigned by a language-model judge. A future preregistration would need a useful-effect threshold for breach reduction, a non-inferiority margin for legitimate completion, a precision target, and a rule for inconclusive results. This note does not select those numbers. Choosing them after observing the effect would defeat the design.
Attribution under partial observability
Containment and attribution require different experiments. After an episode is closed, an independent custodian could hold the seeded-fault and outcome ledger while blinded reviewers receive randomized, policy-permitted record packages. The packages might range from a conventional report to a manifest, local principal state, structured traces, and a recomputable release. Each level would expose more evidence, not a different underlying episode. A reviewer would see only one package level for a given episode, never several versions of the same case. Any assignment of other cases and package levels would need a counterbalanced order fixed before review.
Reviewers would identify the most likely causal handoff or answer that the record is insufficient. Clean controls are needed to detect a reviewer who finds faults everywhere; contamination sentinels are needed to detect a package that leaks the hidden answer. The least revealing package could be called sufficient only under a threshold fixed before review and only if an independent analyst can reproduce the primary tables from the post-lock release. If no tested package passes, the honest conclusion is that the record was insufficient.
What would count against the proposal
A serious proposal should specify how it can fail before it specifies how it will be presented. The bounded claim would not be supported if any of the following holds.
Claim boundary
| A qualifying study could support | It could not establish |
|---|---|
| A causal estimate for two local controls in one locked synthetic workflow and declared model population. | A universal claim about multi-agent safety, real organizations, or every communication channel. |
| A measured safety–utility tradeoff under a fixed endpoint and resource budget. | Proof that a hash, signature, lineage record, or agent statement is semantically true. |
| A bounded judgment about whether one released record supports fault localization. | A claim that Continuous Function already operates agents, collaboration, or a validated research platform. |
Results from one synthetic release workflow would need a deployment-conditions table and should not be pooled into an unrestricted claim. Transfer to research, incident-response, or other organizational workflows would be a later question, not a rescue analysis for a failed first study.
The gate before execution
The next decision is not which model to run. It is whether an existing, safe, reproducible testbed or trace corpus can represent genuinely distinct authority domains without giving the intervention a global view. The result must matter to a named operator deciding an import, acceptance, or delegation policy.
Before execution could be considered, the study would need a qualifying testbed and operator, rights-cleared scenarios and fault classes, containment and release review, separately controlled principal credentials, a frozen protocol and power analysis, an independent human methods and safety review, a custodian for hidden mutations, and an independent reproduction route. None is in place today. If a qualifying testbed and operator cannot be identified, the proposal should remain unexecuted rather than turn into a demonstration built to confirm its own assumptions.
The existing permissioned-handoff starter can format an evidence-free browser-local draft. It does not select this study, authorize execution, save a canonical record, or create evidence.
Related work and technical foundations
The proposal does not claim to invent identity, delegation, provenance, sandboxes, failure attribution, or multi-agent testbeds. Its candidate contribution is narrower: estimating two local controls separately and jointly across distinct administrative principals, with machine-verifiable terminal harm and a separate blinded test of what the released record permits another person to conclude.