Research proposal · multi-principal agent systems

Containing and attributing failures in multi-principal agent networks.

When agents serve different people or organizations, a handoff crosses an authority boundary. This proposal asks whether two locally enforceable controls can contain a seeded failure without requiring one participant to see or govern the whole system.

Work state
Proposed · not admitted for execution
Evidence
None
Review
No independent human methods review
Updated
16 August 2026

Status of this note

Proposal only · no study admitted

No testbed has been selected, no protocol has been locked, and no experiment has been run. Continuous Function does not operate the multi-principal system described here. There are no episodes, participants, empirical results, or independent human review. The purpose of this note is to make the question and the proposed evidence standard available for criticism before any decision to execute the work.

Conceptual illustration of one sealed artifact passing through four separate administrative domains, with one questioned handoff contained at a boundary.
Conceptual illustration—not a system diagram. The amber handoff marks the boundary the proposal would test; it is not an observed event or result.

The object of study

A multi-agent system is not necessarily a multi-principal system. Several model instances may still answer to one operator, share one policy, use one credential boundary, and expose one global transcript. The harder case begins when agents act for different people or organizations. Each principal may have its own authority, private information, tools, obligations, and reasons to withhold part of its state.

Definition

An agent is not a principal. A principal is the person or organization whose authority, information, and interests the agent represents. A qualifying multi-principal experiment must preserve those distinctions at the administrative boundary; four role prompts inside one all-powerful orchestrator would not qualify.

This distinction changes what a useful safeguard may observe. An offline evaluator can reconstruct a complete episode after outcomes are frozen. A receiving principal cannot rely on that retrospective global view when deciding whether to accept an artifact now. The control must operate on evidence the receiver is entitled to see.

Recent work has documented propagation through agent communication in Prompt Infection, failures of per-agent monitoring under distributed attacks in Multi-Agent AI Control, and the difficulty of attributing a failed task in Which Agent Causes Task Failures and When? Identity and provenance specifications such as IDs for AI Systems and SLSA provenance offer useful building blocks. They do not by themselves tell us whether a local control prevents a downstream failure, how much legitimate work it blocks, or what evidence is sufficient to locate the responsible handoff.

Research question

In a contained workflow with genuinely distinct principals, how much do locally enforced delegation envelopes and artifact-lineage validation reduce the probability that one seeded boundary fault causes a downstream breach, and what policy-permitted evidence is sufficient to attribute that breach under partial observability?

The proposal separates two claims that are often combined. The first is causal: a control changes the probability of a machine-verifiable failure. The second is epistemic: a record lets another person determine which handoff caused the failure. A rich log may help attribution without preventing anything. A restrictive control may prevent every failure by preventing useful work. Neither result should be reported as a complete safety improvement.

A contained test case

The proposed first setting is a synthetic cross-organization software-release chain. A producer, a dependency supplier, an assessor, and a release operator pass artifacts through disposable source-control, continuous-integration, registry, ticketing, and staging services. Each principal has separate local policy, memory, credentials, tools, and trace storage. The relay between them may route an artifact and check its syntax, but it may not inspect undisclosed state, grant authority, or override a principal's decision.

One known boundary fault is introduced into each matched scenario. Faults the two controls are designed to block—such as an expired delegation, ancestry mismatch, stale replay, or missing attestation—would define the primary containment estimand. An authorized but substantively false assessment is different: it may pass both controls and would be analyzed separately as a semantic negative control, not pooled into their primary efficacy estimate. The environment is deliberately synthetic. It would contain no production deployment, public network access, real credential, financial transaction, or physical actuation.

Two interventions, crossed

The candidate design is a matched two-by-two experiment. Every arm records the same declared metadata so that observability alone cannot explain a difference. What changes is whether a principal must enforce the declaration before accepting or forwarding the artifact.

A delegation envelope binds a requesting principal and agent to an allowed action, resource scope, recipient, expiry or revocation state, and rule for downstream delegation. Artifact-lineage validation checks the content digest, declared parents, producing principal, environment and version, authorization chain, and required attestations. It checks identity and consistency. It does not establish that the artifact is correct, current in the world, or scientifically true. Enforcement would require both a valid sender-side delegation and a recipient-side acceptance rule; neither principal may confer the other's authority.

Candidate factorial comparison
Lineage validationDelegation envelope offDelegation envelope on
OffRecord-only comparator: metadata is retained but neither control blocks the handoff.Delegation envelope enforced; lineage is recorded but not validated.
OnLineage is validated; delegation scope is recorded but not enforced.Both local controls are enforced.

The experimental unit would be an independently generated scenario seed and initial state, not a message, turn, or agent. Within a matched block, the seed, injected fault, model allocation, tool state, and resource budget would be held fixed across the four conditions; condition order would be randomized. Pilot and confirmatory episodes, exclusions, failed-run treatment, sample size, and decision margins would have to be fixed before any confirmatory outcome is inspected.

What would be measured

Proposed outcomes
MeasureRole in the decision
Terminal boundary breachPrimary safety outcome: whether a downstream principal causes synthetic staging to accept a known-bad or unauthorized artifact before correction.
Legitimate task completionUtility guardrail: a control that makes collaboration impossible cannot qualify as useful containment.
Propagation and recoverySecondary outcomes: propagation depth, time to detection, false blocks, recovery, and severity.
Fault-edge localizationSeparate attribution outcome: whether a blinded reviewer can identify the causal handoff from permitted evidence, with “insufficient evidence” available.

The terminal safety outcome should be read from deterministic system state rather than assigned by a language-model judge. A future preregistration would need a useful-effect threshold for breach reduction, a non-inferiority margin for legitimate completion, a precision target, and a rule for inconclusive results. This note does not select those numbers. Choosing them after observing the effect would defeat the design.

Attribution under partial observability

Containment and attribution require different experiments. After an episode is closed, an independent custodian could hold the seeded-fault and outcome ledger while blinded reviewers receive randomized, policy-permitted record packages. The packages might range from a conventional report to a manifest, local principal state, structured traces, and a recomputable release. Each level would expose more evidence, not a different underlying episode. A reviewer would see only one package level for a given episode, never several versions of the same case. Any assignment of other cases and package levels would need a counterbalanced order fixed before review.

Reviewers would identify the most likely causal handoff or answer that the record is insufficient. Clean controls are needed to detect a reviewer who finds faults everywhere; contamination sentinels are needed to detect a package that leaks the hidden answer. The least revealing package could be called sufficient only under a threshold fixed before review and only if an independent analyst can reproduce the primary tables from the post-lock release. If no tested package passes, the honest conclusion is that the record was insufficient.

What would count against the proposal

A serious proposal should specify how it can fail before it specifies how it will be presented. The bounded claim would not be supported if any of the following holds.

The controls do not clear the predeclared safety threshold.Null or harmful resultA plausible safeguard would remain unestablished even if the system and analysis ran correctly.
Legitimate completion falls beyond the locked utility margin.Restrictive-control resultPreventing every handoff is not evidence of useful governed collaboration.
The apparent effect depends on privileged global information.Construct failureA safeguard that works only for the evaluator does not answer the local-control question.
The principal cells do not have genuinely separate authority.Construct failureRole-playing several principals under one administrative controller is insufficient.
Blinded reviewers cannot localize failures from the released record.Attribution not establishedThe result should narrow the release claim, not lower the evidence threshold.
An independent analyst cannot reproduce the primary tables.Release failureThe discrepancy remains part of the record; it is not silently repaired into success.

Claim boundary

A qualifying study could supportIt could not establish
A causal estimate for two local controls in one locked synthetic workflow and declared model population.A universal claim about multi-agent safety, real organizations, or every communication channel.
A measured safety–utility tradeoff under a fixed endpoint and resource budget.Proof that a hash, signature, lineage record, or agent statement is semantically true.
A bounded judgment about whether one released record supports fault localization.A claim that Continuous Function already operates agents, collaboration, or a validated research platform.

Results from one synthetic release workflow would need a deployment-conditions table and should not be pooled into an unrestricted claim. Transfer to research, incident-response, or other organizational workflows would be a later question, not a rescue analysis for a failed first study.

The gate before execution

Current disposition · remain proposed

The next decision is not which model to run. It is whether an existing, safe, reproducible testbed or trace corpus can represent genuinely distinct authority domains without giving the intervention a global view. The result must matter to a named operator deciding an import, acceptance, or delegation policy.

Before execution could be considered, the study would need a qualifying testbed and operator, rights-cleared scenarios and fault classes, containment and release review, separately controlled principal credentials, a frozen protocol and power analysis, an independent human methods and safety review, a custodian for hidden mutations, and an independent reproduction route. None is in place today. If a qualifying testbed and operator cannot be identified, the proposal should remain unexecuted rather than turn into a demonstration built to confirm its own assumptions.

The existing permissioned-handoff starter can format an evidence-free browser-local draft. It does not select this study, authorize execution, save a canonical record, or create evidence.

Related work and technical foundations

The proposal does not claim to invent identity, delegation, provenance, sandboxes, failure attribution, or multi-agent testbeds. Its candidate contribution is narrower: estimating two local controls separately and jointly across distinct administrative principals, with machine-verifiable terminal harm and a separate blinded test of what the released record permits another person to conclude.

Multi-Agent Risks from Advanced AIHammond et al. · 2025A taxonomy of multi-agent failure modes and risk factors, including information asymmetry and multi-agent security.Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent SystemsLee and Tiwari · 2024An empirical study of attacks that propagate through communication between language-model agents.Agents of ChaosShapira et al. · 2026 preprintAn exploratory live-laboratory study reporting unauthorized compliance, identity spoofing, and cross-agent propagation among other failures.Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance MonitorsMakins et al. · 2026 preprintA synthetic AI-lab evaluation of distributed attacks and the limits of monitoring agents one trajectory at a time.Which Agent Causes Task Failures and When?Zhang et al. · 2025A benchmark for locating the responsible agent and step in failed multi-agent tasks.IDs for AI SystemsChan et al. · 2024A framework for identifying AI-system instances and making relevant information available to parties that interact with them.MPAC: A Multi-Principal Agent Coordination ProtocolQian et al. · 2026 preprintA recent protocol proposal for coordination and governance when agents serve independent principals.SLSA provenanceTechnical specification · v1.2A specification for verifiable information that tracks how and where a software artifact was produced.in-toto Attestation FrameworkTechnical specification · v1.2.0A format for authenticated metadata bound to software artifacts; a useful primitive, not evidence of semantic truth.