Safeguarded AI (ARIA / Zeroth / Heron)
Safety claims are conditional on the declared system boundary; misspecified controllers, emergent coalitions, or agents outside the certified cut can void otherwise correct proofs.
Introduction
The Safeguarded AI cluster—ARIA, Zeroth Research, Heron AI Security, and containment-verification allies—targets an end-to-end stack pairing world models with formal specifications, machine-checkable proof certificates, and runtime action control. Safety claims are conditional on the declared system boundary; misspecified controllers or agents outside the certified cut can void proofs. The programme engages Embedded Agency, Grounding Drift, and Deployment Safety in whole-system assurance form.
Who carries it: ARIA Safeguarded AI; Zeroth Research (Luca Arnaboldi, Pascal Berrang); Heron AI Security; Moon & Varshney (Containment Verification)
What they aim to do. Assure embedded model–tool–environment systems using formal methods, runtime controls, and privacy-preserving verification so that safety claims can be checked against an explicit system boundary.
The hard question. Safety claims are conditional on the declared system boundary; misspecified controllers, emergent coalitions, or agents outside the certified cut can void otherwise correct proofs.
What they produce. The cluster targets an end-to-end safeguarded-AI stack: world models paired with formal specifications and machine-checkable proof certificates, runtime action control via the AARM specification, and zero-knowledge attestation for privacy-preserving verification of deployed behavior.
Key terms. Core terms include safeguarded AI, proof certificates, containment verification, AARM runtime action control, zero-knowledge attestation, and architecture-level multi-agent security.
Related field cruxes. Embedded Agency; Audit Independence; Inner Alignment; Grounding Drift; Deployment Safety
What they contribute. Whole-system formal assurance for agentic deployments; cryptographic attestations of runtime behavior; explicit agentic action boundaries; and architecture-level multi-agent security analysis—directly engaging Embedded Agency, Audit Independence, Inner Alignment, Grounding Drift, and Deployment Safety.
How this project treats it. A green proof certificate does not imply that corrections are actually taken up in the live loop; attestation of declared behavior does not imply that the discovered effective agent matches the certified unit—a cousin tension to Guaranteed-Safe AI’s completeness wall versus CIRIS-style operational stacks.
Links
- ARIA Safeguarded AI
- Zeroth Research
- Containment Verification
- Heron AI Security
- AARM specification
- ZK attestation (Berrang)
Map clustering
AISafety.com map listings that roll up to this agenda:
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.