Safeguarded AI (ARIA / Zeroth / Heron)

Safety claims are conditional on the declared system boundary; misspecified controllers, emergent coalitions, or agents outside the certified cut can void otherwise correct proofs.

Introduction

The Safeguarded AI cluster—ARIA, Zeroth Research, Heron AI Security, and containment-verification allies—targets an end-to-end stack pairing world models with formal specifications, machine-checkable proof certificates, and runtime action control. Safety claims are conditional on the declared system boundary; misspecified controllers or agents outside the certified cut can void proofs. The programme engages Embedded Agency, Grounding Drift, and Deployment Safety in whole-system assurance form.

Who carries it: ARIA Safeguarded AI; Zeroth Research (Luca Arnaboldi, Pascal Berrang); Heron AI Security; Moon & Varshney (Containment Verification)

What they aim to do. Assure embedded model–tool–environment systems using formal methods, runtime controls, and privacy-preserving verification so that safety claims can be checked against an explicit system boundary.

The hard question. Safety claims are conditional on the declared system boundary; misspecified controllers, emergent coalitions, or agents outside the certified cut can void otherwise correct proofs.

What they produce. The cluster targets an end-to-end safeguarded-AI stack: world models paired with formal specifications and machine-checkable proof certificates, runtime action control via the AARM specification, and zero-knowledge attestation for privacy-preserving verification of deployed behavior.

Key terms. Core terms include safeguarded AI, proof certificates, containment verification, AARM runtime action control, zero-knowledge attestation, and architecture-level multi-agent security.

Related field cruxes. Embedded Agency; Audit Independence; Inner Alignment; Grounding Drift; Deployment Safety

What they contribute. Whole-system formal assurance for agentic deployments; cryptographic attestations of runtime behavior; explicit agentic action boundaries; and architecture-level multi-agent security analysis—directly engaging Embedded Agency, Audit Independence, Inner Alignment, Grounding Drift, and Deployment Safety.

How this project treats it. A green proof certificate does not imply that corrections are actually taken up in the live loop; attestation of declared behavior does not imply that the discovered effective agent matches the certified unit—a cousin tension to Guaranteed-Safe AI’s completeness wall versus CIRIS-style operational stacks.

Map clustering

AISafety.com map listings that roll up to this agenda:

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.