Research and evidence
The evidence layer for AI-agent security.
CapitalGuard publishes the controls, schemas, test vectors, privacy boundary, and verification protocol. Official score calibration, threat intelligence, benchmark eligibility data, and signing keys remain protected.
Published by CapitalGuard Security Research · Updated July 28, 2026
Citable primary sources
Start with the source, not a marketing summary.
AI Agent Intent Conformance Benchmark
Fifteen deterministic signed-plan, target-binding, sequence, replay, and uncertain-outcome cases with fixed fixtures and a public verifier.
Open sourceAI Agent Delegation Boundary Benchmark
Twelve deterministic signed-chain, attenuation, actor, policy, parent, expiry, depth, and single-use cases with fixed fixtures and a public verifier.
Open sourceMCP Tool Integrity Gate Benchmark
Eight synthetic identity, tool-definition, approval, replay, quarantine, and response-boundary cases publisher-run through the tested CapitalGuard gate implementation.
Open sourceAI Agent Permission Drift Benchmark
One signed synthetic baseline comparison, 36 observed changes, eleven verification cases, privacy-reduced evidence, and an independent verifier.
Open sourceAgent Security License Verification Benchmark
One valid signed chain, eight rejection mutations, a standalone verifier, deterministic results, and release-file digests.
Open sourceGitHub Actions AI-Agent Security Gate
Eight cited controls, a SHA-pinned workflow, exact dependency revisions, explicit limitations, and a versioned machine-readable profile.
Open sourceGitHub Actions Gate Benchmark
Ten inert workflow fixtures, eight static controls, exact failed-control lists, deterministic results, and release-file digests.
Open sourceAgent Runtime Gate Benchmark
Twenty-three synthetic cases across workspace integrity, filtered projection, OCI containment, broker integrity, gateway enforcement, and evidence privacy.
Open sourceOWASP Agentic Top 10 Evidence Map
All ten risk areas mapped to hash-bound CapitalGuard control releases, runtime cases, source bindings, and explicit claim limits.
Open sourceAI Agent Security Evidence Index
A fixed signed challenge with visible publisher, vendor, and independent-lab provenance and no score for untested products.
Open sourceRepository Exposure Benchmark
A six-case synthetic repository, real scanner run, redaction check, reconstructable fixture pack, and digest-bound results.
Open sourceClient Code AI Access Register
A cited pre-flight workflow and blank JSON/CSV register for authorized AI access to client repositories.
Open sourceAI Security Evidence Catalog
The complete 550-record evidence index as human-readable pages, JSON, and CSV with canonical URLs and source citations.
Open sourceCapitalGuard Standard 1.0.0
The normative AI-agent repository security baseline and conformance model.
Open sourceControl catalog
Machine-readable definitions for authorization, inventory, secrets, data, injection, privilege, policy, drift, reporting, traceability, and attestation.
Open sourceTest vectors
Canonical examples for deterministic Guardprint material, output, and verification behavior.
Open sourceAgent security evidence contracts
Versioned A-BOM and tamper-evident event schemas, canonical vectors, and release digests for interoperable evidence.
Open sourcePolicy Language 0.1
A strict default-deny policy schema, privacy-safe requests, stable control IDs, and reproducible decision vectors.
Open sourceEnforced Mode 0.7
Full-hash workspace verification, policy-filtered projection, hardened OCI isolation, signed data boundaries, quarantined agent writes, sequence enforcement, and authenticated broker evidence.
Open sourceFree GitHub Action
A privacy-safe checker that runs on the repository owner's runner and links detected exposure to the full review path.
Open sourceAttack Lab methodology
A non-destructive, architecture-aware challenge model that retains exposed control paths as regression tests.
Open sourceEvidence boundary
Public enough to verify. Private enough to protect.
Public evidence contains aggregate categories, control status, scan limits, timestamps, cryptographic digests, and signatures. It excludes repository names, customer identity, source contents, and secret values.
Benchmark publication rule
CapitalGuard will not publish an exposure rate until at least 30 manually eligible assessments exist, and no reported segment will contain fewer than 10 records.
Until that threshold is met, the methodology is public and the benchmark remains unpublished. This prevents tiny samples, customer inference, and invented statistics.
Preferred citation
CapitalGuard. CapitalGuard Standard 1.0.0: AI-Agent Repository Security Baseline and Guardprint Protocol. 2026.
