Research and evidence

The evidence layer for AI-agent security.

CapitalGuard publishes the controls, schemas, test vectors, privacy boundary, and verification protocol. Official score calibration, threat intelligence, benchmark eligibility data, and signing keys remain protected.

Published by CapitalGuard Security Research · Updated July 28, 2026

Citable primary sources

Start with the source, not a marketing summary.

AI Agent Intent Conformance Benchmark

Fifteen deterministic signed-plan, target-binding, sequence, replay, and uncertain-outcome cases with fixed fixtures and a public verifier.

Open source

AI Agent Delegation Boundary Benchmark

Twelve deterministic signed-chain, attenuation, actor, policy, parent, expiry, depth, and single-use cases with fixed fixtures and a public verifier.

Open source

MCP Tool Integrity Gate Benchmark

Eight synthetic identity, tool-definition, approval, replay, quarantine, and response-boundary cases publisher-run through the tested CapitalGuard gate implementation.

Open source

AI Agent Permission Drift Benchmark

One signed synthetic baseline comparison, 36 observed changes, eleven verification cases, privacy-reduced evidence, and an independent verifier.

Open source

Agent Security License Verification Benchmark

One valid signed chain, eight rejection mutations, a standalone verifier, deterministic results, and release-file digests.

Open source

GitHub Actions AI-Agent Security Gate

Eight cited controls, a SHA-pinned workflow, exact dependency revisions, explicit limitations, and a versioned machine-readable profile.

Open source

GitHub Actions Gate Benchmark

Ten inert workflow fixtures, eight static controls, exact failed-control lists, deterministic results, and release-file digests.

Open source

Agent Runtime Gate Benchmark

Twenty-three synthetic cases across workspace integrity, filtered projection, OCI containment, broker integrity, gateway enforcement, and evidence privacy.

Open source

OWASP Agentic Top 10 Evidence Map

All ten risk areas mapped to hash-bound CapitalGuard control releases, runtime cases, source bindings, and explicit claim limits.

Open source

AI Agent Security Evidence Index

A fixed signed challenge with visible publisher, vendor, and independent-lab provenance and no score for untested products.

Open source

Repository Exposure Benchmark

A six-case synthetic repository, real scanner run, redaction check, reconstructable fixture pack, and digest-bound results.

Open source

Client Code AI Access Register

A cited pre-flight workflow and blank JSON/CSV register for authorized AI access to client repositories.

Open source

AI Security Evidence Catalog

The complete 550-record evidence index as human-readable pages, JSON, and CSV with canonical URLs and source citations.

Open source

CapitalGuard Standard 1.0.0

The normative AI-agent repository security baseline and conformance model.

Open source

Control catalog

Machine-readable definitions for authorization, inventory, secrets, data, injection, privilege, policy, drift, reporting, traceability, and attestation.

Open source

Test vectors

Canonical examples for deterministic Guardprint material, output, and verification behavior.

Open source

Agent security evidence contracts

Versioned A-BOM and tamper-evident event schemas, canonical vectors, and release digests for interoperable evidence.

Open source

Policy Language 0.1

A strict default-deny policy schema, privacy-safe requests, stable control IDs, and reproducible decision vectors.

Open source

Enforced Mode 0.7

Full-hash workspace verification, policy-filtered projection, hardened OCI isolation, signed data boundaries, quarantined agent writes, sequence enforcement, and authenticated broker evidence.

Open source

Free GitHub Action

A privacy-safe checker that runs on the repository owner's runner and links detected exposure to the full review path.

Open source

Attack Lab methodology

A non-destructive, architecture-aware challenge model that retains exposed control paths as regression tests.

Open source

Evidence boundary

Public enough to verify. Private enough to protect.

Public evidence contains aggregate categories, control status, scan limits, timestamps, cryptographic digests, and signatures. It excludes repository names, customer identity, source contents, and secret values.

Benchmark publication rule

CapitalGuard will not publish an exposure rate until at least 30 manually eligible assessments exist, and no reported segment will contain fewer than 10 records.

Until that threshold is met, the methodology is public and the benchmark remains unpublished. This prevents tiny samples, customer inference, and invented statistics.

Preferred citation

CapitalGuard. CapitalGuard Standard 1.0.0: AI-Agent Repository Security Baseline and Guardprint Protocol. 2026.

Verify Evidence