Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
-
Updated
Aug 31, 2026 - Python
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Action-graded severity scoring (L0-L6) for tool-using AI agents, computed from red-team execution traces.
Reproducible evaluation of deterministic sequence rules for AI-agent tool-call traces.
Measure how well an agent tool-call authorization layer separates legitimate actions from injected ones, on AgentDojo ground truth. No agent runs, no model calls, seconds to run. Ships the controls that make a block rate meaningful.
Formal runtime verification for tool-using LLM agents: MFOTL/MonPoly replayed offline on AgentDojo, STAC and R-Judge. Paper, MFOTL specifications, experiments, and the benchmark-readiness audit (CPSIoTSec 2026).
Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
LangChain-native AgentDojo benchmark: utility + ASR evaluation across banking, slack, travel, and workspace suites.
A local-first research scaffold for evaluating models, agent harnesses, and complete agent products on realistic, stateful tasks.
Measuring prompt-injection defences against policy-conformant attacks
Evaluating provenance-gated tool calls as a prompt injection defense on AgentDojo. Reproducible runs, per-case analysis, and published results.
Security audit of LLM-based multi-agent systems with indirect prompt-injection PoCs and mitigations.
Systematic evaluation of prompt injection defenses in LLM agents, across seven models and two benchmarks, plus a mechanistic analysis of why the defenses fail.
Personal research project — solo, unaffiliated. Inspect AI evaluation framework for LLM agent security: ASR, benign utility, and Transparency Rate across prompt injection, tool poisoning, and psych attacks.
To associate your repository with the agentdojo topic, visit your repo's landing page and select "manage topics."