statements ยท statement-BATCH-2026-006-023
source-BATCH-2026-005-006
AgentDojo version 3 contains four simulated environments, 70 tools, 97 user tasks, 27 injection targets, and 629 security cases.
Statement context
- Statement type
- methodological claim
- Exact source locator
- arXiv:2406.13352v3 PDF pp. 1 and 6, Abstract and Table 1
- Source date
- 2024-06-19
- Role at source time
- joint authors of the AgentDojo paper
- Evidence character
- primary empirical result
- Scope
- Version 3 benchmark composition.
Recorded uncertainty: The cases are curated simulations rather than a probability sample of deployed agents.
Indexed source: AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Evidence lineage and transparency
| Relationship field | Linked identifiers |
|---|---|
| topic ids | topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity |
Machine review: machine checked. Human review: approved. Workflow: published.
Complete structured record
- statement id
- statement-BATCH-2026-006-023
- source id
- source-BATCH-2026-005-006
- person id
- Unknown
- institutional author
- Debenedetti et al. (joint paper authors)
- book edition id
- Unknown
- speaker role at source time
- joint authors of the AgentDojo paper
- organization at source time
- ETH Zurich and Invariant Labs
- statement type
- methodological_claim
- neutral paraphrase
- AgentDojo version 3 contains four simulated environments, 70 tools, 97 user tasks, 27 injection targets, and 629 security cases.
- direct quote
- Unknown
- direct quote rights note
- Unknown
- exact locator
- arXiv:2406.13352v3 PDF pp. 1 and 6, Abstract and Table 1
- locator type
- pdf_page
- source date
- 2024-06-19
- topic ids
- topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
- executive role context
- role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
- industry context
- cross-sector
- geographic context
- synthetic environments; no human geography
- evidence character
- primary_empirical_result
- factual verification status
- source_reported_not_independently_verified
- statement scope
- Version 3 benchmark composition.
- uncertainty
- The cases are curated simulations rather than a probability sample of deployed agents.
- extraction method
- Machine-assisted close reading of the exact verified source version; original neutral paraphrase; human review pending.
- machine extraction confidence
- 0.94
- independent agent review status
- not_started
- publication status
- published
- workflow status
- published
- machine review status
- machine_checked
- human review status
- approved
- reviewed by
- Murray Newlands
- reviewed at
- 2026-08-16T23:17:22Z
Provenance and revision history
{
"provenance": [
{
"source_url": "https://arxiv.org/abs/2406.13352",
"accessed_at": "2026-08-15",
"retrieval_method": "Exact verified source version inherited from BATCH-2026-005 and statement-level close reading",
"exact_locator": "arXiv:2406.13352v3 PDF pp. 1 and 6, Abstract and Table 1",
"content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
"batch_id": "BATCH-2026-006",
"prompt_id": "OEII-STATEMENT-CODING",
"prompt_version": "2.0",
"notes": "Source eligibility inherited from source-BATCH-2026-005-006; human attribution and publication review pending."
}
],
"revision_history": [
{
"changed_at": "2026-08-16T23:17:22Z",
"changed_by": "Murray Newlands",
"summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
"batch_id": "BATCH-2026-006"
}
]
}