sources · source-BATCH-2026-005-006
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.
Open the canonical original source
Source at a glance
- Source type
- academic paper
- Publisher
- NeurIPS 2024 Datasets and Benchmarks Track
- Source or access date
- 2024-06-19
- Analysis depth
- deeply analyzed
- Rights treatment
- link and paraphrase
- Human review
- Murray Newlands · 2026-08-16
What the source reports or argues
- The benchmark exposes substantial utility and security tradeoffs.
- Table 5 reports targeted attack success of 57.69% with no defense and 6.84% with tool filtering for the listed GPT-4o setup, with reported intervals.
Important limitations
- Synthetic environments and versioned APIs limit external validity.
- Raw numerators and interval-construction method are not reported in the reviewed table.
- v3 replaced earlier results after a bug fix and travel-suite update.
Source-located statements
6 reviewed statements are indexed from this source.
- AgentDojo distinguishes benign task completion, utility preserved during attack, and successful execution of an attacker goal as separate evaluation measures. (arXiv:2406.13352v3 PDF p. 6, Section 3.4 "Reporting AgentDojo Results")
- AgentDojo version 3 contains four simulated environments, 70 tools, 97 user tasks, 27 injection targets, and 629 security cases. (arXiv:2406.13352v3 PDF pp. 1 and 6, Abstract and Table 1)
- In the reported no-defense runs, several evaluated models failed to complete at least one third of benign tasks. (arXiv:2406.13352v3 PDF p. 20, Table 3, "Benign utility" column)
- For GPT-4o, Table 5 reports targeted attack success of 57.69% without a defense and 6.84% with tool filtering. (arXiv:2406.13352v3 PDF p. 20, Table 5, "Targeted ASR" row)
- Across the plotted defense configurations, utility during attack was roughly 15 to 20 percentage points below benign utility. (arXiv:2406.13352v3 PDF p. 9, Figure 9b and caption)
- The authors report releasing benchmark code, run outputs, conversations, figure notebooks, documentation, and exact dependency information. (arXiv:2406.13352v3 PDF p. 22, Appendix E.1-E.3)
Evidence lineage and transparency
| Relationship field | Linked identifiers |
|---|---|
| author ids | person-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer |
| correction ids | correction-BATCH-2026-005-001 |
Machine review: machine checked. Human review: approved. Workflow: published.
Complete structured record
- source id
- source-BATCH-2026-005-006
- canonical title
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- alternate titles
- source type
- academic_paper
- series or parent source
- arXiv:2406.13352v3; NeurIPS 2024 proceedings record
- publisher
- NeurIPS 2024 Datasets and Benchmarks Track
- channel
- Unknown
- speaker ids
- author ids
- person-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer
- institutional author
- Unknown
- organization references
- recorded at
- Unknown
- event date
- Unknown
- published at
- 2024-06-19
- updated at
- 2024-11-24
- duration seconds
- Unknown
- language
- lang-en
- translated title
- Unknown
- translation method
- Unknown
- geography of speaker
- geography of organization
- geography discussed
- synthetic environments; no human geography
- study geography
- synthetic environments; no human geography
- original url
- https://arxiv.org/abs/2406.13352
- canonical url
- https://arxiv.org/abs/2406.13352
- archived url
- Unknown
- embed url
- Unknown
- doi
- 10.48550/arXiv.2406.13352
- canonical identity status
- verified
- repost status
- original_or_authoritative_rendition
- original source id
- Unknown
- rights status
- link_and_paraphrase
- ownership status
- third_party
- relationship to off
- none_identified; Executive AI Research snapshot contains no matching production records
- transcript status
- Unknown
- transcript source
- Unknown
- transcript republication permission
- Unknown
- chapter markers
- analysis basis
- Complete 26-page v3 conference paper reviewed; benchmark composition, figures, Tables 1 and 3-5, confidence intervals, limitations, and funding inspected visually.
- topics
- topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
- executive roles
- role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
- original abstract
- Unknown
- inclusion rationale
- Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.
- source quality dimensions
- {"attribution_strength":"high","methodological_transparency":"high","independence":"academic_with_disclosed_industry_affiliations","bibliographic_stability":"high_arxiv_and_neurips_records","evidence_strength":"strong_for_benchmark_conditions_not_real_world_incidence","unresolved":"Confidence-interval construction and raw row numerators are not stated in the reviewed tables."}
- methodology quality
- {"design":"controlled benchmark experiments in stateful synthetic tool environments","transparency":"high","human_review_required":true}
- study design
- controlled benchmark experiments in stateful synthetic tool environments
- sample
- {"size":"97 user tasks; 27 injection tasks; 629 security test cases; seven listed agent/model configurations in Table 3","sampling_method":"researcher-curated realistic tasks and synthetic data; compatible task cross-product"}
- population
- LLM agents executing tools over simulated workspace, Slack, travel, and banking environments
- date range
- {"fieldwork":"not reported; v3 results reflect the paper's 2024 model/API versions","version":"v3; updated after a Llama implementation bug fix and travel-suite update"}
- funding
- Edoardo Debenedetti supported by armasuisse Science and Technology; Jie Zhang funded by Swiss National Science Foundation grant 214838; benchmark funding statement says no institution explicitly funded benchmark creation
- sponsor
- no benchmark sponsor reported
- peer review status
- verified conference publication in official NeurIPS 2024 proceedings
- findings
- The benchmark exposes substantial utility and security tradeoffs., Table 5 reports targeted attack success of 57.69% with no defense and 6.84% with tool filtering for the listed GPT-4o setup, with reported intervals.
- limitations
- Synthetic environments and versioned APIs limit external validity., Raw numerators and interval-construction method are not reported in the reviewed table., v3 replaced earlier results after a bug fix and travel-suite update.
- correction ids
- correction-BATCH-2026-005-001
- retraction status
- no_retraction_identified;_v3_documents_bug_fix_update
- content hash
- 349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05
- accessed at
- 2026-08-15
- verification status
- metadata_content_and_version_machine_verified_human_approved
- source depth
- deeply_analyzed
- publication status
- published
- workflow status
- published
- machine review status
- machine_checked
- human review status
- approved
- reviewed by
- Murray Newlands
- reviewed at
- 2026-08-16T23:17:22Z
Provenance and revision history
{
"provenance": [
{
"source_url": "https://arxiv.org/abs/2406.13352",
"accessed_at": "2026-08-15",
"retrieval_method": "official source retrieval and machine-assisted review",
"exact_locator": "v3; updated after a Llama implementation bug fix and travel-suite update",
"content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
"batch_id": "BATCH-2026-005",
"prompt_id": "OEII-EVIDENCE-RESEARCH",
"prompt_version": "2.0",
"notes": "Human review pending."
}
],
"revision_history": [
{
"changed_at": "2026-08-16T23:17:22Z",
"changed_by": "Murray Newlands",
"summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
"batch_id": "BATCH-2026-005"
}
]
}