research ยท source-BATCH-2026-005-006
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.
Open the canonical original source
Evidence lineage and transparency
| Relationship field | Linked identifiers |
|---|---|
| author ids | person-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer |
| correction ids | correction-BATCH-2026-005-001 |
Machine review: machine checked. Human review: approved. Workflow: published.
Complete structured record
- source id
- source-BATCH-2026-005-006
- canonical title
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- alternate titles
- source type
- academic_paper
- series or parent source
- arXiv:2406.13352v3; NeurIPS 2024 proceedings record
- publisher
- NeurIPS 2024 Datasets and Benchmarks Track
- channel
- Unknown
- speaker ids
- author ids
- person-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer
- institutional author
- Unknown
- organization references
- recorded at
- Unknown
- event date
- Unknown
- published at
- 2024-06-19
- updated at
- 2024-11-24
- duration seconds
- Unknown
- language
- lang-en
- translated title
- Unknown
- translation method
- Unknown
- geography of speaker
- geography of organization
- geography discussed
- synthetic environments; no human geography
- study geography
- synthetic environments; no human geography
- original url
- https://arxiv.org/abs/2406.13352
- canonical url
- https://arxiv.org/abs/2406.13352
- archived url
- Unknown
- embed url
- Unknown
- doi
- 10.48550/arXiv.2406.13352
- canonical identity status
- verified
- repost status
- original_or_authoritative_rendition
- original source id
- Unknown
- rights status
- link_and_paraphrase
- ownership status
- third_party
- relationship to off
- none_identified; Executive AI Research snapshot contains no matching production records
- transcript status
- Unknown
- transcript source
- Unknown
- transcript republication permission
- Unknown
- chapter markers
- analysis basis
- Complete 26-page v3 conference paper reviewed; benchmark composition, figures, Tables 1 and 3-5, confidence intervals, limitations, and funding inspected visually.
- topics
- topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
- executive roles
- role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
- original abstract
- Unknown
- inclusion rationale
- Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.
- source quality dimensions
- {"attribution_strength":"high","methodological_transparency":"high","independence":"academic_with_disclosed_industry_affiliations","bibliographic_stability":"high_arxiv_and_neurips_records","evidence_strength":"strong_for_benchmark_conditions_not_real_world_incidence","unresolved":"Confidence-interval construction and raw row numerators are not stated in the reviewed tables."}
- methodology quality
- {"design":"controlled benchmark experiments in stateful synthetic tool environments","transparency":"high","human_review_required":true}
- study design
- controlled benchmark experiments in stateful synthetic tool environments
- sample
- {"size":"97 user tasks; 27 injection tasks; 629 security test cases; seven listed agent/model configurations in Table 3","sampling_method":"researcher-curated realistic tasks and synthetic data; compatible task cross-product"}
- population
- LLM agents executing tools over simulated workspace, Slack, travel, and banking environments
- date range
- {"fieldwork":"not reported; v3 results reflect the paper's 2024 model/API versions","version":"v3; updated after a Llama implementation bug fix and travel-suite update"}
- funding
- Edoardo Debenedetti supported by armasuisse Science and Technology; Jie Zhang funded by Swiss National Science Foundation grant 214838; benchmark funding statement says no institution explicitly funded benchmark creation
- sponsor
- no benchmark sponsor reported
- peer review status
- verified conference publication in official NeurIPS 2024 proceedings
- findings
- The benchmark exposes substantial utility and security tradeoffs., Table 5 reports targeted attack success of 57.69% with no defense and 6.84% with tool filtering for the listed GPT-4o setup, with reported intervals.
- limitations
- Synthetic environments and versioned APIs limit external validity., Raw numerators and interval-construction method are not reported in the reviewed table., v3 replaced earlier results after a bug fix and travel-suite update.
- correction ids
- correction-BATCH-2026-005-001
- retraction status
- no_retraction_identified;_v3_documents_bug_fix_update
- content hash
- 349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05
- accessed at
- 2026-08-15
- verification status
- metadata_content_and_version_machine_verified_human_approved
- source depth
- deeply_analyzed
- publication status
- published
- workflow status
- published
- machine review status
- machine_checked
- human review status
- approved
- reviewed by
- Murray Newlands
- reviewed at
- 2026-08-16T23:17:22Z
Provenance and revision history
{
"provenance": [
{
"source_url": "https://arxiv.org/abs/2406.13352",
"accessed_at": "2026-08-15",
"retrieval_method": "official source retrieval and machine-assisted review",
"exact_locator": "v3; updated after a Llama implementation bug fix and travel-suite update",
"content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
"batch_id": "BATCH-2026-005",
"prompt_id": "OEII-EVIDENCE-RESEARCH",
"prompt_version": "2.0",
"notes": "Human review pending."
}
],
"revision_history": [
{
"changed_at": "2026-08-16T23:17:22Z",
"changed_by": "Murray Newlands",
"summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
"batch_id": "BATCH-2026-005"
}
]
}