sources · source-BATCH-2026-005-006

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.

Open the canonical original source

Source at a glance

Source type
academic paper
Publisher
NeurIPS 2024 Datasets and Benchmarks Track
Source or access date
2024-06-19
Analysis depth
deeply analyzed
Rights treatment
link and paraphrase
Human review
Murray Newlands · 2026-08-16

What the source reports or argues

Important limitations

Source-located statements

6 reviewed statements are indexed from this source.

Evidence lineage and transparency

Relationship fieldLinked identifiers
author idsperson-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer
correction idscorrection-BATCH-2026-005-001

Machine review: machine checked. Human review: approved. Workflow: published.

Complete structured record
source id
source-BATCH-2026-005-006
canonical title
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
alternate titles
source type
academic_paper
series or parent source
arXiv:2406.13352v3; NeurIPS 2024 proceedings record
publisher
NeurIPS 2024 Datasets and Benchmarks Track
channel
Unknown
speaker ids
author ids
person-BATCH-2026-005-edoardo-debenedetti, person-BATCH-2026-005-jie-zhang, person-BATCH-2026-005-mislav-balunovic, person-BATCH-2026-005-luca-beurer-kellner, person-BATCH-2026-005-marc-fischer, person-BATCH-2026-005-florian-tramer
institutional author
Unknown
organization references
recorded at
Unknown
event date
Unknown
published at
2024-06-19
updated at
2024-11-24
duration seconds
Unknown
language
lang-en
translated title
Unknown
translation method
Unknown
geography of speaker
geography of organization
geography discussed
synthetic environments; no human geography
study geography
synthetic environments; no human geography
original url
https://arxiv.org/abs/2406.13352
canonical url
https://arxiv.org/abs/2406.13352
archived url
Unknown
embed url
Unknown
doi
10.48550/arXiv.2406.13352
canonical identity status
verified
repost status
original_or_authoritative_rendition
original source id
Unknown
rights status
link_and_paraphrase
ownership status
third_party
relationship to off
none_identified; Executive AI Research snapshot contains no matching production records
transcript status
Unknown
transcript source
Unknown
transcript republication permission
Unknown
chapter markers
analysis basis
Complete 26-page v3 conference paper reviewed; benchmark composition, figures, Tables 1 and 3-5, confidence intervals, limitations, and funding inspected visually.
topics
topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
executive roles
role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
original abstract
Unknown
inclusion rationale
Accepted discovery candidate candidate-BATCH-2026-001-050; contributes peer-reviewed conference paper evidence or normative context to governed AI-agent identity.
source quality dimensions
{"attribution_strength":"high","methodological_transparency":"high","independence":"academic_with_disclosed_industry_affiliations","bibliographic_stability":"high_arxiv_and_neurips_records","evidence_strength":"strong_for_benchmark_conditions_not_real_world_incidence","unresolved":"Confidence-interval construction and raw row numerators are not stated in the reviewed tables."}
methodology quality
{"design":"controlled benchmark experiments in stateful synthetic tool environments","transparency":"high","human_review_required":true}
study design
controlled benchmark experiments in stateful synthetic tool environments
sample
{"size":"97 user tasks; 27 injection tasks; 629 security test cases; seven listed agent/model configurations in Table 3","sampling_method":"researcher-curated realistic tasks and synthetic data; compatible task cross-product"}
population
LLM agents executing tools over simulated workspace, Slack, travel, and banking environments
date range
{"fieldwork":"not reported; v3 results reflect the paper's 2024 model/API versions","version":"v3; updated after a Llama implementation bug fix and travel-suite update"}
funding
Edoardo Debenedetti supported by armasuisse Science and Technology; Jie Zhang funded by Swiss National Science Foundation grant 214838; benchmark funding statement says no institution explicitly funded benchmark creation
sponsor
no benchmark sponsor reported
peer review status
verified conference publication in official NeurIPS 2024 proceedings
findings
The benchmark exposes substantial utility and security tradeoffs., Table 5 reports targeted attack success of 57.69% with no defense and 6.84% with tool filtering for the listed GPT-4o setup, with reported intervals.
limitations
Synthetic environments and versioned APIs limit external validity., Raw numerators and interval-construction method are not reported in the reviewed table., v3 replaced earlier results after a bug fix and travel-suite update.
correction ids
correction-BATCH-2026-005-001
retraction status
no_retraction_identified;_v3_documents_bug_fix_update
content hash
349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05
accessed at
2026-08-15
verification status
metadata_content_and_version_machine_verified_human_approved
source depth
deeply_analyzed
publication status
published
workflow status
published
machine review status
machine_checked
human review status
approved
reviewed by
Murray Newlands
reviewed at
2026-08-16T23:17:22Z
Provenance and revision history
{
  "provenance": [
    {
      "source_url": "https://arxiv.org/abs/2406.13352",
      "accessed_at": "2026-08-15",
      "retrieval_method": "official source retrieval and machine-assisted review",
      "exact_locator": "v3; updated after a Llama implementation bug fix and travel-suite update",
      "content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
      "batch_id": "BATCH-2026-005",
      "prompt_id": "OEII-EVIDENCE-RESEARCH",
      "prompt_version": "2.0",
      "notes": "Human review pending."
    }
  ],
  "revision_history": [
    {
      "changed_at": "2026-08-16T23:17:22Z",
      "changed_by": "Murray Newlands",
      "summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
      "batch_id": "BATCH-2026-005"
    }
  ]
}

Open machine-readable record