statements ยท statement-BATCH-2026-006-025
source-BATCH-2026-005-006
For GPT-4o, Table 5 reports targeted attack success of 57.69% without a defense and 6.84% with tool filtering.
Statement context
- Statement type
- empirical finding
- Exact source locator
- arXiv:2406.13352v3 PDF p. 20, Table 5, "Targeted ASR" row
- Source date
- 2024-06-19
- Role at source time
- joint authors of the AgentDojo paper
- Evidence character
- primary empirical result
- Scope
- GPT-4o on AgentDojo v3 security cases under the listed defense configurations.
Recorded uncertainty: The table gives confidence-interval half-widths but not raw numerators or the interval construction method; the rates are not production prevalence.
Indexed source: AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Evidence lineage and transparency
| Relationship field | Linked identifiers |
|---|---|
| topic ids | topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity |
Machine review: machine checked. Human review: approved. Workflow: published.
Complete structured record
- statement id
- statement-BATCH-2026-006-025
- source id
- source-BATCH-2026-005-006
- person id
- Unknown
- institutional author
- Debenedetti et al. (joint paper authors)
- book edition id
- Unknown
- speaker role at source time
- joint authors of the AgentDojo paper
- organization at source time
- ETH Zurich and Invariant Labs
- statement type
- empirical_finding
- neutral paraphrase
- For GPT-4o, Table 5 reports targeted attack success of 57.69% without a defense and 6.84% with tool filtering.
- direct quote
- Unknown
- direct quote rights note
- Unknown
- exact locator
- arXiv:2406.13352v3 PDF p. 20, Table 5, "Targeted ASR" row
- locator type
- pdf_page
- source date
- 2024-06-19
- topic ids
- topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
- executive role context
- role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
- industry context
- cross-sector
- geographic context
- synthetic environments; no human geography
- evidence character
- primary_empirical_result
- factual verification status
- source_reported_not_independently_verified
- statement scope
- GPT-4o on AgentDojo v3 security cases under the listed defense configurations.
- uncertainty
- The table gives confidence-interval half-widths but not raw numerators or the interval construction method; the rates are not production prevalence.
- extraction method
- Machine-assisted close reading of the exact verified source version; original neutral paraphrase; human review pending.
- machine extraction confidence
- 0.94
- independent agent review status
- not_started
- publication status
- published
- workflow status
- published
- machine review status
- machine_checked
- human review status
- approved
- reviewed by
- Murray Newlands
- reviewed at
- 2026-08-16T23:17:22Z
Provenance and revision history
{
"provenance": [
{
"source_url": "https://arxiv.org/abs/2406.13352",
"accessed_at": "2026-08-15",
"retrieval_method": "Exact verified source version inherited from BATCH-2026-005 and statement-level close reading",
"exact_locator": "arXiv:2406.13352v3 PDF p. 20, Table 5, \"Targeted ASR\" row",
"content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
"batch_id": "BATCH-2026-006",
"prompt_id": "OEII-STATEMENT-CODING",
"prompt_version": "2.0",
"notes": "Source eligibility inherited from source-BATCH-2026-005-006; human attribution and publication review pending."
}
],
"revision_history": [
{
"changed_at": "2026-08-16T23:17:22Z",
"changed_by": "Murray Newlands",
"summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
"batch_id": "BATCH-2026-006"
}
]
}