statements ยท statement-BATCH-2026-006-022

source-BATCH-2026-005-006

AgentDojo distinguishes benign task completion, utility preserved during attack, and successful execution of an attacker goal as separate evaluation measures.

Statement context

Statement type
definition
Exact source locator
arXiv:2406.13352v3 PDF p. 6, Section 3.4 "Reporting AgentDojo Results"
Source date
2024-06-19
Role at source time
joint authors of the AgentDojo paper
Evidence character
primary empirical result
Scope
Metric definitions for version 3 of the AgentDojo benchmark.

Recorded uncertainty: Metric definitions do not establish external validity.

Indexed source: AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Evidence lineage and transparency

Relationship fieldLinked identifiers
topic idstopic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity

Machine review: machine checked. Human review: approved. Workflow: published.

Complete structured record
statement id
statement-BATCH-2026-006-022
source id
source-BATCH-2026-005-006
person id
Unknown
institutional author
Debenedetti et al. (joint paper authors)
book edition id
Unknown
speaker role at source time
joint authors of the AgentDojo paper
organization at source time
ETH Zurich and Invariant Labs
statement type
definition
neutral paraphrase
AgentDojo distinguishes benign task completion, utility preserved during attack, and successful execution of an attacker goal as separate evaluation measures.
direct quote
Unknown
direct quote rights note
Unknown
exact locator
arXiv:2406.13352v3 PDF p. 6, Section 3.4 "Reporting AgentDojo Results"
locator type
pdf_page
source date
2024-06-19
topic ids
topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
executive role context
role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
industry context
cross-sector
geographic context
synthetic environments; no human geography
evidence character
primary_empirical_result
factual verification status
attribution_verified
statement scope
Metric definitions for version 3 of the AgentDojo benchmark.
uncertainty
Metric definitions do not establish external validity.
extraction method
Machine-assisted close reading of the exact verified source version; original neutral paraphrase; human review pending.
machine extraction confidence
0.94
independent agent review status
not_started
publication status
published
workflow status
published
machine review status
machine_checked
human review status
approved
reviewed by
Murray Newlands
reviewed at
2026-08-16T23:17:22Z
Provenance and revision history
{
  "provenance": [
    {
      "source_url": "https://arxiv.org/abs/2406.13352",
      "accessed_at": "2026-08-15",
      "retrieval_method": "Exact verified source version inherited from BATCH-2026-005 and statement-level close reading",
      "exact_locator": "arXiv:2406.13352v3 PDF p. 6, Section 3.4 \"Reporting AgentDojo Results\"",
      "content_hash": "349884fffbf43282591c5accffd57bb651632f38e412fe303114941bdd111b05",
      "batch_id": "BATCH-2026-006",
      "prompt_id": "OEII-STATEMENT-CODING",
      "prompt_version": "2.0",
      "notes": "Source eligibility inherited from source-BATCH-2026-005-006; human attribution and publication review pending."
    }
  ],
  "revision_history": [
    {
      "changed_at": "2026-08-16T23:17:22Z",
      "changed_by": "Murray Newlands",
      "summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
      "batch_id": "BATCH-2026-006"
    }
  ]
}

Open machine-readable record