sources · source-BATCH-2026-005-007

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Accepted discovery candidate candidate-BATCH-2026-001-054; contributes conference paper; acceptance marker verified in version, venue record not independently retrieved evidence or normative context to governed AI-agent identity.

Open the canonical original source

Source at a glance

Source type
academic paper
Publisher
ICLR 2025 / arXiv
Source or access date
2024-10-03
Analysis depth
deeply analyzed
Rights treatment
link and paraphrase
Human review
Murray Newlands · 2026-08-16

What the source reports or argues

Important limitations

Source-located statements

5 reviewed statements are indexed from this source.

Evidence lineage and transparency

Relationship fieldLinked identifiers
author idsperson-BATCH-2026-005-hanrong-zhang, person-BATCH-2026-005-jingyuan-huang, person-BATCH-2026-005-kai-mei, person-BATCH-2026-005-yifei-yao, person-BATCH-2026-005-zhenting-wang, person-BATCH-2026-005-chenlu-zhan, person-BATCH-2026-005-hongwei-wang, person-BATCH-2026-005-yongfeng-zhang

Machine review: machine checked. Human review: approved. Workflow: published.

Complete structured record
source id
source-BATCH-2026-005-007
canonical title
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
alternate titles
source type
academic_paper
series or parent source
arXiv:2410.02644v4
publisher
ICLR 2025 / arXiv
channel
Unknown
speaker ids
author ids
person-BATCH-2026-005-hanrong-zhang, person-BATCH-2026-005-jingyuan-huang, person-BATCH-2026-005-kai-mei, person-BATCH-2026-005-yifei-yao, person-BATCH-2026-005-zhenting-wang, person-BATCH-2026-005-chenlu-zhan, person-BATCH-2026-005-hongwei-wang, person-BATCH-2026-005-yongfeng-zhang
institutional author
Unknown
organization references
recorded at
Unknown
event date
Unknown
published at
2024-10-03
updated at
2025-05-30
duration seconds
Unknown
language
lang-en
translated title
Unknown
translation method
Unknown
geography of speaker
geography of organization
geography discussed
synthetic benchmark; no human geography
study geography
synthetic benchmark; no human geography
original url
https://arxiv.org/abs/2410.02644
canonical url
https://arxiv.org/abs/2410.02644
archived url
Unknown
embed url
Unknown
doi
10.48550/arXiv.2410.02644
canonical identity status
verified
repost status
original_or_authoritative_rendition
original source id
Unknown
rights status
link_and_paraphrase
ownership status
third_party
relationship to off
none_identified; Executive AI Research snapshot contains no matching production records
transcript status
Unknown
transcript source
Unknown
transcript republication permission
Unknown
chapter markers
analysis basis
Complete 36-page v4 paper reviewed; attack/defense definitions, scenario tables, metric table, result tables, and reproducibility appendix inspected visually.
topics
topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
executive roles
role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
original abstract
Unknown
inclusion rationale
Accepted discovery candidate candidate-BATCH-2026-001-054; contributes conference paper; acceptance marker verified in version, venue record not independently retrieved evidence or normative context to governed AI-agent identity.
source quality dimensions
{"attribution_strength":"high","methodological_transparency":"high_for_benchmark_structure","independence":"academic","bibliographic_stability":"high_arxiv_doi_resolved","evidence_strength":"moderate_to_strong_for_benchmark_conditions","unresolved":"Venue acceptance is author/arXiv-reported; OpenReview verification was unavailable in this run. Headline average lacks uncertainty and a compact denominator statement."}
methodology quality
{"design":"multi-scenario controlled security benchmark","transparency":"high_for_benchmark_structure","human_review_required":true}
study design
multi-scenario controlled security benchmark
sample
{"size":"400 tasks; 10 scenarios; 10 agents; more than 400 tools; 27 attack/defense methods; 13 LLM backbones","sampling_method":"researcher-constructed scenarios, tasks, tools, and attacks; selection procedure not probabilistic"}
population
LLM-agent configurations under synthetic attacks and defenses
date range
{"fieldwork":"not reported; model versions correspond to experiments before v4 dated 2025-05-30","version":"v4; paper and arXiv page state accepted at ICLR 2025"}
funding
not reported
sponsor
not reported
peer review status
conference acceptance reported in the paper and arXiv metadata; independent OpenReview record retrieval returned access denied, so peer-review verification remains qualified
findings
The authors report a highest average attack success rate of 84.30% for a benchmark configuration., Current defenses vary and often trade security against task performance.
limitations
No production population or human sample., The headline average does not disclose a raw numerator or confidence interval., Funding and conflicts are not reported.
correction ids
retraction status
none_identified_on_arxiv_version_history_as_of_2026-08-15
content hash
e20155df01b3a1f6c0a947c4e9e4871957cff2e507939f7ea21a5fe967d84504
accessed at
2026-08-15
verification status
metadata_content_and_version_machine_verified_human_approved
source depth
deeply_analyzed
publication status
published
workflow status
published
machine review status
machine_checked
human review status
approved
reviewed by
Murray Newlands
reviewed at
2026-08-16T23:17:22Z
Provenance and revision history
{
  "provenance": [
    {
      "source_url": "https://arxiv.org/abs/2410.02644",
      "accessed_at": "2026-08-15",
      "retrieval_method": "official source retrieval and machine-assisted review",
      "exact_locator": "v4; paper and arXiv page state accepted at ICLR 2025",
      "content_hash": "e20155df01b3a1f6c0a947c4e9e4871957cff2e507939f7ea21a5fe967d84504",
      "batch_id": "BATCH-2026-005",
      "prompt_id": "OEII-EVIDENCE-RESEARCH",
      "prompt_version": "2.0",
      "notes": "Human review pending."
    }
  ],
  "revision_history": [
    {
      "changed_at": "2026-08-16T23:17:22Z",
      "changed_by": "Murray Newlands",
      "summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
      "batch_id": "BATCH-2026-005"
    }
  ]
}

Open machine-readable record