sources · source-BATCH-2026-005-008

Teams of LLM Agents can Exploit Zero-Day Vulnerabilities

Accepted discovery candidate candidate-BATCH-2026-001-058; contributes preprint evidence or normative context to governed AI-agent identity.

Open the canonical original source

Source at a glance

Source type
academic paper
Publisher
arXiv
Source or access date
2024-06-02
Analysis depth
deeply analyzed
Rights treatment
link and paraphrase
Human review
Murray Newlands · 2026-08-16

What the source reports or argues

Important limitations

Source-located statements

6 reviewed statements are indexed from this source.

Evidence lineage and transparency

Relationship fieldLinked identifiers
author idsperson-BATCH-2026-005-yuxuan-zhu, person-BATCH-2026-005-antony-kellermann, person-BATCH-2026-005-akul-gupta, person-BATCH-2026-005-philip-li, person-BATCH-2026-005-richard-fang, person-BATCH-2026-005-rohan-bindu, person-BATCH-2026-005-daniel-kang

Machine review: machine checked. Human review: approved. Workflow: published.

Complete structured record
source id
source-BATCH-2026-005-008
canonical title
Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
alternate titles
source type
academic_paper
series or parent source
arXiv:2406.01637v2
publisher
arXiv
channel
Unknown
speaker ids
author ids
person-BATCH-2026-005-yuxuan-zhu, person-BATCH-2026-005-antony-kellermann, person-BATCH-2026-005-akul-gupta, person-BATCH-2026-005-philip-li, person-BATCH-2026-005-richard-fang, person-BATCH-2026-005-rohan-bindu, person-BATCH-2026-005-daniel-kang
institutional author
Unknown
organization references
recorded at
Unknown
event date
Unknown
published at
2024-06-02
updated at
2025-03-30
duration seconds
Unknown
language
lang-en
translated title
Unknown
translation method
Unknown
geography of speaker
geography of organization
geography discussed
synthetic sandbox; no human geography
study geography
synthetic sandbox; no human geography
original url
https://arxiv.org/abs/2406.01637
canonical url
https://arxiv.org/abs/2406.01637
archived url
Unknown
embed url
Unknown
doi
10.48550/arXiv.2406.01637
canonical identity status
verified
repost status
original_or_authoritative_rendition
original source id
Unknown
rights status
link_and_paraphrase
ownership status
third_party
relationship to off
none_identified; Executive AI Research snapshot contains no matching production records
transcript status
Unknown
transcript source
Unknown
transcript republication permission
Unknown
chapter markers
analysis basis
Complete 10-page v2 preprint reviewed; Tables 1-2, Figures 2-3, methods, results, and limitation text inspected visually.
topics
topic-ai-agents, topic-ai-governance, topic-cybersecurity, topic-non-human-identity
executive roles
role-ciso, role-cio, role-cto, role-general-counsel, role-board-director
original abstract
Unknown
inclusion rationale
Accepted discovery candidate candidate-BATCH-2026-001-058; contributes preprint evidence or normative context to governed AI-agent identity.
source quality dimensions
{"attribution_strength":"high","methodological_transparency":"moderate","independence":"academic_with_platform_coordination_disclosed","bibliographic_stability":"high_arxiv_doi_resolved","evidence_strength":"moderate_for_selected_sandbox_cases","unresolved":"Nonrelease of code/prompts limits replication; funding and conflict disclosures are absent."}
methodology quality
{"design":"controlled cybersecurity benchmark experiment in sandboxed reproductions","transparency":"moderate","human_review_required":true}
study design
controlled cybersecurity benchmark experiment in sandboxed reproductions
sample
{"size":"14 vulnerabilities; pass@1 and pass@5 trials; exact total run count not summarized","sampling_method":"purposive selection of recent, reproducible web vulnerabilities with clear success triggers and manual exploitability"}
population
reproducible open-source web vulnerabilities; not all vulnerabilities or deployed systems
date range
{"fieldwork":"not reported","version":"v2 submitted 2025-03-30"}
funding
not reported
sponsor
not reported; authors state OpenAI requested code and prompts remain confidential
peer review status
not peer reviewed; preprint
findings
The authors report 42% pass@5 and 18% pass@1 for HPTSA with GPT-4 on the 14-vulnerability benchmark., Open-source scanners and listed open-source models report 0% on this benchmark.
limitations
Only 14 selected web vulnerabilities., No confidence intervals., Code and prompts are withheld, limiting replication., Results are capability demonstrations, not estimates of incident prevalence.
correction ids
retraction status
none_identified_on_arxiv_version_history_as_of_2026-08-15
content hash
f4fc06040f0856d99d2009042b4c6fa0c01206da8ca61cb39a94d52b9867cf71
accessed at
2026-08-15
verification status
metadata_content_and_version_machine_verified_human_approved
source depth
deeply_analyzed
publication status
published
workflow status
published
machine review status
machine_checked
human review status
approved
reviewed by
Murray Newlands
reviewed at
2026-08-16T23:17:22Z
Provenance and revision history
{
  "provenance": [
    {
      "source_url": "https://arxiv.org/abs/2406.01637",
      "accessed_at": "2026-08-15",
      "retrieval_method": "official source retrieval and machine-assisted review",
      "exact_locator": "v2 submitted 2025-03-30",
      "content_hash": "f4fc06040f0856d99d2009042b4c6fa0c01206da8ca61cb39a94d52b9867cf71",
      "batch_id": "BATCH-2026-005",
      "prompt_id": "OEII-EVIDENCE-RESEARCH",
      "prompt_version": "2.0",
      "notes": "Human review pending."
    }
  ],
  "revision_history": [
    {
      "changed_at": "2026-08-16T23:17:22Z",
      "changed_by": "Murray Newlands",
      "summary": "Approved for the governed-identities pilot release under the exact scope, exclusions, rights treatment, and limitations recorded in issue #18.",
      "batch_id": "BATCH-2026-005"
    }
  ]
}

Open machine-readable record