SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Abstract
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first benchmark for evaluating LLM agents on the post-compromise incident-response workflow. Given a forensic disk snapshot of a compromised host together with the alerts, vulnerability scans, and baseline checks reported by a host security product, agents are required to produce forensic reports on intrusions, baseline risks, and vulnerability risks, together with a remediation plan. We instantiate this task across 10 cyber ranges, each constructed from a distinct compromised cloud host, spanning 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. We evaluate 23 frontier LLMs on the OpenCode agent harness. Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range. This reveals a fundamental bottleneck in building agents for real-world incident response. The benchmark is publicly available at https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond.
Community
We introduce SecRespond, a benchmark for evaluating AI agents on real-world post-compromise incident response. It includes 10 reproducible cyber ranges across Linux and Windows, frozen forensic disk snapshots, synthetic security-product evidence, and expert-authored evaluation checklists that separately assess detection and remediation planning.
Paper: https://arxiv.org/abs/2607.26791
Code and data: https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond
Dataset: https://huggingface.co/datasets/Alibaba-NLP/SecRespond
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges (2026)
- Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens (2026)
- OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills (2026)
- AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments (2026)
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents (2026)
- GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines (2026)
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper