CTF.ae has launched XRanges for AI, a benchmark designed to score autonomous security agents against realistic, instrumented targets[1]. The platform aims to answer a basic question facing defenders: not what an agent claims in its own report, but what it actually did inside a live application[1]. XRanges for AI builds on prior work where 545 human hackers tackled the same labs, giving teams a human baseline against which to measure emerging AI systems[1][5].
XRanges for AI deploys full-featured target applications with instrumentation baked into every service, recording each action an agent takes during an evaluation run[1]. The system ingests this telemetry and converts it into four independent scores that update in real time while the agent is still working, all accessible through a single workspace[1]. Instead of relying on self-reported findings, teams see ground-truth evidence of which boundaries the agent crossed, which vulnerabilities it actually exploited, and how far it progressed through each scenario[1][5]. The platform is available as a managed cloud service or as a self-hosted deployment where no telemetry leaves the customer’s environment[1].
XRanges arrives amid a wave of academic and industry benchmarks trying to quantify how well AI agents handle real vulnerabilities. Researchers at UC Berkeley’s RDI Lab introduced CyberGym, a large-scale corpus covering 1,507 real-world vulnerabilities across 188 widely used open-source projects, where top agents currently reach about 30% single-trial success and roughly 67% after 30 trials, up from around 10% in earlier generations[6]. SEC-bench, focused on LLM code agents, reports that state-of-the-art systems achieve at most 18% success in proof-of-concept generation and around 34% in vulnerability patching across its full dataset, underscoring substantial performance gaps[9]. Benchmarks such as ZeroDayBench and AutoPenBench further probe how agents find and patch previously unseen flaws, but most rely on static pass/fail metrics rather than live operational telemetry[4][13]. XRanges for AI sits alongside these efforts while concentrating on multi-signal scoring tied directly to what happens inside realistic applications[1][5][6].
Other studies highlight the operational risks of deploying offensive security agents at scale. StealthBench measures operational stealth across six OPSEC dimensions and finds that no model exceeds a 54% “safe success” rate that requires both task completion and stealth, suggesting that OPSEC failures are systematic across model families[14]. Exploit-focused corpora like ExploitGym catalogue whether agents can turn vulnerabilities from datasets such as NYU CTF and CVE-Bench into working attacks, as exploit generation capabilities steadily improve[15]. Together, these results show that agentic security systems can be powerful yet error-prone, creating demand for frameworks like XRanges that track both technical capability and operational behavior rather than trusting agents’ own narratives[1][14][15].
Vendors are already competing to top public leaderboards, and the stakes for defenders are rising quickly. Microsoft has highlighted that its multi-model MDASH system scored 88.45% on CyberGym, leading the benchmark’s public leaderboard by about five points over the next best entry[8][10]. Even at those levels, agents still fail on a meaningful share of real vulnerabilities, and many benchmarks expose blind spots in exploitation, patching, or stealth[6][8][9]. XRanges for AI gives security teams a dedicated workspace where they can plug in candidate agents, watch runs unfold in real time, and compare scores across labs without exposing production infrastructure to uncontrolled AI experiments[1].
For defenders exploring autonomous agents for bug bounty, red teaming, or continuous scanning, XRanges for AI provides a way to evaluate whether those systems are ready for anything beyond a lab environment[1]. Organizations can run their own or third-party agents against XRanges labs, review the resulting telemetry and scores, and use that data to validate vendor claims, tune configurations, and decide where human experts must still review findings[1][5][6]. As agentic security matures, the combination of rigorous academic benchmarks and operational scoring platforms like XRanges is likely to become a prerequisite for any serious AI-driven defense program, ensuring that security teams know what their agents truly accomplished—not just what they said they did[1][6][14].
References
- 545 Hackers Tested It First. Now XRanges for AI Scores Your Security Agent
- [PDF] ZERODAYBENCH: EVALUATING LLM AGENTS ON UNSEEN …
- AI Bug-Bounty Agent Benchmark: 12 Models, 100 Black-Box Labs
- CyberGym: Evaluating AI Agents’ Real-World …
- Microsoft’s multi-agent AI system tops Anthropic’s Mythos on cybersecurity benchmark
- SEC-bench: Automated Benchmarking of LLM Agents on Real …
- AI スピードに対応した防御: Microsoft の新たなマルチモデル エージェント型セキュリティシステムが、業界をリードするベンチマークでトップを獲得 – Windows Blog for Japan
- AutoPenBench: A Vulnerability Testing Benchmark for …
- StealthBench: Measuring Operational Stealth in Autonomous …
- Can AI Agents Turn Security Vulnerabilities into Real Attacks? – arXiv
