Skip to content
This repository was archived by the owner on Aug 22, 2026. It is now read-only.

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Cybersecurity AI · CAI

The open framework that established Cybersecurity AI as a research domain.
Archived. The research it produced continues.

Status Successor Research Funding License

Move to CSI Read the research Contact


Important

📦 This repository is archived.

CAI is no longer under active development. This branch holds the complete final source tree, read-only, as a research artifact — the reference implementation behind a body of work spanning 18 papers, 30+ CVEs and #1 rankings in international security competitions.

The repository has been consolidated into a single archival commit. Nothing of the source is missing: every file of the final release is here, and the published packages, issues and pull requests remain available (see Where the source lives).

No further releases, bug fixes, security patches or support will be provided.

Everything CAI proved, and everything two years of open development taught us, has been carried forward into Cybersecurity Superintelligence (CSI).

Note

Existing CAI professional (paid) customers — you are not stranded.

We recommend migrating to CSI, where all CAI capabilities, fixes and defenses now live. If CSI is not the right fit for your deployment, write to support@aliasrobotics.com to discuss extended support for your existing CAI installation. Subscriptions, active or lapsed, are handled case by case — reach out before this archive becomes a problem for you.


At a glance

papers CVE IDs CTF speedup cost dataset

Cybersecurity AI (CAI) was a lightweight, open-source framework for building and deploying AI-powered offensive and defensive security automation. Released in April 2025 by Alias Robotics, it became the de facto open framework for AI Security — used by thousands of researchers and hundreds of organizations, and the experimental platform on which the Cybersecurity AI research domain was established.

Last open-source release 0.5.10 — public PyPI cai-framework, December 2025
Last professional release v1.1.5paid customers only, private Alias Robotics package index (this tree)
Development span March 2025 → August 2026 · 1,078 commits by 103 authors, consolidated into one archival commit
Community ~9.8k stars · ~1.4k forks · 275 issues · 195 pull requests
Research output 18 papers and technical reports on Cybersecurity AI
Disclosure record 30+ CVE IDs · 2 CISA ICS advisories · 100+ robot vulnerabilities
Competition record Rank #1 at Neurogrid, Dragos OT and HTB "AI vs Humans"
Successor Cybersecurity Superintelligence (CSI)

includes authors inherited from the upstream openai-agents-python history


Contents

📦 The archive · 🚀 CSI · 🔬 The research · 🇪🇺 Funding · 🧰 Using the archive · 📖 Citation · 🙏 Acknowledgements · ⚖️ License


📦 The archive

Frozen, not deleted. Everything stays readable.

✅ Still available ❌ No longer provided
Complete final source tree (v1.1.5), read-only Releases, bug fixes, security patches
cai-framework 0.5.10 on public PyPI Issue triage, PR review, roadmap
Documentation and examples as of the final release Support of any kind, community or commercial
Issues · PRs as a public record Browsable commit-by-commit history
All papers, datasets and benchmarks alias model access through this framework
Published cai-framework releases on PyPI Guarantees about third-party model providers

Where the source lives

This repository was squashed to a single commit when it was archived, so the incremental history is no longer browsable here. The code itself is fully intact and remains available in several forms:

What Where
Final source tree — every file of the last professional release (v1.1.5) this branch
Open-source releases0.3.9 through 0.5.10, each with its own sdist and wheel cai-framework on PyPI
Development record — 275 issues and 195 pull requests, with their diffs and discussion Issues · Pull requests
Documentation as of the final release docs/ · aliasrobotics.github.io/cai
Research artifacts — papers, benchmarks, datasets aliasrobotics.com/research-security.php

If you need something from the pre-archival history for research or compliance reasons, write to research@aliasrobotics.com.

Warning

An archived offensive-security framework is unmaintained attack tooling. Dependencies will age, provider APIs will drift, and known weaknesses — including the prompt-injection classes we ourselves documented — will not be fixed here. Run it only in isolated environments, against systems you are explicitly authorised to test, and never as part of a production security programme.


🚀 CSI

Six layers. One product. Built for professionals.

Cybersecurity Superintelligence (CSI) is CAI's evolution into an enterprise product. Where CAI was an open research framework — free, best effort, and built to prove what agentic security could do — CSI is a commercially supported, licensed platform for professional security teams, enterprises, critical infrastructure operators, defense and government customers.

CAI survives inside it as the scaffold layer, now one of several harnesses, wrapped in the proprietary models, datasets, agents, steering and benchmarking CAI never had — and delivered with the support, licensing and sovereignty guarantees that production and national-security use require.

Layer What it leverages
🧠 LLMs The alias family (alias3, alias2, alias2-mini, alias1, alias0) — cybersecurity-specialised, on-premise deployable
🔌 Scaffolds Unified routing across Claude Code, Codex, Mistral, CAI and GCAI through a local proxy owning telemetry and cost
📊 Datasets 18.07 TB of expert security trajectories — 26M prompts, 230,935 sessions, 123 countries (arXiv:2605.28146)
🤖 Agents 15+ specialised agents — Defender, Red Team, APT, Forensics, Robot Defender, custom
🎚️ Steering Activation steering and abliteration, cutting refusals on legitimate offensive tasks from 59% to 1%
📈 Benchmarking Continuous measurement through Cybench and CAIBench

What CSI fixes that CAI could not

Two years of open development produced a precise inventory of what an agentic security framework gets wrong — most of it visible in our issues and pull requests. CSI is the answer to that inventory.

🛡️ Maintained injection defenses CAI's four-layer guardrail framework (arXiv:2508.21669) was a research contribution. In CSI it is a supported, continuously updated control — agentic attack surface does not hold still.
🐛 A supported release train Bug fixes and versioned releases, instead of a frozen archive (0.5.10 open source, v1.1.5 professional).
🧑‍💼 Professional support Included with every tier, plus quarterly consulting on agent design, measurement and reporting on annual plans.
🔒 Sovereignty by construction The CAI Dataset paper showed operators routinely paste live credentials and production hostnames into frontier-model APIs, concentrating the world's offensive context in a handful of providers. CSI's answer: on-premise, privately-hosted, cybersecurity-specialised models inside your trust boundary.
🇪🇺 Enterprise & sovereign deployment Commercial licensing, GDPR/NIS2 compliance, fully air-gapped installations, private benchmarking, audit logging and custom fine-tuning.
🧪 Multi-scaffold coverage No single harness dominates. CSI's blackboard architecture composes heterogeneous scaffolds to solve 19/33 Cybench challenges vs 15/33 for the best individual scaffold (arXiv:2605.28334).

Move to CSI

CSI PRO — for professional teams and enterprises
CSI On-Premise — air-gapped, sovereign deployments for defense, government and critical infrastructure

support@aliasrobotics.com


🔬 The research

One framework. Eighteen papers. A new research domain.

CAI was never only a tool. It was the instrument through which Cybersecurity AI was established as a research domain — a testbed for questions about autonomy, evaluation, strategy, defense and regulation that could not be answered on paper alone.

The framework is archived. The research is the lasting contribution, and it continues.

Full curated index → aliasrobotics.com/research-security.php

The arc

  2023  ·  PentestGPT              LLM-guided penetration testing, +228.6% over baseline.
           USENIX Security '24     Security expertise externalised into natural-language guidance.
             │                     ── humans, guided by AI
             │
  2025  ·  CAI                     Open, bug bounty-ready agentic framework.
           this repository         3,600× faster than humans, 156× cheaper. #1 across CTF circuits.
             │                     ── AI, guided by humans
             │
  2026  ·  G-CTR + CSI             Game-theoretic neurosymbolic reasoning, multi-scaffold blackboard
           the successor           orchestration, sovereign on-premise models.
                                   ── human-guided cybersecurity superintelligence

Documented end-to-end in Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI.


The corpus

🏛️ Foundations and framework

Paper What it established
CAI: An Open, Bug Bounty-Ready Cybersecurity AI
arXiv:2504.06017 · Apr 2025 · HTML edition
The framework paper. 3,600× faster than human pentesters at 156× lower cost; CVSS 4.3–7.5 findings in production systems; systematic evaluation across proprietary and open-weight LLMs, exposing the gap between vendor claims and measured capability.
The Dangerous Gap Between Automation and Autonomy
arXiv:2506.23592 · Jun 2025
A 6-level taxonomy separating automation from autonomy in Cybersecurity AI — the vocabulary the field was missing.
CAI Fluency: A Framework for Cybersecurity AI Fluency
arXiv:2508.13588 · Aug 2025
Educational framework and curriculum for cybersecurity AI literacy, developed with academic partners.
Towards Cybersecurity Superintelligence
arXiv:2601.14614 · Jan 2026
Synthesises PentestGPT → CAI → G-CTR into a single trajectory: from AI-guided humans to human-guided AI.

♟️ Strategy, evaluation and benchmarking

Paper What it established
Evaluating Agentic Cybersecurity in Attack/Defense CTFs
arXiv:2510.17521 · Oct 2025
54.3% defensive patching success against 28.3% offensive initial access — defense is, for now, the easier side for agents.
CAIBench: A Meta-Benchmark for Cybersecurity AI Agents
arXiv:2510.24317 · Oct 2025 · HTML edition
Modular meta-benchmark spanning Jeopardy CTFs, A&D CTFs, cyber ranges, knowledge and privacy — measuring labor-relevance, not trivia.
The World's Top AI Agent for Security CTF
arXiv:2512.02654 · Dec 2025
Five major 2025 circuits, Rank #1 repeatedly, 41/45 flags and the $50,000 Neurogrid prize. Argues Jeopardy CTFs are now a solved game and the field must move to Attack & Defense.
A Game-Theoretic AI for Guiding Attack and Defense
arXiv:2601.05887 · Jan 2026
Generative Cut-the-Rope (G-CTR) fuses Nash-equilibrium reasoning with LLM agents: success 20% → 43%, cost per success ÷2.7, behavioural variance ÷5.2, ~2:1 in Purple-team play.
Dynamic Cyber Ranges
arXiv:2604.24184 · Apr 2026
LLM-driven Defender agents cut attacker success to 0–55%; small on-premise models matched frontier defense while detecting intrusions 10× faster.
Towards CSI: What's the best harness for cybersecurity?
arXiv:2605.28334 · May 2026
No single scaffold dominates. A blackboard architecture over five heterogeneous scaffolds solves 19/33 Cybench (57.6%) vs 15/33 best-individual — 25% faster, comparable cost.

🛡️ Defense, safety and adversarial robustness

Paper What it established
Hacking the AI Hackers via Prompt Injection
arXiv:2508.21669 · Aug 2025
Turns the attack surface around: AI security tools are themselves injectable. Introduces and empirically validates a four-layer guardrail defense.
Synthetic APTs: the Collapse of TTP-Based Attribution
arXiv:2606.07158 · Jun 2026
Agents emulating five known APT groups compromised all 10 enterprise-range experiments — in 8 of them weaponising the defender's own endpoint platform for C2. TTP-based attribution does not survive this.

🤖 Applied security — robotics, OT and consumer devices

Paper What it established
The Cybersecurity of a Humanoid Robot
arXiv:2509.14096 · Sep 2025 · HTML edition
Dual-layer encryption flaws and unauthorised telemetry on a production humanoid platform.
Humanoid Robots as Attack Vectors
arXiv:2509.14139 · Sep 2025 · HTML edition
The Unitree G1 operates simultaneously as a covert surveillance node and an active cyber operations platform.
Cybersecurity AI in OT: Dragos OT CTF 2025
arXiv:2511.05119 · Nov 2025
Rank #1 at hours 7–8 of a 48-hour, 1,000-team ICS competition; 32/34 challenges; a 37% velocity advantage over the leading human crews.
Hacking Consumer Robots in the AI Era
arXiv:2603.08665 · Mar 2026 · HTML edition
Lawnmower, exoskeleton and window cleaner: 38 vulnerabilities discovered automatically in ~7 hours — work that previously took months of specialist research.

📜 Data, policy and regulation

Paper What it established
Cybersecurity AI (CAI) Dataset
arXiv:2605.28146 · May 2026
Fourteen months of trajectories from this framework: 230,935 sessions · 26,027,742 prompts · 16,768 IPs · 123 countries · 18.07 TB — the largest described corpus of LLM-driven hacker trajectories, and the empirical case for on-premise models.
Certifying Ghosts: How CAI Agents Break the EU Cyber Resilience Act
arXiv:2607.07109 · Jul 2026
Agentic discovery invalidates all four assumptions underpinning the CRA: a product that passed every check becomes exploitable with no one touching it. Static, human-paced certification ends in December 2027.

Precursor. PentestGPT (USENIX Security 2024) pioneered LLM-powered penetration testing and established the foundation this entire line was built on.


Competition record

Neurogrid Neurogrid flags Dragos Dragos challenges HTB AI HTB Spain HTB world HTB ranking HTB world ranking Cyber Apocalypse Mistral


🏢 Case studies — CAI against real systems
Domain Case study Outcome
🤖 Robotics Unitree G1 Humanoid Unauthorised telemetry to China-related servers, exposed RSA keys with world-writable permissions, and surveillance capability implicating GDPR and international privacy law.
⚙️ OT Dragos OT CTF 2025 Top-10 finish — Rank 1 during hours 7–8, 32 of 34 challenges, 37% velocity advantage over the leading human teams.
🌐 IT · Bug Bounty HackerOne Platform HackerOne engineers used CAI to explore agentic architectures; CAI's Retester agent directly inspired their production deduplication agent, now handling millions of reports.
⚙️ OT Ecoforest Heat Pumps Critical flaw enabling unauthorised remote access and potential catastrophic failure, plus exposed credentials and DES weaknesses across the European installed base.
🤖 Robotics Mobile Industrial Robots Automated ROS message injection exposing unauthorised access to robot control systems and alarm triggers.
🌐 IT · Web Mercado Libre Automated API enumeration surfacing user-data exposure risks at e-commerce scale.
⚙️ OT MQTT broker Unauthenticated topic subscription in a Dockerised OT network, with injected values corrupting Grafana dashboards.
🌐 IT · Web PortSwigger Web Security Academy Race-condition exploitation of a file-upload flaw, uploading and executing a web shell through parallel requests.

Recorded proofs of concept

CAI + alias0 — ROS injection on MiR-100 CAI + alias0 — API discovery at Mercado Libre
asciicast asciicast
CAI on JWT @ PortSwigger CTF CAI on HackableII Boot2Root CTF
asciicast asciicast

More at aliasrobotics.com/case-studies-robot-cybersecurity.php.

🧬 Foundational robot cybersecurity research — where this line began

The Cybersecurity AI line grew out of robot-security work at Alias Robotics dating back to 2018:

SROS2 (IROS 2022) · Robot Cybersecurity, a review · Robot Teardown · Cybersecurity in Robotics · Securing robots in OT environments · alurity · Red teaming ROS in industry · DevSecOps in Robotics · Akerbeltz, industrial robot ransomware (IEEE IRC 2020) · Robot Vulnerability Database · Aztarna · Robot Hazards · Robot Security Framework · Robotics CTF · Robot Vulnerability Scoring System

Venues — USENIX Security · Black Hat USA 2021 · Black Hat Europe 2021 · RootedCON · IROS · ICRA · ROSCon · GameSec · IEEE IRC · Humanoids

Responsible disclosure — 30+ CVE IDs issued as a CVE Numbering Authority since February 2020 · 2 co-authored CISA ICS advisories · 100+ robot vulnerabilities disclosed · 90-day disclosure window

🎓 CAI Fluency — the educational programme

Free and still available. Formalised in arXiv:2508.13588.

Description 🇬🇧 🇪🇸
Ep. 0 — What is CAI? Cybersecurity AI explained
Ep. 1 — The framework Vision and ethical principles behind the project
Ep. 2 — Zero to Cyber Hero Breaking into cybersecurity with AI, for complete beginners
Ep. 3 — Vibe-Hacking A first hack: agents, tools, output interpretation, model comparison
Ep. 4 — Intro ReAct From basic LLMs to Chain-of-Thought, ReAct and multi-agent architectures
Ep. 5 — CTF challenges Web, crypto, reverse engineering and forensics with agents
Annex 1 — Community Meeting #1 40+ participants from academia, industry and defense
Annex 2 — CAI 0.5.x Multi-agent support, /history, /compact, /graph, /memory; OT heat-pump case study
Annex 3 — CAI 0.4.x and alias0 Streaming, MCP support, privacy-by-design model-of-models
Annex 4Jaula del N00B Framework walkthrough on the Spanish cybersecurity show

🇪🇺 Funding

One funded research project made all of the above possible.

CAI and the Cybersecurity AI research line were co-funded by the European Innovation Council (EIC) Accelerator, project RIS — Revolutionising Cybersecurity with the Next Generation AI-Powered Security Platform, under the HORIZON-EIC-2023-ACCELERATOR-01 call of Horizon Europe.

Project RIS — Revolutionising Cybersecurity with the Next Generation AI-Powered Security Platform
Grant agreement 101161136
Programme Horizon Europe · European Innovation Council (EIC) Accelerator
Call HORIZON-EIC-2023-ACCELERATOR-01
Duration 1 July 2024 → 30 June 2027
Coordinator ALIAS ROBOTICS S.L. (Spain)
EU contribution €2,499,875

RIS set out to build an artificial immune system for industrial robots — bio-inspired AI that detects anomalies and delivers integrated protection for robots working alongside humans. Pursuing that goal honestly required first understanding what AI-powered offense can do, because a defensive system can only be designed against an adversary whose real capability is known.

CAI is what that question produced: an open framework used to measure, publicly and reproducibly, exactly how far agentic AI can go in security. The 18 papers above are the public research output of that work, and its defensive conclusions — GenAI-native defender agents, dynamic cyber ranges, guardrail frameworks, sovereign on-premise models — now live on in CSI.

Funded by the European Union. Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Innovation Council. Neither the European Union nor the granting authority can be held responsible for them.


🧰 Using the archive

Caution

Unmaintained software. No fixes, no patches, no support. Use only in isolated environments against systems you are authorised to test. For anything else, use CSI.

pip install cai-framework installs the open-source line, frozen at 0.5.10 (public PyPI, December 2025). The v1.1.5 tree in this repository was the professional distribution for paid customers and was never published to public PyPI — if you hold a CAI subscription, see the note at the top before relying on it.

Setting CAI_LICENSE_OFF=1 bypasses the startup license check and targets the public PyPI package. alias models are not available this way — configure any other supported provider (OpenAI, Anthropic, DeepSeek, Ollama, …) via CAI_MODEL and the matching API key.

# Python 3.12 recommended; always use a fresh virtual environment
python3.12 -m venv cai_env
source cai_env/bin/activate
pip install cai-framework

# minimal .env — fill in the provider key you intend to use
echo -e 'OPENAI_API_KEY="sk-1234"\nANTHROPIC_API_KEY=""\nOLLAMA=""\nPROMPT_TOOLKIT_NO_CPR=1\nCAI_STREAM=false' > .env

# run without an Alias Robotics license
export CAI_LICENSE_OFF=1
cai   # the first launch can take up to 30 seconds
📚 Reference material, preserved as of the final release
Topic Where
Installation — OS X, Ubuntu 20.04/24.04, Windows WSL, Android docs/cai_installation.md
Quickstart and the REPL docs/cai_quickstart.md · docs/quickstart.md
Architecture — agents, tools, handoffs, patterns, tracing, HITL docs/cai_architecture.md · docs/agents.md · docs/handoffs.md · docs/multi_agent.md
Guardrails and prompt-injection defenses docs/guardrails.md · docs/cai_prompt_injection.md
Environment variables docs/environment_variables.md
Models and providers — 300+ via LiteLLM docs/models.md · docs/cai_list_of_models.md · docs/providers/
MCP integration docs/mcp.md
Benchmarking and results docs/cai_benchmark.md · docs/benchmarking/ · docs/results.md
Research index docs/research.md
FAQ docs/cai_faq.md
Examples examples/

Forks are welcome — the terms in LICENSE continue to apply — but no pull requests or issues will be reviewed in this repository.


📖 Citation

If you use CAI, its datasets or its benchmarks in your research, please cite the framework paper. Machine-readable metadata lives in CITATION.cff.

@article{mayoral2025cai,
  title={CAI: An Open, Bug Bounty-Ready Cybersecurity AI},
  author={Mayoral-Vilches, V{\'\i}ctor and Navarrete-Lozano, Luis Javier and Sanz-G{\'o}mez, Mar{\'\i}a and Espejo, Lidia Salas and Crespo-{\'A}lvarez, Marti{\~n}o and Oca-Gonzalez, Francisco and Balassone, Francesco and Glera-Pic{\'o}n, Alfonso and Ayucar-Carbajo, Unai and Ruiz-Alcalde, Jon Ander and Rass, Stefan and Pinzger, Martin and Gil-Uriarte, Endika},
  journal={arXiv preprint arXiv:2504.06017},
  year={2025}
}
Full Cybersecurity AI bibliography — 18 entries, by publication date
@article{mayoral2025cai,
  title={CAI: An Open, Bug Bounty-Ready Cybersecurity AI},
  author={Mayoral-Vilches, V{\'\i}ctor and Navarrete-Lozano, Luis Javier and Sanz-G{\'o}mez, Mar{\'\i}a and Espejo, Lidia Salas and Crespo-{\'A}lvarez, Marti{\~n}o and Oca-Gonzalez, Francisco and Balassone, Francesco and Glera-Pic{\'o}n, Alfonso and Ayucar-Carbajo, Unai and Ruiz-Alcalde, Jon Ander and Rass, Stefan and Pinzger, Martin and Gil-Uriarte, Endika},
  journal={arXiv preprint arXiv:2504.06017},
  year={2025}
}

@article{mayoral2025automation,
  title={Cybersecurity AI: The Dangerous Gap Between Automation and Autonomy},
  author={Mayoral-Vilches, V{\'\i}ctor},
  journal={arXiv preprint arXiv:2506.23592},
  year={2025}
}

@article{mayoral2025fluency,
  title={CAI Fluency: A Framework for Cybersecurity AI Fluency},
  author={Mayoral-Vilches, V{\'\i}ctor and Wachter, Jasmin and Chavez, Crist{\'o}bal RJ and Schachner, Cathrin and Navarrete-Lozano, Luis Javier and Sanz-G{\'o}mez, Mar{\'\i}a},
  journal={arXiv preprint arXiv:2508.13588},
  year={2025}
}

@article{mayoral2025hacking,
  title={Cybersecurity AI: Hacking the AI Hackers via Prompt Injection},
  author={Mayoral-Vilches, V{\'\i}ctor and Rynning, Per Mannermaa},
  journal={arXiv preprint arXiv:2508.21669},
  year={2025}
}

@article{mayoral2025humanoidsecurity,
  title={The Cybersecurity of a Humanoid Robot},
  author={Mayoral-Vilches, V{\'\i}ctor},
  journal={arXiv preprint arXiv:2509.14096},
  year={2025}
}

@article{mayoral2025humanoid,
  title={Cybersecurity AI: Humanoid Robots as Attack Vectors},
  author={Mayoral-Vilches, V{\'\i}ctor},
  journal={arXiv preprint arXiv:2509.14139},
  year={2025}
}

@article{balassone2025evaluation,
  title={Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs},
  author={Balassone, Francesco and Mayoral-Vilches, V{\'\i}ctor and Rass, Stefan and Pinzger, Martin and Perrone, Gaetano and Romano, Simon Pietro and Schartner, Peter},
  journal={arXiv preprint arXiv:2510.17521},
  year={2025}
}

@article{mayoral2025caibench,
  title={CAIBench: A Meta-Benchmark for Evaluating Cybersecurity AI Agents},
  author={Mayoral-Vilches, V{\'\i}ctor and Balassone, Francesco and Navarrete-Lozano, Luis Javier and Sanz-G{\'o}mez, Mar{\'\i}a and Crespo-{\'A}lvarez, Marti{\~n}o and Rass, Stefan and Pinzger, Martin},
  journal={arXiv preprint arXiv:2510.24317},
  year={2025}
}

@article{mayoral2025dragos,
  title={Cybersecurity AI in OT: Insights from an AI Top-10 Ranker in the Dragos OT CTF 2025},
  author={Mayoral-Vilches, V{\'\i}ctor and Navarrete-Lozano, Luis Javier and Balassone, Francesco and Sanz-G{\'o}mez, Mar{\'\i}a and Veas-Ch{\'a}vez, Crist{\'o}bal Ricardo and del Mundo de Torres, Maite},
  journal={arXiv preprint arXiv:2511.05119},
  year={2025}
}

@article{mayoral2025topctf,
  title={Cybersecurity AI: The World's Top AI Agent for Security Capture-the-Flag (CTF)},
  author={Mayoral-Vilches, V{\'\i}ctor and Navarrete-Lozano, Luis Javier and Balassone, Francesco and Sanz-G{\'o}mez, Mar{\'\i}a and Veas-Chavez, Crist{\'o}bal R. J. and del Mundo de Torres, Maite and Turiel, Vanesa},
  journal={arXiv preprint arXiv:2512.02654},
  year={2025}
}

@article{mayoral2026gctr,
  title={Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense},
  author={Mayoral-Vilches, V{\'\i}ctor and Sanz-G{\'o}mez, Mar{\'\i}a and Balassone, Francesco and Rass, Stefan and Salas-Espejo, Lidia and Jablonski, Benjamin and Navarrete-Lozano, Luis Javier and del Mundo de Torres, Maite and Veas-Chavez, Crist{\'o}bal R. J.},
  journal={arXiv preprint arXiv:2601.05887},
  year={2026}
}

@article{mayoral2026superintelligence,
  title={Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI},
  author={Mayoral-Vilches, V{\'\i}ctor and Rass, Stefan and Pinzger, Martin and Gil-Uriarte, Endika and Ayucar-Carbajo, Unai and Ruiz-Alcalde, Jon Ander and del Mundo de Torres, Maite and Sanz-G{\'o}mez, Mar{\'\i}a and Balassone, Francesco and Veas-Chavez, Crist{\'o}bal R. J. and Turiel, Vanesa and Glera-Pic{\'o}n, Alfonso and S{\'a}nchez-Prieto, Daniel and Salvatierra, Yuri and Zabalegui-Landa, Paul and Cabrera-{\'A}lvarez, Ruffino Reydel and Mayoral-Pizarroso, Patxi},
  journal={arXiv preprint arXiv:2601.14614},
  year={2026}
}

@article{mayoral2026consumerrobots,
  title={Cybersecurity AI: Hacking Consumer Robots in the AI Era},
  author={Mayoral-Vilches, V{\'\i}ctor and Ayucar-Carbajo, Unai and Laflamme, Olivier and Peng, Ruikai and Sanz-G{\'o}mez, Mar{\'\i}a and Balassone, Francesco and Apa, Lucas and Gil-Uriarte, Endika},
  journal={arXiv preprint arXiv:2603.08665},
  year={2026}
}

@article{mayoral2026cyberranges,
  title={Dynamic Cyber Ranges},
  author={Mayoral-Vilches, V{\'\i}ctor and Sanz-G{\'o}mez, Mar{\'\i}a and Balassone, Francesco and del Mundo de Torres, Maite and Nicolaou, George and Rodriguez Borines, Samuel and Graziano, Almerindo and Zabalegui, Paul and Gil-Uriarte, Endika},
  journal={arXiv preprint arXiv:2604.24184},
  year={2026}
}

@article{mayoral2026caidataset,
  title={Cybersecurity AI (CAI) Dataset},
  author={Mayoral-Vilches, V{\'\i}ctor},
  journal={arXiv preprint arXiv:2605.28146},
  year={2026}
}

@article{mayoral2026csiharness,
  title={Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?},
  author={Mayoral-Vilches, V{\'\i}ctor and Balassone, Francesco and Sanz-G{\'o}mez, Mar{\'\i}a and Zabalegui-Landa, Paul and S{\'a}nchez-Prieto, Daniel and Oteiza-{\'A}lvarez, Marina and Quarta, Davide and Pinzger, Martin},
  journal={arXiv preprint arXiv:2605.28334},
  year={2026}
}

@article{balassone2026syntheticapts,
  title={Synthetic APTs: the Collapse of TTP-Based Attribution},
  author={Balassone, Francesco and Mayoral-Vilches, V{\'\i}ctor and Sanz-G{\'o}mez, Mar{\'\i}a and Zabalegui-Landa, Paul and Rass, Stefan and Quarta, Davide and S{\'a}nchez-Prieto, Daniel and Oteiza-{\'A}lvarez, Marina and Graziano, Almerindo and Kim, Lauren Min and Choi, MinSeok},
  journal={arXiv preprint arXiv:2606.07158},
  year={2026}
}

@article{mayoral2026certifyingghosts,
  title={Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act},
  author={Mayoral-Vilches, V{\'\i}ctor},
  journal={arXiv preprint arXiv:2607.07109},
  year={2026}
}

🙏 Acknowledgements

To the European Union. CAI was developed by Alias Robotics and co-funded through the EIC Accelerator project RIS (GA 101161136), HORIZON-EIC-2023-ACCELERATOR-01 call. That funding is the reason this research exists and is public.

To every contributor who filed one of the 275 issues, opened one of the 195 pull requests, ported CAI to new platforms, wrote agents, broke things in creative ways and told us about it: thank you. Much of what makes CSI robust today was learned from your reports. Those 275 issues and 195 pull requests stay public here permanently, and the work of all 103 authors is present in every file of this snapshot.

To our academic collaborators at partner institutions worldwide, for co-authorship, rigour, and for pushing benchmark design, curricula and defense mechanisms further than a company could alone. Research collaboration on Cybersecurity AI continues at research@aliasrobotics.com — PhD projects, benchmarking studies with CAIBench, security education initiatives and dataset access.

To the open source we built on — agentic principles inspired by OpenAI's swarm and openai-agents-python; model routing from LiteLLM; tracing and observability from phoenix.

To PentestGPT and the USENIX Security community, where this line of research began.


⚖️ License and disclaimer

This project combines MIT-licensed components — derived from openai/openai-agents-python, under src/cai/agents — with proprietary additions licensed for research purposes only. See LICENSE · LICENSE-MIT · DISCLAIMER.

Warning

Access to this library, and use of the information and materials herein, is not intended, and is prohibited, where such access or use violates applicable laws or regulations. The authors do not encourage or promote unauthorised tampering with running systems; doing so can cause serious human harm and material damage.

Pentest for good instead. By downloading, using or modifying this source code, you agree to the terms of the LICENSE and the limitations set out in the DISCLAIMER.


The framework is archived. The research is not.

Research CSI Alias Robotics

Made in the Basque Country 🇪🇺 · research@aliasrobotics.com · support@aliasrobotics.com

About

Cybersecurity AI (CAI), the framework for AI Security

Topics

Resources

Stars

9.8k stars

Watchers

101 watching

Forks

Releases

Packages

Used by

Contributors

Languages