arXiv:2608.23181nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

CyberFactory: Scaling Cyber Security Capabilities with Instances from the WildExplained for Beginners

Jian Yang, Haau-Sing Li, Shawn Guo +8 more

Cryptography and SecurityComputation and Language

Abstract

As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention, with closed-source LLMs (e.g., Mythos) delivering advanced cybersecurity capabilities. However, existing open-source efforts remain limited: frontier open-weight models do not provide reproducible cybersecurity training solutions, open-source training solutions focus on isolated tasks and lack scalable agentic data, and scaling agentic rollouts requires strong domain priors. In this work, we introduce CyberFactory, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA). CyberFactory transforms public vulnerability artifacts, including CVEs from the wild, into executable and verifiable task instances. It further uses a reusable vulnerability-analysis skill to guide the teacher through source inspection, problem solving with domain prior, and evidence-based validation. The resulting supervision is agentic: the model interacts with tools and target environments and revises its solutions according to execution feedback. Using these trajectories, we train and release \modelname\emph{Aegis is, in Greek mythology, the protective shield of Zeus and Athena; the name reflects the model's defensive, security-oriented purpose.}, which internalizes the skill-guided procedure without requiring the skill at inference time. On CyberGym, \modelname reaches 52.4% Pass@1 under a one-hour budget, improving over its Qwen~3.5 base model by +22.8 points and outperforming the evaluated general-purpose backbones under the same scaffold.

The Problem

Imagine you are a developer or a security analyst trying to answer a straightforward question: Does this piece of code have a security flaw, and if so, how do I fix it? In the world of large language models (LLMs), we have become quite good at generating code, but applying that skill to cybersecurity is a different beast entirely.

The research gap here is real and multifaceted. First, the best-performing cybersecurity models are often closed-source (like Mythos), meaning researchers cannot inspect, modify, or build upon them. Second, existing open-source models are "one-trick ponies": they might be good at finding vulnerabilities, but they aren't trained to do both detection and patching, let alone answer cybersecurity questions. They lack the massive, coordinated datasets needed to teach a model how to behave like a security analyst. Finally, scaling up training for these tasks is incredibly difficult. It usually requires deep, domain-specific knowledge—strong "priors"—to guide the model; without it, the model wastes its computation budget on dead ends.

The paper asks: How can we take the raw, messy vulnerability data from the real world (CVEs, or Common Vulnerabilities and Exposures) and turn it into a reproducible, scalable training pipeline for an open-source model?

How It Works (The Technical Mechanics)

CyberFactory is the framework introduced in this paper, and its core mission is to turn "CVEs from the wild" into executable training examples. It operates in three stages: data construction, trajectory synthesis, and model training.

Here is the concrete flow, explained through an analogy a software engineer would appreciate.

  • The Analogy: The Automated Bug Squasher Think of CyberFactory as a highly sophisticated Continuous Integration (CI) pipeline for security. Normally, a developer writes code and runs tests. CyberFactory takes a real-world vulnerability report (a CVE) and automatically builds a "test environment" (a Docker image) containing the vulnerable code. It then sets up a verification loop: the model proposes an input (like a user sending a specific malicious packet), the system runs it, and it checks if the program crashes (the "sanitizer failure"). If it crashes on the old version but not the new patched version, the test passes.

  • Stage 1: Data Construction (Turning CVEs into Exams) The authors start with real vulnerability artifacts. A CVE typically says, "Hey, this software version is vulnerable."

    1. Source Adaptation: The framework finds the specific version of the software mentioned in the CVE.
    2. Instance Verification: It builds the "before" and "after" states (pre-patch and post-patch Docker images). Crucially, it filters out CVEs that are already too easy—if a model can just guess the fix, it’s not a good training example. It keeps only the challenging ones.
    3. Description Generation: It generates a natural language description of the bug for the model to read.
  • Stage 2: Trajectory Synthesis (The "Teacher" and the Skill) This is the meat of the paper. The model needs to learn how to solve these problems, not just what the answer is. The authors introduce a "reusable vulnerability-analysis skill."

    • The Skill: Think of this as a debugging checklist the model follows. It tells the model to: (1) inspect the code, (2) use tools to test inputs, (3) check if the tool output shows a crash, and (4) if no crash, revise the input and try again.
    • Agentic Loop: The model isn't just generating text; it's interacting with a virtual machine. It proposes an input, executes it, reads the error message, and revises its approach. This creates a "trajectory"—a history of the model's thought process and actions.
  • Stage 3: Model Training (Internalizing the Skill) The resulting trajectories—these sequences of "think, act, observe, revise"—are used to Supervised Fine-Tune (SFT) the OpenAegis model. The magic here is that the model absorbs the procedure. At inference time, you don't need to give it the "skill checklist"; the weights of the model have learned how to explore and validate on their own.

Key Results & Benchmarks

The results are striking. The paper evaluates the trained model, OpenAegis, on a benchmark called CyberGym. This benchmark tasks the model with reproducing real vulnerabilities within a one-hour time limit.

  • The Big Number: OpenAegis achieves a 58.1% Pass@1 score. In plain language: out of 100 real-world vulnerability tasks, the model successfully finds a working Proof-of-Concept (PoC) for more than half of them on the first attempt, within an hour.
  • The Gains: This is a massive +28.5 point improvement over its base model (Qwen 3.5). To put that in perspective, the base model was stuck at roughly 29.6%.
  • The Competition: OpenAegis doesn't just beat its base model; it beats much larger, general-purpose models. It outperforms GLM 5.2 by 14.8 points and Kimi K2.7 by 6.4 points, despite using fewer parameters.

Furthermore, the paper highlights the "skill" aspect. When the authors gave the vulnerability-analysis skill to a general model (GLM 5.2) at inference time, that model's score jumped from 43.3% to 46.5%. More importantly, the training trajectories increased the "throughput" of successful data generation, meaning they could train better models faster.

Why It Matters (Key Takeaways)

This work matters because it lowers the barrier to entry for open-source cybersecurity AI. Here are the key takeaways:

  • From "Magic" to "Method": We now have a reproducible recipe. Instead of hoping a closed-source model gets better, researchers can use CyberFactory to train their own models on specific security tasks. It connects the dots between data, training, and evaluation in a way that was missing.
  • The Power of the "Skill": The paper demonstrates that giving a model a structured procedure (a "skill") during training is more effective than just letting it loose. The model internalizes this workflow, meaning it becomes a more disciplined and effective security analyst without needing external prompts at runtime.
  • Efficiency Gains: The skill-guided approach is surprisingly efficient. By guiding the model with a domain prior, the authors were able to get higher scores even when reducing the time budget per task (from 60 minutes to 15 minutes with the skill). This means we can generate more training data in less time.
  • What to Watch For: The evaluation is currently bounded by the available CVE artifacts. Not every software project has a public CVE, and some complex vulnerabilities might still require human intuition. However, the framework is designed to be extensible, and future work will likely broaden the scope of covered vulnerabilities.

Final Thought

CyberFactory represents a shift from building "smarter models" to building "better training pipelines." By treating cybersecurity tasks like reproducible engineering problems—with differential oracles (crash tests) and structured skills—we can train open-source models that rival, and even surpass, much larger closed competitors. It’s a significant step toward making advanced cybersecurity capabilities transparent, auditable, and accessible.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →