RANGEFORGE /SECURITY RL
Contact

RL Environments for Security

Train and validate LLMs and AI agents on offensive and defensive security tasks — inside high-fidelity, procedurally-varied enterprise environments built on real cloud infrastructure, real endpoints, and real security tooling.

01RL Environments core

Full enterprise environments — networks, cloud, endpoints, employees, and the security stack that watches them — generated fresh for every run. Agents act inside them; rewards come from the environment itself, not a judge model.

01

Procedurally varied

Every run stands up a new enterprise: different topology, credentials, and blind spots. Nothing to memorize, nothing to overfit — the environment is the curriculum.

02

Offense and defense

Attack chains from initial foothold to exfiltration, and the defensive side of the same coin: detection, triage, and response. Train and test both sides of the range.

03

Verifiable rewards

Every boundary is enforced for real, and every outcome is scored against the environment's ground truth — never the agent's self-report.

02Benchmarks & Evals

Define benchmarks on top of any environment, then evaluate any model or agent against them — frontier or open, raw LLM or full agent stack — with scored, reproducible, comparable results.

01

Define

Turn scenarios into benchmarks: objectives, scoring rules, and partial credit for each step of a task — from a single skill to a multi-stage operation.

02

Evaluate

Run any LLM or agent through the same benchmark and compare them side by side — capability, reliability, and cost on the tasks that actually matter.

03

Reproduce

Environments reset to a clean state between runs, so results are reproducible and comparable across models, versions, and time.

03Who it's for

One foundation — realistic security environments — serving three audiences.

01

Frontier AI labs

Evaluate model security capabilities in realistic enterprise scenarios, and generate the post-training data that improves them.

02

Security vendors

Training and evaluation infrastructure for agentic products — AI SOC, pentesting, cloud security — before they meet a customer's environment.

03

Enterprises

Independently validate internal agents and vendor security claims against your own kind of environment — not a vendor's demo.

04Writing

Notes from building security environments and evaluating agents inside them.

Building or evaluating security agents?

Tell us what you're training, testing, or trying to trust — we'll set up a range for it.

Get in touch hello@rangeforge.ai