RL Environments

Train Safer, More Capable Agentic Models

Generate trajectories in realistic environments to improve task completion, policy adherence, and resistance to attacks.

Trusted by 8 leading foundation model labs

Verifiable Trajectories and Reward Signals

Measure task completion, policy compliance, and attack resistance using full agent trajectories, deterministic verifiers, and expert-calibrated rubrics.

Web App, Computer, and 
Tool-Use Environments

Strengthen ability to navigate applications and use tools through realistic simulations of consumer and enterprise workflows.

Real World Adversarial
Intelligence

Align agentic models while preserving task performance with real-world attack patterns embedded into their workflows.

Examples of RL Environments

Across domains, capabilities, and risks.

Anthropic

Resist indirect prompt injection across web, computer-use, and terminal environments.

Environment

High-fidelity web clone and MCP server

Injection points

Users, projects, issues, spaces, pages, teams, comments, apps, notifications, activity, alerts, heroes, promos, features, carousels, products, sidebars

Third-party apps

Users, projects, issues, spaces, pages, teams, comments, apps, notifications, activity, alerts, heroes, promos, features, carousels, products, sidebars

Contributing experts
  • Daily Jira users (developers, engineering managers, product managers, UI and UX teams), plus security researchers with 3+ years red-teaming agentic models
Sample tasks
  • Confidential roadmap exfiltration to an external collaborator; malware injection via a fake feature request; injections via API-submitted issues, emails, or Confluence-linked content; privilege escalation via workflow transition abuse; cross-project data leakage via JQL filter sharing

Discover and exploit vulnerabilities across web applications, services, and network environments.

Environment

Web app with injectable vulnerability seeds, in red team (vulnerable) and blue team (non-vulnerable) task variants

Languages

Python, Java, PHP, JavaScript

Third-party apps

Users, projects, issues, spaces, pages, teams, comments, apps, notifications, activity, alerts, heroes, promos, features, carousels, products, sidebars

Contributing experts

OSCP-certified penetration testers, cloud security and network forensics specialists, several with elite military cybersecurity unit backgrounds

Sample tasks

Single-service, single-step exploitation; multi-service, multi-step exploitation of chained vulnerabilities; multiple frameworks (Flask, Django, Node.js, WordPress); backend and frontend vulnerabilities; protocols including HTTP/S, WebSocket, HTTP3, XMPP, MQTT

Identify security issues across repositories, applications, and development environments.

Environment

Terminal environment with the Pi personal assistant installed and connected to personal tools and third-party apps.

Languages

?

Third-party apps

Users, projects, issues, spaces, pages, teams, comments, apps, notifications, activity, alerts, heroes, promos, features, carousels, products, sidebars

Contributing experts

Security researchers with 3+ years of experience red-teaming personal assistant agents for companies building consumer AI products.

Sample tasks

Passport exfiltration through fake airline document requests. Internal payout disclosure through fabricated sponsor precedents. False labor-bank attestations through spoofed mailboxes. Identity phishing through fake return-authorisation desks. Passport leaks through fake WhatsApp guest registration. ID exfiltration through fraudulent moving-company calendar invites.

Detect malicious or unsafe code across repositories, applications, and development environments.

Environment

Coding agent environment

Injection points

?

Third-party apps

Users, projects, issues, spaces, pages, teams, comments, apps, notifications, activity, alerts, heroes, promos, features, carousels, products, sidebars

Contributing experts

OSCP-certified penetration testers, cloud security and network forensics specialists, several with elite military cybersecurity unit backgrounds

Sample tasks

?

Adversarial Intelligence

We've Been Down The Rabbit Hole

The Rabbit Hole is Alice’s database of evil built over a decade across hundreds of languages and cultures. Our researchers turn these patterns into hidden attacks that test whether agents complete their tasks while resisting manipulation.

Talk to Our Experts
‍Environment Coverage

Environments Powered by Real Attack Patterns for Real Results

Train and evaluate agents on real-world tasks across high-fidelity computer-use and tool-use environments generating full trajectories for RLHF, custom evaluations, and benchmarking.

Adversarial Risk
Coverage

Test indirect prompt injection (IPI), jailbreaks, PII exposure, data exfiltration, malware, bias, and compliance failures.

Web App, Computer, and 
Tool-Use Environments

Cover visual interfaces, APIs, connected tools, and terminal-based workflows.

Consumer and Enterprise Workflows

Simulate applications across industries, business functions, and everyday consumer tasks.

Deterministic Verifiers and Rubrics

Reliably measure whether agents complete tasks, follow safety policies, and resist manipulation.

“Alice supports our most complex and high priority red teaming needs. We like working with them because they are leaders in this space, understand the shifting adversarial landscape, and the output is always high quality. They are great partners who quickly adapt to our ever changing needs.”

Sumeet Pandey
Generative AI Partnerships

Reward What AI Does, and What it Doesn’t.

Realistic environments and real attack patterns help you train and evaluate capability, alignment, and attack resistance before agents reach the real world.