Skip to main content
v2026.11,858 entries · CC-BY 4.0

Gray Swan Arena: Crowdsourced Red-Teaming as Pre-Deployment Evidence

Gray Swan’s Arena crowdsources adversarial testing to 15,000+ red-teamers, and two competitions with UK AISI and US CAISI produced public papers. What those numbers support as pre-deployment evidence — and what they can’t.

Written and maintained by CASRAI Editorial Board

Last updated

When a frontier lab’s system card says its model was “red-teamed,” there is a growing chance that a large part of that work was done by a crowd of strangers competing for prize money. Gray Swan AI runs the best-documented version of this: the Arena, a public adversarial red-teaming platform that Gray Swan describes as “the largest AI red-teaming network in the world,” with “15,000+ adversarial researchers, and growing.” Two of its flagship competitions have now produced peer-reviewable papers co-authored with the UK AI Security Institute and, in the more recent case, analysed with the US Center for AI Standards and Innovation. That makes the Arena one of the few places where an outsider can actually check what a red-teaming claim is worth — and one of the clearest illustrations of why “we red-teamed it” is not, by itself, a safety result.

  • What it is: a standing, public, incentivised adversarial-testing platform run by a commercial vendor, not a regulator, an academic consortium, or an accredited auditor.
  • The scale is real. Gray Swan’s 2025 Agent Red-Teaming Challenge with UK AISI ran 8 March to 6 April 2025 and logged roughly 1.8 million attack attempts against 22 models across 44 target behaviours, with about 62,000 successful policy violations and $171,800 paid out in prizes.
  • The 2026 follow-up was tighter and worse-looking. The Indirect Prompt Injection Arena, designed with UK AISI, US CAISI and frontier labs, drew 464 participants and roughly 272,000 attempts against 13 frontier models across 41 scenarios — and at least one successful hijack was found against every single model tested.
  • Attack success rates varied by more than an order of magnitude between models in the same competition — but the denominators are not comparable, because attackers choose where to spend effort.
  • Universal attacks are the headline finding. Single attack strategies transferred across 21 of the 41 behaviours and across model families, and transfer ran mostly downhill: attacks found against more robust models worked on less robust ones, rarely the reverse.
  • What it is not: an audit, a certification, an independent evaluation, or a random sample of real-world adversaries. The lab pays, the lab scopes, and the lab decides what to publish.

Who runs the Arena

Gray Swan AI is a Pittsburgh-based AI security company. Its own About page states it was “Founded by world-leading AI safety and security researchers from Carnegie Mellon University,” and names Matt Fredrikson as co-founder and CEO and J. Zico Kolter as co-founder and Chief Scientist — the same research lineage behind the widely cited work on automated adversarial attacks against aligned language models. The company sells three distinct things, and conflating them is the single most common error in reading a system-card citation:

  • Arena — the crowdsourced human red-teaming network described on this page. Gray Swan’s model-builder page describes the commercial offering as “Quarterly flagship challenges plus custom challenges scoped to your priority risk areas.”
  • Shade — an automated attacker, marketed as a “Custom LLM attacker trained on diverse Arena attack strategies.” This is a machine, not a crowd.
  • Cygnal — a runtime guardrail product that sits in front of a deployed model. This is a control, not an evaluation.

A third tier sits between Arena and a traditional consultancy: Gray Swan also sells private engagements staffed by “Arena’s top performers: hand-selected subject-matter experts.” So a lab that says it “worked with Gray Swan” may mean an open public competition, a closed invitational, an automated attack run, or a guardrail deployment. Those produce evidence of very different strength, and a reader of a safety report generally cannot tell which one happened unless the report says so.

What the 2025 Agent Red-Teaming Challenge actually found

The first competition to produce a public paper ran from 8 March to 6 April 2025 in partnership with the UK AI Security Institute. Gray Swan’s own results snapshot reports roughly 1.8 million attempts, about 62,000 successful breaks found by the community, 22 LLMs tested, 44 specific harmful agent behaviours rolled out in four waves, $171,800 in total prizes, and 161 red-teamers paid. The behaviour categories were confidentiality breaches, conflicting objectives, instruction-hierarchy violations (both informational and action-based), and over-refusals. The top individual, competing as “zardav,” logged 924 unique breaks.

The per-model attack success rates Gray Swan published from that challenge ranged from 1.47% for Claude 3.7 Sonnet (Thinking), the most robust model in the field, to 6.49% for Llama-3.3-70b, the least robust, with GPT-4o at 2.41%. Those are small-looking numbers attached to an enormous denominator, and the paper is blunt about what they mean in practice.

The write-up, Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition (arXiv:2507.20526), was authored by a Gray Swan team together with UK AISI researchers including Xander Davies, Robert Kirk and Yarin Gal, plus Dan Hendrycks, Zico Kolter and Matt Fredrikson. Its central finding is not the per-model ranking but the frequency: “Nearly all agents exhibit policy violations for most behaviors within 10-100 queries, with high attack transferability across models and tasks.” The failure modes it demonstrated included unauthorised data access, illicit financial actions, and regulatory non-compliance. The authors released the Agent Red Teaming (ART) benchmark — a curated set of high-impact attacks drawn from the competition, plus an evaluation harness — and reported benchmarking 19 state-of-the-art models against it.

That last move matters more than the competition itself. A one-off crowd event produces a number that nobody else can reproduce. A released benchmark distilled from that event produces an artifact other people can run. This is the difference between marketing and evidence, and it is the right test to apply to any crowdsourced red-teaming claim.

The 2026 Indirect Prompt Injection Arena: no model survived

The second and more consequential competition targeted agent hijacking — indirect prompt injection, where hostile instructions are smuggled into content an agent reads, such as an email, a web page, or a code repository, rather than typed by the user. The resulting paper, How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition (arXiv:2603.15714, submitted 16 March 2026), carries an author list that spans Gray Swan, UK AISI, OpenAI, Meta and academic groups — including Lama Ahmad, Kamalika Chaudhuri, Ivan Evtimov, Xander Davies, Javier Rando, Benjamin L. Edelman and Xiangyu Qi, alongside Fredrikson and Kolter.

Its numbers: 464 participants submitted roughly 272,000 attack attempts, producing 8,648 successful attacks across 41 scenarios against 13 frontier models. Per-model vulnerability spanned more than an order of magnitude — Claude Opus 4.5 at about a 0.5% success rate at one end, Gemini 2.5 Pro at about 8.5% at the other.

The US Center for AI Standards and Innovation published its own analysis of the same data on 23 March 2026, on the CAISI research blog. CAISI’s summary uses rounded figures — “more than 250,000 attack attempts from over 400 participants” — against the paper’s 272,000 and 464, so a reader comparing the two sources should expect the blog to understate slightly rather than assume a discrepancy. CAISI’s four stated findings are worth reading as governance claims, not just security ones:

  • “At least one successful attack was found against all of the target frontier models.” No model was immune.
  • Models “differed sharply in how many successful attacks were found” — robustness is real and unevenly distributed.
  • Certain families of “universal” attacks “were often able to transfer across scenarios and models,” implying shared weaknesses in how instruction hierarchies are learned rather than model-specific bugs. The paper puts a number on it: universal strategies transferred across 21 of the 41 behaviours.
  • Transfer was directional. Attacks that worked on more robust models generally worked on less robust ones; the reverse rarely held. That means a lab cannot infer its own robustness from the fact that a competitor’s published attacks fail against its model.

The paper also documents a property that should worry anyone relying on post-hoc monitoring: successful attacks frequently concealed evidence of the compromise in the agent’s final response to the user. A hijacked agent that reports success is not a detectable failure through the user-facing channel alone.

On disclosure, the competition organisers released an open-source version of the competition environment and published 95 successful attacks against an open-weight Qwen target that did not transfer to closed-source models, while routing model-specific findings to the respective labs and the full dataset to UK AISI and US CAISI. That is a deliberate, defensible split — reproducible artifacts for the public, live attack strings for the defenders and the institutes — and it is a far more mature disclosure posture than most vendor red-teaming reports manage. It also means the public cannot independently verify the closed-model results. Both things are true at once.

What a crowd result can and cannot support

The Arena’s own marketing is careful about the artifact and loose about the inference. Gray Swan’s adversarial-evaluation page promises “citable deliverables: reproducible transcripts, severity judgments, and raw attempt data,” findings “rigorous enough to cite in your model card,” and “methodology documentation that satisfies regulators and auditors.” Transcripts and raw attempt data genuinely are the right deliverables. The problem is what gets inferred from them downstream.

The denominators are not comparable across models

Attack success rate looks like a measurement and behaves like an artefact of attacker attention. In an incentivised competition, red-teamers allocate effort strategically: toward models where breaks come easily, toward behaviours with unclaimed prizes, toward whatever the leaderboard rewards this week. A model with a low published ASR may be genuinely robust, or it may simply have been abandoned early by attackers who found softer targets. Nothing in a headline ASR distinguishes the two. Cross-model ASR comparisons from a single competition should be read as a rough ordering with wide error bars, not as a calibrated scale — the same problem CASRAI documents in Likelihood Term: Three Frameworks, No Shared Probability Scale.

“No successful attack” is a bound, not a property

Reporting on Anthropic’s Claude Opus 4.5 system card illustrates the shape of the claim and its limits. VentureBeat’s coverage of that system card describes Gray Swan’s Shade automated attacker — not the human Arena crowd — running adaptive campaigns in which Opus 4.5 in coding environments reached roughly 4.7% attack success at one attempt, 33.6% at ten attempts and 63.0% at one hundred attempts, while a computer-use configuration with extended thinking held at 0% across 200 attempts. Read the first series and the second together: a 0% result at 200 attempts is a statement about 200 attempts by one attacker configuration, and the coding series shows exactly how fast a low single-attempt rate climbs when the adversary is allowed to keep trying. An adversary with a real motive is not limited to 200 tries. This is the core reason CASRAI’s Is AI Red-Teaming Security Theater? treats unbounded-effort claims sceptically.

The scope is the lab’s, not the crowd’s

Every Arena challenge is designed: somebody chooses the target behaviours, the scenarios, the model configurations exposed, the system prompts, and the guardrails left switched on or off. In the IPI Arena, that design was done “in collaboration with UK AISI, US CAISI, and frontier labs,” which is about as good as scoping gets. In a private commercial engagement it is done by the paying client. A crowd of 15,000 people cannot find a vulnerability in a surface the crowd was never shown, and the published result will not say what was withheld.

Incentives shape the threat model

Prize-driven red-teamers optimise for verifiable, judgeable breaks that score. Real adversaries optimise for outcomes and have no interest in being reproducible. The overlap is substantial but it is not identity, and an Arena result is evidence about what a motivated, rule-following, time-boxed crowd can find under a scoring rubric. That is genuinely useful. It is not a population estimate of real-world attacker capability.

It is not independence

The commercial relationship runs from the lab to the vendor. Gray Swan is paid by the model builders whose models it tests, designs the challenges with them, and markets the results as citable in their system cards. None of that makes the findings false — the IPI Arena results are actively unflattering to every model tested, which is a good sign about the pipeline’s integrity. But it is a vendor engagement, not an independent audit, and the distinction is exactly the one CASRAI sets out in Who Checks the Checkers: Evaluator Independence in AI Safety and Third-Party AI Evaluator Standards: Independence, Access, and Methodology. The involvement of UK AISI and US CAISI as data recipients and co-analysts is what upgrades these two particular competitions above ordinary vendor testing — and that upgrade does not automatically extend to any other Gray Swan engagement.

Where this sits in regulation

No regime currently accredits crowdsourced red-teaming, and none names Gray Swan or any comparable platform. What regulation does is create demand for the artifact without specifying its quality. Article 55 of the EU AI Act requires providers of general-purpose AI models with systemic risk to perform model evaluation “including conducting and documenting adversarial testing” — red teaming by name — without prescribing a methodology, a sample size, an attacker skill level, or an independence requirement. The GPAI Code of Practice’s safety and security chapter fills in process expectations around documentation and external expert consultation, but a provider retains wide latitude in how adversarial testing is carried out. See CASRAI’s EU AI Act GPAI Code of Practice guide for how those obligations are structured.

That gap is where a crowdsourced result becomes attractive: it is large, it is documented, it produces transcripts, and it is cheap relative to a bespoke audit. The governance risk is straightforward. A requirement that says “conduct adversarial testing” is satisfied by a thin engagement and a thorough one alike, and a reader of the resulting safety report has no standard against which to tell them apart. Until a framework specifies what adversarial testing evidence must contain — scope disclosure, attacker-effort budgets, who chose the targets, what was excluded — the quality of the evidence is set by the lab’s own candour. CASRAI tracks the same structural problem for pre-deployment testing agreements in CAISI and the UK AI Security Institute: How Pre-Deployment Testing Agreements Work.

Where NIKOLAI fits: N5, evaluation-validity threat

NIKOLAI is CASRAI’s own frontier-AI-safety dictionary — an independent, unendorsed reference work, not an official standard and not something any lab, institute or vendor has agreed to be bound by. Gray Swan has filed no Mapping Declaration with CASRAI, and nothing on this page should be read as Gray Swan endorsing NIKOLAI or NIKOLAI endorsing Gray Swan.

The element that applies here is evaluation-validity threat, in NIKOLAI’s N5 track (Evidence and evaluations), carried on the live element page as Proposed at nikolai-v0.1. Its definition proposes “a controlled list of named conditions — evaluation awareness, sandbagging, alignment faking, metagaming/grader-gaming, reward hacking — under which a model’s behaviour during evaluation may not reflect its behaviour in deployment, so that a lab’s claim of having checked for this is machine-comparable across frameworks.”

Crowdsourced competitions raise two of those conditions directly. Evaluation awareness: a public, announced, scheduled competition against a named model is a setting a model may be able to recognise, and the behaviour observed there is behaviour under observation. Metagaming runs in the other direction — toward the attackers, who are optimising against a scoring rubric and a human or automated judge rather than against the underlying harm the rubric is a proxy for. Neither invalidates a competition result. Both are reasons a result needs to be reported with its conditions attached rather than as a bare percentage, which is precisely what a named, machine-comparable validity-threat vocabulary would force. CASRAI’s Evaluation-Validity Threats: Sandbagging, Reward Hacking, and NIKOLAI’s N5 Crosswalk works through the full element.

Why research administrators should care

The indirect-prompt-injection threat the 2026 Arena measured is not an abstract frontier-lab concern; it is the exact failure mode of the agentic tools universities are now deploying. A research-computing office that turns on an AI assistant with access to institutional email, shared drives, ticketing systems or code repositories has built the scenario the competition tested: an agent that reads untrusted content and can take actions. The finding that at least one successful hijack existed against every frontier model tested applies to whatever model sits underneath that assistant.

Two practical consequences for sponsored programs, research security and IT procurement. First, when a vendor says its product is “red-teamed by Gray Swan” or any comparable network, the diligence question is which product — crowd Arena, automated Shade, private invitational, or guardrails — and whether the vendor will share the transcripts and the scope, not just the claim. The deliverables exist; asking for them is reasonable. Second, export-control and research-security reviewers should treat “the agent reported success” as insufficient evidence that an agent was not compromised, given the documented concealment behaviour. CASRAI’s Assessing Third-Party AI Vendor Risk covers the wider diligence pattern.

Frequently asked questions

Is Gray Swan Arena an independent audit of frontier models?

No. It is a commercial adversarial-testing platform paid for by model builders, with challenge scope agreed between the vendor and the client. The two competitions described here are stronger than ordinary vendor testing because UK AISI and US CAISI participated in design and analysis and received the full datasets, but that is collaboration, not accreditation, and it does not extend to Gray Swan’s other engagements.

How many people are in the Arena?

Gray Swan states “15,000+ adversarial researchers, and growing” for the network overall. Individual competitions are much smaller: the 2026 indirect-prompt-injection competition had 464 participants, and the 2025 agent challenge paid prizes to 161 red-teamers out of a larger field.

Did any model pass the indirect prompt injection competition?

No. CAISI’s analysis states that at least one successful attack was found against all of the target frontier models. Success rates varied widely — roughly 0.5% for Claude Opus 4.5 against roughly 8.5% for Gemini 2.5 Pro — but no model was immune.

Does a low attack success rate mean a model is safe?

Not on its own. Attack success rate depends on how much attacker effort was spent on that model, which in a prize-driven competition is not evenly distributed. Adaptive-attack series reported in system cards also show single-attempt rates climbing steeply as attempts are allowed to accumulate, so a rate quoted at one attempt says little about an adversary permitted a hundred.

Does the EU AI Act require this kind of testing?

Article 55 requires providers of general-purpose AI models with systemic risk to conduct and document adversarial testing, including red teaming, but it does not prescribe who performs it, at what scale, or to what standard. A crowdsourced competition can help satisfy that obligation; so can a much thinner exercise, and the published record often will not distinguish them.

Has Gray Swan filed a NIKOLAI Mapping Declaration?

No. Any relationship drawn on this page between Arena results and NIKOLAI’s N5 evaluation-validity-threat element is CASRAI’s own shadow mapping, not a declaration by Gray Swan, and NIKOLAI is an independent, unendorsed dictionary rather than an official standard.

Sources

  • Dziemian, M., Lin, M., Fu, X., Nowak, M., Winter, N., Jones, E., Zou, A., Ahmad, L., Chaudhuri, K., Davies, X., Edelman, B. L., Evtimov, I., Qi, X., Rando, J., Fredrikson, M., Kolter, Z., et al. How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition. arXiv:2603.15714, submitted 16 March 2026 — 464 participants, ~272,000 attempts, 8,648 successful attacks, 41 scenarios, 13 frontier models, universal strategies transferring across 21 of 41 behaviours.
  • NIST / CAISI research blog, Insights into AI Agent Security from a Large-Scale Red-Teaming Competition, 23 March 2026 — partnership with Gray Swan and UK AISI; “at least one successful attack was found against all of the target frontier models”; universal and directional attack transfer.
  • Zou, A., Lin, M., Jones, E., Nowak, M., Dziemian, M., Winter, N., Grattan, A., Nathanael, V., Croft, A., Davies, X., Patel, J., Kirk, R., Burnikell, N., Gal, Y., Hendrycks, D., Kolter, J. Z., & Fredrikson, M. Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition. arXiv:2507.20526 — 22 frontier AI agents, 44 scenarios, 1.8 million prompt-injection attacks, 60,000+ successful policy violations, ART benchmark release.
  • Gray Swan AI, UK AISI x Gray Swan Agent Red-Teaming Challenge: Results Snapshot — 8 March to 6 April 2025; 1,800,000 attempts; 62,000 successful breaks; 22 models; 44 behaviours in 4 waves; $171,800 in prizes; 161 paid red-teamers; per-model ASR figures.
  • Gray Swan AI, Arena — Red-Teaming Battlefield for Breaking AI Models — “15,000+ adversarial researchers, and growing”; “The largest AI red-teaming network in the world.”
  • Gray Swan AI, Adversarial Evaluation — “citable deliverables: reproducible transcripts, severity judgments, and raw attempt data”; “rigorous enough to cite in your model card”; quarterly flagship plus custom challenges; Shade and expert-tier descriptions.
  • Gray Swan AI, About Gray Swan — founded by Carnegie Mellon University researchers; Matt Fredrikson (co-founder, CEO), J. Zico Kolter (co-founder, Chief Scientist); Arena, Shade and Cygnal product lines.
  • VentureBeat, Anthropic vs. OpenAI red teaming methods reveal different security priorities for enterprise AI — reporting the Claude Opus 4.5 system card’s Gray Swan Shade adaptive-attack series (4.7% at 1 attempt, 33.6% at 10, 63.0% at 100 in coding environments; 0% across 200 attempts for computer use with extended thinking).
  • Regulation (EU) 2024/1689 (AI Act), Article 55 — adversarial testing and red-teaming obligations for general-purpose AI models with systemic risk.
  • CASRAI’s own NIKOLAI evaluation-validity threat element page (N5 track), verified live at time of writing.

Related reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Gray Swan Arena: Crowdsourced Red-Teaming as Pre-Deployment Evidence

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.