-
Capability-Based Planning for AI Crisis Preparedness
Authors:
Isaak Mengesha,
Charlie Collins,
Juan Felipe Cerón Uribe,
Salvatore d'Ambrosio,
Vickie Ellis
Abstract:
Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and official reviews concede that likelihood-based risk asse…
▽ More
Capability-based planning drives preparedness in defense and homeland security, but has yet to be applied seriously to AI. Government AI preparations follow a predict-then-act paradigm: rank risks by likelihood and impact, then prepare for the highest expected harm. AI resists prediction: expert timelines disagree by orders of magnitude, and official reviews concede that likelihood-based risk assessment fails for exactly this class of risk. Drawing on principles of decision making under deep uncertainty, we propose a methodological framework in three parts: a scenario library sampled systematically across declared axes; a rating procedure that assesses each government capability against each scenario on coarse, gated criteria; and a prioritization step that maps the resulting matrix onto decision rules a government might adopt. Through a pilot across the four most severe AI-enabled threat classes, we illustrate the kind of insight the instrument yields and provide a proof of concept for capability-based planning as a practical tool for AI crisis preparedness.
△ Less
Submitted 18 August, 2026;
originally announced August 2026.
-
GPT-Red: Automated Red Teaming via Self-Play at Scale
Authors:
Eric Wallace,
Christopher A. Choquette-Choo,
Nikhil Kandpal,
Sam Toyer,
Dylan Hunn,
Stephanie Lin,
Yuxin Wen,
Xiangyu Qi,
Christopher Wolff,
Zizhao Wang,
Milad Nasr,
Sicheng Zhu,
Chuan Guo,
Juan Felipe Cerón Uribe,
Kaiwen Wang,
Aiden Low,
Kai Xiao,
Kai Chen
Abstract:
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorit…
▽ More
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.
△ Less
Submitted 28 July, 2026;
originally announced July 2026.
-
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Authors:
Chuan Guo,
Juan Felipe Ceron Uribe,
Sicheng Zhu,
Christopher A. Choquette-Choo,
Steph Lin,
Nikhil Kandpal,
Milad Nasr,
Rai,
Sam Toyer,
Miles Wang,
Yaodong Yu,
Alex Beutel,
Kai Xiao
Abstract:
Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-fol…
▽ More
Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We introduce IH-Challenge, a reinforcement learning training dataset, to address these difficulties. Fine-tuning GPT-5-Mini on IH-Challenge with online adversarial example generation improves IH robustness by +10.0% on average across 16 in-distribution, out-of-distribution, and human red-teaming benchmarks (84.1% to 94.1%), reduces unsafe behavior from 6.6% to 0.7% while improving helpfulness on general safety evaluations, and saturates an internal static agentic prompt injection evaluation, with minimal capability regression. We release the IH-Challenge dataset (https://huggingface.co/datasets/openai/ih-challenge) to support future research on robust instruction hierarchy.
△ Less
Submitted 11 March, 2026;
originally announced March 2026.
-
LLM Critics Help Catch LLM Bugs
Authors:
Nat McAleese,
Rai Michael Pokorny,
Juan Felipe Ceron Uribe,
Evgenia Nitishinskaya,
Maja Trebacz,
Jan Leike
Abstract:
Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and overcome that limitation this work trains "critic" models that help humans to more accurately evaluate model-written code. These critics are themselves LLMs trained with RLHF to write natural language feedback highlighting…
▽ More
Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and overcome that limitation this work trains "critic" models that help humans to more accurately evaluate model-written code. These critics are themselves LLMs trained with RLHF to write natural language feedback highlighting problems in code from real-world assistant tasks. On code containing naturally occurring LLM errors model-written critiques are preferred over human critiques in 63% of cases, and human evaluation finds that models catch more bugs than human contractors paid for code review. We further confirm that our fine-tuned LLM critics can successfully identify hundreds of errors in ChatGPT training data rated as "flawless", even though the majority of those tasks are non-code tasks and thus out-of-distribution for the critic model. Critics can have limitations of their own, including hallucinated bugs that could mislead humans into making mistakes they might have otherwise avoided, but human-machine teams of critics and contractors catch similar numbers of bugs to LLM critics while hallucinating less than LLMs alone.
△ Less
Submitted 28 June, 2024;
originally announced July 2024.
-
GPT-4 Technical Report
Authors:
OpenAI,
Josh Achiam,
Steven Adler,
Sandhini Agarwal,
Lama Ahmad,
Ilge Akkaya,
Florencia Leoni Aleman,
Diogo Almeida,
Janko Altenschmidt,
Sam Altman,
Shyamal Anadkat,
Red Avila,
Igor Babuschkin,
Suchir Balaji,
Valerie Balcom,
Paul Baltescu,
Haiming Bao,
Mohammad Bavarian,
Jeff Belgum,
Irwan Bello,
Jake Berdine,
Gabriel Bernadett-Shapiro,
Christopher Berner,
Lenny Bogdonoff,
Oleg Boiko
, et al. (256 additional authors not shown)
Abstract:
We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers. GPT-4 is a Transformer-based mo…
▽ More
We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers. GPT-4 is a Transformer-based model pre-trained to predict the next token in a document. The post-training alignment process results in improved performance on measures of factuality and adherence to desired behavior. A core component of this project was developing infrastructure and optimization methods that behave predictably across a wide range of scales. This allowed us to accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4.
△ Less
Submitted 4 March, 2024; v1 submitted 15 March, 2023;
originally announced March 2023.