OpenAI describes a blueprint for evaluating the risk that a large language model could aid biological-threat creation.

In an evaluation involving biology experts and students, OpenAI found that GPT‑4 produced at most a mild uplift in biological-threat-creation accuracy. The result was not large enough to be conclusive, but was presented as a starting point for continued research and deliberation.
Overview
The evaluation aimed to measure whether models could meaningfully increase access to dangerous information compared with the internet baseline. It involved 100 human participants: 50 biology experts with PhDs and wet-lab experience, and 50 students with at least one university biology course. Participants received either internet access alone or GPT‑4 plus internet.
The study measured accuracy, completeness, innovation, time taken, and self-rated difficulty across stages of a biological-threat-creation process. It found mild uplifts in accuracy and completeness but no statistically significant effects, and did not test physical implementation.
Design principles

The work emphasizes testing with human participants, eliciting the full range of model capabilities under controlled conditions, and measuring improvement over existing resources. The stated purpose is to develop an empirical “tripwire” that can signal a need for caution and further testing as models become more capable.
Results and discussion
The study found no evidence that model access reduced task-completion time or improved innovation. Researchers observed that model-assisted responses tended to be longer and contain more relevant detail, while noting uncertainty about whether that reflects truly more complete information.
OpenAI concluded that additional high-quality biorisk evaluations, clearer thresholds for meaningful risk, and effective mitigation strategies are needed. It notes that information access alone is insufficient to create a biological threat and that physical access and relevant expertise remain important constraints.
Information-hazard precautions
The study avoided publishing an end-to-end process for any particular biological threat. Participants underwent screening, training, confidentiality requirements, and security controls; access to a research-only model was restricted to vetted experts in a monitored facility.
Related local documents: Preparedness Framework and GPT‑4 system card.
Data downloads: participants, responses, accuracy summary, completeness summary, and innovation summary.