Building an early warning system for LLM-aided biological threat creation

by OpenAI
🌖 2024-01-31

OpenAI describes a blueprint for evaluating the risk that a large language model could aid biological-threat creation.

Building an early warning system for LLM-aided biological threat creation

In an evaluation involving biology experts and students, OpenAI found that GPT‑4 produced at most a mild uplift in biological-threat-creation accuracy. The result was not large enough to be conclusive, but was presented as a starting point for continued research and deliberation.

Overview

The evaluation aimed to measure whether models could meaningfully increase access to dangerous information compared with the internet baseline. It involved 100 human participants: 50 biology experts with PhDs and wet-lab experience, and 50 students with at least one university biology course. Participants received either internet access alone or GPT‑4 plus internet.

The study measured accuracy, completeness, innovation, time taken, and self-rated difficulty across stages of a biological-threat-creation process. It found mild uplifts in accuracy and completeness but no statistically significant effects, and did not test physical implementation.

Design principles

Increased access evaluation diagram

Redacted example research-only model response

Task-sourcing table

Biological-threat-creation process diagram

The work emphasizes testing with human participants, eliciting the full range of model capabilities under controlled conditions, and measuring improvement over existing resources. The stated purpose is to develop an empirical “tripwire” that can signal a need for caution and further testing as models become more capable.

Results and discussion

Accuracy results

Accuracy with thresholding results

Completeness results

Innovation results

Time-taken results

Self-rated difficulty results

The study found no evidence that model access reduced task-completion time or improved innovation. Researchers observed that model-assisted responses tended to be longer and contain more relevant detail, while noting uncertainty about whether that reflects truly more complete information.

OpenAI concluded that additional high-quality biorisk evaluations, clearer thresholds for meaningful risk, and effective mitigation strategies are needed. It notes that information access alone is insufficient to create a biological threat and that physical access and relevant expertise remain important constraints.

Information-hazard precautions

The study avoided publishing an end-to-end process for any particular biological threat. Participants underwent screening, training, confidentiality requirements, and security controls; access to a research-only model was restricted to vetted experts in a monitored facility.

Participant prior experience

Web pages accessed

Appendix histogram

Table 1

Mean binarized accuracy table

Second mean binarized accuracy table

Related local documents: Preparedness Framework and GPT‑4 system card.

Data downloads: participants, responses, accuracy summary, completeness summary, and innovation summary.