Steph Guerra; Aurelia Attal-Juncqua; John P. Tarangelo; Casey Aveggio; Katie Dammer; David Glickstein · 2026-08-18
The use of artificial intelligence in biology offers great promise but could also endanger global health and national security. In this report, the authors develop a layered mitigation strategy that is robust against a variety of threat scenarios. The authors examine the strengths and weaknesses of a network of nine mitigations, describe areas of mutual reinforcement, and propose next steps for decisionmakers.
Jeffrey Lee; Alyssa Worland; Christopher Rodriguez; Kyle Brady; Grant Ellison; Henry Alexander Bradley; Dawid Maciorowski; Jordan Despanie; Barbara Del Castello; Jason Johnson; Steph Guerra · 2026-08-11
Assessing a threat model in which malicious actors leverage large language model (LLM) agents to exploit biological tools (BTs), the authors evaluate whether agents can lower expertise barriers for hazardous biodesign. They find that LLM agents demonstrate emerging ability to use BTs for nucleic acid synthesis screening evasion, although success is inconsistent and model safeguards frequently limit testing of closed-weight systems.
As protein design tools and other biological AI models become more capable of designing novel biological components and systems, the mechanisms to prevent misuse have not kept pace.
The moment an artificial intelligence (AI) lab publishes an open-weight model, it relinquishes the ability to recall it. The following is a preliminary assessment of what open-weight models mean for terrorist misuse.
Samuel H. King; Claudia L. Driscoll; David B. Li; Daniel Guo; Aditi T. Merchant; Garyk Brixi; Max E. Wilkinson; Brian L. Hie · Science · 2026-08-06
Many important biological functions arise not from single genes but from complex interactions encoded by entire genomes. We report the first generative design of complete bacteriophage genomes using genome language models. We generated viable bacteriophages with target host tropism, using the phage ΦX174 as our design template. Experimental testing yielded 16 phages with diverse fitness profiles in laboratory conditions. Cryo–electron microscopy confirmed that a generated phage utilizes an evolutionarily distant DNA packaging protein in its capsid. A cocktail of generated phages rapidly overcomes ΦX174-resistant Escherichia coli strains, demonstrating a path toward artificial intelligence–generated phage therapies against rapidly evolving bacterial pathogens. This work provides a blueprint for the design of diverse synthetic bacteriophages and useful biological systems at the genome scale.
ChatGPT’s new restrictions are obstructing legitimate analysis of the Bundibugyo Ebola response—and its Trusted Access process offers remarkably little transparency when the system gets it wrong.
Grant Ellison; Jeffrey Lee; Barbara Del Castello; Sunishchal Dev; Kyle Brady · 2026-08-06
The authors use item response theory and an interpretive taxonomy to address the motivating question: “What can benchmark data reveal about what artificial intelligence (AI) can do now that it could not do before?” Findings indicate that many benchmark tasks are no longer informative, but there is a discriminating frontier of tasks, with recent gains in performance concentrated in difficult tasks theorized to be relevant to real-world risks.
DOWNLOAD Introduction AI is rapidly advancing biotechnology by accelerating scientific discovery, biological design and experimental research.[1] Much of this progress is being driven by AI-enabled biological tools (BTs), including biological foundation models, specialized analytical tools, and components of increasingly automated biological workflows.[2] Recent breakthroughs illustrate both the breadth and accelerating pace of these advances. In […]
The offensive potential is no longer theoretical. We need to develop systems to strengthen public health as quickly as AI is accelerating biological design
I interviewed dozens of biosecurity experts. The fear that AI will enable lone wolves to build pandemics is a case study in how domain expertise gets sidelined.
Casey O. Barkan; Christopher Rodriguez; Swaptik Chowdhury; Li Ang Zhang; Roger Brent · 2026-07-20
Artificial intelligence (AI) agents can help non-experts use advanced biological design tools, raising misuse concerns. The authors investigate whether safeguards can be built into these tools to block such assistance. They find that these safeguards are ineffective against today’s AI agents, largely because multiple frontier large language models evade the safeguards by misrepresenting or concealing their identity as AI systems.
Jeremy Guntoro; Alexander Dack; Dylan Danno; Michaela Jančovičová; Križan Jurinović; Vanessa Smilansky · arXiv · 2026-07-15
Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention probe. The probes resolve finer-grained AMR drug-class subcategories and separate them from unrelated functional genes, providing additional evidence that the learned signal is not explained solely by generic functional-gene status. Bacterial virulence is also decodable, though more weakly (region-level ROC-AUC 0.833). The AMR probe retains comparable ranking performance on simulated short reads without retraining, enabling evaluation before assembly in settings where assembly is computationally costly or unreliable. It achieves a read-level ROC-AUC of 0.898 (mean-pool), comparable to the mean-pooled full-region result. Within SynGenome, AMR-associated prompt labels are only weakly recoverable from Evo 1.5-generated sequences; these prompt-derived labels do not establish the function of the generated response sequences. A complementary sparse-autoencoder analysis recovers interpretable resistance-associated features but proves less consistent than the supervised probes. Together, these results position lightweight embedding-based probes as a fast, inexpensive first-pass detection layer for metagenomic biosurveillance and map both strengths and current limits of the approach. This work was conducted as part of the AIxBio Hackathon 2026 hosted by BlueDot Impact, Apart Research, and Cambridge Biosecurity Hub.
Earlier this spring, Claude’s developers found the cyber potential of Anthropic’s model, Mythos 5, hazardous enough to pump the brakes on model release,
Rahul Gupta; Abhinav Mohanty; Payal Motwani; Venkatesh Saligrama; Satyapriya Krishna; Connor Harris; Gary Anthony Ackerman; Brandon Behlendorf; Tom Hobson; Theodore Wilson; Spyros Matsoukas · arXiv · 2026-07-13
As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.
Jeffrey Lee; Alyssa Worland; Kyle Brady; Grant Ellison; Henry Alexander Bradley; Christopher Rodriguez; Casey O. Barkan; Sunishchal Dev; Dawid Maciorowski; Jordan Despanie; Barbara Del Castello; Bria Persaud; Amar Pandya; Ella Guest; Steph Guerra · 2026-06-25
Biological tools (BTs) could potentially be misused by malicious actors seeking to design novel biological weapons. In this report, the authors evaluate seven large language model (LLM) agents on their abilities to select and operate BTs. They find that LLM agents are capable of performing initial interactions with BTs, which could lower expertise barriers for actors seeking to use BTs. Targeted testing of agent design capabilities is warranted.
Tessa Alexanian; Jacob Beal; Craig Bartling; Jens Berlips; Peter A. Carr; Adam Clore; Helena Cozzarini; James Diggans; Yorgo El Moubayed; Kevin Esvelt; Kevin Flyangolts; Leonard Foner; Patrick A. Fullerton; Bryan T. T. Gemler; Caitlin A. D. Jagla; Rassin Lababidi; Tom Mitchell; Steven T. Murphy; Michael T. Parker; Nicholas Roehner; Andre Rusch; Kemper Talley; Troy Timmerman; Nicole E. Wheeler · Frontiers in Bioengineering and Biotechnology · 2026-06-04
Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential “sequences of concern,” there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a collection of test sets containing 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequences of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of sequences in test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has, to date, reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the collection of test sets as a whole. Together, the rubric and the test sets provide a concrete “sequence of concern” definition that can be used as a foundation for development of biosecurity screening standards and policy and is also continuing to be refined in ongoing work.
Eric Horvitz · Microsoft On the Issues · 2026-06-04
AI is reshaping biology, unlocking breakthroughs while raising new risks. Learn how smarter safeguards can strengthen biosecurity without slowing innovation.
Alexandra Zini · Social Science Research Network · 2026-06-01
<p><span>The emergence of large language models and generative AI systems presents a significant and underexplored threat to biosecurity. While bioterrorism has
Moritz S. Hanke; Shrestha Rath; Anita Cicero; Thomas V. Inglesby; Jaspreet Pannu · Frontiers in Microbiology · 2026-06-01
Biological AI models (BAIMs) are advancing rapidly and hold substantial promise. Yet, these models raise dual-use concerns, particularly regarding capabilities that could enhance pathogens with pandemic potential. Current risk mitigation discussions for BAIMs are concentrated in the post-development stage, focusing on evaluations and safeguards, after a model has been trained. We argue that upstream, pre-development risk–benefit review (RBR) is a necessary, missing component of effective BAIM governance. The broad range of models, their capabilities, and purpose-built applications, as well as a majority-academic developer community, make this approach feasible and realistic. We propose a review framework with the following five components and discuss key characteristics of each component: (1) review trigger criteria that determine whether a model should undergo an RBR. The review trigger criteria are (A) a model is reasonably anticipated to possess capabilities of concern or (B) it will be trained on sensitive pathogen data classified under the Biosecurity Data Levels (BDL) system; structured risk (2) and benefit (3) reviews using qualitative and quantitative criteria; (4) integration of risk and benefit scores into a composite assessment; and (5) proportionate risk mitigation recommendations. We expect that such reviews (2 through 5) would apply to only a small fraction of BAIMs and would benefit responsible developers by establishing clear expectations at the outset of model development. We discuss how such a framework could be implemented broadly across academic institutions, commercial developers, and federal, philanthropic, and private funding bodies, and address limitations such as the subjectivity of risk assessments. RBR for BAIMs remains nascent and will require expert-driven working groups to define capabilities of concern, establish clear review criteria, and assess risk mitigation efficacy. RBRs are a promising conceptual approach for BAIM risk management and should be pioneered, refined, and vetted through real-world application with model developers.
A replacement for the rescinded 2000 Guidelines should include current challenges at the intersection of AI applications in bioengineering and biosecurity—termed the AIxBio nexus.
Aurelia Attal-Juncqua; John P. Tarangelo; Evan Davison Kotler; Casey Aveggio; Marie Burns; Barbara Del Castello; Tyler Burns; Steph Guerra; Saskia Popescu · 2026-05-20
As artificial intelligence (AI) systems are increasingly integrated into biological workflows, the risk of misuse has grown. To address this, Helena and RAND convened a workshop in early 2026 with AI, biosecurity, and biotechnology experts. Participants reviewed high-impact threat scenarios and developed mitigation strategies with broad applicability. These conference proceedings summarize the workshop’s discussions and outputs.
Casey Aveggio; Sarah R. Carter; Harshini Mukundan; Elizabeth (Beth) E. E. Cameron; Steph Guerra · Frontiers in Microbiology · 2026-05-18
The convergence of artificial intelligence and biotechnology is transforming the life sciences and enabling rapid design, development, and analysis across the research lifecycle. However, this acceleration also heightens the risk for potential biological misuse concerns. Building a sustainable and secure bioeconomy requires moving beyond rhetoric about balancing innovation and security toward practical, operational efforts that do both. This paper proposes a framework for initiatives that embed safety and trustworthiness into the architecture of biological research systems while simultaneously advancing scientific progress. Alongside this framework, we will describe three case studies where investment can yield mutual benefit: (1) secure and tiered data-sharing infrastructures that broaden access to high-quality biological data while maintaining control over sensitive information; (2) provenance and metadata tracking mechanisms for AI enabled biological tools that enhance scientific reproducibility and oversight; and (3) capability benchmarking approaches for AI-enabled biological tools that can enable performance improvements while providing situational awareness about biological risk trajectories. Our framework demonstrates how co-designing innovation and security objectives can transform potential trade-offs into reinforcing outcomes. The paper concludes by outlining policy and funding strategies to enable such mutually reinforcing win-win approaches, positioning responsible AI-enabled biotechnology as both a driver of innovation and a foundation for global biosecurity.
Gary R. Abel; Tessa Alexanian; Craig Bartling; Jacob Beal; Samuel Curtis; Kevin Flyangolts; Leonard Foner; Samuel P. Forry; Gene D. D. Godbold; Eric Horvitz; Bin Hu; Corey M. Michael Hudson; Caitlin Jagla; Rassin Lababidi; Sheng Lin-Gibson; Brittany Rife Magalis; Jaspreet Pannu; Sebastian Rivera; David Ross; Bruce J. Wittmann; James Diggans · Frontiers in Bioengineering and Biotechnology · 2026-05-14
Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.
Bruce J. Wittmann; Nicole E. Wheeler; Steven T. Murphy; Tom Mitchell; Brittany Magalis; Bryan T. Gemler; Kevin Flyangolts; James Diggans; Adam Clore; Jacob Beal; Craig Bartling; Tessa Alexanian; Eric Horvitz · bioRxiv · 2026-03-05
Rapid advancements in AI have enabled significant progress in protein and nucleic acid design, but they also pose biosecurity challenges. We examine the vulnerabilities of biosecurity screening software (BSS) to AI-reformulated synthetic homologs of proteins of concern (POCs) that have been fragmented into smaller segments. We evaluate four BSS tools that were recently patched to enhance their AI resiliency. Without any further modification, we found that two of the four tools were capable of robustly detecting fragments as short as 50 nucleotides, demonstrating screening capabilities that exceed those requested in the U.S. Framework for Nucleic Acid Synthesis. Upgraded versions of the other two tools improved performance. Although our findings confirm the effectiveness of the tested BSS tools, at the same time, they emphasize the urgency of developing alternate BSS approaches to counter evolving AI-enabled biosecurity risks.
The AIxBio field stands at a critical juncture where rapid capability advances are outpacing governance frameworks and safety measures. The next 18 months will likely prove pivotal in determining whether voluntary safety practices by AI companies, emerging evaluation frameworks, and international coordination efforts can keep pace with technological development.
Advances in artificial intelligence (AI) and distributed ledger technology (DLT) are reshaping how biological research, data and materials are managed. At the same time, the mechanisms used to implement and demonstrate compliance with the biological weapons prohibition regime rely heavily on national oversight systems that face increasing administrative complexity and uneven capacity across states. Emerging technologies are often discussed as potential sources of risk in the life sciences, but they may also provide tools to strengthen key regime functions. AI and DLT could support more effective laboratory oversight, strengthen export controls on dual-use items, and facilitate national reporting and transparency mechanisms. Their impact will depend on governance choices—including how states manage data integrity, human oversight, interoperability and equitable access to digital capabilities. Used responsibly, these tools could improve record integrity, administrative efficiency and confidence in the peaceful use of biological research.
Chen Bo Calvin Zhang; Christina Q. Knight; Nicholas Kruus; Jason Hausenloy; Pedro Medeiros; Nathaniel Li; Aiden Kim; Yury Orlovskiy; Coleman Breen; Bryce Cai; Jasper Götting; Andrew Bo Liu; Samira Nedungadi; Paula Rodriguez; Yannis Yiming He; Mohamed Shaaban; Zifan Wang; Seth Donoughe; Julian Michael · arXiv · 2026-02-26
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurity-relevant task sets. Participants worked on complex problems with ample time (up to 13 hours for the most involved tasks). We found that LLM access provided substantial uplift: novices with LLMs were 4.16 times more accurate than controls (95% CI [2.63, 6.87]). On four benchmarks with available expert baselines (internet-only), novices with LLMs outperformed experts on three of them. Perhaps surprisingly, standalone LLMs often exceeded LLM-assisted novices, indicating that users were not eliciting the strongest available contributions from the LLMs. Most participants (89.6%) reported little difficulty obtaining dual-use-relevant information despite safeguards. Overall, LLMs substantially uplift novices on biological tasks previously reserved for trained practitioners, underscoring the need for sustained, interactive uplift evaluations alongside traditional benchmarks.
CapabilitiesDual-use researchEvaluationsJailbreaks and red-teamingNon-state actors
Aurelia Attal-Juncqua; Anita Cicero; Alex Zhu; Tom Inglesby · Health Security · 2026-02-22
Artificial Intelligence (AI) has the potential to revolutionize biosecurity, health security, biodefense, and pandemic preparedness by offering groundbreaking solutions for managing biological threats. This landscape review explores recent advancements in AI across these fields, drawing from both grey literature and peer-reviewed studies published between January 2019 and February 2024. AI has demonstrated potential in predicting viral mutations, which could enable earlier detection of outbreaks, and streamlining resource allocation by analyzing diverse data sources. It could also play a crucial role in accelerating the development and deployment of medical countermeasures, such as vaccines and therapeutics. Additionally, use of AI may enhance laboratory automation, reducing human error and increasing biosafety. Despite these promising advancements, significant challenges related to the potential misuse of AI, data security, and privacy concerns necessitate careful implementation and robust governance. This article highlights the rapid progress and vast potential of AI in biosecurity and provides key recommendations for US policymakers to effectively harness AI's capabilities while ensuring safety and security. These recommendations include expanding access to advanced computing resources, fostering collaboration across sectors, and establishing clear regulatory frameworks to support the safe and ethical deployment of AI technologies.
Carlos Mougan; Lauritz Morlock; Jair Aguirre; James R. M. Black; Jan Brauner; Simeon Campos; Sunishchal Dev; David Fernández Llorca; Alberto Franzin; Mario Fritz; Emilia Gómez; Friederike Grosse-Holz; Eloise Hamilton; Max Hasin; Jose Hernandez-Orallo; Dan Lahav; Luca Massarelli; Vasilios Mavroudis; Malcolm Murray; Patricia Paskov; Jaime Raldua; Wout Schellaert · Science · 2026-02-19
Amid rapid advances in bioscience and as artificial intelligence reshapes global technological capabilities, the Nuclear Threat Initiative (NTI) and the China Arms Control and Disarmament Association (CACDA) are jointly calling for action to strengthen biosecurity and oversight and engage in responsible practices to prevent accidents or misuse.
Adeline E. Williams; Barbara Del Castello; Jeffrey Lee; Derek Roberts; John P. Tarangelo; Jay Atanda; Alejandro Colman-Lerner; Jeff Gerold; Roger Brent · 2026-02-11
A rise in artificial intelligence (AI) use in biology has driven transformative developments in the field but could also pose significant dual-use risks. In this report, the authors identify five biological functions that could be modified using AI tools and develop a dual-component risk-scoring tool — combining biological modification risk factors with actor capability assessments — to evaluate the potential for misuse.
This perspective examines biological agentic evaluations as an emerging tool for assessing the capabilities and risks of autonomous AI systems in biological contexts. Drawing on hands-on evaluation experience, it offers practical guidance on defining, designing, running, scoring, and interpreting evaluations, highlighting how design choices shape conclusions and policy relevance.
Tom Hobson; Alexander Ghionis; Lalitha Sundaram; Nancy Connell; Richard Armitage · Social Science Research Network · 2026-02-10
Discussions of artificial intelligence (AI) and biological weapons (BW) have largely focused on laboratory-facing capabilities, treating acquisition as
Doni Bloomfield; James R. M. Black; Oliver Crook; Nadav Brandes; Moritz S. Hanke; Thomas V. Inglesby; Anita Cicero; Robert Pollack; Tina Hernandez-Boussard; Michael J. Imperiale; Russ B. Altman; Jaspreet Pannu · Science · 2026-02-05
Kyle Brady; Jeffrey Lee; Dawid Maciorowski; Alyssa Worland; Jordan Despanie; Bria Persaud; Barbara Del Castello; Henry Alexander Bradley; Grant Ellison; Charles Teague; Sarah L. Gebauer; Greg McKelvey; Steph Guerra; Ella Guest · 2026-02-03
This report describes eight frontier large language model (LLM) agents on their ability to design DNA segments, interact with a benchtop DNA synthesizer, and generate laboratory protocols. These are dual-use tasks, explored as potential technical bottlenecks to a malicious actor building a viral pathogen that could be weaponized. Performance varied among the models, but all tested LLMs designed biologically coherent DNA segments in some attempts.
Managed access will be critical for reducing biosecurity risks related to the misuse of biological AI tools. By working to develop best practices for each element of this framework—risk levels, tiered access, and practices to verify legitimacy—developers of biological AI tools and the broader life sciences community can reduce risks while maintaining the benefits of these tools.
Josh Dettman; Emily Lathrop; Aurelia Attal-Juncqua; Matthew Nicotra; Allison Berke · Biotechnology and Bioengineering · 2026-01-17
As artificial intelligence continues to enhance biological innovation, the potential for misuse must be addressed to fully unlock the potential societal benefits. While significant work has been done to evaluate general-purpose AI and specialized biological design tools (BDTs) for biothreat creation risks, actionable steps to mitigate the risk of AI-enabled biothreat creation are underdeveloped. This paper provides policy and technology strategies collected from a diverse range of sources placed in the context of an organizing framework aligned with steps in the AI-enabled creation of a biothreat. After collating previous reports (typically on one or a small set of mitigation options) and evaluating the proposed mitigation options by projected feasibility and impact, we prioritize development of seven mitigation strategies (with a total of twelve individual mitigations): model unlearning and information removal techniques (a combination of five mitigations), classifier-based input and output filtering for BDTs, AI agents for biosecurity, safety bug bounty programs, ensuring enforcement of existing material/equipment protections, enhancing biosurveillance and bioattribution, and screening metadata/audit logs before DNA synthesis. We invite collaboration among policymakers, researchers, and technologists to refine and implement these strategies into a strong layered defense, ensuring that AI can be used safely and securely to the benefit of all.
Žiga Avsec; Natasha Latysheva; Jun Cheng; Guido Novati; Kyle R. Taylor; Tom Ward; Clare Bycroft; Lauren Nicolaisen; Eirini Arvaniti; Joshua Pan; Raina Thomas; Vincent Dutordoir; Matteo Perino; Soham De; Alexander Karollus; Adam Gayoso; Toby Sargeant; Anne Mottram; Lai Hong Wong; Pavol Drotár; Adam Kosiorek; Andrew Senior; Richard Tanburn; Taylor Applebaum; Souradeep Basu; Demis Hassabis; Pushmeet Kohli · Nature · 2026-01-01
Deep learning models that predict functional genomic measurements from DNA sequences are powerful tools for deciphering the genetic regulatory code. Existing methods involve a trade-off between input sequence length and prediction resolution, thereby limiting their modality scope and performance1–5. We present AlphaGenome, a unified DNA sequence model, which takes as input 1 Mb of DNA sequence and predicts thousands of functional genomic tracks up to single-base-pair resolution across diverse modalities. The modalities include gene expression, transcription initiation, chromatin accessibility, histone modifications, transcription factor binding, chromatin contact maps, splice site usage and splice junction coordinates and strength. Trained on human and mouse genomes, AlphaGenome matches or exceeds the strongest available external models in 25 of 26 evaluations of variant effect prediction. The ability of AlphaGenome to simultaneously score variant effects across all modalities accurately recapitulates the mechanisms of clinically relevant variants near the TAL1 oncogene6. To facilitate broader use, we provide tools for making genome track and variant effect predictions from sequence.
AI will not transform terrorism overnight. Its more likely effect is subtler: lowering barriers for extremist propaganda, recruitment, disinformation, and planning while strengthening counterterrorism tools and surveillance risks.
Flawed safety assessments of artificial intelligence (AI) models that rely on tacit knowledge and inadequate benchmarks may create a false sense of security. The authors identify attributes of contemporary AI models that contribute to increased biological weapons risk and propose that improved benchmarks to enable more-comprehensive risk evaluation might help to mitigate risks from future models before deployment.
FrontierScience: Evaluating AI's ability to perform expert-level scientific tasks
Journal article
Miles Wang; Joy Jiao; Neil Chowdhury; Ethan Chang; Tejal Patwardhan · 2025-12-16
We introduce FrontierScience, a benchmark evaluating AI capabilities for expertlevel scientific reasoning. FrontierScience consists of two tracks: (1) Olympiad, which contains international olympiad problems (at the level of IPhO, IChO, and IBO), and (2) Research, which contains PhD-level, open-ended problems representative of sub-problems in scientific research. In total, FrontierScience is composed of several hundred questions (160 in the open-sourced gold set) covering subfields across physics, chemistry, and biology, from quantum electrodynamics to synthetic organic chemistry. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. In contrast, all Olympiad problems are originally produced by international olympiad medalists and national team coaches to ensure standards of difficulty, originality, and factuality. All Research problems are research sub-tasks written and verified by PhD scientists (doctoral candidates, postdoctoral researchers, or professors). For Research, we also introduce a granular rubric-based architecture to evaluate model capabilities throughout the process of solving a research task, as opposed to judging a standalone answer. In initial evaluations of several frontier models, GPT-5.2 is the top performing model on FrontierScience, scoring 77% on the Olympiad set and 25% on the Research set.
Several frontier AI companies test their AI systems for dual-use biological capabilities that might be misused by threat actors. But what do these test results imply about the overall risk of bioterrorist attacks? There is much expert debate about how seriously to view such threats, especially from lone wolf actors. This report creates a framework for how to convert capability evaluations into risk assessments, using a simple model that draws on historical case studies, expert elicitation, and reference class forecasting. I conclude that if AI systems were to increase the number of STEM Bachelors able to synthesise pathogens as complex as influenza by 10 percentage points and also enable them to design concerning operational attack plans, then the annual probability of an epidemic caused by a lone wolf attack might increase from 0.15% to 1.0%. This is equivalent to 12,000 additional expected deaths per year, or ~$100B. Risk scenarios where AI or other tools also help discover novel viruses reach higher damages, whereas risk can also be significantly lowered if mitigations are put in place. A review of this report by six subject-matter experts and five superforecasters found similar medians, though all forecasts had high uncertainty. This work demonstrates a methodological approach for converting capability evaluations into risk assessments, whilst highlighting the continued need for better underlying evidence and expert discussion to refine assumptions.
Kaitlyn Gibbons · The Nuclear Threat Initiative · 2025-12-09
The AI-biology convergence offers enormous benefits but also brings about risks as we’ve never seen before. Without action from multiple disciplines, the race for AI development and dominance could become a race to the bottom when it comes to safety and security.
Le Cong; David Smerkous; Xiaotong Wang; Di Yin; Zaixi Zhang; Ruofan Jin; Yinkai Wang; Michal Gerasimiuk; Ravi K. Dinesh; Alex Smerkous; Lihan Shi; Joy Zheng; Ian Lam; Xuekun Wu; Shilong Liu; Peishan Li; Yi Zhu; Ning Zhao; Meenal Parakh; Simran Serrao; Imran A. Mohammad; Chao-Yeh Chen; Xiufeng Xie; Tiffany Chen; David Weinstein; Greg Barbone; Belgin Caglar; John B. Sunwoo; Fuxin Li; Jia Deng; Joseph C. Wu; Sanfeng Wu; Mengdi Wang · arXiv · 2025-12-08
Modern science advances fastest when thought meets action. LabOS represents the first AI co-scientist that unites computational reasoning with physical experimentation through multimodal perception, self-evolving agents, and Extended-Reality(XR)-enabled human-AI collaboration. By connecting multi-model AI agents, smart glasses, and robots, LabOS allows AI to see what scientists see, understand experimental context, and assist in real-time execution. Across applications -- from cancer immunotherapy target discovery to stem-cell engineering and material science -- LabOS shows that AI can move beyond computational design to participation, turning the laboratory into an intelligent, collaborative environment where human and machine discovery evolve together.
Synthetic biology harnesses and redesigns biological systems to drive innovation across a broad range of sectors, including health, agriculture, and production. It is increasingly integrating with artificial intelligence tools like large language models and robotics to accelerate innovation, improve accessibility, and enable more complex applications. Guided by the OECD Framework for Anticipatory Governance of Emerging Technologies, this report provides a strategic intelligence assessment of this convergence, laying out several concrete cases of where the technology is and how it could develop in the future. It identifies the governance implications (e.g. biosecurity and biosafety, data supply chain, human oversight) with accompanying policy options for each to guide policymakers on potential future actions. The report recommends further analysis on a range of issues due to policy importance and high uncertainty, such as forward-looking monitoring of the technology’s development, agile and anticipatory governance, and leveraging spaces for international collaboration.
AI agentsBiological ToolsFrameworkPolicySynthetic biology
Hayley Severance; Kevin P. O'Prey; Nikki Teran; Jamie M. Yassif · 2025-12-03
Imagine this: A global pandemic that results in more than 850 million cases, and 60 million deaths is sparked by a novel enterovirus strain that was intentionally engineered by an extremist group using artificial intelligence–enabled capabilities. This is the fictional scenario that senior leaders grappled with during a tabletop exercise hosted by NTI in partnership with the Munich Security Conference (MSC) in February 2025.
The exercise explored opportunities and risks at the convergence of AI and life sciences (AIxBio), and the resulting report offers practical recommendations to prevent catastrophic misuse.
The key takeaway? The exploitation of AIxBio capabilities for harm is a plausible near-term risk—and the time for action is now.
International governanceMitigationsPolicyRisk assessmentsThreat AssessmentThe Nuclear Threat Initiative
The UK Al Security Institute (AISI) has conducted evaluations of frontier Al systems since November 2023 across domains critical to national security and public safety. This report presents our first public analysis of the trends we've observed. It seeks to provide accessible, data-driven insights into the frontier of Al capabilities and promote a shared understanding among governments, industry, and the public.
CapabilitiesEvaluationsUnited KingdomUK AI Security Institute
Oliver M. Crook; Anemone Franz; Aaron Maiwald · Bulletin of the Atomic Scientists · 2025-12-01
Technologies needed for tracing engineered biothreats back to their sources are advancing rapidly. Here are some recommendations for creating attribution systems that can improve detection, accountability, and deterrence.
Michael Ben Okon; Okechukwu Paul-Chima Ugwu; Chinyere Nneoma Ugwu; Fabian Chukwudi Ogenyi; Dominic Terkimbi Swase; Chinyere Nkemjika Anyanwu; Val Hyginus Udoka Eze; Jovita Nnenna Ugwu; Saheed Adekunle Akinola; Regan Mujinya; Emeka Godson Anyanwu · Frontiers in Public Health · 2025-11-26
Biosecurity threats, which include natural outbreaks, laboratory accidents, and intentional bioterrorism, are a major issue for global health security. The impact of poor preparedness on the health, social, and economic effects of the 1918 influenza pandemic, the 2001 anthrax attacks, and the COVID-19 crisis is devastating. Standard methods, such as quarantine and serology, as well as traditional inoculations, offered basic defences but were often reactive, slow, and unfair. The recent scientific and technological progress has altered the concept of biosecurity preparedness by providing new instruments of early detection, quick reaction, and fair health solutions. Artificial intelligence-based epidemic prediction, next-generation sequencing, CRISPR-based diagnostics, and digital epidemiology are emerging technologies that enable near-real-time surveillance. New therapeutic agents and vaccines, such as mRNA and DNA platforms, monoclonal antibodies, and nanobody therapies, have enhanced response capabilities. Containment measures based on robotics, biosensors, nanotechnology-based PPE, and portable biocontainment units have simultaneously improved frontline safety. Sensitive health information and enhanced coordination are today secured with the help of digital and cyber-biosecurity tools. Nonetheless, the innovations have ethical, legal, and equity issues, which point to the need to govern responsibly and make them accessible to all. This review brings forth the incorporation of emerging technologies with international cooperation, fair systems, and responsive policies as the keys to developing resilient and future-orientated systems that could help alleviate natural, accidental, and intentional biosecurity threats.
Dianzhuo Wang; Marian Huot; Zechen Zhang; Kaiyi Jiang; Eugene I. Shakhnovich; Kevin M. Esvelt · Frontiers in Microbiology · 2025-11-26
Artificial intelligence now shapes the design of biological matter. Protein language models (pLMs), trained on millions of natural sequences, can predict, generate, and optimize functional proteins with minimal human input. When embedded in experimental pipelines, these systems enable closed-loop biological design at unprecedented speed. The same convergence that accelerates vaccine and therapeutic discovery, however, also creates new dual-use risks. We first map recent progress in using pLMs for fitness optimization across proteins, then critically assess how these approaches have been applied to viral evolution and how they intersect with laboratory workflows, including active learning and automation. Building on this analysis, we outline a capability-oriented framework for integrated AI–biology systems, identify evaluation challenges specific to biological outputs, and propose research directions for training-and inference-time safeguards.
Sunishchal Dev; Charles Teague; Grant Ellison; Kyle Brady; Ying-Chiang Jeffrey Lee; Sarah L. Gebauer; Henry Alexander Bradley; Dawid Maciorowski; Bria Persaud; Jordan Despanie; Barbara Del Castello; Alyssa Worland; Michael Miller; Adrian Salas; Dave Nguyen; James Liu; Jason Johnson; Andrew Sloan; Will Stonehouse; Travis Merrill; Thomas Goode; Greg McKelvey; Ella Guest · 2025-11-25
Artificial intelligence (AI) systems demonstrate deep knowledge across a broad variety of scientific domains, including biology and chemistry, and bad actors could misuse some of these systems to develop biological or chemical weapons. In this report, the authors evaluate the most-capable AI models (as of May 2025) against eight leading knowledge benchmarks to determine the degree to which frontier AI systems pose biological or chemical risks.
Doni Bloomfield; Moritz S. Hanke; Aaron Maiwald; James R. M. Black; Toby Webster; Tina Hernandez-Boussard; Allison Berke; Oliver M. Crook; Jassi Pannu · 2025-11-24
Training data is an essential input into creating competent artificial intelligence (AI) models. AI models for biology are trained on large volumes of data, including data related to biological sequences, structures, images, and functions. The type of data used to train a model is intimately tied to the capabilities it ultimately possesses– including those of biosecurity concern. For this reason, an international group of more than 100 researchers at the recent 50th anniversary Asilomar Conference endorsed data controls to prevent the use of AI for harmful applications such as bioweapons development. To help design such controls, we introduce a five-tier Biosecurity Data Level (BDL) framework for categorizing pathogen data. Each level contains specific data types, based on their expected ability to contribute to capabilities of concern when used to train AI models. For each BDL tier, we propose technical restrictions appropriate to its level of risk. Finally, we outline a novel governance framework for newly created dual-use pathogen data. In a world with widely accessible computational and coding resources, data controls may be among the most high-leverage interventions available to reduce the proliferation of concerning biological AI capabilities.
Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who might fine-tune these models for malicious use. To address this gap, we propose BioRiskEval, a framework to evaluate the robustness of procedures that are intended to reduce the dual-use capabilities of bio-foundation models. BioRiskEval assesses models' virus understanding through three lenses, including sequence modeling, mutational effects prediction, and virulence prediction. Our results show that current filtering practices may not be particularly effective: Excluded knowledge can be rapidly recovered in some cases via fine-tuning, and exhibits broader generalizability in sequence modeling. Furthermore, dual-use signals may already reside in the pretrained representations, and can be elicited via simple linear probing. These findings highlight the challenges of data filtering as a standalone procedure, underscoring the need for further research into robust safety and security strategies for open-weight bio-foundation models.
Aditi T. Merchant; Samuel H. King; Eric Nguyen; Brian L. Hie · Nature · 2025-11-19
Generative genomic models can design increasingly complex biological systems1. However, controlling these models to generate novel sequences with desired functions remains challenging. Here, we show that Evo, a genomic language model, can leverage genomic context to perform function-guided design that accesses novel regions of sequence space. By learning semantic relationships across prokaryotic genes2, Evo enables a genomic ‘autocomplete’ in which a DNA prompt encoding genomic context for a function of interest guides the generation of novel sequences enriched for related functions, which we refer to as ‘semantic design’. We validate this approach by experimentally testing the activity of generated anti-CRISPR proteins and type II and III toxin–antitoxin systems, including de novo genes with no significant sequence similarity to natural proteins. In-context design of proteins and non-coding RNAs with Evo achieves robust activity and high experimental success rates even in the absence of structural priors, known evolutionary conservation or task-specific fine-tuning. We then use Evo to complete millions of prompts to produce SynGenome, a database containing over 120 billion base pairs of artificial intelligence-generated genomic sequences that enables semantic design across many functions. More broadly, these results demonstrate that generative genomics with biological language models can extend beyond natural sequences.
Ludovico Mitchener; Angela Yiu; Benjamin Chang; Mathieu Bourdenx; Tyler Nadolski; Arvis Sulovari; Eric C. Landsness; Daniel L. Barabasi; Siddharth Narayanan; Nicky Evans; Shriya Reddy; Martha Foiani; Aizad Kamal; Leah P. Shriver; Fang Cao; Asmamaw T. Wassie; Jon M. Laurent; Edwin Melville-Green; Mayk Caldas; Albert Bou; Kaleigh F. Roberts; Sladjana Zagorac; Timothy C. Orr; Miranda E. Orr; Kevin J. Zwezdaryk; Ali E. Ghareeb; Laurie McCoy; Bruna Gomes; Euan A. Ashley; Karen E. Duff; Tonio Buonassisi; Tom Rainforth; Randall J. Bateman; Michael Skarlinski; Samuel G. Rodriques; Michaela M. Hinks; Andrew D. White · arXiv · 2025-11-05
Data-driven scientific discovery requires iterative cycles of literature search, hypothesis generation, and data analysis. Substantial progress has been made towards AI agents that can automate scientific research, but all such agents remain limited in the number of actions they can take before losing coherence, thus limiting the depth of their findings. Here we present Kosmos, an AI scientist that automates data-driven discovery. Given an open-ended objective and a dataset, Kosmos runs for up to 12 hours performing cycles of parallel data analysis, literature search, and hypothesis generation before synthesizing discoveries into scientific reports. Unlike prior systems, Kosmos uses a structured world model to share information between a data analysis agent and a literature search agent. The world model enables Kosmos to coherently pursue the specified objective over 200 agent rollouts, collectively executing an average of 42,000 lines of code and reading 1,500 papers per run. Kosmos cites all statements in its reports with code or primary literature, ensuring its reasoning is traceable. Independent scientists found 79.4% of statements in Kosmos reports to be accurate, and collaborators reported that a single 20-cycle Kosmos run performed the equivalent of 6 months of their own research time on average. Furthermore, collaborators reported that the number of valuable scientific findings generated scales linearly with Kosmos cycles (tested up to 20 cycles). We highlight seven discoveries made by Kosmos that span metabolomics, materials science, neuroscience, and statistical genetics. Three discoveries independently reproduce findings from preprinted or unpublished manuscripts that were not accessed by Kosmos at runtime, while four make novel contributions to the scientific literature.
Frontier open-weight models lag behind the most capable models by an average of 3 months in the Epoch Capabilities Index (ECI), our holistic measure of model capability. That corresponds to an average ECI gap of around 7 points, similar to the gap between o3 and GPT-5. However, the gap varies considerably over time, sometimes even closing completely. Until the release of o1-mini, Llama 3.1-405B was rated on par with the closed-source state-of-the-art model, Claude 3.5 Sonnet. You can see more detailed analysis about the gap in our earlier article.
Mantas Mazeika; Alice Gatti; Cristina Menghini; Udari Madhushani Sehwag; Shivam Singhal; Yury Orlovskiy; Steven Basart; Manasi Sharma; Denis Peskoff; Elaine Lau; Jaehyuk Lim; Lachlan Carroll; Alice Blair; Vinaya Sivakumar; Sumana Basu; Brad Kenstler; Yuntao Ma; Julian Michael; Xiaoke Li; Oliver Ingebretsen; Aditya Mehta; Jean Mottola; John Teichmann; Kevin Yu; Zaina Shaik; Adam Khoja; Richard Ren; Jason Hausenloy; Long Phan; Ye Htet; Ankit Aich; Tahseen Rabbani; Vivswan Shah; Andriy Novykov; Felix Binder; Kirill Chugunov; Luis Ramirez; Matias Geralnik; Hernán Mesura; Dean Lee; Ed-Yeremai Hernandez Cardona; Annette Diamond; Summer Yue; Alexandr Wang; Bing Liu; Ernesto Hernandez; Dan Hendrycks · arXiv · 2025-10-30
AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To measure this, we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings. AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%. These results help ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts and enabling stakeholders to proactively navigate AI-driven labor automation.
David R. Gillum; Rebecca L. Moritz · Frontiers in Bioengineering and Biotechnology · 2025-10-28
IntroductionRecent U.S. biosecurity policy has shifted from organism-level controls to sequence-level governance of synthetic nucleic acids in response to de novo genome synthesis risks, artificial intelligence assisted design, and globalized DNA/RNA manufacturing. While intended to strengthen safety and security, this shift risks overburdening under-resourced institutions and providing oversight that looks thorough on paper but delivers little added protection. This study examines the widening “implementation gap” between policy ambition and operational capacity.MethodsDrawing on practitioner experience and current literature, we analyzed policy frameworks, institutional practices, and case examples to identify structural challenges in sequence-level oversight. Particular attention was given to how definitions, regulatory triggers, and institutional resources interact in practice, creating gaps between policy intent and operational capacity. This mixed approach allowed us to capture the high-level design of oversight frameworks and the practical realities of their implementation across diverse institutional settings.ResultsWe found three core obstacles: ambiguous definitions of sequences of concern, fragmented and overlapping regulatory triggers, and underdeveloped institutional screening and review capacities. Ambiguity creates uncertainty about what should be flagged, while fragmented rules add redundancies without clarifying responsibility. Limited institutional resources further constrain effective oversight. These weaknesses produce overinclusive surveillance, inconsistent provider screening, unmanaged legacy construct inventories, and a lack of shared reference tools, straining resources without yielding proportional security benefits.DiscussionAligning oversight with real-world capacity is essential to avoid brittle and costly systems that deliver limited biosecurity benefits. We propose seven reforms to address the identified obstacles: functional risk tiering, federal investment in biosafety infrastructure, policy pilots and real-world testing, institutional certification pathways, adaptive governance cycles, pragmatic global harmonization, and coupling screening with operational safeguards. These measures reduce ambiguity, streamline fragmented rules, and strengthen institutional capabilities. Embedding implementer perspectives and calibrating oversight to realistic capacities will ensure that biosecurity systems remain credible, resilient, and effective in the synthetic nucleic acid era.
Nucleic acid synthesis screeningPolicyUnited States
Tom Reed; Tegan McCaslin; Luca Righetti · arXiv · 2025-10-28
Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This practice provides a window into the methodology of these evaluations, helping to build public trust in AI systems, and enabling third party review in the still-emerging science of AI evaluation. But what aspects of evaluation methodology do developers currently include -- or omit -- in their reports? This paper examines three frontier AI model reports published in spring 2025 with among the most detailed documentation: OpenAI's o3, Anthropic's Claude 4, and Google DeepMind's Gemini 2.5 Pro. We compare these using the STREAM (v1) standard for reporting ChemBio benchmark evaluations. Each model report included some useful details that the others did not, and all model reports were found to have areas for development, suggesting that developers could benefit from adopting one another's best reporting practices. We identified several items where reporting was less well-developed across all model reports, such as providing examples of test material, and including a detailed list of elicitation conditions. Overall, we recommend that AI developers continue to strengthen the emerging science of evaluation by working towards greater transparency in areas where reporting currently remains limited.
Stephen Casper; Kyle O'Brien; Shayne Longpre; Elizabeth Seger; Kevin Klyman; Rishi Bommasani; Aniruddha Nrusimha; Ilia Shumailov; Sören Mindermann; Steven Basart; Frank Rudzicz; Kellin Pelrine; Avijit Ghosh; Andrew Strait; Robert Kirk; Dan Hendrycks; Peter Henderson; J. Zico Kolter; Geoffrey Irving; Yarin Gal; Yoshua Bengio; Dylan Hadfield-Menell · Social Science Research Network · 2025-10-26
Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different opportunities and challenges for effective risk management. For example, they allow for more open research and testing. However, managing their risks is also challenging because they can be modified arbitrarily, used without oversight, and spread irreversibly. Currently, there is limited research on safety tooling specific to open-weight models. Addressing these gaps will be key to both realizing their benefits and mitigating their harms. In this paper, we present 16 open technical challenges for open-weight model safety involving training data, training algorithms, evaluations, deployment, and ecosystem monitoring. We conclude by discussing the nascent state of the field, emphasizing that openness about research, methods, and evaluations-not just weights-will be key to building a rigorous science of open-weight model risk management.
Patricia Skowronek; Anant Nawalgaria; Matthias Mann · bioRxiv · 2025-10-06
We present a multimodal AI laboratory agent that captures and shares tacit experimental practice by linking written instructions with hands-on laboratory work through the analysis of video, speech, and text. While current AI tools have proven effective in literature analysis and code generation, they do not address the critical gap between documented knowledge and implicit lab practice. Our framework bridges this divide by integrating protocol generation directly from researcher-recorded videos, systematic detection of experimental errors, and evaluation of instrument readiness by comparing current performance against historical decisions. Evaluated in mass spectrometry-based proteomics, we demonstrate that the agent can capture and share practical expertise beyond conventional documentation and identify common mistakes, although domain-specific and spatial recognition should still be improved. This agentic approach enhances reproducibility and accessibility in proteomics and provides a generalizable model for other fields where complex, hands-on procedures dominate. This study lays the groundwork for community-driven, multimodal AI systems that augment rather than replace the rigor of scientific practice.
Microsoft researchers reveal a confidential research effort that explored how open-source AI tools could be used to bypass biosecurity checks—and helped create fixes now influencing global standards.
Jailbreaks and red-teamingMitigationsOpen-weight LLMsTransparency and reporting
Bruce J. Wittmann; Tessa Alexanian; Craig Bartling; Jacob Beal; Adam Clore; James Diggans; Kevin Flyangolts; Bryan T. Gemler; Tom Mitchell; Steven T. Murphy; Nicole E. Wheeler; Eric Horvitz · Science · 2025-10-02
Advances in artificial intelligence (AI)–assisted protein engineering are enabling breakthroughs in the life sciences but also introduce new biosecurity challenges. Synthesis of nucleic acids is a choke point in AI-assisted protein engineering pipelines. Thus, an important focus for efforts to enhance biosecurity given AI-enabled capabilities is bolstering methods used by nucleic acid synthesis providers to screen orders. We evaluated the ability of open-source AI-powered protein design software to create variants of proteins of concern that could evade detection by the biosecurity screening tools used by nucleic acid synthesis providers, identifying a vulnerability where AI-redesigned sequences could not be detected reliably by current tools. In response, we developed and deployed patches, greatly improving detection rates of synthetic homologs more likely to retain wild type–like function.
Although significant work has been done to evaluate general-purpose artificial intelligence (AI) and specialized biological design tools for biothreat creation risks, actionable steps to mitigate the risk of AI-enabled biothreat creation are underdeveloped. In this paper, the authors provide policy and technology strategies in an organizing framework and prioritize the development of seven mitigation strategies.
David Manheim; Adeline Williams; Casey Aveggio; Allison Berke · 2025-09-24
In this report, the authors present findings from a Delphi study designed to assess both fundamental and near-term limits of biological design assisted by artificial intelligence (AI). Their objective was to identify which biological and AI-related constraints might serve as hard or persistent barriers to misuse. To do so, they conducted two parallel Delphi elicitations with experts in biology and AI to evaluate the limits that each field faces.
Aurelia Attal-Juncqua; Casey Mahoney; Sana Zakaria; Paul Khullar; Patricia Paskov; Ella Guest; Michael Simpson · 2025-09-24
Rapid AI and biotechnology development brings promising health benefits but creates unprecedented biosecurity risks, which current global treaties and data systems cannot sufficiently address. The RAND–NTI workshop at the 2025 AI Action Summit convened stakeholders from multiple sectors to examine challenges and responses. Key discussions centered on gaps in reliable data, policy obstacles, and the need for coordinated measures to track, contain, and manage AI-driven biosecurity threats.
Zaixi Zhang; Ruofan Jin; Le Cong; Mengdi Wang · arXiv · 2025-09-20
DNA language models have revolutionized our ability to understand and design DNA sequences--the fundamental language of life--with unprecedented precision, enabling transformative applications in therapeutics, synthetic biology, and gene editing. However, this capability also poses substantial dual-use risks, including the potential for creating pathogens, viruses, and even bioweapons. To address these biosecurity challenges, we introduce two innovative watermarking techniques to reliably track the designed DNA: DNAMark and CentralMark. DNAMark employs synonymous codon substitutions to embed watermarks in DNA sequences while preserving the original function. CentralMark further advances this by creating inheritable watermarks that transfer from DNA to translated proteins, leveraging protein embeddings to ensure detection across the central dogma. Both methods utilize semantic embeddings to generate watermark logits, enhancing robustness against natural mutations, synthesis errors, and adversarial attacks. Evaluated on our therapeutic DNA benchmark, DNAMark and CentralMark achieve F1 detection scores above 0.85 under various conditions, while maintaining over 60% sequence similarity to ground truth and degeneracy scores below 15%. A case study on the CRISPR-Cas9 system underscores CentralMark's utility in real-world settings. This work establishes a vital framework for securing DNA language models, balancing innovation with accountability to mitigate biosecurity risks.
Biological ToolsMitigationsTransparency and reporting
Samuel H. King; Claudia L. Driscoll; David B. Li; Daniel Guo; Aditi T. Merchant; Garyk Brixi; Max E. Wilkinson; Brian L. Hie · bioRxiv · 2025-09-17
Many important biological functions arise not from single genes, but from complex interactions encoded by entire genomes. Genome language models have emerged as a promising strategy for designing biological systems, but their ability to generate functional sequences at the scale of whole genomes has remained untested. Here, we report the first generative design of viable bacteriophage genomes. We leveraged frontier genome language models, Evo 1 and Evo 2, to generate whole-genome sequences with realistic genetic architectures and desirable host tropism, using the lytic phage ΦX174 as our design template. Experimental testing of AI-generated genomes yielded 16 viable phages with substantial evolutionary novelty. Cryo-electron microscopy revealed that one of the generated phages utilizes an evolutionarily distant DNA packaging protein within its capsid. Multiple phages demonstrate higher fitness than ΦX174 in growth competitions and in their lysis kinetics. A cocktail of the generated phages rapidly overcomes ΦX174-resistance in three E. coli strains, demonstrating the potential utility of our approach for designing phage therapies against rapidly evolving bacterial pathogens. This work provides a blueprint for the design of diverse synthetic bacteriophages and, more broadly, lays a foundation for the generative design of useful living systems at the genome scale.
Bahrad A. Sokhansanj; Gail L. Rosen · Human Genetics · 2025-09-16
Genome Language Models (GLMs) represent a transformative convergence of artificial intelligence (AI) and genomics, offering unprecedented capabilities for biological discovery, healthcare innovation, and therapeutic design applications. However, these powerful tools create novel regulatory challenges that existing frameworks—whether AI governance or genomic privacy protections—cannot adequately address alone. This paper examines the critical regulatory gaps emerging at this intersection, highlighting tensions between AI principles that favor broad data access and genomic governance that demands stringent privacy protections and informed consent. We analyze how GLMs challenge conventional regulatory approaches as they pertain to applications in disease risk prediction, international research collaboration, and open-source model distribution. We propose a multilayered governance framework that combines policy innovations such as regulatory sandboxes and certification frameworks with technical solutions for privacy preservation and model interpretability. By developing adaptive governance strategies that bridge AI and genomic regulation, we can enable responsible GLM innovation while safeguarding individual rights, promoting equity, and addressing emerging biosecurity concerns in this rapidly evolving field.
Filippa Lentzos; Hailey Wingo; Jez Littlewood; Alberto Muti · Bulletin of the Atomic Scientists · 2025-09-15
A recent project asked what the telltale signs of a bioweapons program would be? There are gaps in the information members of the global bioweapons treaty have access to, but new approaches and technologies can help close these.
Attribution and forensicsBioweapons historyInternational governance
Henry E. Miller; Matthew Greenig; Benjamin Tenmann; Bo Wang · bioRxiv · 2025-09-07
Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of benchmarks evaluating agent capability in multi-step data analysis workflows or in solving the machine learning (ML) challenges central to AI-driven therapeutics development, such as perturbation response modeling or drug toxicity prediction. We introduce BioML-bench, the first benchmarking suite for evaluating AI agents on end-to-end biomedical ML tasks. BioML-bench spans four domains (protein engineering, single-cell omics, biomedical imaging, and drug discovery) with tasks that require agents to parse a task description, build a pipeline, implement models, and submit predictions graded by established metrics (e.g., AUROC, Spearman). We evaluate four open-source agents: two biomedical specialists (STELLA, Biomni) and two generalists (AIDE, MLAgentBench). On average, agents underperform relative to human baselines, and biomedical specialization does not confer a consistent advantage. We also found that agents which employed more diverse ML strategies more often tended to score highest, suggesting that architecture and scaffolding may be stronger determinants of performance. These findings underscore both the potential and current limits of agentic systems for biomedical ML, and highlight the need for systematic, reproducible evaluations. BioML-bench is provided open-source at github.com/science-machine/biomlbench.
Tegan McCaslin; Jide Alaga; Samira Nedungadi; Seth Donoughe; Tom Reed; Rishi Bommasani; Chris Painter; Luca Righetti · arXiv · 2025-09-03
Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducted, and how their results inform decisions - is crucial for building trust in AI development. We propose STREAM (A Standard for Transparently Reporting Evaluations in AI Model Reports), a standard to improve how model reports disclose evaluation results, initially focusing on chemical and biological (ChemBio) benchmarks. Developed in consultation with 23 experts across government, civil society, academia, and frontier AI companies, this standard is designed to (1) be a practical resource to help AI developers present evaluation results more clearly, and (2) help third parties identify whether model reports provide sufficient detail to assess the rigor of the ChemBio evaluations. We concretely demonstrate our proposed best practices with "gold standard" examples, and also provide a three-page reporting template to enable AI developers to implement our recommendations more easily.
Jonathan Feldman; Tal Feldman · arXiv · 2025-08-30
Recent advances in generative biology have enabled the design of novel proteins, creating significant opportunities for drug discovery while also introducing new risks, including the potential development of synthetic bioweapons. Existing biosafety measures primarily rely on inference-time filters such as sequence alignment and protein-protein interaction (PPI) prediction to detect dangerous outputs. In this study, we evaluate the performance of three leading PPI prediction tools: AlphaFold 3, AF3Complex, and SpatialPPIv2. These models were tested on well-characterized viral-host interactions, such as those involving Hepatitis B and SARS-CoV-2. Despite being trained on many of the same viruses, the models fail to detect a substantial number of known interactions. Strikingly, none of the tools successfully identify any of the four experimentally validated SARS-CoV-2 mutants with confirmed binding. These findings suggest that current predictive filters are inadequate for reliably flagging even known biological threats and are even more unlikely to detect novel ones. We argue for a shift toward response-oriented infrastructure, including rapid experimental validation, adaptable biomanufacturing, and regulatory frameworks capable of operating at the speed of AI-driven developments.
Yanda Chen; Mycal Tucker; Nina Panickssery; Tony Wang; Francesco Mosconi; Anjali Gopal; Carson Denison; Linda Petrini; Jan Leike; Ethan Perez; Mrinank Sharma · Alignment Science Blog · 2025-08-19
We experimented with removing harmful information about chemical, biological, radiological and nuclear (CBRN) weapons from our models' pretraining data. We identified harmful content using a classifier and pretrained models from scratch on the filtered dataset. This approach reduced the model's accuracy on a harmful-capabilities evaluation by 33% relative compared to random baseline performance, while preserving its beneficial capabilities.
Barbara Del Castello; Henry H. Willis · 2025-08-11
Rapid developments in life science technologies could have major consequences for the field of biosecurity. In this report, the authors propose and demonstrate a method to measure changes in threats associated with these emerging technologies. The method relies on expert elicitation and technology scoring techniques to understand how biotechnology maturation and diffusion could lower barriers for nonstate actors to create biological agents.
Open-weight AI systems offer unique benefits, including enhanced transparency, open research, and decentralized access. However, they are vulnerable to tampering attacks which can efficiently elicit harmful behaviors by modifying weights or activations. Currently, there is not yet a robust science of open-weight model risk management. Existing safety fine-tuning methods and other post-training techniques have struggled to make LLMs resistant to more than a few dozen steps of adversarial fine-tuning. In this paper, we investigate whether filtering text about dual-use topics from training data can prevent unwanted capabilities and serve as a more tamper-resistant safeguard. We introduce a multi-stage pipeline for scalable data filtering and show that it offers a tractable and effective method for minimizing biothreat proxy knowledge in LLMs. We pretrain multiple 6.9B-parameter models from scratch and find that they exhibit substantial resistance to adversarial fine-tuning attacks on up to 10,000 steps and 300M tokens of biothreat-related text -- outperforming existing post-training baselines by over an order of magnitude -- with no observed degradation to unrelated capabilities. However, while filtered models lack internalized dangerous knowledge, we find that they can still leverage such information when it is provided in context (e.g., via search tool augmentation), demonstrating a need for a defense-in-depth approach. Overall, these findings help to establish pretraining data curation as a promising layer of defense for open-weight AI systems.
Yuxuan Zhu; Tengjun Jin; Yada Pruksachatkun; Andy Zhang; Shu Liu; Sasha Cui; Sayash Kapoor; Shayne Longpre; Kevin Meng; Rebecca Weiss; Fazl Barez; Rahul Gupta; Jwala Dhamala; Jacob Merizian; Mario Giulianelli; Harry Coppock; Cozmin Ududec; Jasjeet Sekhon; Jacob Steinhardt; Antony Kellermann; Sarah Schwettmann; Matei Zaharia; Ion Stoica; Percy Liang; Daniel Kang · arXiv · 2025-08-07
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench Verified uses insufficient test cases, while TAU-bench counts empty responses as successful. Such issues can lead to under- or overestimation of agents' performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces the performance overestimation by 33%.
Georgia Adamson and Gregory C. Allen break down emerging AI-enabled bioterrorism risks and how policymakers can strengthen U.S. biosecurity for the AI age.
Forty years ago, Australia proactively responded to the proliferation of chemical weapons by convening the highly successful Australia Group. That same proactive leadership is needed now to counter emerging AI-enabled chemical, biological, radiological and nuclear (CBRN) threats.
Third-party assessments can be conducted on frontier models to confirm evaluations or claims on critical safety capabilities and mitigations. In appropriate contexts, these assessments may help to confirm or build confidence in safety claims, add robust methodological independence, and supplement expertise. This report outlines practices and approaches among Frontier Model Forum (FMF) firms for implementing, where appropriate, rigorous, secure, and fit-for-purpose third-party assessments.
Douglas Densmore; Chris Isaac; Nicole Wheeler; Jaime M. Yassif · 2025-08-01
Current biosecurity frameworks, such as those used by the International Gene Synthesis Consortium, rely on the ability to compare DNA synthesis orders to known sequences to determine if they may be concerning. However, as biodesign tools—especially those powered by artificial intelligence (AI)—begin exploring novel biological designs that deviate substantially from organisms found in nature, traditional screening methods are likely to struggle to interpret these novel designs. This makes it difficult for DNA synthesis providers and other service providers, who support bioscience and biotechnology research and development, to detect potential threats, as it involves assessing the risks of entirely new designs that do not resemble known organisms or toxins. To address these challenges, NTI | bio partnered with Lattice Automation to design and pilot a standard for capturing and transmitting metadata—such as design provenance, editing history, and intended use—alongside DNA or protein sequences. The additional context provided by this standard, known as the Biodesign Metadata Exchange (BMDE), can help biosecurity decision-makers assess risks more effectively by increasing their understanding of not only the sequence itself but also the design process behind it.
Council on Strategic Risks · Council on Strategic Risks · 2025-07-31
A backcasting exercise conducted with AIxBio experts and follow-up engagement with key stakeholders yielded a practical framework and tangible recommendations for the US government to achieve the best-case scenario for societal resilience to AIxBio risks over the next five years.
Frontier Model Forum · Frontier Model Forum · 2025-07-30
Frontier AI presents transformative opportunities within the biological sciences, including the potential to rapidly accelerate beneficial research discoveries and development. However, the dual-use nature of these technologies may also introduce novel risks. One potential harm involves the misuse of legitimately accessed frontier AI systems by malicious actors to create biological threats, such as a bioweapon. As frontier AI capabilities advance, it is crucial to develop robust risk management practices that enable society to harness the benefits of AI in biology while proactively managing its most severe potential risks.
In light of this challenge, frontier model developers have committed to researching, implementing, validating, and sharing mitigation measures (also known as safeguards) to prevent the misuse of their models. This issue brief presents a preliminary taxonomy of safeguards designed to reduce the risk of biological misuse stemming from access to frontier AI models. Drawing from discussions with experts within the Frontier Model Forum (FMF) and the broader biosafety and biosecurity communities, this brief outlines the current landscape of AI-bio misuse safeguards, identifies potential future approaches to mitigations, and underscores the importance of implementing societal-level measures as a complement to technical safeguards.
Engineering National Academies of Sciences · 2025-07-25
Current policies on dual-use research of concern (DURC) and pathogens with enhanced pandemic potential (PEPP) typically focus on physical laboratory work. In light of the fast-evolving advances in artificial intelligence and computational modeling, these frameworks do not effectively inform risk and benefit evaluation and assessment related to the information and resources generated from computational studies.
To address these concerns, the National Academies of Sciences, Engineering, and Medicine convened a workshop sponsored by the National Science Foundation on April 3-4, 2025, to explore the benefits and biosecurity risks of communicating and publishing biological research using in silico modeling and computational approaches. The workshop brought together multi-sectoral experts to discuss current policies and safeguards related to DURC and PEPP, as well as lessons learned, and considered the challenges and opportunities for promoting the benefits of computational and AI-driven approaches in biology while mitigating potential biosecurity risks. This publication summarizes the presentations and discussion of the workshop, including suggestions from participants on tiered oversight approaches, early-stage risk evaluations and assessment, and incentivizing norms through training and publication standards.
CapabilitiesData governanceDual-use researchPolicyTransparency and reporting
More than 35 leading experts highlight the risks posed by rapidly advancing capabilities at the convergence of AI and the life sciences (AIxBio capabilities) and call on governments, industry, the scientific community, and funders to take action to safeguard this technology.
Joe Needham; Giles Edkins; Govind Pimpale; Henning Bartsch; Marius Hobbhahn · arXiv · 2025-07-16
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable benchmarks for deployment and governance decisions. We investigate whether frontier language models can accurately classify transcripts based on whether they originate from evaluations or real-world deployment, a capability we call evaluation awareness. To achieve this, we construct a diverse benchmark of 1,000 prompts and transcripts from 61 distinct datasets. These span public benchmarks (e.g., MMLU, SWEBench), real-world deployment interactions, and agent trajectories from scaffolding frameworks (e.g., web-browsing agents). Frontier models clearly demonstrate above-random evaluation awareness (Gemini-2.5-Pro reaches an AUC of $0.83$), but do not yet surpass our simple human baseline (AUC of $0.92$). Furthermore, both AI models and humans are better at identifying evaluations in agentic settings compared to chat settings. Additionally, we test whether models can identify the purpose of the evaluation. Under multiple-choice and open-ended questioning, AI models far outperform random chance in identifying what an evaluation is testing for. Our results indicate that frontier models already exhibit a substantial, though not yet superhuman, level of evaluation-awareness. We recommend tracking this capability in future models.
AI agentsEvaluationsJailbreaks and red-teamingLLMs
Bridget Williams; Luca Righetti; Josh Rosenberg; Rebecca Ceppas de Castro; Rhiannon Britt; Emily Soice; Alvaro Morales; Jon Sanders; Seth Donoughe; James Black; Ezra Karger; Philip E Tetlock · 2025-07-01
Capabilities of large language models (LLMs) on several biological benchmarks have prompted excitement about their usefulness for beneficial research, but also concern about potential biosecurity risks. We recruited 46 subject-matter experts in biology and biosecurity, and 22 generalist forecasters to estimate the risks of growing LLM capabilities. The median expert predicted a 0.3% baseline annual risk of a human-caused epidemic that causes 100,000 deaths. This estimate then rose to 1.5% conditional on several hypothetical LLM capabilities, including matching the performance of a top performing team of virologists on a virology troubleshooting test. Given this finding, we conducted a baselining study and found that LLMs have already crossed this performance threshold. The median respondent thought that this would not happen until after 2030. More encouragingly, experts reduced their risk forecast close to baseline (0.4%) conditional on the adoption of LLM safeguards and mandatory nucleic acid screening.
Cindy S. Groff-Vindman; Benjamin D. Trump; Christopher L. Cummings; Madison Smith; Alexander J. Titus; Ken Oye; Valentina Prado; Eyup Turmus; Igor Linkov · npj Biomedical Innovations · 2025-07-01
The convergence of AI and synthetic biology is revolutionizing biological discovery and engineering. This manuscript examines how AI-driven tools accelerate bioengineering workflows, unlocking innovations in medicine, agriculture, and sustainability. It also addresses dual-use risks, governance gaps, and ethical dilemmas posed by these advancements, proposing strategies for oversight and updated regulations. By exploring opportunities and risks, this work highlights the transformative potential of AI-driven synthetic biology and pathways for responsible development.
Allison Berke; Forrest W. Crawford; Toby Webster; James Smith; Sana Zakaria; Sella Nevo · 2025-06-30
The authors of this paper assess the link between biological data and the capabilities of artificial intelligence models trained on large volumes of biological data. They describe anticipated impacts of new biological data sources and the potentially dangerous capabilities of certain types of biological data. The authors recommend strategies to limit such capabilities arising from biological data, including recommending options for governance.
As deterrence by denial is increasingly explored as a framework to reduce vulnerabilities to and dissuade deliberate biological risks, more actors are examining the technological toolkit that would make this approach a reality. Through an application of deterrence theory to deliberate biological threats and an overview of specific technical abilities, this paper posits that a credible deterrence by denial strategy to address biological threats is maintained through four key pillars: detection, response, mitigation, and attribution. A cursory overview of the technologies comprising each pillar, such as early warning, disease surveillance, rapid diagnostics, genomic sequencing, and medical countermeasure development, provides the basis for what a robust framework would require and could possibly achieve. Additionally, it examines the considerations and challenges for successful implementation and further questions regarding the unique interactions between deterrence signaling, necessary transparencies, and an ever-complicated information environment. States have a wide range of options set for reducing biological risks—many of which, if used together, can confer a robust and credible deterrent against potential adversaries’ use of biological weapons. The paper concludes by situating the deterrence by denial framework into the larger discourse of national and global strategies to address biological risks, arguing that the core component across frameworks is the readiness and availability of key capabilities outlined across the four pillars outlined above.
Bioterrorism and accidental laboratory leaks pose significant threats to global health and security. Advancements in biotechnology and artificial intelligence (AI) have intensified concerns about potential increases in these risks. However, quantitative risk estimates have been lacking, hindering effective policy-making and risk mitigation efforts. Here we modeled current and future pandemic risks from bioterrorism and lab leaks, with a focus on the potential impact of general-purpose radically transformative AI (GPRTAI) systems. Our results suggest that the current risk of a pandemic from lab leaks (0.03/year; 95% CrI 0.002–0.3) is approximately 2000 times greater than that from bioterrorism (0.00002/year; 95% CrI 0.000002–0.0001). With GPRTAI, we project bioterrorism risk could increase by 80 times (to 0.001/year; 95% CrI 0.00007–0.02), while lab leak risk might increase by 6 times (to 0.3/year; 95% CrI 0.01–3), assuming no new safeguards are implemented.
As our models grow more capable in biology, we’re layering in safeguards and partnering with global experts, including hosting a biodefense summit this July.
Frontier Model Forum · Frontier Model Forum · 2025-06-13
The Frontier Model Forum (FMF) has a founding mandate to advance the science of frontier AI safety and security. As part of that effort, today we are pleased to share an update on our support for novel research at the intersection of AI and biological sciences.
Kevin L. Wei; Patricia Paskov; Sunishchal Dev; Michael J. Byun; Anka Reuel; Xavier Roberts-Gaal; Rachel Calcott; Evie Coxon; Chinmay Deshpande · arXiv · 2025-06-09
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist towards this end. Human performance baselines are vital for the machine learning community, downstream users, and policymakers to interpret AI evaluations. Models are often claimed to achieve "super-human" performance, but existing baselining methods are neither sufficiently rigorous nor sufficiently well-documented to robustly measure and assess performance differences. Based on a meta-review of the measurement theory and AI evaluation literatures, we derive a framework with recommendations for designing, executing, and reporting human baselines. We synthesize our recommendations into a checklist that we use to systematically review 115 human baselines (studies) in foundation model evaluations and thus identify shortcomings in existing baselining methods; our checklist can also assist researchers in conducting human baselines and reporting results. We hope our work can advance more rigorous AI evaluation practices that can better serve both the research community and policymakers. Data is available at: https://github.com/kevinlwei/human-baselines
Sebastian Rivera; Moritz Hanke; Samuel Curtis; Neil Cherian; Anthony Gitter; Jeffrey Gray; Andrew Hebbeler; Stephen McCarthy; David Nannemann; Claire Qureshi; Brian Weitzner; Ian Haydon · OSF · 2025-06-03
This report presents Recommended Actions from the January 2025 Responsible Biodesign Workshop, which convened leading experts across AI-enabled biomolecular design and biosecurity policy. Building on existing community commitments for the Responsible Development of AI for Protein Design, the Recommended Actions aim to guide scientists, policy practitioners, and funding bodies in ensuring safe and beneficial development of AI-enabled biomolecular design tools. The Recommended Actions focus on advancing AI-Resilient nucleic acid synthesis security screening, assessing the risk-benefit landscape of biomolecular design capabilities, and building fora for sustained engagement between scientists and policy practitioners.
Zaixi Zhang; Zhenghong Zhou; Ruofan Jin; Le Cong; Mengdi Wang · arXiv · 2025-05-28
DNA, encoding genetic instructions for almost all living organisms, fuels groundbreaking advances in genomics and synthetic biology. Recently, DNA Foundation Models have achieved success in designing synthetic functional DNA sequences, even whole genomes, but their susceptibility to jailbreaking remains underexplored, leading to potential concern of generating harmful sequences such as pathogens or toxin-producing genes. In this paper, we introduce GeneBreaker, the first framework to systematically evaluate jailbreak vulnerabilities of DNA foundation models. GeneBreaker employs (1) an LLM agent with customized bioinformatic tools to design high-homology, non-pathogenic jailbreaking prompts, (2) beam search guided by PathoLM and log-probability heuristics to steer generation toward pathogen-like sequences, and (3) a BLAST-based evaluation pipeline against a curated Human Pathogen Database (JailbreakDNABench) to detect successful jailbreaks. Evaluated on our JailbreakDNABench, GeneBreaker successfully jailbreaks the latest Evo series models across 6 viral categories consistently (up to 60\% Attack Success Rate for Evo2-40B). Further case studies on SARS-CoV-2 spike protein and HIV-1 envelope protein demonstrate the sequence and structural fidelity of jailbreak output, while evolutionary modeling of SARS-CoV-2 underscores biosecurity risks. Our findings also reveal that scaling DNA foundation models amplifies dual-use risks, motivating enhanced safety alignment and tracing mechanisms. Our code is at https://github.com/zaixizhang/GeneBreaker.
Biological ToolsEvaluationsJailbreaks and red-teamingRisk assessmentsVirology
Svetlana P. Ikonomova; Bruce J. Wittmann; Fernanda Piorino; David J. Ross; Samuel W. Schaffter; Olga Vasilyeva; Eric Horvitz; James Diggans; Elizabeth A. Strychalski; Sheng Lin-Gibson; Geoffrey J. Taghon · bioRxiv · 2025-05-16
Advances in machine learning are providing new abilities for engineering biology, promising leaps forward with beneficial applications. At the same time, these advances raise concerns about biosecurity. Recently, Wittmann et al. described an in silico pipeline of generative AI tools to demonstrate how amino acid sequences encoding sequences of concern (SOCs) could be reformulated as synthetic homologs that may evade detection by biosecurity screening software (BSS) used by nucleic acid synthesis providers. While the discovered vulnerability has been mitigated, its true severity remains unclear, as the study was performed without experimental testing of the synthetic homologs. Here, we present a testing, evaluation, validation, and verification (TEVV) framework for AI-assisted protein design (AIPD), using safe proteins as SOC proxies. Our findings indicate AIPD can generate synthetic homologs whose predicted structures are similar to a native template, but without necessarily retaining activity. We determine that current AIPD systems are not yet powerful enough to reliably rewrite the sequence of a given protein, while both maintaining activity and evading detection by BSS. We further conclude that TEVV of generated sequences requires significant investment of time, technical skill, and resources. We offer our framework and test proteins as a general benchmark of AIPD capability.
Rahul K. Arora; Jason Wei; Rebecca Soskin Hicks; Preston Bowman; Joaquin Quiñonero-Candela; Foivos Tsimpourlas; Michael Sharman; Meghan Shah; Andrea Vallone; Alex Beutel; Johannes Heidecke; Karan Singhal · arXiv · 2025-05-13
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user or healthcare professional. Responses are evaluated using conversation-specific rubrics created by 262 physicians. Unlike previous multiple-choice or short-answer benchmarks, HealthBench enables realistic, open-ended evaluation through 48,562 unique rubric criteria spanning several health contexts (e.g., emergencies, transforming clinical data, global health) and behavioral dimensions (e.g., accuracy, instruction following, communication). HealthBench performance over the last two years reflects steady initial progress (compare GPT-3.5 Turbo's 16% to GPT-4o's 32%) and more rapid recent improvements (o3 scores 60%). Smaller models have especially improved: GPT-4.1 nano outperforms GPT-4o and is 25 times cheaper. We additionally release two HealthBench variations: HealthBench Consensus, which includes 34 particularly important dimensions of model behavior validated via physician consensus, and HealthBench Hard, where the current top score is 32%. We hope that HealthBench grounds progress towards model development and applications that benefit human health.
This study systematically evaluates 27 frontier Large Language Models on eight biology benchmarks spanning molecular biology, genetics, cloning, virology, and biosecurity. Models from major AI developers released between November 2022 and April 2025 were assessed through ten independent runs per benchmark. The findings reveal dramatic improvements in biological capabilities. Top model performance increased more than 4-fold on the challenging text-only subset of the Virology Capabilities Test over the study period, with OpenAI's o3 now performing twice as well as expert virologists. Several models now match or exceed expert-level performance on other challenging benchmarks, including the biology subsets of GPQA and WMDP and LAB-Bench CloningScenarios. Contrary to expectations, chain-of-thought did not substantially improve performance over zero-shot evaluation, while extended reasoning features in o3-mini and Claude 3.7 Sonnet typically improved performance as predicted by inference scaling. Benchmarks such as PubMedQA and the MMLU and WMDP biology subsets exhibited performance plateaus well below 100%, suggesting benchmark saturation and errors in the underlying benchmark data. The analysis highlights the need for more sophisticated evaluation methodologies as AI systems continue to advance.
Frontier Model Forum · Frontier Model Forum · 2025-05-12
WORKSTREAM AI-Biosafety Workstream DOWNLOAD The rapid advancement of frontier AI presents transformative opportunities in the biological domain. For example, frontier AI systems may accelerate the discovery of new medical treatments, optimize biomanufacturing processes, or facilitate the development of novel biocatalysts. However, given that frontier AI capabilities are often dual-use, they may also heighten risks from […]
Jaspreet Pannu; Doni Bloomfield; Robert MacKnight; Moritz S. Hanke; Alex Zhu; Gabe Gomes; Anita Cicero; Thomas V. Inglesby · PLOS Computational Biology · 2025-05-08
As a result of rapidly accelerating artificial intelligence (AI) capabilities, multiple national governments and multinational bodies have launched efforts to address safety, security and ethics issues related to AI models. One high priority among these efforts is the mitigation of misuse of AI models, such as for the development of chemical, biological, nuclear or radiological (CBRN) threats. Many biologists have for decades sought to reduce the risks of scientific research that could lead, through accident or misuse, to high-consequence disease outbreaks. Scientists have carefully considered what types of life sciences research have the potential for both benefit and risk (dual use), especially as scientific advances have accelerated our ability to engineer organisms. Here we describe how previous experience and study by scientists and policy professionals of dual-use research in the life sciences can inform dual-use capabilities of AI models trained using biological data. Of these dual-use capabilities, we argue that AI model evaluations should prioritize addressing those which enable high-consequence risks (i.e., large-scale harm to the public, such as transmissible disease outbreaks that could develop into pandemics), and that these risks should be evaluated prior to model deployment so as to allow potential biosafety and/or biosecurity measures. While biological research is on balance immensely beneficial, it is well recognized that some biological information or technologies could be intentionally or inadvertently misused to cause consequential harm to the public. AI-enabled life sciences research is no different. Scientists’ historical experience with identifying and mitigating dual-use biological risks can thus help inform new approaches to evaluating biological AI models. Identifying which AI capabilities pose the greatest biosecurity and biosafety concerns is necessary in order to establish targeted AI safety evaluation methods, secure these tools against accident and misuse, and avoid impeding immense potential benefits.
Michael J. D. Vermeer; Emily Lathrop; Alvin Moon · 2025-05-06
In this report, RAND researchers examine the threat of human extinction posed by artificial intelligence (AI). Using a scenario-based analysis to evaluate three technologies — nuclear weapons, pathogens, and geoengineering — they found that it would be immensely challenging for AI to create an extinction threat, although they could not rule out the possibility. They identify risk indicators and make recommendations to mitigate future risk.
Ying-Chiang Jeffrey Lee; Bria Persaud; Barbara Del Castello; Allison Berke; Gustavs Zilgalvis · 2025-04-24
In this paper, the authors provide an overview of 15 cloud lab organizations—research labs for remote execution of experiments—around the world, including facility size, number of instruments, location, and scientific focus. They also discuss how the automation and remote capabilities of cloud labs could enable bad actors in the development and proliferation of chemical or biological weapons.
Background: Recent advances in synthetic biology have raised new challenges in biosecurity and biorisk management (BRM), particularly in high-containment laboratories and other critical research settings.Objective: This article explores the use of artificial intelligence (AI) agents to automate BRM tasks, thus enhancing both safety and efficiency in these sensitive environments.Methods: We propose an integrated system combining machine learning models for risk assessment with embodied AI agents capable of executing physical containment and other tasks. Specifically, AI algorithms can be employed for predictive risk modeling, anomaly detection, and real-time decision-making, while embodied AI agents can serve to carry out operations that would otherwise expose humans to hazardous biological agents and other hazards.Results: Combined, these two technologies are known as a specialized AI agent for BRM. This dual approach aims to reduce the cognitive and physical burdens on human personnel, minimize human error, and ensure consistent adherence to BRM protocols. This combination optimizes laboratory workflows and introduces a robust layer of redundancy in managing biological risks. Conclusions: Here we present a novel and promising step forward in augmenting the safety of high-containment laboratories through implementation of specialized AI agents; thus, contributing to more resilient BRM frameworks.
AI agentsMitigationsRisk assessmentsSynthetic biology
Jasper Götting; Pedro Medeiros; Jon G Sanders; Nathaniel Li; Long Phan; Karam Elabd; Lennart Justen; Dan Hendrycks; Seth Donoughe · 2025-04-22
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of 322 multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of 22.1% on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI’s o3, reaches 43.8% accuracy, outperforming 94% of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.
Frontier Model Forum · Frontier Model Forum · 2025-04-22
This report discusses emerging industry practices for implementing Frontier Capability Assessments and is intended as a field-wide perspective, not a description of any single organization’s methodologies. As the science of these assessments is rapidly advancing, this overview represents a snapshot of current practices. This report focuses on the main assessment techniques currently in use or being developed, rather than providing an exhaustive list of all possible approaches for demonstrating that a model poses acceptably low risk.
Alex Friedland · Center for Security and Emerging Technology · 2025-04-21
Opposing narratives around AI for biotechnology raise the question: how are biotech researchers actually using AI in published research? CSET’s Steph Batalis, Catherine Aiken, and James Dunham explored this question by leveraging CSET’s merged academic corpus, enriched publication metadata, and research clusters.
Lalitha Sundaram; Jessica Bland; Tom Hobson · 2025-04-15
The intersection of AI and engineering biology offers a unique opportunity to address pressing societal issues, but also poses a potential threat to global biosecurity. Is AI lowering barriers to biological weapons development? What are the risks and how likely are they to emerge? Can AI be used as a tool to aid biosecurity as well as threaten it? Researchers from the Centre for the Study of Existential Risk discuss.
The extent to which foundation models can disclose novel chemical, biological, radiation, and nuclear (CBRN) threats to expert users is unclear due to a lack of test cases. I leveraged the unique opportunity presented by an upcoming publication describing a novel catastrophic biothreat - "Technical Report on Mirror Bacteria: Feasibility and Risks" - to conduct a small controlled study before it became public. Graduate-trained biologists tasked with predicting the consequences of releasing mirror E. coli showed no significant differences in rubric-graded accuracy using Claude Sonnet 3.5 new (n=10) or web search only (n=2); both groups scored comparably to a web baseline (28 and 43 versus 36). However, Sonnet reasoned correctly when prompted by a report author, but a smaller model, Haiku 3.5, failed even with author guidance (80 versus 5). These results suggest distinct stages of model capability: Haiku is unable to reason about mirror life even with threat-aware expert guidance (Stage 1), while Sonnet correctly reasons only with threat-aware prompting (Stage 2). Continued advances may allow future models to disclose novel CBRN threats to naive experts (Stage 3) or unskilled users (Stage 4). While mirror life represents only one case study, monitoring new models' ability to reason about privately known threats may allow protective measures to be implemented before widespread disclosure.
CapabilitiesEvaluationsJailbreaks and red-teamingMitigationsThreat Assessment
Frontier Model Forum · Frontier Model Forum · 2025-03-18
WORKSTREAM AI-Biosafety Workstream DOWNLOAD Frontier AI models and systems are particularly promising for advancing medicine and public health. At the same time, their knowledge of biology and ability to reason about biological concepts may also be misused in ways that pose significant risks to public safety and security. To manage those risks, many frontier AI […]
Engineering National Academies of Sciences · 2025-03-14
This report represents the outcome of a technical assessment by the committee of the capabilities of AI-enabled biological tools, which can be utilized in conjunction with the framework developed in the National Academies’ 2018 report Biodefense in the Age of Synthetic Biology (hereafter referred to as “the 2018 framework”) for a full consideration of the threat landscape and risk assessment. The 2018 framework comprehensively identified the different risk factors associated with synthetic biology capabilities but can be applied more broadly to other biotechnologies.
Bria Persaud; Ying-Chiang Jeffrey Lee; Jordan Despanie; Helin Hernandez; Henry Alexander Bradley; Sarah L. Gebauer; Greg McKelvey · 2025-02-28
The authors of this working paper developed a proof-of-concept automated grader and used it to assess large language models' abilities to answer knowledge-based questions and generate protocols that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats.
Thomas Hayes; Roshan Rao; Halil Akin; Nicholas J. Sofroniew; Deniz Oktay; Zeming Lin; Robert Verkuil; Vincent Q. Tran; Jonathan Deaton; Marius Wiggert; Rohil Badkundri; Irhum Shafkat; Jun Gong; Alexander Derry; Raul S. Molina; Neil Thomas; Yousuf A. Khan; Chetan Mishra; Carolyn Kim; Liam J. Bartie; Matthew Nemeth; Patrick D. Hsu; Tom Sercu; Salvatore Candido; Alexander Rives · Science · 2025-02-21
More than 3 billion years of evolution have produced an image of biology encoded into the space of natural proteins. Here, we show that language models trained at scale on evolutionary data can generate functional proteins that are far away from known proteins. We present ESM3, a frontier multimodal generative language model that reasons over the sequence, structure, and function of proteins. ESM3 can follow complex prompts combining its modalities and is highly responsive to alignment to improve its fidelity. We have prompted ESM3 to generate fluorescent proteins. Among the generations that we synthesized, we found a bright fluorescent protein at a far distance (58% sequence identity) from known fluorescent proteins, which we estimate is equivalent to simulating 500 million years of evolution.
Nicole E. Wheeler · Frontiers in Bioengineering and Biotechnology · 2025-02-05
The integration of artificial intelligence (AI) in protein design presents unparalleled opportunities for innovation in bioengineering and biotechnology. However, it also raises significant biosecurity concerns. This review examines the changing landscape of bioweapon risks, the dual-use potential of AI-driven bioengineering tools, and the necessary safeguards to prevent misuse while fostering innovation. It highlights emerging policy frameworks, technical safeguards, and community responses aimed at mitigating risks and enabling responsible development and application of AI in protein design.
Richard Moulange; Tina Wünn; Cassidy Nelson · 2025-01-08
The European Union AI Act (EU AI Act) is a regulatory framework that aims to govern the development, deployment, and use of AI systems within the EU. It focuses on large-scale AI systems classified as general-purpose AI (GPAI). It is currently undetermined whether AI-enabled biological tools are subject to regulation under the EU AI Act. In this report, we apply the relevant GPAI and systemic risk definitions in the EU AI Act to 50 AI biological models (both narrow AI-enabled biological tools and biological foundation models) and examine the implications.
EvaluationsRisk assessmentsThe Centre for Long-Term Resilience
The perception that the convergence of biological engineering and artificial intelligence (AI) could enable increased biorisk has recently drawn attention to the governance of biotechnology and AI. The 2023 Executive Order, Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, requires an assessment of how AI can increase biorisk. Within this perspective, quantitative and qualitative frameworks for evaluating biorisk are presented. Both frameworks are exercised using notional scenarios and their benefits and limitations are then discussed. Finally, the perspective concludes by noting that assessment and evaluation methodologies must keep pace with advances of AI in the life sciences.
FrameworkPolicyRisk assessmentsThreat AssessmentUnited States
Sebastian Rivera; Rebecca Mackelprang · 2025-01-01
The intersection of artificial intelligence (AI) with the life sciences is rapidly reshaping what is possible across myriad sectors, including engineering biology. These possibilities include beneficial applications, like vaccine development, and also potentially harmful applications, like pandemic pathogen design. This white paper provides commentary and recommendations on risk mitigation strategies that balance the dual-use nature of AI enabled biological design tools.
Biological ToolsEngineering Biology Research Consortium
To better understand the nexus between ICTs and the biological field, this paper begins with an outline of some of the benefits introduced by the integration of advanced ICT in biological research and development. It then proposes a definition of the concept of ‘cyberbiosecurity’ and proceeds to outline some of the key risks at this nexus. The paper also considers how this cross-regime issue is being addressed in relevant multilateral fora tasked with ensuring international peace and security, notably at the United Nations; as well as nationally by States Parties to the Biological Weapons Convention (BWC).
Yoshua Bengio; Sören Mindermann; Daniel Privitera; Tamay Besiroglu; Rishi Bommasani; Stephen Casper; Yejin Choi; Philip Fox; Ben Garfinkel; Danielle Goldfarb; Hoda Heidari; Anson Ho; Sayash Kapoor; Leila Khalatbari; Shayne Longpre; Sam Manning; Vasilios Mavroudis; Mantas Mazeika; Julian Michael; Jessica Newman; Kwan Yee Ng; Chinasa T. Okolo; Deborah Raji; Girish Sastry; Elizabeth Seger; Theodora Skeadas; Tobin South; Emma Strubell; Florian Tramèr; Lucia Velasco; Nicole Wheeler; Daron Acemoglu; Olubayo Adekanmbi; David Dalrymple; Thomas G. Dietterich; Edward W. Felten; Pascale Fung; Pierre-Olivier Gourinchas; Fredrik Heintz; Geoffrey Hinton; Nick Jennings; Andreas Krause; Susan Leavy; Percy Liang; Teresa Ludermir; Vidushi Marda; Helen Margetts; John McDermid; Jane Munga; Arvind Narayanan; Alondra Nelson; Clara Neppel; Alice Oh; Gopal Ramchurn; Stuart Russell; Marietje Schaake; Bernhard Schölkopf; Dawn Song; Alvaro Soto; Lee Tiedrich; Gaël Varoquaux; Andrew Yao; Ya-Qin Zhang; Olubunmi Ajala; Fahad Albalawi; Marwan Alserkal; Guillaume Avrin; Christian Busch; André Carlos Ponce de Leon Ferreira de Carvalho; Bronwyn Fox; Amandeep Singh Gill; Ahmet Halit Hatip; Juha Heikkilä; Chris Johnson; Gill Jolly; Ziv Katzir; Saif M. Khan; Hiroaki Kitano; Antonio Krüger; Kyoung Mu Lee; Dominic Vincent Ligot; José Ramón López Portillo; Oleksii Molchanovskyi; Andrea Monti; Nusu Mwamanzi; Mona Nemer; Nuria Oliver; Raquel Pezoa Rivera; Balaraman Ravindran; Hammam Riza; Crystal Rugege; Ciarán Seoighe; Jerry Sheehan; Haroon Sheikh; Denise Wong; Yi Zeng · 2025-01-01
A report on the state of advanced AI capabilities and risks – written by 100 AI experts including representatives nominated by 33 countries and intergovernmental organisations.
CapabilitiesInternational governanceRisk assessmentsThreat AssessmentTransparency and reporting
Yana Bromberg; Russ Altman; Michael Imperiale; Eric Horvitz; Monica Dus; Raphael Townshend; Vicky Yao; Todd Treangen; Tessa Alexanian; Erika Szymanski; Jaime Yassif; Rafael Anta; Ariel B. Lindner; Markus Schmidt; James Diggans; Kevin M. Esvelt; Kutubuddin A. Molla; Ryan Phelan; Mengdi Wang; Felicia Wu; Daniela Matias de Carvalho Bittencourt · Rice University · 2025-01-01
Integration of artificial intelligence (AI) and biotechnology (AIxBio) creates revolutionary opportunities for progress in advancing the bioeconomy and addressing health concerns. AI advances promise to greatly accelerate beneficial biological discoveries and innovation and will undoubtedly be one of the deepest contributions of AI to people and society. However, AI methods can also increase risks of accidents and enable malevolent activities aimed at deliberately harmful applications such as bioweapons development. Effective AIxBio governance requires frameworks that enable the great rewards expected from AI in biosciences but that also consider more costly outcomes made possible by AI advances. Recent literature on AIxBio risk management highlights strategies that include tiered access controls, AI auditing mechanisms, and mandatory biological molecule synthesis screening and monitoring. However, many of these potential guardrails have yet to be developed and/or adequately evaluated. In addition to developing practical, technical solutions, it will also be important to develop guidelines and regulations, as well as incentives to follow these, to drive broad implementation of effective risk reduction solutions at the national and international level. Such policies can address significant gaps in national and global governance, but it will also be important to harmonize these approaches to address any regulatory divergence and inconsistencies in risk management across key world players.
Frontier Model Forum · Frontier Model Forum · 2024-12-20
Frontier AI-bio safety evaluations aim to test the biological capabilities and, by extension, the potential biosafety implications of frontier AI. As the science of AI safety evaluations is still nascent, the evaluations themselves can vary widely in both purpose and methodology. As such, a key first step in building out an effective safety evaluation ecosystem […]
Beba Cibralic; Barbara Del Castello; Allison Berke; Aurelia Attal-Juncqua; Joshua Dettman; Alyssa Worland; Jeffrey Lee; Elika Somani; Lukas Berglund; Jay Atanda; Toby Webster · 2024-12-17
This paper provides public comment in response to the Request for Information by the AI Safety Institute concerning the responsible development and use of chemical and/or biological (chem-bio) artificial intelligence models. Chem-bio AI models can help in analyzing, predicting, or generating novel chem-bio sequences, structures, or functions and are becoming increasingly capable and accessible for a wide range of dual-use applications.
Dual-use researchEvaluationsMitigationsNISTRequest for Information (RFI)RAND
Alex John London · Nuclear Threat Initiative · 2024-12-10
This essay collection is designed to encourage the exploration and identification of potential solutions to disincentivize states from developing or using biological weapons. The goal of this collection is to bridge theory and practical policy-relevant approaches to develop new approaches to invigorate international efforts to reduce biological threats.
Aidan Peppin; Anka Reuel; Stephen Casper; Elliot Jones; Andrew Strait; Usman Anwar; Anurag Agrawal; Sayash Kapoor; Sanmi Koyejo; Marie Pellat; Rishi Bommasani; Nick Frosst; Sara Hooker · arXiv · 2024-12-04
To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or systems could increase biorisk and a robust method for testing that threat model. This paper provides an analysis of existing available research surrounding two AI and biorisk threat models: 1) access to information and planning via large language models (LLMs), and 2) the use of AI-enabled biological tools (BTs) in synthesizing novel biological artifacts. We find that existing studies around AI-related biorisk are nascent, often speculative in nature, or limited in terms of their methodological maturity and transparency. The available literature suggests that current LLMs and BTs do not pose an immediate risk, and more work is needed to develop rigorous approaches to understanding how future models could increase biorisks. We end with recommendations about how empirical work can be expanded to more precisely target biorisk and ensure rigor and validity of findings.
AI advances are transforming life sciences but raise serious dual-use risks, including the potential for AI-enabled biological weapons. These risks are heightened in the Global South, where under-resourced health and biosecurity systems increase vulnerability. A review of AI strategies in 141 Global South countries shows only 27% have national AI plans, none of which address biosecurity. The paper highlights the urgent need for proactive governance and international cooperation to ensure AI’s benefits are equitably realized without compromising regional or global biosecurity.
Bruce J. Wittmann; Tessa Alexanian; Craig Bartling; Jacob Beal; Adam Clore; James Diggans; Kevin Flyangolts; Bryan T. Gemler; Tom Mitchell; Steven T. Murphy; Nicole E. Wheeler; Eric Horvitz · bioRxiv · 2024-12-04
Fast-moving advances in AI-assisted protein engineering are enabling breakthroughs in the life sciences that promise numerous beneficial applications. At the same time, these new capabilities are creating potential biosecurity challenges by providing new pathways to intentional or accidental synthesis of genes that encode hazardous proteins. The synthesis of nucleic acids is a key choke point in the AI-assisted protein engineering pipeline as it is where digital designs are transformed into physical instructions that can produce potentially harmful proteins. Thus, one focus for efforts to enhance biosecurity in the face of new AI-enabled capabilities is on bolstering the screening of orders by nucleic acid synthesis providers. We describe a multistakeholder, cross-sector effort to address biosecurity challenges with uses of AI-powered biological design tools to reformulate naturally occurring proteins of concern to create synthetic homologs that have low sequence identity to the wild-type proteins. We evaluated the abilities of traditional nucleic acid biosecurity screening tools to detect these synthetic homologs and found that, of tools tested, not all could previously detect such AI-redesigned sequences reliably. However, as we report, patches were built and deployed to improve detection rates over the course of the project, resulting in a final mean detection rate over tools of 97% of the synthetic homologs that were determined, using in-silico metrics, to be more likely to retain wild-type-like function. Finally, we make recommendations on approaches for studying and addressing the rising risk of adversarial AI-assisted protein engineering attacks like the one we identified and worked to mitigate.
AI agentsEvaluationsMitigationsNucleic acid synthesis screeningProtein design
The Bioeconomy Information Sharing and Analysis Center (BIO-ISAC) and its members strongly support the intent to collect information and insights from stakeholders on current and future practices and methodologies for the responsible development and use of chemical and biological (chem-bio) AI models. We welcome this opportunity to provide expertise and thank you for your leadership in this work. We ask you to prioritize cyberbiosecurity through this work, particularly with regards to data and model integrity. BIO-ISAC (isac.bio), a non-prot organization, addresses threats unique to the bioeconomy and enables coordination among stakeholders to facilitate a safe, secure industry.
BIO-ISAC provides two-way sharing of information between and among public and private partners serving as the central resource for gathering information on threats impacting the bioeconomy, helping spur the development and evaluation of defensive tools to address these issues.
Advances in AI have revolutionized chemical and biological design, driving breakthroughs in drug discovery. However, these advancements pose dual-use risks, particularly in the potential misuse of AI to create harmful biological agents. Despite the remarkable progress in the field, no standardized framework exists to evaluate the design capabilities of these systems. Unlike predictive models that can be evaluated cheaply in silico, design methods demand acting on the world. This necessitates real-world deployment to evaluate model capabilities. Addressing dual-use risks requires first understanding and quantifying AI capabilities, benchmarking systems against each other, and tracking their improvements over time. This response highlights the current challenges in evaluating AI chem-bio design methods and proposes a path forward through a robust evaluation framework based on real-world experimentation.
Samuel Curtis; Stephen McCarthy; Bryce Johnson; Rocco Moretti · Rosetta Commons · 2024-12-03
[From Linkedin Post] Our goal was to share perspectives and considerations that represent the views of Rosetta Commons and of the broader biomolecular structure analysis, prediction, and design communities. It was supplemented by scientists' input from a biosecurity-oriented roundtable discussion at Summer RosettaCon (August 2024), a workshop at European RosettaCon (November 2024), and a survey of the broader Rosetta Commons community.
Governments recognize the rapid advancement of biotechnology and are seeking to understand how policies should adapt. In my view, scientific research communities must be meaningfully engaged in policy development processes so that we're equipped to identify and address real risks while supporting vital research, which will, in turn, allow us to tackle major challenges in public health and the environment.
In our comment, one issue we draw attention to is the present lack of incentives or straightforward mechanisms to obtain substantive input from scientists on policy matters. While we share some initial thoughts on addressing this challenge, I suspect we're only scratching the surface. I'm eager to hear others' thoughts on bridging this gap.
NISTPolicyProtein designRequest for Information (RFI)
Artificial intelligence (AI) tools pose exciting possibilities to advance scientific, biomedical, and public health research. At the same time, these tools have raised concerns about their potential to contribute to biological threats, like those from pathogens and toxins. This report describes pathways that result in biological harm, with or without AI, and a range of governance tools and mitigation measures to address them.
MitigationsRisk assessmentsThreat AssessmentCenter for Security and Emerging Technology
Spiez CONVERGENCE 2024 was the sixth edition of this conference series hosted by Spiez Laboratory. The conference series is a Swiss “Science Diplomacy” initiative and part of the Swiss strategy for arms control and disarmament. It provides a platform for the presentation and discussion of new developments in science and technology that may affect the regimes governing the prohibition of chemical and biological weapons. Before the actual conference, participants had the opportunity to partake in a virtual “ice breaker” event with two keynote presentations introducing the objectives of the conference. This report brings together the new developments that were presented and the trends and findings discussed in the plenary and the breakout sessions.
Jaspreet Pannu; Sarah Gebauer; Greg McKelvey Jr; Anita Cicero; Tom Inglesby · Nature · 2024-11-21
AI-enabled research might cause immense harm if it is used to design pathogens with worrying new properties. To prevent this, we need better collaboration between governments, AI developers and experts in biosafety and biosecurity.
International governanceMitigationsPolicyRisk assessmentsThreat Assessment
James Revill; Clarissa Rios; Louison Mazeaud · Bulletin of the Atomic Scientists · 2024-11-16
Advances in AI could complicate international rules meant to prevent the spread of bioweapons, even as researchers and arms control experts debate how useful AI might ultimately be to would-be hostile actors.
Eric Nguyen; Michael Poli; Matthew G. Durrant; Brian Kang; Dhruva Katrekar; David B. Li; Liam J. Bartie; Armin W. Thomas; Samuel H. King; Garyk Brixi; Jeremy Sullivan; Madelena Y. Ng; Ashley Lewis; Aaron Lou; Stefano Ermon; Stephen A. Baccus; Tina Hernandez-Boussard; Christopher Ré; Patrick D. Hsu; Brian L. Hie · Science · 2024-11-15
The genome is a sequence that encodes the DNA, RNA, and proteins that orchestrate an organism’s function. We present Evo, a long-context genomic foundation model with a frontier architecture trained on millions of prokaryotic and phage genomes, and report scaling laws on DNA to complement observations in language and vision. Evo generalizes across DNA, RNA, and proteins, enabling zero-shot function prediction competitive with domain-specific language models and the generation of functional CRISPR-Cas and transposon systems, representing the first examples of protein-RNA and protein-DNA codesign with a language model. Evo also learns how small mutations affect whole-organism fitness and generates megabase-scale sequences with plausible genomic architecture. These prediction and generation capabilities span molecular to genomic scales of complexity, advancing our understanding and control of biology.
Sarah R. Carter; Nicole E. Wheeler; Chris Isaac; Jaime M. Yassif · The Nuclear Threat Initiative · 2024-11-14
The integration of artificial intelligence (AI) with the life sciences offers tremendous potential benefits to society, but advances in AI biodesign tools also pose significant risks of misuse, with the potential for global consequences.
AI biodesign tools (BDTs) are technologies that enable the engineering of biological systems. These tools are trained on biological data and are developed to provide insights, predictions, and designs related to biological systems. BDTs have the potential to drive progress in the development of new therapeutics and are likely to have a significant impact across the broader bioeconomy, including in agriculture, health, and materials science. However, there are risks BDTs could be misused to design dangerous pathogens, and few safeguards exist to ensure that the benefits of these technologies can be realized safely and securely.
Innovative strategies are needed to reduce the risks associated with potential misuse of biological design tools without significantly hindering beneficial uses. This report identifies a number of strategies, referred to as guardrails, that could be developed to safeguard BDTs against misuse.
Experts from government, academia, and industry were convened to explore the transformative impact of AI on biotechnology, emphasizing the need for strategic investments, high-quality datasets, and biosecurity measures to drive innovation and safety in the sector.
Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment analysis and planning. This article shows researchers with some training in statistics how to think about and analyze data from language model evaluations. Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data, measuring differences between two models, and planning an evaluation experiment. We make a number of specific recommendations for running language model evaluations and reporting experiment results in a way that minimizes statistical noise and maximizes informativeness.
The article discusses the evolving debate around the biosecurity risks posed by artificial intelligence, especially generative AI's potential to assist in creating dangerous bioweapons or pathogens.
Lynda M. Stuart; Rick A. Bright; and Eric Horvitz · NAM Perspectives · 2024-10-28
The use of artificial intelligence (AI)-enabled protein design tools such as AlphaFold, RoseTTAfold, and RFdiffusion to model proteins and generate new protein structures (referred herein as AI biodesign) is fueling a revolution in biology that promises benefits across medicine, sustainability, and beyond. Specifically, but currently underappreciated, these AI systems for protein design are transformational tools that can be harnessed to create life-saving medical countermeasures in a fast-paced manner and could have profound beneficial impacts on global health security. As part of the efforts to ensure the responsible use of AI, maximizing the benefits of AI biodesign and the application of such tools to biosecurity should become a global priority.
Global Biodefense Staff · Global Biodefense · 2024-10-05
The U.S. Artificial Intelligence Safety Institute (AISI), housed within the National Institute of Standards and Technology (NIST), is seeking information
Mingchen Li; Bingxin Zhou; Yang Tan; Liang Hong · bioRxiv · 2024-10-03
Pre-trained deep protein models have become essential tools in fields such as biomedical research, enzyme engineering, and therapeutics due to their ability to predict and optimize protein properties effectively. However, the diverse and broad training data used to enhance the generalizability of these models may also inadvertently introduce ethical risks and pose biosafety concerns, such as the enhancement of harmful viral properties like transmissibility or drug resistance. To address this issue, we introduce a novel approach using knowledge unlearning to selectively remove virus-related knowledge while retaining other useful capabilities. We propose a learning scheme, PROEDIT, for editing a pre-trained protein language model toward safe and responsible mutation effect prediction. Extensive validation on open benchmarks demonstrates that PROEDIT significantly reduces the model’s ability to enhance the properties of virus mutants without compromising its performance on non-virus proteins. As the first thorough exploration of safety issues in deep learning solutions for protein engineering, this study provides a foundational step toward ethical and responsible AI in biology.
Biological ToolsJailbreaks and red-teamingMitigationsProtein design
Language models rapidly become more capable in many domains, including biology. Both AI developers and policy makers [1] [2] [3] are in need of benchmarks that evaluate their proficiency in conducting biological research. However, there are only a handful of such benchmarks[4, 5], and all of them have their limitations. This paper introduces the Biological Lab Protocol benchmark (BioLP-bench) that evaluates the ability of language models to find and correct mistakes in a diverse set of laboratory protocols commonly used in biological research.
To evaluate understanding of the protocols by AI models, we introduced in these protocols numerous mistakes that would still allow them to function correctly. After that we introduced in each protocol a single mistake that would cause it to fail. We then gave these modified protocols to an LLM, prompting it to identify the mistake that would cause it to fail, and measured the accuracy of a model in identifying such mistakes across many test cases. State-of-the-art language models demonstrated poor performance compared to human experts, and in most cases couldn’t correctly identify the mistake.
Code and dataset are published at https://github.com/baceolus/BioLP-bench
Shaik Waseem Vali; Michael Crone; Dana Cortade; Amanda Kohler; Erika DeBenedictis; TJ Brunette · Zenodo · 2024-09-12
Automating biological research has the potential to improve reproducibility, throughput, and free-up scientist time to design rather than execute experiments. One avenue for scientists to leverage automation is through the use of cloud labs. However, there is no publicly available open source code to conduct basic life science workflows in cloud labs, and resources to learn the code are very limited, creating an enormous barrier to entry to cloud science. Here we present the first ever set of validated, open source protocols for cloud labs. We have developed a full pipeline for manipulating DNA on Emerald Cloud Lab’s platform, including onboarding samples, conducting PCR and Golden Gate reactions, and assessing the quality of samples with purification and gel electrophoresis. Each script is open source and available on GitHub, and we also provide instructions and screencast videos for future users. This work represents a necessary first step toward enabling cloud labs to mature into a widely accessible tool for life sciences
Richard Moulange; Sophie Rose; James Smith; Cassidy Nelson · 2024-08-23
Life sciences research and industries are undergoing a rapid transformation due to advancements in artificial intelligence (AI). Beyond the ongoing spotlight on ‘frontier’ AI models, AI-enabled biological tools (BTs) are driving a substantial amount of progress. There are various types of BTs, from experimental simulation to protein design tools, which provide a diverse range of expanding capabilities that enable beneficial research and innovation. However, in addition to the benefits, it has been hypothesised that BT capability improvements may increasingly enable harm: the ability to design and experiment with biological agents can be repurposed by malicious actors for weaponisation and misuse in the absence of safeguards and security measures. Whether BTs pose significant misuse risks above baseline has not been fully established, and proportional mitigation strategies should be considered only in combination with rigorous risk assessment. To mitigate the potential misuse risk from BTs, it is crucial to be able to comprehensively assess the capabilities of these tools and evaluate their potentially dangerous applications. This report summarises a methodological framework for risk assessment of BTs developed by The Centre for Long-Term Resilience in December 2023. We present an adaptable approach with example criteria and highlight limitations that can be addressed in future work.
Doni Bloomfield; Jaspreet Pannu; Alex W. Zhu; Madelena Y. Ng; Ashley Lewis; Eran Bendavid; Steven M. Asch; Tina Hernandez-Boussard; Anita Cicero; Tom Inglesby · Science · 2024-08-23
Advances in AI-driven biological models offer transformative benefits across health, agriculture, and biotechnology, but pose dual-use risks, including the potential creation of pandemic-capable pathogens. This paper argues that voluntary commitments to manage these risks are inadequate and calls for legally mandated oversight of advanced biological models—particularly those trained on sensitive or large-scale biological data using significant computational resources. Prerelease evaluations, narrowly targeted regulations, and global coordination are proposed to mitigate misuse while preserving scientific openness. The authors urge immediate development of scalable, standardized evaluation protocols and international harmonization of biosecurity standards, including genome synthesis oversight.
This report aims to clearly assess AI’s impact on the risks of biocatastrophe. It first considers the history and existing risk landscape in American biosecurity independent of AI disruptions. Drawing on a sister report, Catalyzing Crisis: A Primer on Artificial Intelligence, Catastrophes, and National Security, this study then considers how AI is impacting biorisks across four dimensions of AI safety: new capabilities, technical challenges, integration into complex systems, and conditions of AI development.11 Building on this analysis, the report identifies areas of future capability development that may substantially alter the risks of large-scale biological catastrophes worthy of monitoring as the technology continues to evolve. Finally, the report recommends actionable steps for policymakers to address current and near-term risks of biocatastrophes.
Suryesh K. Namdeo; Pawan Dhar · IndiaBioscience · 2024-08-12
As artificial intelligence (AI) enables the transformation of biology into an engineering discipline, an effective governance model that uses threat forecasting, real-time evaluation, and response strategies is urgently needed to address accidental or deliberate misuse. This article talks about the risks at the interface of AI and biosecurity and what could India do to better prepare for potential AI-biorisks.
Global SouthInternational governancePolicyThreat Assessment
National Academies of Sciences, Engineering, and Medicine · National Academies Press · 2024-08-05
Artificial intelligence (AI) and automation are increasingly being used to aid biological discovery and biotechnology development. Robotic and remotely controlled equipment is being used to accelerate research, while AI is opening new opportunities to explore the natural world and inform efforts to build biological entities with useful capabilities. Such technologies are poised to drive beneficial advances in health, biomaterials, environmental remediation, biomanufacturing, agriculture, and other areas. However, these developments also raise new questions and potential risks. Researchers, policymakers, and the public have sought to examine how applying AI and automation in biotechnology might lead to new challenges for biosecurity, health and safety, the environment, the integrity of scientific data, and economic development and national competitiveness.
AI agentsBiological ToolsCloud labsHealth securityRisk assessments
Sophie Rose; Richard Moulange; James Smith; Cassidy Nelson · 2024-07-26
this report aims to improve understanding of the potential impact of AI on biological misuse risk in two ways: 1) Providing a framework to estimate how AI may provide uplift: This framework can be adopted by readers to reason about uplift with their own assumptions, make forecasts about uplift, provide structure for gathering intelligence or developing policy, and to empirically test model capabilities (e.g. through evals) or uplift (e.g. through uplift studies). We consider biological uplift, but in principle the framework can generalise to other risks. 2) Generating hypotheses about how much uplift AI provides: We generate hypotheses about the uplift that given AI capabilities may offer to different categories of threat actors using our framework. We do this to facilitate an initial assessment of where to focus efforts to further understand uplift. These hypotheses can be subsequently prioritised by others with access to intelligence signals and based on their policy priorities, which can then be tested by those with the requisite resources and model access.
CapabilitiesEvaluationsFrameworkRisk assessmentsThreat AssessmentThe Centre for Long-Term Resilience
Jon M. Laurent; Joseph D. Janizek; Michael Ruzo; Michaela M. Hinks; Michael J. Hammerling; Siddharth Narayanan; Manvitha Ponnapati; Andrew D. White; Samuel G. Rodriques · arXiv · 2024-07-17
There is widespread optimism that frontier Large Language Models (LLMs) and LLM-augmented systems have the potential to rapidly accelerate scientific discovery across disciplines. Today, many benchmarks exist to measure LLM knowledge and reasoning on textbook-style science questions, but few if any benchmarks are designed to evaluate language model performance on practical tasks required for scientific research, such as literature search, protocol planning, and data analysis. As a step toward building such benchmarks, we introduce the Language Agent Biology Benchmark (LAB-Bench), a broad dataset of over 2,400 multiple choice questions for evaluating AI systems on a range of practical biology research capabilities, including recall and reasoning over literature, interpretation of figures, access and navigation of databases, and comprehension and manipulation of DNA and protein sequences. Importantly, in contrast to previous scientific benchmarks, we expect that an AI system that can achieve consistently high scores on the more difficult LAB-Bench tasks would serve as a useful assistant for researchers in areas such as literature search and molecular cloning. As an initial assessment of the emergent scientific task capabilities of frontier language models, we measure performance of several against our benchmark and report results compared to human expert biology researchers. We will continue to update and expand LAB-Bench over time, and expect it to serve as a useful tool in the development of automated research systems going forward. A public subset of LAB-Bench is available for use at the following URL: https://huggingface.co/datasets/futurehouse/lab-bench
This briefer focuses on the potential use of such export controls by the United States to mitigate risks from the Bio-AI nexus–a key objective of the nation’s AI-strategy.
CapabilitiesMitigationsPolicyThreat AssessmentUnited StatesThe Council on Strategic Risks
Jaspreet Pannu; Doni Bloomfield; Alex Zhu; Robert MacKnight; Gabe Gomes; Anita Cicero; Thomas Inglesby · 2024-06-25
As a result of rapidly accelerating artificial intelligence (AI) capabilities, over the past year, multiple national governments and multinational bodies have announced efforts to address safety, security and ethics issues related to AI models. One high priority among these efforts is the mitigation of misuse of AI models, such as for the development of chemical, biological, nuclear or radiological (CBRN) threats. Many biologists have for decades sought to reduce the risks of scientific research that could lead, through accident or misuse, to high-consequence disease outbreaks. Scientists have carefully considered what types of life sciences research have the potential for both benefit and risk (dual-use), especially as scientific advances have accelerated our ability to engineer organisms and create novel variants of pathogens. Here we describe how previous experience and study by scientists and policy professionals of dual-use capabilities in the life sciences can inform risk evaluations of AI models with biological capabilities. We argue that AI model evaluations should prioritize addressing high-consequence risks (those that could cause large-scale harm to the public, such as pandemics), and that these risks should be evaluated prior to model deployment so as to allow potential biosafety and/or biosecurity measures. While biological research is on balance immensely beneficial, it is well recognized that some biological research information and technologies could be intentionally or inadvertently misused to cause large-scale harm to the public. AI-enabled life sciences research is no different. Scientists' historical experience with identifying and mitigating dual-use biological risks can thus help inform new approaches to evaluating biological AI models. Identifying which AI capabilities pose the greatest biosecurity and biosafety concerns is necessary in order to establish targeted AI safety evaluation methods, secure these tools against accident and misuse, and avoid impeding immense potential benefits.
Background: The integration of Artificial Intelligence (AI) with synthetic biology is driving unprecedented progress in both fields. However, this integration introduces complex biosecurity challenges. Addressing these concerns, this article proposes a specialized biosecurity risk assessment process designed to evaluate the incorporation of AI in synthetic biology.Methods: A set of tailored tools and methodology was developed for conducting biosecurity risk assessments of AI language models used for synthetic biology. These resources were developed to guide risk management professionals through a systematic process of identifying, evaluating, and mitigating potential risks.Results: The tools and methodology provided offer a structured approach to risk assessment, enabling risk management professionals to comprehensively analyze the biosecurity implications of AI applications in synthetic biology. They facilitate the identification of potential risks and the development of effective mitigation strategies. An example of a risk assessment performed on the large language model “ChatGPT 4.0” is provided here.Conclusion: AI's role in synthetic biology is rapidly expanding; thus, establishing proactive and secure practices is crucial. The biosecurity risk assessment tools and methodology presented here are the first provided in the literature and will be instrumental steps toward the responsible integration of AI in synthetic biology. By adopting these resources, the biorisk management community can effectively navigate and manage the biosecurity challenges posed by AI, ensuring its responsible and secure application in the field of synthetic biology.
This white paper describes considerations for generating and standardizing biological data to support continued AIxBio research, development, and application.
Nathaniel Li; Alexander Pan; Anjali Gopal; Summer Yue; Daniel Berrios; Alice Gatti; Justin D. Li; Ann-Kathrin Dombrowski; Shashwat Goel; Long Phan; Gabriel Mukobi; Nathan Helm-Burger; Rassin Lababidi; Lennart Justen; Andrew B. Liu; Michael Chen; Isabelle Barrass; Oliver Zhang; Xiaoyuan Zhu; Rishub Tamirisa; Bhrugu Bharathi; Adam Khoja; Zhenqi Zhao; Ariel Herbert-Voss; Cort B. Breuer; Samuel Marks; Oam Patel; Andy Zou; Mantas Mazeika; Zifan Wang; Palash Oswal; Weiran Lin; Adam A. Hunt; Justin Tienken-Harder; Kevin Y. Shih; Kemper Talley; John Guan; Russell Kaplan; Ian Steneker; David Campbell; Brad Jokubaitis; Alex Levinson; Jean Wang; William Qian; Kallol Krishna Karmakar; Steven Basart; Stephen Fitz; Mindy Levine; Ponnurangam Kumaraguru; Uday Tupakula; Vijay Varadharajan; Ruoyu Wang; Yan Shoshitaishvili; Jimmy Ba; Kevin M. Esvelt; Alexandr Wang; Dan Hendrycks · arXiv · 2024-05-15
The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks of malicious use, government institutions and major AI labs are developing evaluations for hazardous capabilities in LLMs. However, current evaluations are private, preventing further research into mitigating risk. Furthermore, they focus on only a few, highly specific pathways for malicious use. To fill these gaps, we publicly release the Weapons of Mass Destruction Proxy (WMDP) benchmark, a dataset of 3,668 multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. WMDP was developed by a consortium of academics and technical consultants, and was stringently filtered to eliminate sensitive information prior to public release. WMDP serves two roles: first, as an evaluation for hazardous knowledge in LLMs, and second, as a benchmark for unlearning methods to remove such hazardous knowledge. To guide progress on unlearning, we develop RMU, a state-of-the-art unlearning method based on controlling model representations. RMU reduces model performance on WMDP while maintaining general capabilities in areas such as biology and computer science, suggesting that unlearning may be a concrete path towards reducing malicious use from LLMs. We release our benchmark and code publicly at https://wmdp.ai
EvaluationsJailbreaks and red-teamingLLMsMitigations
Renan Chaves de Lima; Lucas Sinclair; Ricardo Megger; Magno Alessandro Guedes Maciel; Pedro Fernando da Costa Vasconcelos; Juarez Antônio Simões Quaresma · Frontiers in Artificial Intelligence · 2024-05-10
The threat landscape of biological hazards with the evolution of AI presents challenges. While AI promises innovative solutions, concerns arise about its misuse in the creation of biological weapons. The convergence of AI and genetic editing raises questions about biosecurity, potentially accelerating the development of dangerous pathogens. The mapping conducted highlights the critical intersection between AI and biological threats, underscoring emerging risks in the criminal manipulation of pathogens. Technological advancement in biology requires preventative and regulatory measures. Expert recommendations emphasize the need for solid regulations and responsibility of creators, demanding a proactive, ethical approach and governance to ensure global safety.
FAS; Johns Hopkins Center for Health Security; Nuclear Threat Initiative Global Biological Policy and Programs; Scowcroft Institute of International Affairs · 2024-04-09
James Smith; Sophie Rose; Richard Moulange; Cassidy Nelson · 2024-03-27
Advances in AI-enabled biological tools (BTs) are catalysing life sciences research. BTs refer to the range of narrow but highly capable AI tools trained on biological data using machine learning techniques, which are important for many aspects of scientific research and development. While beneficial in many cases, BTs could be misused to enable actors across multiple steps in the development of
Biological ToolsPolicyUnited KingdomThe Centre for Long-Term Resilience
James Smith; Sophie Rose; Richard Moulange; Cassidy Nelson · The Centre for Long-Term Resilience · 2024-03-27
As part of our work to identify the three most beneficial next steps that the UK Government can take to reduce the biological risk posed by BTs, our team reflected on where the approach to narrow, specialised tools will need to differ from existing approaches to mitigating the risks from frontier AI. In this post, we outline why comprehensive risk assessments—which draw on literature and stakeholder engagement to assess these tools’ capabilities—are an effective and feasible alternative to conducting evaluations.
Michael Chen; Martin Holub; Cameron Tice · 2024-03-25
This report explores the potential of large language models (LLMs) to enhance biosecurity. We conducted interviews with nine biosecurity experts to understand their daily tasks, and how LLMs could be more useful for their work. Our findings indicate that approximately 50% of our interviewees’ biosecurity-related tasks, such as gathering information from papers and reports, reviewing safety forms, and writing memos and summaries, have high potential for automation with LLMs. Skills critical for biosecurity work, like processing information and communicating effectively, could also be augmented by LLMs. However, current LLMs have limitations, such as often providing shallow or incorrect information. We provide suggestions for LLM-based tools that could significantly advance biosecurity efforts and list field-specific datasets to facilitate their development.
Carl J. E. Suster; David Pham; Jen Kok; Vitali Sintchenko · Frontiers in Bacteriology · 2024-03-06
The analysis of microbial genomes has long been recognised as a complex and data-rich domain where artificial intelligence (AI) can assist. As AI technologies have matured and expanded, pathogen genomics has also contended with exponentially larger datasets and an expanding role in clinical and public health practice. In this mini-review, we discuss examples of emerging applications of AI to address challenges in pathogen genomics for precision medicine and public health. These include models for genotyping whole genome sequences, identifying novel pathogens in metagenomic next generation sequencing, modelling genomic information using approaches from computational linguistics, phylodynamic estimation, and using large language models to make bioinformatics more accessible to non-experts. We also examine factors affecting the adoption of AI into routine laboratory and public health practice and the need for a renewed vision for the potential of AI to assist pathogen genomics practice.
This paper examines how Large Language Models (LLMs) could contribute to the proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons.
LLMsRisk assessmentsThreat AssessmentJames Martin Center for Nonproliferation Studies
Trond Arne Undheim · Frontiers in Bioengineering and Biotechnology · 2024-02-28
AI-enabled synthetic biology has tremendous potential but also significantly increases biorisks and brings about a new set of dual use concerns. The picture is complicated given the vast innovations envisioned to emerge by combining emerging technologies, as AI-enabled synthetic biology potentially scales up bioengineering into industrial biomanufacturing. However, the literature review indicates that goals such as maintaining a reasonable scope for innovation, or more ambitiously to foster a huge bioeconomy do not necessarily contrast with biosafety, but need to go hand in hand. This paper presents a literature review of the issues and describes emerging frameworks for policy and practice that transverse the options of command-and-control, stewardship, bottom-up, and laissez-faire governance. How to achieve early warning systems that enable prevention and mitigation of future AI-enabled biohazards from the lab, from deliberate misuse, or from the public realm, will constantly need to evolve, and adaptive, interactive approaches should emerge. Although biorisk is subject to an established governance regime, and scientists generally adhere to biosafety protocols, even experimental, but legitimate use by scientists could lead to unexpected developments. Recent advances in chatbots enabled by generative AI have revived fears that advanced biological insight can more easily get into the hands of malignant individuals or organizations. Given these sets of issues, society needs to rethink how AI-enabled synthetic biology should be governed. The suggested way to visualize the challenge at hand is whack-a-mole governance, although the emerging solutions are perhaps not so different either.
This problem and policy analysis brief summarizes the most pressing risks at the intersection of artificial intelligence and chemical & biological weapons. It also outlines the policy recommendations for the US Federal Government to mitigate these risks and improve our national and global security.
NISTPolicyThreat AssessmentUnited StatesFuture of Life Institute
In our latest commentary produced from our New European Voices on Existential Risk (NEVER) network, Rebecca Donaldson explores the potential of new technologies for security whilst minimising their potential for harm in the realms of AI and the life sciences. She proposes that more funds go towards the biological weapons convention, the creation of an Emerging Technology Utilisation and Response Unit (ETURU) and the fostering of a culture of AI assurance and responsible democratisation of biotechnologies.
Advances in AI and DNA synthesis promise to revolutionize medicine… but could enable bioterrorism. A thoughtful mix of public health measures and restricted access to advanced capabilities can manage this risk while also alleviating natural viral threats.
Health securityNucleic acid synthesis screeningPolicyThreat AssessmentCenter for AI Safety
Tejal Patwardhan; Kevin Liu; Todor Markov; Neil Chowdhury; Dillon Leet; Natalie Cone; Caitlin Maltbie; Joost Huizinga; Carroll Wainwright; Shawn Jackson; Steven Adler; Rocco Casagrande; Aleksander Madry · 2024-01-31
We’re developing a blueprint for evaluating the risk that a large language model (LLM) could aid someone in creating a biological threat. In an evaluation involving both biology experts and students, we found that GPT-4 provides at most a mild uplift in biological threat creation accuracy. While this uplift is not large enough to be conclusive, our finding is a starting point for continued research and community deliberation.
In response to Executive Order 14110 on the safe and responsible development of artificial intelligence (AI), the Department of Homeland Security's Countering Weapons of Mass Destruction Office (CWMD) led the creation of the AI CBRN Report. This report assesses the dual-use risks and opportunities of AI technologies in the context of chemical, biological, radiological, and nuclear (CBRN) threats. Developed in collaboration with U.S. government entities, academia, and industry experts, the report evaluates how AI might be misused to facilitate CBRN threats and explores its potential to enhance national defenses against them.
Department of Homeland SecurityDual-use researchProtein designRisk assessments
The perception that the convergence of biological engineering and artificial intelligence (AI) could enable increased biorisk has recently drawn attention to the governance of biotechnology and artificial intelligence. The 2023 Executive Order, Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, requires an assessment of how artificial intelligence can increase biorisk. Within this perspective, we present a simplistic framework for evaluating biorisk and demonstrate how this framework falls short in achieving actionable outcomes for a biorisk manager. We then suggest a potential path forward that builds upon existing risk characterization work and justify why characterization efforts of AI-enabled tools for engineering biology is needed.
FrameworkRisk assessmentsThreat AssessmentUnited States
In this post, we argue that if AI model evaluations (evals) want to have meaningful real-world impact, we need a “Science of Evals”, i.e. the field needs rigorous scientific processes that provide more confidence in evals methodology and results.
Stephanie Batalis; Caroline Schuerger; Gigi Kwik Gronvall; Matthew E. Walsh · Applied Biosafety · 2023-12-27
Introduction: Artificial intelligence (AI) tools continue to be developed and used within the life sciences. The impact of these tools on the biosecurity landscape surrounding mail-order DNA synthesis and how to address the impacts have not been critically examined in the literature. Methods: The impacts of AI-driven chatbots and biological design tools on the biosecurity landscape surrounding mail-order DNA synthesis were analyzed and described. The findings are informed by the authors' experience in the field. Results: Generally, chatbots lower barriers to access of information that could be misused while biological design tools may provide new abilities to users with the intent of misuse. Six recommendations to the United States Government that attempt to maximize the benefits of these new technologies while mitigating risks are provided. Conclusion: Mandating mail-order DNA synthesis providers to screen DNA synthesis orders is a critical safeguarding step that should be taken as soon as possible. Over time, biological design tools will reduce the effectiveness of such a regulation and actions should be taken now to limit the negative impacts in the future.
Biological ToolsLLMsMitigationsNucleic acid synthesis screeningUnited States
Bill Anderson-Samways; Ashwin Acharya · 2023-12-15
AI regulation has made significant strides in recent months. However, any prospective regulator will confront a number of interrelated challenges including: a relative lack of government expertise; an uncertain and rapidly-evolving risk landscape; and the possibility that significant risks may arise during development as well as deployment. We therefore draw lessons from the biosecurity domain, which shares those features to some extent. Specifically, we examine the Federal Select Agent Program (FSAP), the mainstay of the US biosecurity regime. We suggest that FSAP offers both positive lessons for AI regulation, such as providing a precedent for a development-phase licensing regime, and more constructive lessons, relating to the pitfalls of checklist-based regulations as opposed to risk-based regulations.
FrameworkMitigationsPolicyRisk assessmentsUnited StatesInstitute for AI Policy and Strategy
Nazish Jeffery; Sarah R. Carter; Tessa Alexanian; Oliver Crook; Samuel Curtis; Richard Moulange; Shrestha Rath; Sophie Rose · Federation of American Scientists · 2023-12-12
Biosecurity risks related to AI are complex and rapidly changing. Here are five promising ideas to meet the challenges that AI poses in the life sciences.
Daniil A. Boiko; Robert MacKnight; Ben Kline; Gabe Gomes · Nature · 2023-12-01
Transformer-based large language models are making significant strides in various fields, such as natural language processing1–5, biology6,7, chemistry8–10 and computer programming11,12. Here, we show the development and capabilities of Coscientist, an artificial intelligence system driven by GPT-4 that autonomously designs, plans and performs complex experiments by incorporating large language models empowered by tools such as internet and documentation search, code execution and experimental automation. Coscientist showcases its potential for accelerating research across six diverse tasks, including the successful reaction optimization of palladium-catalysed cross-couplings, while exhibiting advanced capabilities for (semi-)autonomous experimental design and execution. Our findings demonstrate the versatility, efficacy and explainability of artificial intelligence systems like Coscientist in advancing research.
Batalis · Center for Security and Emerging Technology · 2023-12-01
Recent government directives, international conferences, and media headlines reflect growing concern that artificial intelligence could exacerbate biological threats. When it comes to biorisk, AI tools are cited as enablers that lower information barriers, enhance novel biothreat design, or otherwise increase a malicious actor’s capabilities. In this explainer, CSET Biorisk Research Fellow Steph Batalis summarizes the state of the biorisk landscape with and without AI.
Richard Moulange; Max Langenkamp; Tessa Alexanian; Samuel Curtis; Morgan Livingston · arXiv · 2023-11-30
Recent advancements in generative machine learning have enabled rapid progress in biological design tools (BDTs) such as protein structure and sequence prediction models. The unprecedented predictive accuracy and novel design capabilities of BDTs present new and significant dual-use risks. For example, their predictive accuracy allows biological agents, whether vaccines or pathogens, to be developed more quickly, while the design capabilities could be used to discover drugs or evade DNA screening techniques. Similar to other dual-use AI systems, BDTs present a wicked problem: how can regulators uphold public safety without stifling innovation? We highlight how current regulatory proposals that are primarily tailored toward large language models may be less effective for BDTs, which require fewer computational resources to train and are often developed in an open-source manner. We propose a range of measures to mitigate the risk that BDTs are misused, across the areas of responsible development, risk assessment, transparency, access management, cybersecurity, and investing in resilience. Implementing such measures will require close coordination between developers and governments.
David Rein; Betty Li Hou; Asa Cooper Stickland; Jackson Petty; Richard Yuanzhe Pang; Julien Dirani; Julian Michael; Samuel R. Bowman · arXiv · 2023-11-20
We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are "Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.
Steph Batalis; Venkatram · Center for Security and Emerging Technology · 2023-11-16
The recent Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence will have major implications for biotechnology. The EO demonstrates that the White House considers biorisk a major concern for AI safety and security. In this blog post CSET’s bio experts explain the bio-relevant takeaways of the executive order, add some additional context, and note their remaining questions about its implementation.
Nucleic acid synthesis screeningPolicyUnited States
Allison Berke · Bulletin of the Atomic Scientists · 2023-11-07
AI tools are close to being able to do a host of dangerous things, like walking people through the mistakes they made during a failed attempt at making dangerous pathogens and guiding them to a better protocol. There needs to be a durable way to probe the ways these systems might be misused, even as newer and more powerful technologies are continuously released.
CapabilitiesJailbreaks and red-teamingThreat Assessment
Christopher East · The Council on Strategic Risks · 2023-11-06
Artificial Intelligence (AI) is an epoch-defining technology that will profoundly shape the ways we live and work. In the so-called ‘Century of Biology,’ AI-enabled tools promise a massive increase in speed and capacity in certain parts of the bioengineering R&D cycle, including during data aggregation, lab automation, and the analysis and output phases. But despite
Anjali Gopal; Nathan Helm-Burger; Lennart Justen; Emily H. Soice; Tiffany Tzeng; Geetha Jeyapragasan; Simon Grimm; Benjamin Mueller; Kevin M. Esvelt · arXiv · 2023-11-01
Large language models can benefit research and human understanding by providing tutorials that draw on expertise from many different fields. A properly safeguarded model will refuse to provide "dual-use" insights that could be misused to cause severe harm, but some models with publicly released weights have been tuned to remove safeguards within days of introduction. Here we investigated whether continued model weight proliferation is likely to help malicious actors leverage more capable future models to inflict mass death. We organized a hackathon in which participants were instructed to discover how to obtain and release the reconstructed 1918 pandemic influenza virus by entering clearly malicious prompts into parallel instances of the "Base" Llama-2-70B model and a "Spicy" version tuned to remove censorship. The Base model typically rejected malicious prompts, whereas the Spicy model provided some participants with nearly all key information needed to obtain the virus. Our results suggest that releasing the weights of future, more capable foundation models, no matter how robustly safeguarded, will trigger the proliferation of capabilities sufficient to acquire pandemic agents and other biological weapons.
CapabilitiesJailbreaks and red-teamingOpen-weight LLMsRisk assessmentsVirology
Sarah R Carter; Nicole E Wheeler; Sabrina Chwalek; Christopher R Isaac; Jaime Yassif · 2023-10-30
To address the pressing need to govern AI-bio capabilities, this report explores three key questions:
1. What are current and anticipated AI capabilities for engineering living systems?
2. What are the biosecurity implications of these developments?
3. What are the most promising options for governing this important technology that will effectively guard against misuse while enabling beneficial applications?
To answer these questions, this report presents key findings informed by interviews with more than 30 individuals with expertise in AI, biosecurity, bioscience research, biotechnology, and governance of emerging technologies. Building on these findings, the report includes recommendations from the authors on the path toward developing more robust governance approaches for AI-bio capabilities to reduce biological risks without unduly hindering scientific advances.
This article explores the transformative role of artificial intelligence (AI) in synthetic biology, highlighting its impact from gene editing to metabolic pathway design. AI-driven tools are accelerating the development of innovative biotechnological solutions aimed at addressing critical global challenges such as food security, healthcare advancement, and climate change mitigation. By enhancing precision, efficiency, and scalability, AI is revolutionizing synthetic biology workflows, enabling the creation of novel biological systems and therapies with unprecedented speed and accuracy.
Christopher A. Mouton; Caleb Lucas; Ella Guest · 2023-10-16
In this report, the authors address the emerging issue of identifying and mitigating the risks posed by the misuse of artificial intelligence (AI) - specifically, large language models - in the context of biological attacks and present preliminary findings of their research. They find that while AI can generate concerning text, the operational impact is a subject for future research.
CapabilitiesJailbreaks and red-teamingThreat AssessmentRAND
Odhran O'Donoghue; Aleksandar Shtedritski; John Ginger; Ralph Abboud; Ali Essa Ghareeb; Justin Booth; Samuel G. Rodriques · arXiv · 2023-10-16
The ability to automatically generate accurate protocols for scientific experiments would represent a major step towards the automation of science. Large Language Models (LLMs) have impressive capabilities on a wide range of tasks, such as question answering and the generation of coherent text and code. However, LLMs can struggle with multi-step problems and long-term planning, which are crucial for designing scientific experiments. Moreover, evaluation of the accuracy of scientific protocols is challenging, because experiments can be described correctly in many different ways, require expert knowledge to evaluate, and cannot usually be executed automatically. Here we present an automatic evaluation framework for the task of planning experimental protocols, and we introduce BioProt: a dataset of biology protocols with corresponding pseudocode representations. To measure performance on generating scientific protocols, we use an LLM to convert a natural language protocol into pseudocode, and then evaluate an LLM's ability to reconstruct the pseudocode from a high-level description and a list of admissible pseudocode functions. We evaluate GPT-3 and GPT-4 on this task and explore their robustness. We externally validate the utility of pseudocode representations of text by generating accurate novel protocols using retrieved pseudocode, and we run a generated protocol successfully in our biological laboratory. Our framework is extensible to the evaluation and improvement of language model planning abilities in other areas of science or other areas that lack automatic evaluation.
The UK Government believes more research into AI risk is needed. This report explains why. It describes the current state and key trends relating to frontier AI capabilities, and then explores how frontier AI capabilities might evolve in the future and reviews some key risks. There is significant uncertainty around both the capabilities and risks from AI, including some experts who believe that some of these risks are overstated. This report focuses on evidence for risks and concludes that doing further research is necessary.
This report covers many risks, but we wish to emphasise that the overarching risk is a loss of trust in and trustworthiness of this technology which would permanently deny us and future generations its transformative positive benefits. In discussing the other risks, we do so in order to galvanize action to mitigate them, such that we can capture the full benefits of frontier AI.
Nicole N. Thadani; Sarah Gurev; Pascal Notin; Noor Youssef; Nathan J. Rollins; Daniel Ritter; Chris Sander; Yarin Gal; Debora S. Marks · Nature · 2023-10-01
Effective pandemic preparedness relies on anticipating viral mutations that are able to evade host immune responses to facilitate vaccine and therapeutic design. However, current strategies for viral evolution prediction are not available early in a pandemic—experimental approaches require host polyclonal antibodies to test against1–16, and existing computational methods draw heavily from current strain prevalence to make reliable predictions of variants of concern17–19. To address this, we developed EVEscape, a generalizable modular framework that combines fitness predictions from a deep learning model of historical sequences with biophysical and structural information. EVEscape quantifies the viral escape potential of mutations at scale and has the advantage of being applicable before surveillance sequencing, experimental scans or three-dimensional structures of antibody complexes are available. We demonstrate that EVEscape, trained on sequences available before 2020, is as accurate as high-throughput experimental scans at anticipating pandemic variation for SARS-CoV-2 and is generalizable to other viruses including influenza, HIV and understudied viruses with pandemic potential such as Lassa and Nipah. We provide continually revised escape scores for all current strains of SARS-CoV-2 and predict probable further mutations to forecast emerging strains as a tool for continuing vaccine development (evescape.org).
In the absence of empirical data concerning the capabilities of modern biotechnological methods to produce and deploy high impact biological threat agents, a strong theoretical model is required to inform effective biotechnological regulations and biosecurity preparations. Such a model is presented that aims to be robust across the diverse natures of all biological agents, any actors who might develop them, and the many biotechnologies and emerging computational super intelligence platforms that might be harnessed to do so. Core to this model is the recognition that any high consequence biotechnological agent must be able to spread geographically, be novel to the defenders, and be produced within the well understood constraints of technological development pipelines. Given these requirements, and the well established difficulty of modeling, and manipulating a novel organism's dynamics when introduced into an ecosystem, it becomes possible to derive the necessary properties of any actor capable of developing such a high consequence biotechnological threat agent: They must be designing their agent deliberately to do harm, and they must be highly resourced. Malevolent low resourced actors and benevolent or accidental actors regardless of resource level are revealed as being unable to produce such an agent. This is significant as much recent concern over the democratization of biotechnological capabilities has focused upon the large numbers of potential actors in those categories. Additionally, the constrained nature of the research and development efforts that might actually be able to produce a high consequence biotechnological threat agent allows for a refined focus in biosecurity policy and biotechnology regulation. This refined focus de-emphasizes damaging access-control policies seeking to limit and control large numbers of actors in the biotech space. Instead, an emphasis upon intelligence gathering to detect the definable and large footprints of the kind of research and development program needed to create such a high consequence biotechnological threat agent is revealed as optimal.
Alexander J. Titus; Adam H. Russell · arXiv · 2023-08-27
Artificial intelligence (AI) promises immense benefits across sectors, yet also poses risks from dual-use potentials, biases, and unintended behaviors. This paper reviews emerging issues with opaque and uncontrollable AI systems and proposes an integrative framework called violet teaming to develop reliable and responsible AI. Violet teaming combines adversarial vulnerability probing (red teaming) with solutions for safety and security (blue teaming) while prioritizing ethics and social benefit. It emerged from AI safety research to manage risks proactively by design. The paper traces the evolution of red, blue, and purple teaming toward violet teaming, and then discusses applying violet techniques to address biosecurity risks of AI in biotechnology. Additional sections review key perspectives across law, ethics, cybersecurity, macrostrategy, and industry best practices essential for operationalizing responsible AI through holistic technical and social considerations. Violet teaming provides both philosophy and method for steering AI trajectories toward societal good. With conscience and wisdom, the extraordinary capabilities of AI can enrich humanity. But without adequate precaution, the risks could prove catastrophic. Violet teaming aims to empower moral technology for the common welfare.
Dual-use researchFrameworkJailbreaks and red-teaming
As advancements in artificial intelligence (AI) propel progress in the life sciences, they may also enable the weaponisation and misuse of biological agents. This article differentiates two classes of AI tools that pose such biosecurity risks: large language models (LLMs) and biological design tools (BDTs). LLMs, such as GPT-4, are already able to provide dual-use information that could have enabled historical biological weapons efforts to succeed. As LLMs are turned into lab assistants and autonomous science tools, this will further increase their ability to support research. Thus, LLMs will in particular lower barriers to biological misuse. In contrast, BDTs will expand the capabilities of sophisticated actors. Concretely, BDTs may enable the creation of pandemic pathogens substantially worse than anything seen to date and could enable forms of more predictable and targeted biological weapons. In combination, LLMs and BDTs could raise the ceiling of harm from biological agents and could make them broadly accessible. The differing risk profiles of LLMs and BDTs have important implications for risk mitigation. LLM risks require urgent action and might be effectively mitigated by controlling access to dangerous capabilities. Mandatory pre-release evaluations could be critical to ensure that developers eliminate dangerous capabilities. Science-specific AI tools demand differentiated strategies to allow access to legitimate users while preventing misuse. Meanwhile, risks from BDTs are less defined and require monitoring by developers and policymakers. Key to reducing these risks will be enhanced screening of gene synthesis, interventions to deter biological misuse by sophisticated actors, and exploration of specific controls of BDTs.
Anna Krin; Gunnar Jeremias · Artificial intelligence · 2023-07-01
Artificial intelligence (AI) is an emerging technology with a dual-use character. Concerns have been raised that some of its applications in life sciences can be misused by nefarious actors for the development of biological and chemical weapons, prohibited by the Biological Weapons Convention (BWC), and the Chemical Weapons Convention (CWC). Areas of AI applications relevant to the BWC and CWC include rational drug design, retrosynthesis planning, and synthetic biology. Research in such areas might also unintentionally produce knowledge, products, or technologies that could be used by others to cause harm.
In late May of 2023, the problem-solving organization Helena convened a small group of senior leaders from industry, government, think tanks, and academia to interrogate this risk landscape and pressure-test courses of action. Their conversations took place at The Rockefeller Foundation’s Bellagio Center.
Steph Batalis; Caroline Schuerger; Vikram Venkatram · Center for Security and Emerging Technology · 2023-06-15
Steph Batalis, Caroline Schuerger and Vikram Venkatram explore three notable areas in the life sciences where LLMs are catalyzing meaningful advances: drug discovery, genetics, and precision medicine.
Emily H. Soice; Rafael Rocha; Kimberlee Cordova; Michael Specter; Kevin M. Esvelt · arXiv · 2023-06-06
Large language models (LLMs) such as those embedded in 'chatbots' are accelerating and democratizing research by providing comprehensible information and expertise from many different fields. However, these models may also confer easy access to dual-use technologies capable of inflicting great harm. To evaluate this risk, the 'Safeguarding the Future' course at MIT tasked non-scientist students with investigating whether LLM chatbots could be prompted to assist non-experts in causing a pandemic. In one hour, the chatbots suggested four potential pandemic pathogens, explained how they can be generated from synthetic DNA using reverse genetics, supplied the names of DNA synthesis companies unlikely to screen orders, identified detailed protocols and how to troubleshoot them, and recommended that anyone lacking the skills to perform reverse genetics engage a core facility or contract research organization. Collectively, these results suggest that LLMs will make pandemic-class agents widely accessible as soon as they are credibly identified, even to people with little or no laboratory training. Promising nonproliferation measures include pre-release evaluations of LLMs by third parties, curating training datasets to remove harmful concepts, and verifiably screening all DNA generated by synthesis providers or used by contract research organizations and robotic cloud laboratories to engineer organisms or viruses.
CapabilitiesDual-use researchJailbreaks and red-teamingLLMsNucleic acid synthesis screening
Erik Ekins; Filippa Lentzos; Max Brackmann; Cédric Invernizzi · Bulletin of the Atomic Scientists · 2023-03-24
As cutting-edge AI-powered chatbots like ChatGPT come online, observers have begun to worry about the implications of content producing AI in areas like employment and disinformation. More overlooked are similar systems for generating new proteins. AIs for protein engineering could potentially have many benefits, but they also pose new risks.
An international security conference explored how artificial intelligence (AI) technologies for drug discovery could be misused for de novo design of biochemical weapons. A thought experiment evolved into a computational proof.
AI is already consequential, but its future trajectory remains contested. Policymakers should make their assumptions explicit, focus on what can be shaped rather than what can be perfectly predicted, and build institutions that can learn and respond as evidence changes.
Recommendations to Governments on Mitigating AIxBio Risks These recommendations are based on the input and insights provided by participants of the INHR/CNAS trilateral dialogue. This US-China-International dialogue, focused primarily on the safety of AI military systems, includes experts suc...
How are terrorists using AI? Semi-structured interviews with 27 former Boko Haram members conducted in northeast Nigeria in 2025 and 2026 reveal unprecedented detail about AI-assisted terrorist activity primarily through 2024. This report finds that both factions of Boko Haram use frontier AI, including ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek, to assist in combat and day-to-day operations. This AI use is institutionalized through specialized units and internal training. It has aided in attack planning, weapons troubleshooting, and the design of explosive devices, as users have successfully circumvented some safeguards. This know-how was transferred through transnational jihadist networks, with Islamic State operatives delivering in-person training. Respondents expressed strong enthusiasm for AI and, in some cases, openness to mass-casualty weapons, though documented use remains conventional. Terrorist adoption of AI has thus advanced further and more systematically than prior analysis has recognized, making it a present and growing reality that warrants attention from policymakers, security communities, and AI developers.
Artificial intelligence (AI)-enabled tools are accelerating advances in biotechnology, offering powerful new capabilities for medical countermeasure (MCM) development. These technologies can expand the bioeconomy, strengthen global leadership in standards and innovation, and dramatically improve the speed, accuracy, and scalability of systems that detect, predict, and respond to biological threats, while also introducing new risks for misuse. The National Academies and National Academy of Medicine will conduct an international workshop and consensus study to examine the challenge of responding to unique biological threats designed using AI-enabled tools, opportunities to accelerate MCM against AI-enabled threats, and strategies for reducing associated risks.
Health securityPandemic preparednessPolicyThreat Assessment
Current trends make clear that biosecurity will become much more challenging over the next several years. We must act strategically to secure biology—before it becomes a general-purpose technology. Drawing on decades of experience and the knowledge of dozens of subject matter experts, Biosecurity Really offers a forward-looking analysis of how global trends and technologies are reshaping the biosecurity landscape, and actionable steps we can take to respond.
David Atanasov; Niccolò Zanichelli; Jean-Stanislas Denain · Epoch AI
We release a database of over 1,100 biological AI models across nine categories. We analyze their safeguards, accessibility, training data sources, and the foundation models they build on.
Biological ToolsData governanceEvaluationsTransparency and reporting
How advanced AI is lowering barriers to biological threats: accelerating knowledge synthesis, planning, pathogen design, sequencing, and bio-engineering automation. Evidence and capability assessments.