Did you know that artificial intelligence (AI) could be used to create biological weapons? Fortunately, tools like OpenAI are constantly evolving to prevent this.

In a recent development, OpenAI has implemented new safeguards in its most advanced AI models, o3 and o4-mini, to prevent them from being used to create biological and chemical threats.

But what exactly are these safeguards and why are they so important? Read on and discover what AI safeguards are, the general risks involved in its use and, in particular, the biorisks that concern experts so much.

What are AI safeguards?

Safeguards in AI function as safety brakes: mechanisms that identify and block uses that could be harmful. In the case of OpenAI, this system is called safety-focused reasoning monitor.

This custom component runs on top of the o3 and o4‑mini models, analyzes each query and rejects those related to biological or chemical risks, returning a warning instead of dangerous instructions (techcrunch.com).

To train this monitor, OpenAI had a team of red teamers who spent nearly 1,000 hours identifying unsafe conversations and labeling thousands of examples of risky requests.

In this way, the system learns to recognize patterns that could indicate the creation of biological weapons, toxins or harmful chemical agents.

General AI Risks

While biorisks attract attention due to their severity, AI presents other large-scale threats that are worth remembering:

  • Bias and discrimination: algorithms trained with partial data can perpetuate existing prejudices and adversely affect vulnerable groups.
  • Job displacement: automation threatens jobs in manufacturing, customer service and other sectors, generating social and economic challenges.
  • Deepfakes and disinformation: the creation of falsified audiovisual content can undermine public trust and facilitate disinformation campaigns.
  • Over-reliance: Blind trust in AI systems can erode human skills, such as critical thinking and problem solving.

Although these risks require regulatory frameworks and constant audits, biorisks add a layer of complexity due to their catastrophic potential.

What are biorisks?

Biorisks encompass the threats derived from the manipulation of biological agents. With AI, even users without specialized training could design biotoxins, optimize infectious agents, or even create biological weapons.

Although it may seem incredible, AI algorithms can analyze genetic sequences and predict mutations that increase the lethality or resistance of pathogens. Or they could identify genetic combinations that increase transmissibility, immune evasion or resistance to treatments.

Finally, tools designed to accelerate medical discoveries can be redirected toward destructive ends. That is, turning something beneficial into a biological weapon.

Safeguards implemented by OpenAI

According to OpenAI’s report, the reasoning monitor has demonstrated a 98.7% rejection rate against high-risk requests during internal testing. When it detects a dangerous pattern, the system stops generating the response and issues an explanatory message, describing why it cannot help with that query.

In addition, OpenAI updated its Preparedness Framework, a set of guidelines to evaluate risk levels and define action protocols:

  • High Threshold: If a model can allow a novice user to replicate known threats, additional controls are required before deploying it to production
  • Critical Threshold: if the model is capable of assisting experts in the development of new highly dangerous threats, its launch is stopped until safeguards are reinforced and exhaustive audits are carried out

With this classification, OpenAI decides whether to proceed with the release, delay the deployment to improve security measures, or discard the project in its current phase.

Continuous evaluation and future challenges

Although the system has been successful in testing, OpenAI recognizes key limitations. First of all, there is the possibility of evading filters. The tests did not simulate persistent users rephrasing their queries to bypass the monitor.

Furthermore, despite its automation, expert monitoring remains essential to identify new malicious tactics.

Finally there is competitive pressure. Some have pointed to very tight deadlines for evaluating models, which could compromise the depth of testing.

To address these challenges, OpenAI collaborates with laboratories such as Los Alamos and third parties specialized in biosafety, continually refining its classifiers and protocols. It also plans to expand its network of independent evaluators and publish security reports with greater transparency.

Towards responsible use of AI

Mass adoption of AI requires a balance between innovation and caution. OpenAI’s measures represent an important advance, but should not be understood as a definitive solution. Governments, academic institutions and the technology industry must:

  • Develop global regulatory frameworks that coordinate biosafety standards.
  • Promote transparency through the publication of security reports and independent audits.
  • Promote training in bioethics and cybersecurity for AI developers and users.

Only with a collaborative and multidisciplinary approach can we harness the potential of AI without exposing humanity to new dangers.

AI offers unprecedented opportunities, from faster medical diagnoses to industrial process optimization. But, as we have seen, it can also facilitate the creation of biological threats.

OpenAI safeguards are a step forward, demonstrating that it is possible to anticipate misuses of technology. However, security is never static: we must remain vigilant and committed to the continuous improvement of these systems to ensure a future where AI is a force for progress and not destruction.

This post is also available in: Español Français Русский Italiano