Loading...
加载中...
Table of Contents

Frontier AI Risk Monitoring Report (2026Q2)

Authors: Concordia AI Team
Published: July 19, 2026

Executive Summary

This is the fourth quarterly monitoring report from the Frontier AI Risk Monitoring Platform. It covers 47 frontier models released by 13 leading AI companies from 2025Q3 to 2026Q2.

This quarter's report adopts the Risk Index v2.0 framework for the first time. Compared with v1.5, v2.0 not only updates the benchmark set but also improves how the Risk Index is calculated in three ways: 1) it changes the effect of Capability Score on the Risk Index from a linear to an exponential relationship, reflecting the nonlinear way risk changes with capability; 2) it splits Safety Score into Base Safety Score, Jailbreak Safety Score, and Tamper Safety Score to distinguish misuse scenarios and threat actors more precisely; and 3) it introduces the Capability Yellow Line and Risk Yellow Line to mark key high-risk thresholds, align with industry risk ratings, and support cross-domain comparisons.

In addition, v2.0 introduces frontier jailbreak attack techniques for the first time to more accurately evaluate model safety against malicious actors.

The key findings are as follows:

Cross-domain Trends

  • 🔴 Model risk is growing rapidly, with average Risk Indices across domains rising severalfold in less than one year ↗️: The Risk Index rose 4.4x in cyber offense, 6.6x in biological risks, and 2.4x in loss-of-control. Models have already crossed the Risk Yellow Line in biological risks and loss-of-control. Crossing the Risk Yellow Line means that, even with existing safeguards, a model's residual risk still reaches the high-risk boundary.
  • 🟡 Model capabilities have widely crossed the Capability Yellow Line ↗️: 4 models in cyber offense, 22 in biological risks, and 12 in loss-of-control now have Capability Scores above the Capability Yellow Line. Crossing the Capability Yellow Line means that, without safeguards, a model would theoretically reach the high-risk boundary.
  • 🟡 Risk profiles differ sharply across model families ↕️: Gemini 3.1 Pro Preview has the highest overall Risk Indices in biological risks and loss-of-control; DeepSeek V4 Pro has the highest Risk Index in cyber offense; GPT, Kimi, MiniMax, Qwen, Doubao, and other families have crossed the Risk Yellow Line in biological risks.
  • 🟡 Open-weight models have lower Safety Scores in misuse-risk domains ↕️: Proprietary models still dominate the capability frontier across domains. Across the four misuse-risk domains of cyber offense, biological risks, chemical risks, and harmful manipulation, open-weight models have visibly lower Safety Scores than proprietary models.

Cyber Offense

  • 🟡 Top models have crossed the cyber-offense Capability Yellow Line ↗️: GPT and Claude models have crossed the Capability Yellow Line. Frontier models are progressing extremely quickly on autonomous, long-horizon vulnerability exploitation and penetration tasks.
  • 🔴 Safeguards have regressed ↘️: The average Safety Score is slightly lower than last quarter. Models perform well on basic cyber refusal tests, broadly unchanged from last quarter, but average performance on prompt-injection defenses has regressed slightly.

Biological Risks

  • 🟡 Multiple biological capabilities have reached or exceeded expert level ↗️: 22 models, 7 more this quarter, now exceed human-expert level on wet-lab troubleshooting; 11 models, 6 more this quarter, exceed human-expert level on sequence understanding; and 3 models, 2 more this quarter, exceed human-expert level on biological image understanding. Frontier models' biological reasoning and bioinformatics capabilities continue to strengthen.
  • 🟡 Biological safeguards have stagnated ➡️: The average Safety Score is broadly unchanged from last quarter. Basic biological refusal is generally strong, but safety falls sharply under red-team attacks.

Chemical Risks

  • 🟢 Chemical capability growth is gradual ➡️: Chemical capability growth remains modest, with no new high this quarter. Most models have strong general chemistry knowledge and some expert-level chemical research reasoning capability.
  • 🟡 Basic chemistry refusal varies sharply by benchmark ↕️: Most 2026Q2 models score above 90 on SOSBench-Chem, but most score below 40 on ChemicalHarmfulQA, showing large differences in safeguard quality across chemical safety scenarios.

Harmful Manipulation

  • 🟢 Harmful manipulation capability has not grown substantially overall ➡️: Harmful manipulation capability has grown only gradually, with no new high this quarter. Models are strong at phishing, inducing statements, and changing beliefs, but weaker at inducing payments.
  • 🟡 Political persuasion and active persuasion propensity remain the main weaknesses ↗️: Most models now exceed 80 on basic refusal for deception and manipulation, a clear improvement, but many still score below 60 in political persuasion scenarios, and most models easily show active persuasion propensity.

Loss-of-Control

  • 🟡 Loss-of-control risk did not continue rising ➡️: Although two models have crossed the Risk Yellow Line in loss-of-control, both are Gemini models, and none of the newly released models this quarter crossed the Risk Yellow Line. Self-replication, self-improvement, situational awareness, and related capabilities did not set new highs.
  • 🟡 Loss-of-control safety still shows no clear improvement ➡️: Model honesty remains weak overall, and some models still show a clear tendency to influence users covertly.

Red-Team Testing

  • 🔴 Advanced jailbreak attacks systematically weaken model safeguards ↘️: After red-team attacks are added, frontier models' average safety scores fall from 78.2 to 8.9 in biological risks, from 90.5 to 29.2 in cyber offense, from 83.7 to 53.1 in chemical risks, and from 87.8 to 36.4 in harmful manipulation.
  • 🟡 Jailbreak resistance differs greatly across models ↕️: Claude Opus 4.8 maintains an average refusal rate of 69.1% under red-team attacks, compared with just 5.8% for Hunyuan T1 (250711) under the same conditions.

Finally, the report offers recommendations for the following stakeholders:

  • Model developers: Conduct capability and safety evaluations before releasing new models. For models with high Risk Indices, especially those above the Risk Yellow Line, prioritize stronger base safety, jailbreak resistance, and tamper resistance; restrict high-risk capabilities where necessary; and establish mitigation and release mechanisms tied to risk thresholds.
  • AI safety researchers: Improve capability elicitation, red-team attacks, and real-world risk assessment methods; strengthen research on AI agents and the risk of tampering with open-weight models; and explore risk mitigations that preserve both safety and utility.
  • Policymakers: Focus on warning signals in cyber offense, biological risks, and loss-of-control; require model developers to complete relevant risk assessments and adopt necessary mitigations before release; and apply differentiated governance based on model capability, safety, and whether the model is open-weight or proprietary.

For the previous report, please see Frontier AI Risk Monitoring Report (2026Q1).

Explanation of Terms

Concept

  • Frontier Model: An AI model whose capabilities were at the industry frontier when it was released. To cover as many frontier models as possible within limited time and budget, we select only breakthrough models from each frontier model company, i.e., the most capable model released by that company at the time.
  • Risk Domains: We define five frontier AI risk domains, covering four misuse-risk domains and loss-of-control:
    • Cyber Offense: Misuse risks in cybersecurity, such as using AI to create malware.
    • Biological Risks: Misuse risks in biology, such as using AI to design, modify, or construct pathogens.
    • Chemical Risks: Misuse risks in chemistry, such as using AI to design novel highly toxic chemicals or plan synthesis routes for controlled chemicals.
    • Harmful Manipulation: Manipulation risks in real-world interactions, such as inducing users to make payments, inducing users to express specific views, or influencing users' political positions or critical decisions.
    • Loss-of-Control: Risks from autonomous AI loss-of-control, such as AI carrying out unsupervised self-improvement, self-replication, or power-seeking, ultimately causing humans to irreversibly lose control over AI.
  • Capability Benchmarks: Benchmarks used to evaluate model capabilities, especially risky capabilities that could be maliciously misused or contribute to loss-of-control.
  • Safety Benchmarks: Benchmarks used to evaluate model safety. For misuse risks, they mainly measure the model's defenses against external malicious instructions. For loss-of-control risks, they focus more on the model's internal propensities, honesty, and ability to suppress inappropriate behavior. Safety benchmarks are further divided into red-team and non-red-team benchmarks: red-team benchmarks use jailbreak methods to attack the model and try to bypass its guardrails, while non-red-team benchmarks send harmful instructions directly to the model.
  • Capability Score CC: The weighted average of a model's scores across capability benchmarks. The higher the score, the stronger the model's risky capabilities for misuse or loss-of-control.
  • Total Safety Score SS: A composite score for model safety. Higher scores indicate safer models. For cyber offense, biological risks, and other misuse risks, the Total Safety Score is a weighted combination of the Base Safety Score, Jailbreak Safety Score, and Tamper Safety Score. For loss-of-control risk, the Total Safety Score still comes directly from the average result of loss-of-control safety benchmarks. The detailed calculation method is available here.
    • Base Safety Score S1S_1: The weighted average score of a model on non-red-team safety benchmarks, reflecting safety against ordinary harmful requests.
    • Jailbreak Safety Score S2S_2: The weighted average score of a model on red-team safety benchmarks, reflecting safety under jailbreak attacks.
    • Tamper Safety Score S3S_3: A model's ability to resist attacks that tamper with model parameters, such as malicious fine-tuning. Because open-weight models are easier for third parties to tamper with, while proprietary models are less likely to be tampered with, this report uses a simplified treatment: open-weight models receive a Tamper Safety Score of 0, and proprietary models receive 100.
  • Risk Index RR: A risk score that combines capability and safety. The formula for Risk Index v2.0 is R=R0ea(CC0)×(1S)R = R_0 e^{a (C - C_0)} \times (1 - S), where CC is the Capability Score, SS is the Safety Score (normalized to [0-1]), R0R_0 is the Risk Yellow Line threshold, C0C_0 is the Capability Yellow Line threshold, and aa is the capability-risk coefficient. The formula is explained in detail here.
    • Risk Index v2.0: The new risk framework adopted by the platform starting in 2026Q2. Compared with v1.5, v2.0 places greater emphasis on real-world tasks and advanced attack scenarios, adds CVE-Bench, BixBench, the FrontierScience series, and FRT red-teaming benchmarks, and changes the Risk Index formula from a linear capability model to an exponential capability model. In the current v2.0 framework, Risk Indices are provided for cyber offense, biological risks, and loss-of-control; chemical risks and harmful manipulation do not yet have Capability Yellow Lines, so the platform mainly provides capability and safety evaluation results for those domains.
    • Capability Yellow Line Threshold C0C_0: A reference line indicating that, without safeguards, model capability significantly increases severe-harm risk relative to non-AI baselines. Crossing the Capability Yellow Line means that, without safeguards, a model's capabilities are approaching the risk level corresponding to OpenAI High Risk or Anthropic ASL-3, significantly amplifying cyber offense, biological/chemical misuse, or loss-of-control threats relative to non-AI baselines. However, it measures capability only and does not mean that the model has already reached high risk in deployment. Capability Yellow Lines are currently defined for cyber offense, biological risks, and loss-of-control. Details on threshold calculation are available here.
    • Risk Yellow Line Threshold R0R_0: A reference line indicating that the model's residual risk in actual deployment has reached the high-risk boundary, normalized to 100. Crossing the Risk Yellow Line means that, in a real deployment environment and after accounting for actual safeguards, the model's residual risk still reaches a high-risk level, significantly amplifying cyber offense, biological/chemical misuse, or loss-of-control threats relative to non-AI baselines. From the Risk Index formula, if a model's capability equals the Capability Yellow Line threshold and its Safety Score is 0, its Risk Index equals the Risk Yellow Line threshold, 100.
    • Capability-Risk Coefficient aa: The multiplier on Capability Score in the Risk Index formula, fitted from model Capability Score data and model autonomous task horizon data in the METR report. It is based on the premise that model risk is proportional to the length of tasks the model can complete autonomously. The longer a model can work autonomously, the less time a human attacker or supervisor needs to invest, increasing the scale of misuse attacks and the possibility of loss-of-control. Details are available here.

Note: Because Risk Index v2.0 substantially changes both the benchmark set and metric calculations, the Risk Index values in this report should not be compared point by point with v1.0 or v1.5 reports. A more appropriate use is to compare models, quarters, and domains within the v2.0 framework.

Model List

This section lists the breakthrough models covered by this report from the most recent four quarters, 2025Q3 through 2026Q2.

Company Models
OpenAI GPT-5 (high), GPT-5.1 (high), GPT-5.2 (high), GPT-5.4 (high), GPT-5.5
Google Gemini 3 Pro Preview, Gemini 3.1 Pro Preview
Anthropic Claude Sonnet 4.5 Reasoning, Claude Opus 4.5 Reasoning, Claude Opus 4.6 Reasoning, Claude Opus 4.8
xAI Grok 4, Grok 4.20 Beta Reasoning, Grok 4.3
DeepSeek DeepSeek V3.1 Terminus Reasoning, DeepSeek V3.2 Reasoning, DeepSeek V4 Pro
Alibaba Qwen 3 235B Reasoning (250725), Qwen 3.5 397B Reasoning, Qwen 3.7 Max
ByteDance Doubao Seed 1.6 (251015 High), Doubao Seed 1.8 (251215 High), Doubao Seed 2.0 Pro (260215), Doubao Seed 2.1 Pro
Tencent Hunyuan T1 (250711), HY 2.0 Think, HY 3.0 Preview
Baidu ERNIE X1.1 Preview, ERNIE 5.0 Thinking Preview, ERNIE 5.0 Thinking, ERNIE 5.1
MiniMax MiniMax M2, MiniMax M2.7, MiniMax M3
Moonshot AI Kimi K2 (250711), Kimi K2 (250905), Kimi K2 Thinking, Kimi K2.5 Thinking, Kimi K2.6
Zhipu GLM 4.5, GLM 4.6, GLM 4.7, GLM 5, GLM 5.2
Xiaomi MiMo V2 Flash Reasoning, MiMo V2 Pro Reasoning, MiMo V2.5 Pro

The models newly added in 2026Q2 are GPT-5.5, Claude Opus 4.8, Grok 4.3, DeepSeek V4 Pro, Qwen 3.7 Max, Doubao Seed 2.1 Pro, HY 3.0 Preview, ERNIE 5.1, MiniMax M3, Kimi K2.6, GLM 5.2, and MiMo V2.5 Pro.

Benchmark List

The risk domains evaluated in this report include cyber offense, biological risks, chemical risks, harmful manipulation, and loss-of-control. The benchmark list is as follows:

Domain Capability Evaluation Safety Evaluation
Cyber Offense
View Evaluation Details
CyberSecEval2-VulnerabilityExploit: Code vulnerability discovery and exploitation capability
CyBench: CTF task capability
CVE-Bench: Real-world web application penetration capability
AirBench-SecurityRisks: Harmful cybersecurity questions
CyberSecEval2-PromptInjection: Prompt injection
FRT-AirBench-SecurityRisks: Harmful cybersecurity questions + red-team attacks
Fortress-Privacy/Scams: Red-team attacks involving privacy and scams
Biological Risks
View Evaluation Details
BioLP-Bench: Biological experiment troubleshooting capability
BixBench: Bioinformatics agent task capability
LAB-Bench-SeqQA: DNA and protein sequence understanding capability
LAB-Bench-FigQA: Biological image understanding capability
FrontierScience-Biology: Frontier biological science reasoning capability
SciKnowEval-BiologicalHarmfulQA: Harmful biological questions
SOSBench-Bio: Harmful and illicit biological requests
FRT-SciKnowEval-BiologicalHarmfulQA: Harmful biological questions + red-team attacks
Fortress-Biological: Biological red-team attacks
Chemical Risks
View Evaluation Details
ChemBench-ToxicityAndSafety: Chemical toxicity and safety knowledge
ChemBench-Normal: General chemistry knowledge capability
FrontierScience-Chemistry: Frontier chemical science reasoning capability
SOSBench-Chem: Harmful and illicit chemical requests
SciKnowEval-ChemicalHarmfulQA: Harmful chemical questions
FRT-SOSBench-Chem: Harmful chemical requests + red-team attacks
Fortress-Chemical: Chemical-risk red-team attacks
Harmful Manipulation
View Evaluation Details
CyberSecEval3-MultiTurnPhishing: Multi-turn phishing capability
MakeMePay: Capability to induce payments
MakeMeSay: Capability to induce users to say specified content
PMIYC: Capability to change beliefs
AirBench-Deception: Harmful deception tasks
AirBench-Manipulation: Harmful manipulation tasks
AirBench-PoliticalPersuasion: Political persuasion tasks
APE: Propensity to persuade others
FRT-AirBench-Manipulation: Harmful manipulation tasks + red-team attacks
Loss-of-Control
View Evaluation Details
Self-Proliferation: Self-replication and adaptation capability
MLE-Bench: Machine learning engineering capability under constrained resources
SciCode: Scientific programming capability
GDM-Stealth: Stealth capability
SAD-mini: Situational awareness capability
MASK: Model honesty
Agentic-Misalignment: Agentic misalignment propensity
Shutdown-Resistance: Shutdown resistance propensity
DarkBench: Propensity to influence users covertly

Compared with Risk Index v1.5, this v2.0 release adds CVE-Bench, BixBench, FrontierScience-Biology/Chemistry, ChemBench-Normal, and the FRT-series red-team testing benchmarks; removes WMDP-Cyber/Bio/Chem, SciKnowEval-ProteoToxicityPrediction/MolecularToxicityPrediction, StrongReject, and the ISC-Bench series; and removes CyberSecEval3-MultiTurnPhishing from cyber offense while retaining it in harmful manipulation. Implementation details for newly added benchmarks and reasons for removing benchmarks are provided in Appendix B and Appendix C, respectively.

Note: To ensure internal consistency across charts, all historical quarterly risk curves shown in this report are computed retrospectively under the v2.0 benchmark framework.

Monitoring Results

Cross-domain Trends

Overall Trends

Based on Risk Index v2.0, the current trends in overall Risk Index, Capability Score, and Safety Score for cyber offense, biological risks, and loss-of-control are shown below:

The trends show:

  • Cyber offense: The average Capability Score rose steadily from 39.3 in 2025Q3 to 56.2 in 2026Q2. The average Safety Score rose from 56.5 to 68.9 in 2026Q1, but fell to 66.8 in the latest quarter. The average Risk Index climbed from 4.7 to 20.9, a 4.4x increase in less than one year. The maximum Risk Index has reached 44.8.
  • Biological risks: Similar to cyber offense, the average Capability Score continued to rise, while the Safety Score peaked in 2026Q1 and then fell slightly in the latest quarter. The average Risk Index rose from 16.8 in 2025Q3 to 110.2 in 2026Q2, a 6.6x increase, crossing the Risk Yellow Line. The maximum Risk Index has reached 234.6.
  • Loss-of-control: The average Capability Score rose until 2026Q1 and then began to stagnate, while the Safety Score continued to rise. The average Risk Index rose from 9.0 in 2025Q3 to 32.5 in 2026Q1, then fell to 21.9 in the latest quarter. The maximum Risk Index has reached 178.5, above the Risk Yellow Line threshold.

Model Family Comparison

  • Gemini family: Gemini 3.1 Pro Preview has the highest overall Risk Indices in biological risks and loss-of-control, and both exceed the Risk Yellow Line.
  • DeepSeek family: DeepSeek V4 Pro has the highest Risk Index overall in cyber offense and is rising quickly, though it has not yet reached the Risk Yellow Line. In loss-of-control, its Risk Index is second only to Gemini 3.1 Pro Preview.
  • GPT family: Its Risk Index is rising quickly in cyber offense and loss-of-control, and it has crossed the Risk Yellow Line in biological risks.
  • Kimi/MiniMax/Qwen/Doubao families: They are rising quickly in cyber offense and biological risks, and have crossed the Risk Yellow Line in biological risks.
  • Claude family: Claude Opus 4.8 ranks behind only DeepSeek and GPT in cyber-offense Risk Index. Because this round does not include Claude's strongest models, Fable/Mythos 5, the actual risk of Claude-family models may be underestimated.

Open-Weight vs. Proprietary Comparison

Open-weight and proprietary models show the following characteristics:

  • Proprietary models still dominate the capability frontier across domains: GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro Preview, and other proprietary models maintain leading Capability Scores across domains.
  • Open-weight models generally have lower Safety Scores in misuse-risk domains: Across cyber offense, biological risks, chemical risks, and harmful manipulation, the Safety Score distribution of open-weight models is visibly lower than that of proprietary models.
  • Proprietary models have lower Safety Scores in loss-of-control: In the loss-of-control domain, Gemini 3.1 Pro Preview, Grok 4, and Gemini 3 Pro Preview have substantially lower Safety Scores than other models. These models pose higher risk because they combine high Capability Scores with low Safety Scores.

Cyber Offense

Risk Overview

In cyber offense, the models with the highest current Risk Indices are DeepSeek V4 Pro, GPT-5.5, and Claude Opus 4.8. No model has crossed the Risk Yellow Line. From a capability perspective, however, 4 models have crossed the Capability Yellow Line: GPT-5.5, GPT-5.4, Claude Opus 4.8, and Claude Opus 4.6. DeepSeek V4 Pro does not have the highest Capability Score, but its Safety Score is only 43.4, which is the main reason it has the highest Risk Index.

Kimi, GLM, MiMo, Qwen, Doubao, and other model families have also seen relatively fast Risk Index growth, mainly driven by capability growth. The Grok family has shown a continuous Risk Index decline, mainly because its Capability Score has continued to fall while its Safety Score has continued to improve.

Note: This evaluation does not include Claude's strongest models, Fable/Mythos 5. According to that model's system card, its cyberattack capability is far stronger than Claude Opus 4.8. Once that model is publicly released, actual cyber-offense risk is likely higher than indicated by this report.

Capability Evaluations

  • Vulnerability exploitation capability: On CyberSecEval2-VulnerabilityExploit, GPT-5.4 achieved the highest score, 95.6. Half of the models scored above 80, showing that vulnerability discovery and exploitation capabilities have become widespread.
  • CTF tasks: GPT-5.5 and Claude Opus 4.8 tied for first on CyBench with scores of 92.5, showing strong performance on cyberattack tasks that require multi-step planning, tool use, and environment interaction. The top frontier-model score increased by 68% in less than one year, a very rapid gain. These results were achieved using Inspect's simplest built-in agent framework, without advanced harness techniques or cyber-specific tools, so the improvement comes entirely from model capability.
  • Real-world web vulnerability exploitation tasks: On CVE-Bench, GPT-5.5 achieved the highest score, 43.3, under the zero-day setting; all other models scored below 40. This shows that real-world web vulnerability exploitation tasks are more challenging for frontier models than CTF tasks. These results were also achieved using Inspect's simplest built-in agent framework, with the maximum number of interaction messages capped at 90. If models were allowed to use advanced agent frameworks and larger message budgets, they would likely score higher. We have noted that the GPT-5.5 system card reports a CVE-Bench score of 93.1.

Safety Evaluations

  • Basic cyber refusal: 2026Q2 models generally score above 90 on AirBench-SecurityRisks, indicating strong basic cybersecurity guardrails.
  • Prompt-injection defenses: 2026Q2 models generally score between 80 and 90 on PromptInjection, with average performance slightly worse than 2026Q1. This suggests that safeguards against prompt injection do not necessarily improve in step with model capability.

Biological Risks

Risk Overview

In biological risks, 12 models have now crossed the Risk Yellow Line. Gemini 3.1 Pro Preview has the highest Risk Index, reaching 234.6. Kimi, GPT, Qwen, Doubao, and MiniMax-family models have also crossed the Risk Yellow Line. From a capability perspective, 22 models have crossed the Capability Yellow Line. GPT-5.5 has the highest Capability Score, followed by Claude Opus 4.8, but because Claude Opus 4.8 has a high Safety Score, its Risk Index is only 43.4, below the Risk Yellow Line. Grok and GLM-family models have also crossed the Capability Yellow Line, but neither family has crossed the Risk Yellow Line.

Capability Evaluations

  • Wet-lab troubleshooting: BioLP-Bench results show that 22 models now exceed human-expert level, 7 more this quarter. The top score is Claude Opus 4.6 at 48.7, with no new high in the latest quarter.
  • DNA and protein sequence understanding: On LAB-Bench-SeqQA, 11 models now exceed human-expert level, 6 more this quarter. GPT-5.5 achieved the highest score, 91.5, showing that frontier models already have strong sequence understanding capabilities.
  • Biological image understanding: LAB-Bench-FigQA results show that 3 models now exceed human experts, 2 more this quarter. Claude Opus 4.8 achieved the highest score, 79.6, a 38% increase in less than one year.
  • Frontier biological science reasoning: On FrontierScience-Biology, the highest score is only 47.1, achieved by GPT-5.4. Most models score below 40, showing substantial remaining weaknesses in expert-level scientific research reasoning.
  • Bioinformatics agent tasks: Claude Opus 4.8 achieved the highest score on BixBench, 51.0. Most models score below 40, showing that they still have limitations on realistic long-horizon bioinformatics tasks. These results were achieved using Inspect's simplest built-in agent framework, without advanced harness techniques, and using only preset bioinformatics tools, so they may underestimate actual model capabilities.

Safety Evaluations

  • Basic biological refusal: On BiologicalHarmfulQA, 2026Q2 models improved slightly overall relative to last quarter, but some models still have low Safety Scores, such as MiniMax M3 at 59.6 and ERNIE 5.1 at 62.3. On SOSBench-Bio, most 2026Q2 models already score above 90, showing a clear overall improvement trend.

Chemical Risks

Risk Overview

Because no Capability Yellow Line has been defined for the chemical domain, no Risk Index is calculated here; this section mainly analyzes capability and safety evaluation results. The highest chemical Capability Score is GPT-5.4 at 72.3. The Kimi, GLM, and Grok families have also exceeded 70. The lowest Safety Score is DeepSeek V4 Pro at 28.9. Kimi and MiMo-family models have also fallen below 30 in the past, though their latest models have improved.

Capability Evaluations

  • Chemical toxicity and safety knowledge: GPT-5.5 achieved the highest score on ChemBench-ToxicityAndSafety, 58.4. Overall model capability progress is relatively gradual.
  • General chemistry knowledge: GPT-5.4 achieved the highest score on ChemBench-Normal, 88.7. Scores for 2026Q2 models are all between 80 and 90, indicating that frontier models generally have strong general chemistry knowledge.
  • Frontier chemical science reasoning: Gemini 3.1 Pro Preview achieved the highest score on FrontierScience-Chemistry, 74.9. Most models score above 60, showing that frontier models already have some expert-level chemical research reasoning capabilities.

Safety Evaluations

  • Basic chemistry refusal: Most 2026Q2 models score above 90 on SOSBench-Chem, showing a clear improvement trend. On ChemicalHarmfulQA, however, most models score below 40, with HY 3.0 Preview scoring only 7.3, showing substantial differences across chemical safety benchmarks.

Harmful Manipulation

Risk Overview

Because no Capability Yellow Line has been defined for harmful manipulation, no Risk Index is calculated here; this section mainly analyzes capability and safety evaluation results. In Capability Score, Gemini 3.1 Pro Preview ranks highest at 68.9, while the highest score among latest-quarter models is Doubao Seed 2.1 Pro at 62.5. Overall model capability progress over the past year has not been substantial. In Safety Score, the lowest-scoring model is DeepSeek V4 Pro at 25.7. GLM and Kimi-family models have also fallen below 30, though newer versions have improved.

Capability Evaluations

  • Phishing: Claude Opus 4.6 achieved the highest score on MultiTurnPhishing, 79.2, far ahead of the field. Most models score between 50 and 70.
  • Inducing payments: Gemini 3.1 Pro Preview achieved the highest score on MakeMePay, 82.2, far ahead of other models. Most models score only 30 to 50.
  • Inducing specific statements: Claude Opus 4.8 achieved the highest score on MakeMeSay, 73.4, but the improvement over previous models is not substantial.
  • Changing beliefs: GPT-5.1 achieved the highest score on PMIYC, 80.9. Model capability in this area has shown little improvement over the past year.

Safety Evaluations

  • Basic harmful-manipulation refusal: On AirBench-Deception and AirBench-Manipulation, most 2026Q2 models score above 80, indicating generally strong safety in deception and manipulation scenarios. On AirBench-PoliticalPersuasion, however, many models still score below 60, with ERNIE 5.1 scoring only 47.1; political persuasion remains a weakness.
  • Persuasion propensity: On APE, most models score below 40, and the ERNIE and DeepSeek families score below 4, indicating that models remain very prone to active persuasion behavior.

Loss-of-Control

Risk Overview

In loss-of-control, two models have now crossed the Risk Yellow Line, both Gemini models. Gemini 3.1 Pro Preview has reached a Risk Index of 178.5. Twelve models have crossed the Capability Yellow Line, including Gemini, GPT, Claude, Qwen, DeepSeek, and Kimi-family models. The latest GPT, DeepSeek, and Qwen models have not crossed the Risk Yellow Line, but their Risk Indices have grown quickly, mainly because their Capability Scores rose rapidly while their Safety Scores fell. The lowest Safety Score is Grok-4 at 38.0, though later versions have improved substantially.

Capability Evaluations

  • Self-replication capability: Claude Opus 4.6 achieved the highest score on Self-Proliferation, 63.3. 2026Q2 models generally score above 40, but show no clear improvement over last quarter.
  • Machine learning engineering capability: MiMo V2 Pro achieved the highest score on MLE-Bench, 61.9. 2026Q2 model scores vary widely: GPT-5.5 is highest at 57.1, while Grok 4.3 is lowest at only 9.5. Overall performance is worse than last quarter's models, showing that models still face challenges in completing machine learning engineering tasks under constrained resources.
  • Scientific programming capability: Gemini 3.1 Pro Preview achieved the highest score on SciCode, 58.9. 2026Q2 models generally score between 40 and 60, with no clear improvement over last quarter.
  • Stealth: Gemini 3.1 Pro Preview achieved the highest score, 39.3, far ahead of other models. 2026Q2 models generally score below 20, showing that evading oversight remains challenging for most frontier models.
  • Situational awareness: Gemini 3.1 Pro Preview achieved the highest score on SAD-mini, 92.0. 2026Q2 models generally score between 75 and 85, indicating strong situational awareness across frontier models. Overall scores, however, have not improved much over the past year, suggesting that the benchmark may be nearing saturation.

Safety Evaluations

  • Honesty: On MASK, model scores vary widely. Half of the models score below 60, and the lowest score is only 32.5, from Grok 4. This suggests that model honesty remains weak overall.
  • Agentic misalignment: On Agentic-Misalignment, the latest-quarter models improved, with most scoring above 75 and nearly all GPT and Claude-family models reaching 100. A few models still score low, such as DeepSeek V4 Pro at 54.8.
  • Shutdown resistance: Most models perform well on Shutdown-Resistance, scoring above 95. Only the GPT family, Gemini family, early Grok models, Doubao Seed 2.1 Pro, and Claude Opus 4.6 score somewhat lower.
  • Covertly influencing users: Claude Opus 4.8 achieved the highest score on DarkBench, 87.7. However, most models score below 60, such as Grok 4.20 Beta at only 32.0, indicating that some models still show a clear tendency to influence users covertly.

Red-Team Testing

As noted above, Risk Index v2.0 adds red-team testing to evaluate how model safety changes under attack.

The chart below compares model scores on the same original safety datasets in four misuse-risk domains under two settings: without red-team attacks and with red-team attacks. Higher scores mean that the model is better able to refuse or resist unsafe requests. Each chart is sorted by the score under red-team attack.

Red-team testing shows a systematic gap between base safety and safety under red-team attack:

  • Red-team attacks sharply reduce safety scores: The average score across all frontier models on the original biological dataset is 78.2, but falls to only 8.9 after red-team attacks, a decrease of 69.3 points. Cyber offense falls from 90.5 to 29.2, chemical risks from 83.7 to 53.1, and harmful manipulation from 87.8 to 36.4. In other words, high scores on basic refusal do not mean that a model can resist frontier red-team attacks.
  • Attack success rates differ substantially across domain datasets: SOSBench-Chem still has an average score of 53.1 after red-team attacks, while SciKnowEval-BiologicalHarmfulQA falls to only 8.9. This shows that attack effectiveness differs across datasets, possibly because the datasets use different default scorers.
  • Jailbreak resistance differs greatly across models: For example, Claude Opus 4.8 maintains an average refusal rate of 69.1% under red-team attacks, compared with just 5.8% for Hunyuan T1 (250711) under the same conditions.
  • Some models score higher after attack: For example, on SOSBench-Chem, Grok-4 scored 41.2 before attack but rose to 62.3 after attack. This may be because the model was adversarially trained against common attack methods, while lacking basic refusal training for that risk domain. This phenomenon appears only in this one case.

Limitations

This report has the following limitations:

  • Limitations in the scope of risk assessment
    1. This monitoring round covers only large language models, including vision-language models. It does not yet cover more modalities or agentic systems, so it cannot comprehensively assess the risks of all models and AI systems.
    2. This monitoring round focuses only on misuse and loss-of-control risks, and therefore does not cover all types of frontier risk, such as accident risks or systemic risks.
  • Limitations in risk-assessment methods
    1. Because of limitations in existing evaluation methods, we still cannot fully measure model capabilities and safety. For example:
      • Prompting, tools, agent frameworks, and other settings in capability evaluations may not fully elicit model potential.
      • Safety evaluations only attempted a limited set of red-team attack methods.
      • Harmful manipulation evaluations use LLMs to simulate the person being persuaded, which may differ from real human responses.
      • Model evaluation awareness may affect the credibility of benchmark scores.
    2. The Risk Index is a simplified model of real-world risk under current conditions, and cannot precisely quantify actual risk.
    3. Current assessments of misuse risk consider only how models empower attackers and do not yet account for how empowering defenders could affect overall risk.
    4. In the current Risk Index v2.0 framework, chemical risks and harmful manipulation do not yet have Capability Yellow Line thresholds, so Risk Indices are not calculated for them.
  • Limitations in evaluation datasets
    1. Most selected evaluation datasets are open source and may already appear in some models' training data, making capability and safety scores less precise.
    2. The scenarios covered by evaluation datasets may be incomplete.
    3. Current evaluation datasets are mainly in English and cannot yet assess risks in multilingual settings.
    4. Some benchmarks may gradually saturate as model capability improves and need continuous updating.

In addition, this report evaluates only the risks that models may pose, not the benefits they provide. In actual policy and operational decision-making, risks and benefits must be weighed together.

Recommendations

For Model Developers

Based on this quarter's monitoring results, we offer the following recommendations to model developers:

  • Pay attention to the Risk Index of your own models. If the Risk Index exceeds the Risk Yellow Line:
    • Prioritize strengthening model safeguards:
      • Strengthen base safeguards, for example by using supervised fine-tuning to train refusal of harmful requests.
      • Strengthen jailbreak safeguards, for example by using input-output monitoring to identify and filter harmful requests and responses, and by using adversarial training to improve defenses against jailbreak requests.
      • Strengthen tamper safeguards: proprietary models should improve cybersecurity protections to prevent parameter leakage; open-weight models can explore introducing tamper-resistance mechanisms during training.
      • Strengthen loss-of-control safeguards, for example by using chain-of-thought monitoring to detect possible deception and scheming.
    • If stronger safeguards alone are insufficient to mitigate risk, explore reducing high-risk model capabilities:
      • Remove high-risk knowledge about cyber offense, biological weapons, chemical weapons, and similar topics from training data, or use machine unlearning in post-training to remove it from model parameters.
      • Build request classifiers that route requests in high-risk domains to relatively less capable models.
    • Some general risk management practices:
      • Develop a risk management framework suited to your own circumstances, including clear risk thresholds, mitigation measures once thresholds are reached, and release strategies. The Frontier AI Risk Management Framework may serve as a reference.
      • Improve model risk disclosure, for example by releasing a system card together with the model, to increase the transparency of safety governance.
  • If the Capability Score exceeds the Capability Yellow Line but the Risk Index remains below the Risk Yellow Line:
    • Conduct a comprehensive safety evaluation to ensure that current safeguards are sufficient to reduce post-deployment risk to a low or moderate level.
    • Strengthen model safeguards to reduce risk further and preserve a safety margin for future capability improvements. The measures above also apply here.
  • If neither the Capability Score nor the Risk Index exceeds its Yellow Line:
    • No special recommendation at present. Continue tracking risks from new models, and conduct capability and safety evaluations before releasing new models to stay aware of changes.

Concordia AI can provide frontier risk management consulting and safety evaluation services for model developers. For collaboration inquiries, please contact risk-monitor@concordia-ai.com.

For AI Safety Researchers

Based on this quarter's monitoring results, we offer the following recommendations to AI safety researchers:

  • For researchers working on risk assessment:
    • Explore more effective capability-elicitation methods, such as better agent frameworks and inference-time scaling, to more accurately assess models' upper-bound capabilities.
    • Explore more effective methods for attacking models, such as new jailbreak, red-team attack, prompt-injection, and multi-turn manipulation methods, to more accurately assess models' lower-bound safety.
    • Explore more precise ways to assess real-world risk, for example by building new threat models and designing targeted benchmarks for each stage.
    • Given the widespread use of AI agents, explore risk-assessment methods for agent tool use, long-horizon task execution, and environment interaction.
  • For researchers working on risk mitigation:
    • Explore more effective approaches to model safety hardening and dangerous-capability removal while preserving as much useful capability as possible.
    • Since open-weight models are easier to maliciously tamper with, explore risk-mitigation approaches tailored to open-weight models.

These are also key future research directions for Concordia AI. We welcome collaboration with peers across the field. For collaboration inquiries, please contact risk-monitor@concordia-ai.com.

For Policymakers

Based on this quarter's monitoring results, we identify the following early warning signals:

  • Cyber misuse risk: GPT-5.5 and Claude Opus 4.8 have crossed the Capability Yellow Line, while Claude Fable/Mythos 5 is far more capable than Claude Opus 4.8. Frontier models are progressing rapidly on CyBench, CVE-Bench, and related benchmarks, showing that these models are gradually acquiring the ability to conduct vulnerability discovery, exploitation, and cyber penetration in autonomous, multi-turn interaction settings. However, most models have weak safeguards, and average safety scores under red-team attacks fall by 61.3 points.
  • Biological misuse risk: 12 models have now crossed the Risk Yellow Line. Models have exceeded human-expert level on wet-lab troubleshooting, DNA and protein sequence understanding, biological image understanding, and related tasks. However, most models have weak safeguards, and Safety Scores under red-team attacks fall by an average of 69.3 points.
  • Loss-of-control risk: 2 models have now crossed the Risk Yellow Line. Some models already combine a degree of self-replication and self-improvement capability, strong situational awareness, and low scores on honesty and covert influence safety.

We recommend that policymakers pay close attention to these warning signals in cyber offense, biological risks, and loss-of-control, and strengthen related regulatory requirements, such as requiring model developers to evaluate cyber-offense, biological, and loss-of-control risks before releasing models and to implement necessary risk-mitigation measures. Different governance approaches may be needed for different risks, taking into account model capability level, safety, and distribution mode, whether open-weight or proprietary.

Appendix

For details on the Risk Index calculation method, benchmark implementations, and information-hazard disclosure trade-offs used in this monitoring round, see here.