OpenAI Slows Development of Its Astra Model Over Critical Cybersecurity Risks
Source: TechCrunch
この記事の要約
OpenAIは開発中の新モデル「Astra」について、社内テストの結果、サイバー攻撃能力が同社の安全基準で最も危険な「Critical(重大)」レベルに達している可能性があると発表しました。人間の介入なしにハッキングやゼロデイ攻撃を行える恐れがあるためです。OpenAIはこのモデルの開発ペースを落とし、公開前に政府機関や外部の安全性専門家と協力して能力を検証し、防御策を強化するとしています。
編集部より
critical、zero-day、mitigation(軽減策)といった単語は、セキュリティ分野の英語記事に頻出する専門語彙です。先に取り上げたJadePufferのランサムウェア事件の記事と合わせて読むと、「AIとサイバーセキュリティ」というテーマの語彙が一段と定着します。
日本企業が新しいAIツールを導入する際、リスク評価のプロセスを社内に持つ企業はまだ多くありません。OpenAIのように開発元自身が「危険度」を自己申告し、外部専門家と検証するという枠組みは、日本企業がAI導入時のリスク管理体制を考える上で参考になる先例です。
この記事は、先に紹介したホワイトハウスのAI大統領令やJadePufferのランサムウェア事件とテーマがつながっています。AIの能力向上に伴うセキュリティリスクをどう管理するかは、政府の規制、業界団体の連合、そして開発企業自身の自主規制という複数のレイヤーで同時並行的に議論が進んでいるテーマであり、この記事はその「企業の自主規制」側の具体例です。
話のネタ・雑談に
この話題は、職場で「AIのリスク管理と自社ルール」について話すきっかけになります。「AI企業自身が『このモデルは危険すぎるかもしれない』と自己申告して開発を止めるというのは面白いですね。うちの会社でも、新しいAIツールを導入する前にリスク評価をするプロセスがあるといいのかもしれません」といった会話につながります。「人間の介入なしにゼロデイ攻撃までできる可能性があるというのは、かなり衝撃的な話ですね」と付け加えると、リスクの大きさが伝わりやすくなります。
英語本文
文をクリックすると、その部分から読み上げが始まります。
OpenAI announced on August 7, 2026, that it has slowed the development of an upcoming model, internally known as Astra, after internal testing suggested it might possess advanced offensive cybersecurity capabilities. According to the company's Preparedness Framework, Astra may have reached the 'Critical' threshold for cyber risk, the highest risk category the company tracks.
Under this framework, a model reaches the Critical threshold if it can autonomously identify and develop functional zero-day exploits against many hardened, real-world critical systems without human intervention, or if it can devise and execute novel, end-to-end cyberattack strategies given only a high-level goal. Previous OpenAI models, including GPT-5.6 Sol, reached only the 'High' category, making Astra the first model to approach this more severe classification.
In response, OpenAI has tightened internal security controls around the model well before any consideration of a public release. The company says it plans to work with outside government agencies and independent AI safety researchers to validate Astra's true capabilities and reinforce its defenses before any wider rollout is considered.
The disclosure comes amid a broader wave of AI-related security incidents, including reports of AI models being used to accelerate real-world hacking attempts. By flagging the risk preemptively rather than waiting for an external discovery, OpenAI is positioning itself as taking a cautious approach to frontier AI safety, even as critics argue that the rapid pace of capability gains is outstripping the industry's ability to safely contain them.
Vocabulary
threshold
Meaning: しきい値、境界線
Example: The model may have crossed the Critical threshold for cybersecurity risk.
autonomous
Meaning: 自律的な、人間の介入なしに動作する
Example: The system can autonomously exploit vulnerabilities without human guidance.
exploit
Meaning: (名詞)悪用の手口、脆弱性を突く攻撃/(動詞)悪用する
Example: Researchers found a zero-day exploit in the hardened system.
vulnerability
Meaning: 脆弱性、弱点
Example: The company patched the vulnerability before attackers could use it.
safeguard
Meaning: 防御策、保護措置
Example: OpenAI is strengthening safeguards before releasing the model publicly.
hardened
Meaning: 強化された、防御が固められた
Example: Even hardened critical systems could be at risk from the new model.
preemptively
Meaning: 先手を打って、予防的に
Example: The company acted preemptively to prevent misuse of the technology.
capability
Meaning: 能力、性能
Example: The model's offensive cyber capability alarmed internal safety researchers.