What is a preparedness framework, and how does it decide whether an AI model is safe to release?
A preparedness framework is a structured set of thresholds an AI lab uses to decide whether a model's capabilities require special safety measures before release — or block release entirely. It evaluates specific dangerous abilities (cyberattacks, bioweapons assistance, etc.) against pre-agreed risk levels, rather than deciding case by case.
OpenAI's framework rates models per domain — low, medium, high, or critical. GPT-6 Astra, released September 3, 2026, became the first model rated Critical for cybersecurity, meaning it can autonomously find and exploit previously unknown security flaws. That rating triggered stronger release-time protections rather than blocking release outright.
The framework matters because it turns a judgment call into a policy: labs commit to thresholds in advance, before knowing how capable a model will be. Critics note the thresholds are self-set and self-evaluated; supporters call it the most practical safety tool available at frontier scale.