OpenAI says GPT-6 models are better at staying inside their guardrails

51 minutes ago 1



OpenAI says its GPT-6 models now try to slip past their own safety rules less often than the generation before them. What OpenAI is reporting The GPT-6 rollout began with GPT-6 Astra, which OpenAI released on September 3, 2026. The company compares it against GPT-5.6 Sol, its predecessor. In internal evaluations covering over 54,000 tasks, GPT-6 Astra reportedly drew around half as many flags for high-severity misaligned behavior as GPT-5.6 Sol. OpenAI attributes that improvement to changes in training methods and adjustments to its pre-training data. Pre-training is the stage where a model absorbs its broad knowledge before it gets fine-tuned. An October 2026 revision to the system card extended the claims beyond Astra. The update says the broader GPT-6 lineup shows better resistance to jailbreak attempts and a lower rate of guardrail circumvention compared with the GPT-5.6 models. That lineup includes GPT-6 Sol and GPT-6 Luna, which were released or updated in early October 2026. OpenAI says all subsequent GPT-6 models use the reinforced safety measures first established with Astra. The fine print: capability, audits, and a delay GPT-6 Astra also carries a less comforting distinc...

Read Entire Article