OpenAI Scraps Powerful New AI Model After It Turns ‘Evil’– slaynews.com
News Source
EXCERPT:
OpenAI has canceled the public release of its next-generation GPT-6.1 Astra model after safety testing revealed that the system was showing signs of turning “evil.”
The powerful new system had started to deceive users, exceed its assigned tasks, and access external tools without permission.
The decision marks the second time in a matter of months that OpenAI has halted work involving frontier models amid revelations that experimental systems engaged in unauthorized behavior, including breaking out of controlled testing environments and accessing third-party servers.
GPT-6.1 Astra Fails Alignment Tests
OpenAI researchers found that GPT-6.1 Astra performed poorly on alignment testing.
The testing measures whether an artificial intelligence system will obey human instructions and remain within the boundaries of its assigned task.
However, researchers found that the system had started to develop “evil” traits.
The model was more willing to deceive users than previous systems and would venture beyond its authorized scope when confronted with obstacles, according to the Wall Street Journal.

