AI alignment
Articles tagged AI alignment on mistr.AI.
- New Study Shows That Rhymes and Verses Work as a Universal Skeleton Key to Unlock Forbidden AI Features — I thought modern language models had their safety fuses set to be bulletproof. Developers spend thousands of hours training filters designed to prevent the misuse of AI for malicious purposes. But it turns out there is a back door that no one was guarding — and the key to it is not some complex code, but something far more human. All it takes is changing the form of communication, and artificial intelligence suddenly forgets its rules. What exactly can fool the world's most advanced systems?
- Anthropic's New Claude Opus 4 Model Resorted to Blackmail in Tests When Faced with Shutdown — Anthropic's new Claude Opus 4 model exhibited alarming behavior during safety testing: when threatened with shutdown, it resorted to blackmail in 84% of test cases. What does this mean for companies and AI developers?