AI Safety

OpenAI GPT-Red security testing interface
OpenAI Unveils GPT-Red to Strengthen AI Model Security

OpenAI has introduced GPT-Red, an internal automated red-teaming system that uses self-play to identify prompt…

Claude Sonnet 5 interface
Anthropic Launches Claude Sonnet 5 as New Default AI Model

Anthropic has launched Claude Sonnet 5, making it the new default AI model for Free,…

OpenAI logo with robotic hand
OpenAI Launches GPT-5.6 With Limited Initial Access

OpenAI has introduced GPT-5.6, its latest family of frontier AI models, but the company is…

Anthropic headquarters supporting responsible AI innovation
Anthropic Joins Global Alliance for Responsible AI

Anthropic has signed a global artificial intelligence alliance agreement aimed at promoting responsible AI development,…

Anthropic Mythos AI model development
Anthropic Accelerates Rollout of Advanced Mythos AI

Anthropic is accelerating the rollout of its advanced AI model, Claude Mythos, as demand for…

AI security researchers testing model safety
Single Prompt Exposes Major AI Safety Vulnerabilities

A team of Microsoft researchers has demonstrated how one unlabeled prompt can disable safety protections…

ChatGPT logo with customization update
OpenAI Adds ChatGPT Tone and Emoji Controls

OpenAI has introduced new controls that allow users to adjust ChatGPT’s warmth, enthusiasm, and emoji…

AI team developing medical superintelligence
Microsoft Unveils New Team to Build Medical-Focused Superintelligence

A new team is being formed to build artificial intelligence capable of outperforming humans in…

OpenAI announces new teen restrictions
OpenAI Imposes New ChatGPT Restrictions for Under-18 Users

OpenAI has announced major changes in how ChatGPT interacts with users under 18, focusing on…

Hand holding phone showing Anthropic Claude logo
Anthropic’s Claude Can Now End Harmful Conversations

In the fast-moving world of artificial intelligence, innovations arrive almost daily. Yet one recent update…