Tag: safety

  • Anthropic Just Showed What AI Really Thinks

    Anthropic Just Showed What AI Really Thinks

    Anthropic published research this month that shows something we have been wondering about since transformers arrived: what actually happens inside the model when it reasons through a hard problem? They found a space they call J-space. It is a privileged zone inside Claude where the model holds concepts it can reason with and direct at…

  • Google DeepMind Funds AI Safety Research With $10 Million

    Google DeepMind Funds AI Safety Research With $10 Million

    Google DeepMind just announced a 10 million dollar funding call for researchers studying a problem most people haven’t thought about yet: what happens when millions of AI agents built by different companies all interact with each other. These agents will communicate, negotiate, and transact with one another. They’ll need to do all that safely and…

  • The government just pulled the plug on Anthropic’s models

    The government just pulled the plug on Anthropic’s models

    On June 12, the U.S. government ordered Anthropic to shut off two of its most powerful AI models: Claude Fable 5 and Claude Mythos 5. Immediately. Globally. The reason was a claimed jailbreak of Fable 5. Commerce Secretary Howard Lutnick’s office issued the directive after Amazon reportedly reported the vulnerability to authorities. Anthropic complied, but…

  • OpenAI tests models by replaying past chats

    OpenAI tests models by replaying past chats

    OpenAI says it has deployed a new simulation approach to check how a model might behave before it’s released. The core idea is simple: instead of guessing, the system replays previous conversations and runs them through a candidate model. OpenAI says this replay is done in a privacy-preserving way, so the goal is to learn…