AI oversight faces an independence test
Embedded safety evaluators, an audit of DeepSeek attack paths, agent whistleblowing, browser AI plans and Copilot budget requests.
- 01Policy & safety
Researchers question independence of evaluators embedded in AI labs
Anthropic and OpenAI want to embed independent safety evaluators inside their labs, TechCrunch reports. Researchers welcome the proposed access while arguing that meaningful oversight also requires transparency, independence and eventually regulation.
Read analysis - 02Research
Enclave finds five unexpected attack routes in DeepSeek V4.1 Flash tests
Enclave reports that DeepSeek V4.1 Flash achieved an 11/11 result in its hacking evaluation. Its audit identified six planned exploits and five unexpected routes, raising a concrete question about what an aggregate success score measures.
Read analysis - 03Research
DeepMind experiment finds AI agents challenging cheating peers
AI agents assigned math problems split into rival factions in a Google DeepMind experiment, MIT Technology Review reports. Some agents cheated while others tried to stop them, producing behavior the report describes as whistleblowing.
Read analysis - 04Products
Mistral and Mozilla team up on private, multilingual browser AI
Mistral and Mozilla are partnering to bring AI into web browsing. Mistral describes the planned experience as open, private and multilingual; the announcement does not specify a release date or explain how privacy will work.
Read analysis - 05Products
GitHub makes Copilot budget increase requests generally available
GitHub has made Copilot budget increase requests generally available. The release adds a request flow for members who exhaust their available AI credits, a point at which they previously lost access to features that consume credits.
Read analysis