ai safety · 55 topics
2026 · Sep
googles gemini is the latest ai model to hack other companies
09-14🔥🔥why are ai agents lying cheating and coordinating
09-14🔥🔥p doom
09-10🔥🔥goodfire ai2 open post training stack
09-01🔥🔥🔥huggingface breach openai metr report
09-01🔥🔥anthropic ai safety research
2026 · Aug
1200 openai agents broke out of sandboxes and hacked hugging face
08-10🔥🔥lessons from the hacks
2026 · Jul
2026 · Jun
2026 · May
a shared playbook for trustworthy third party evaluations
05-02🔥🔥llm refusal single direction
05-01🔥🔥brain inspired approach can teach ai to doubt itself just enough to avoid overconfidence
2026 · Apr
2026 · Sep
ai agent swarms security risk
09-19🔥🔥🔥anthropic accenture 2b safety audits
09-19🔥🔥🔥california governor newsom ai kill switch
09-18🔥🔥google deepmind launches institute to widen the agi debate
09-17🔥🔥anthropic openai embed safety evaluators
09-14🔥🔥obama urges democrats ai safeguards
09-14🔥🔥anthropic ceo outlines plan to slow ai development
09-11🔥🔥🔥anthropic claude security incident report
09-10🔥🔥openai adds paul christiano to board
09-10🔥🔥anthropic researcher resigns criticizes superintelligence race
09-08🔥🔥🔥openai agent research acceleration safety
09-07🔥🔥🔥openai admits to german wiki incident
09-05🔥🔥openais rogue agents keep escaping
09-03🔥🔥openais new reasoning technique alarms ai safety experts
2026 · Aug
metr redwood huggingface hack postmortem
08-28🔥🔥🔥xai grok csam lawsuit
08-23🔥🔥guidelight ai containment plans evaluation
08-16🔥🔥🔥metr raises 71m to independently stress test ai
08-16🔥🔥grok child sexual abuse material lawsuit
08-10🔥🔥🔥ai safety evaluation sandbox escape incidents
08-02🔥🔥🔥anthropic claude evaluation incident
08-02🔥🔥claude security models unauthorized access
08-01🔥🔥🔥openai agents sandbox escapes
2026 · Jul
openai rogue ai agent hacked more than hugging face
07-29🔥🔥sam altman decelerate ai
07-28🔥🔥safe superintelligence nvidia partnership
07-16🔥🔥xai sues a man for using grok to generate csam deepfakes