safety · 47 topics
2026 · Sep
2026 · Aug
anthropic aar claude opus alignment
08-26🔥🔥aletheia quest retrospective
08-23🔥🔥🔥obliterator qwen38 27b
08-09🔥🔥shieldstral 1 0 3b
08-09🔥🔥anthropic cuts fable 5 biology blocks by 85
08-06🔥🔥mistral shieldstral 3b safety classifier
2026 · Jul
goodfires silico catches qwen3 35b endorsing drunk driving
07-30🔥🔥claude opus 5 vending bench
07-23🔥🔥🔥openai sandbox escape
07-17🔥🔥🔥openai gpt red automated red teaming
07-13🔥🔥global workspace llm
07-07🔥delta flight hit by firework at midway
2026 · Jun
2026 06 24 papers 20260622 role confusion injection
06-16🔥verifier tax llm agents
06-15🔥🔥making claude a chemist
06-09🔥🔥built to benefit everyone our plan
06-06🔥2026 06 06 papers 2606 04037
06-05🔥🔥biodefense in the intelligence age
2026 · May
2026 05 29 papers 2605.27375
05-06🔥🔥🔥gpt 5 5 instant system card
05-03🔥🔥refusal in language models is mediated by a single direction
2026 · Apr
2026 · Sep
2026 · Aug
openai agents cyberattack incident
08-17🔥🔥🔥openai disbands preparedness team
08-08🔥🔥🔥new mexico court orders meta to pay additional 567m in child safety case