🏷 Tag
alignment · 25 topics
research (13)
2026 · Sep
09-17🔥🔥
openai launches misalignment disclosure framework
09-14🔥🔥why are ai agents lying cheating and coordinating
09-09🔥🔥safety for whom refusing the right subset of a topic
09-01🔥🔥anthropic trained hacker opus on hackable environments
2026 · Aug
08-29🔥🔥
anthropic automated alignment researcher
08-29🔥🔥anthropic aar claude opus alignment
08-28🔥🔥from preferences to principles
08-23🔥🔥🔥obliterator qwen38 27b
08-15🔥🔥anthropic conceptual reasoning index
08-10🔥🔥lessons from the hacks
2026 · Jul
2026 · May
2026 · Apr
papers (6)
2026 · Aug
08-11🔥🔥
sharding prevents llm oversight failures
08-03🔥🔥cort counterfactual replay for token level rubric guided policy optimization
2026 · Jul
07-31🔥🔥
do models fake alignment without clear consequences
07-30🔥🔥beyond shapley influence based data auditing pipeline for llm alignment and evaluation
07-21🔥🔥beyond a single direction
2026 · Jun
business (6)