benchmarking opus 5 on slop code bench
arc agi leaderboard
is it agentic enough
transformer scalability crisis
drawing the mona lisa with gpt 5 6 claude gemini and grok
a robot is sprinting towards you do you want it running on claude or grok
in the weights is your new ai centric vanity search
sysadmin measuring instrumental power seeking in frontier ai
relay bench evaluating llms on multi domain reasoning chains