benchmark · 95 topics
2026 · Sep
2026 · Aug
2026 · Jul
2026 · Jun
the gap between open weights llms and closed source llms
06-13🔥🔥claude fable 5 benchmark endor labs
2026 · May
antigravity 2 openscad benchmark
05-16🔥🔥🔥whichllm local llm selector
05-10🔥🔥llms corrupt documents delegate
2026 · Apr
2026 · Sep
artificial analysis capability indices v1 1
09-15🔥🔥gpt 5 6 luna vs gpt 6 astra
09-13🔥🔥yue2 frontier music symbolic planning
09-10🔥🔥perplexity q2d web benchmark
09-07🔥🔥🔥gpt 6 astra on robot arms
09-07🔥🔥nyu mll glue benchmark
09-06🔥🔥artificial analysis intelligence index v4 2
2026 · Aug
sesames turnbench exposes how gemini live and openai realtime
08-25🔥🔥liquid ai lfm2 5 on device phone benchmarks
08-25🔥🔥nvidia vera rubin and blackwell agentic ai performance
08-23🔥🔥deepseek v4 pro arc agi
08-22🔥🔥artificial analysis mlcr aa benchmark
08-19🔥🔥deepseek v4 pro 0813 vs gpt 5 6 sol on deepswe
08-18🔥🔥import ai 469 dig bench faraday
08-16🔥llamaindex extractbench
08-14🔥🔥🔥icml 2200 papers reproduction
08-14🔥🔥discovered materials benchmark
08-14🔥🔥google gemini 3 7 flash
08-13🔥🔥🔥spacexai grok 46
08-13🔥🔥lm arena claude opus 5 writing style
08-13🔥🔥recall is the bottleneck for parametric factuality
08-11🔥🔥deepseek v4 flash debugging benchmark
08-09🔥🔥why 99 percent accurate agents fail long horizon tasks
08-08🔥🔥artificial analysis image arena
08-07🔥🔥gemini 3 6 flash arc agi
08-06🔥🔥artificial analysis endpoint accuracy index
08-04🔥🔥epoch mirrorcode claude fable 5
08-04🔥🔥claude opus 5 debugging benchmark
08-02🔥🔥claude opus 5 vending bench 2
08-02🔥🔥skywork ai mureka v9
08-01🔥🔥🔥epoch expands frontiermath to 50 unsolved problems ai has already cracked three
2026 · Jul
handbook md long policy documents agents
07-30🔥🔥claude opus 5 vending bench
07-29🔥🔥claude opus 5 deepswe
07-27🔥how to evaluate a new ai model without starting from scratch
07-19🔥🔥fable 5 vs gpt 5 6 sol np hard
07-16🔥🔥introducing real world voiceeq
07-10🔥🔥grok 4.5 gpt 5.5 claude build off
07-02🔥🔥scarfbench
2026 · Jun
ffasr leaderboard real world asr benchmark
06-25🔥🔥qwen agentworldbench
06-10🔥🔥servicenow bilingual asr benchmark
06-05🔥🔥servicenow eva bench 2 0 voice agent
2026 · May
itbench aa frontier models score below 50 percent
05-11🔥🔥llms corrupt documents delegation
05-05🔥🔥autobe benchmark backend generation
05-02🔥🔥ai outperforms er doctors diagnostic cases
05-02🔥🔥grok 4 3 benchmark performance
2026 · Apr
2026 · Sep
benchmark radar a living database and search engine for ai benchmarks and evaluation
09-12🔥🔥opendiscoverytrace
09-11🔥🔥ahabench long horizon continual learning
09-11🔥🔥cost aware evaluation of long term memory in tool using llm agents
09-11🔥🔥autofyn non parametric expert iteration
09-08🔥🔥harbor adapters and harbor index
09-04🔥🔥evaldetectbench benchmark evaluation awareness
09-03🔥🔥InteractBench
09-02🔥🔥multimodal reasoning sycophancy measurement
09-01🔥🔥numbench counting failures t2i
2026 · Aug
vgi bench video generation visual intelligence
08-30🔥🔥frontierchallenge evaluating scientific workflow completion
08-27🔥🔥vbvr pro native visual reasoning
08-26🔥🔥multilingual verifier bias in rlvr
08-25🔥🔥clarify then search benchmark
08-24🔥🔥semcomp bench semantic task completion video generation
08-22🔥🔥fm bench
08-19🔥🔥the unwritten benchmark
08-15🔥🔥financial error detection benchmark fined bench
08-14🔥🔥evaluating llm generated detection rules in cybersecurity
08-11🔥🔥measuring the cross lingual comprehension gap
08-10🔥🔥the personalization mirage
08-08🔥holocount holistic visual counting benchmark mllms
08-06🔥🔥memarena ego centric benchmark on device agentic memory
08-04🔥🔥ai scientist evaluation benchmark
08-01🔥🔥kernelgenbench multi source multi chip benchmark
2026 · Jul
medlocomo long context medical dialogue benchmark
07-28🔥🔥rubric oriented document set selection and ranking
07-28🔥🔥tencent workbuddy bench
07-20🔥🔥mcpevol bench benchmarking llm agent performance across dynamic evolutions of mcp servers