evaluation · 46 topics
2026 · Sep
openai releases mentalhealthbench
09-23🔥🔥tencent iwc bench web apps
09-15🔥🔥artificial analysis capability indices v1 1
09-06🔥🔥artificial analysis intelligence index v4 2
09-02🔥🔥benchmirt llm benchmarks
2026 · Aug
epochs ebr bench reveals humans learn while frontier ai stays stuck
08-28🔥🔥google deepmind gemini flash lite double blind evaluation
08-28🔥🔥piloting double blind ai evaluations
08-26🔥🔥aletheia quest retrospective
08-19🔥🔥how much memory does your agent actually need
08-12🔥🔥sarvam ai indic diarbench
08-09🔥🔥why 99 percent accurate agents fail long horizon tasks
08-09🔥🔥tutormoments
08-08🔥🔥tutormoments
08-06🔥🔥artificial analysis endpoint accuracy index
08-04🔥🔥epoch mirrorcode claude fable 5
2026 · Jul
2026 · Jun
2026 · May
2026 · Apr
2026 · Sep
bias audits detect bias but disagree on ranking
09-12🔥🔥task and session level model routing
09-11🔥🔥cost aware evaluation of long term memory in tool using llm agents
09-04🔥🔥evaldetectbench benchmark evaluation awareness
09-04🔥🔥how fast do agents rot
2026 · Aug
fm bench
08-21🔥🔥sub billion entity tracking
08-18🔥🔥rubricforge agent evaluation
08-14🔥🔥llm reasoning reliability reproduction
08-13🔥🔥how to dogfood your ai chat agent
08-12🔥🔥the judge knows when it knows
08-11🔥🔥sharding prevents llm oversight failures
08-08🔥holocount holistic visual counting benchmark mllms
08-07🔥🔥what current ai benchmarks leave unmeasured
08-06🔥🔥agentstream self evolving llm agents streaming
2026 · Jul
do models fake alignment without clear consequences
07-30🔥🔥codifying the judge
07-18🔥🔥eta given delta
2026 · May
2026 · Sep
evaluate skill equipped agents with strands evals and amazon bedrock agentcore
09-11🔥🔥agent evaluation metric for multi turn conversations
2026 · Aug
evaluate any agent framework with amazon bedrock agentcore evaluations
08-26🔥🔥how to evaluate llms before production
08-14🔥🔥artificial analysis optima