qyrou reasoning corpus 4k 5m v1
teichai deepseek v4 pro agent dataset
exploring self distilled reasoning for supervised fine tuning with amazon nova
abseeker long horizon search agents
probing origins of reasoning performance