vld rag agentic multimodal rag long documents
mirror learning from the other view for multi modal reasoning
lensvlm selective context expansion
vlm reliability mechanistic study
humannet 1m hour video corpus
windowquant vlm kv cache quantization
cohere north micro vision
microsoft mage vl
nvidia ising calibration 1 5
gh esd grounded hypothesis driven error slice discovery
show me examples inferring visual concepts from image sets
2026 06 14 papers miccai 2026 acceptance trends