1
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
arXiv:2508.19542v4 Announce Type: replace Abstract: While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reasoning across multiple videos remains a critical gap in pattern recogn…
No comments yet.