1TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMsTimeLens2:面向通用视频时序定位的多模态大模型TimeLens2:汎用ビデオ時間グラウンディング137 UPVOTES · YUHAN ZHU ET AL. · MULTIMEDIA COMPUTING GROUP-NANJING UNIVERSITY · ARXIV 2607.17423Video MLLMs can describe what happens in a video but rarely when the supporting evidence occurs. TimeLens2 targets generalist temporal grounding - locating the moments that back an answer across video tasks.视频多模态大模型能描述视频里发生了什么,却很少能说出证据出现在何时。TimeLens2 面向通用时序定位:在各类视频任务中找到支撑答案的具体时刻。動画MLLMは何が起きたかは説明できても、その根拠がいつ現れるかはほとんど示せない。TimeLens2は汎用の時間グラウンディング、つまり答えを裏付ける瞬間の特定に挑む。
5BadWAM: When World-Action Models Dream Right but Act WrongBadWAM:当世界-行动模型梦得对却做得错BadWAM:世界行動モデルが正しく夢を見て誤って動くとき37 UPVOTES · QI LI ET AL. · ARXIV 2607.15207Examines world-action models that predict the future correctly yet still act wrongly - an embodied-control failure mode where dreaming right does not mean acting right.研究“世界-行动模型”预测未来正确却行动错误的现象——一种具身控制的失效模式:梦得对不代表做得对。世界行動モデル(WAM)が未来を正しく予測しながら誤った行動を取る現象を分析。「正しく夢を見ても正しく動ける」とは限らないという身体性制御の失敗モードを示す。