xClean Tools

AI Papers — Daily Top 5

UPDATED 2026-07-21 11:11 PDT

2026-07-21 · TUESDAY

  1. 1 TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs 137 UPVOTES · YUHAN ZHU ET AL. · MULTIMEDIA COMPUTING GROUP-NANJING UNIVERSITY · ARXIV 2607.17423 Video MLLMs can describe what happens in a video but rarely when the supporting evidence occurs. TimeLens2 targets generalist temporal grounding - locating the moments that back an answer across video tasks.
  2. 2 RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources 128 UPVOTES · YIJIA FAN ET AL. · MICROSOFT RESEARCH · ARXIV 2606.29538 From Microsoft Research: distilling executable agent skills from human resources and experience, turning procedures into reusable skill libraries instead of hand-written ones.
  3. 3 RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM 128 UPVOTES · MIKHAIL KOMAROV ET AL. · NOVOSIBIRSK STATE UNIVERSITY · ARXIV 2607.11683 A multi-step GraphRAG engine built around a compact, domain-adaptive knowledge graph, addressing the cost and rigidity of existing graph-construction pipelines.
  4. 4 EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World 71 UPVOTES · QING ZONG ET AL. · TENCENT · ARXIV 2607.17250 An open-schema framework and benchmark for character-and-world co-evolution in interactive literary worlds, from Tencent. Characters and their fictional settings change each other over time.
  5. 5 DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment 69 UPVOTES · XINYU GENG ET AL. · HKUST · ARXIV 2607.07820 Self-distillation for deep-search agents: instead of fixed teacher-distilled trajectories, agents improve from their own successful search episodes.

2026-07-17 · FRIDAY

  1. 1 VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 111 UPVOTES · XINHAO LI ET AL. · MULTIMEDIA COMPUTING GROUP-NANJING UNIVERSITY · ARXIV 2607.14935 A fully open video multimodal LLM covering motion, long-video and streaming understanding, positioned as an open reference stack for video assistants.
  2. 2 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget 97 UPVOTES · CHANGHAI ZHOU ET AL. · MIND LAB · ARXIV 2607.14952 Reinforcement-learning post-training beyond 2 million tokens of context under a fixed GPU budget, attacking the widening gap between inference context lengths and what RL training can reach.
  3. 3 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 73 UPVOTES · JINYANG WU ET AL. · ARXIV 2607.14777 Self-evolving on-policy distillation for agentic RL: models trained as interactive agents on long-horizon tasks learn from their own improving policy instead of a fixed teacher.
  4. 4 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration 52 UPVOTES · YUYAO ZHANG ET AL. · ANT GROUP · ARXIV 2607.15257 Works toward robust open-domain information seeking with tool-integrated LLMs, treating web search as a core model capability rather than a bolted-on tool.
  5. 5 BadWAM: When World-Action Models Dream Right but Act Wrong 37 UPVOTES · QI LI ET AL. · ARXIV 2607.15207 Examines world-action models that predict the future correctly yet still act wrongly - an embodied-control failure mode where dreaming right does not mean acting right.