👍 84
07/08 08:00
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through sc
中文介绍 提出Deep Native Structural Reasoning方法,直接对原始结构证据进行深度推理,实现材料、化学、生物领域结构-性质关系的准确、跨学科、透明理解,无需人工特征工程,提升可解释性与泛化能力。
👍 54
07/02 08:00
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video unde
中文介绍 Video-Oasis重新审视视频理解评估,指出当前基准难以区分性能来源是视觉感知、语言推理还是知识先验,提出共享评价标准以更公平衡量Video-LLM的真实理解能力。
👍 51
07/08 08:00
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory b
中文介绍 针对VLA模型在长时任务中的马尔可夫假设局限,提出Dual Latent Memory,融合扩展观察窗口与历史检索,增强机器人操控对时间依赖任务的记忆与推理能力。
👍 31
07/08 08:00
We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our model achieves an unbounded interaction horizon while maintaining consistent output quality, benefiting from a carefully crafted causal pretraining pa
中文介绍 LingBot-World 2.0实现无限交互世界,通过因果预训练范式在无限长交互中保持输出一致性,支持多样化场景生成,适用于大规模虚拟环境与交互仿真。
👍 30
07/09 08:00
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Be
中文介绍 提出IdeaGene-Bench基准,评估AI系统追踪科学思想遗传谱系的能力,要求模型基于前人工作推理并生成有谱系依据的新想法,衡量科学推理深度。
👍 27
07/09 08:00
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-ba
中文介绍 LongE2V利用预训练视频扩散先验,联合处理事件流视频重建、预测与帧插值,在长时间跨度下保持稳定,避免回归方法的纹理模糊,实现高质量事件视频恢复。
👍 27
07/09 08:00
The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they
中文介绍 UniClawBench提供通用基准,评估能操作工具、协助用户的主动智能体在真实世界任务中的表现,弥补现有基准在任务多样性和环境真实性方面的不足。
👍 27
07/05 08:00
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corp
中文介绍 提出策略感知提示自适应方法,在非可验证指令跟随的RL中,根据策略演化动态调整LLM评判的评分细则与训练提示,提升奖励信号的准确性与对齐效果。
👍 17
07/09 08:00
In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas36
中文介绍 Canvas360两阶段框架:先通过几何感知预训练获得大尺度全景数据表征,再微调实现上下文全景生成,有效缓解训练数据不足问题,提升生成质量。
👍 15
07/08 08:00
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context e
中文介绍 Jet-Long提出动态双焦点RoPE,高效扩展LLM上下文窗口至远超预训练长度,通过聚焦重要位置与稀疏化注意力,在RAG、代码仓库等任务中实现零样本长上下文扩展。
👍 15
06/29 08:00
JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of tens of billions of SKUs. At this scale, high-quality, structured item knowledge underpins a better consumer experience, lower management costs, and higher opera
中文介绍 京东推出Oxygen AIIC V1,以LLM/VLM为核心的工业级商品理解管理系统,在数十亿SKU规模下自动提取结构化商品知识,提升用户体验与管理效率、降低成本。
👍 14
07/08 08:00
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient
中文介绍 提出单轨迹异步优化用于智能体强化学习,相比同步批处理方法,在长程智能体任务中显著提升训练效率,适用于LLM后训练阶段的agentic RL场景。
👍 13
07/09 08:00
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known
中文介绍 OpenCoF通过视频生成实现推理,推理过程在时序帧中展开,模型学习生成因果连贯的视频帧序列,以视觉方式呈现逻辑推理,提升可解释性。
👍 12
07/04 08:00
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterati
中文介绍 CineMobile面向移动端图像生成视频,实现子弹时间、滑动变焦等电影级运镜,通过轻量Diffusion Transformer与多步迭代优化在设备上高效生成高质量动态效果。
👍 11
07/07 08:00
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the
中文介绍 RoboDojo提供统一的仿真与真实基准,系统评估通用机器人操控策略,覆盖多技能、长时域任务,在两种环境中全面衡量泛化能力,弥补现有基准局限。
👍 10
07/09 08:00
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed bey
中文介绍 提出Proactive Memory Agent,主动管理长时任务中的决策相关状态,通过记忆检索与上下文整合防止关键信息被长轨迹淹没,提升智能体在复杂长程任务中的表现。
👍 10
07/08 08:00
Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention models fall behind in long-context recall compared to softmax-attention-based transformer architectures. Increasing the state size of linear attention
中文介绍 Sparse Delta Memory利用稀疏机制扩展线性RNN的状态容量,在不增加计算量的前提下提升长上下文召回能力,弥合线性注意力与softmax注意力在长程记忆上的差距。
👍 10
07/03 08:00
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems still rely on researcher intuition to choose where information is stored, how observations are proces
中文介绍 提出自动化设计具身智能体架构的方法,搜索感知、记忆、规划、行动等模块的最优组合,替代手工设计,自动探索更大架构空间,提升任务适应性与性能。
👍 9
07/07 08:00
Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models have advanced scene reconstruction and embodied intelligence, scaling to entire ci
中文介绍 WildCity提供城市尺度的真实世界测试台,覆盖数平方公里场景,支持渲染、仿真与空间智能研究,用于评估AI在大型陌生环境中构建空间表征、导航与推理的能力。
👍 8
07/04 08:00
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability has focused on individual layers and circuits (where), leaving the token-level dynamics of multimodal computation during
中文介绍 研究MLLM在自回归生成中每个token的多模态计算动态,通过细粒度注意力分析揭示视觉与语言信息逐步整合的token级机制,深化对多模态生成过程的理解。