The highest-impact AI research papers published today on arXiv and Hugging Face. Curated with fast AI summaries, community discussions, and open-source GitHub code.
Mayank Singh, Michele Stoppa, +8 more
Luce unifies geometry and PBR materials in a voxelized Gaussian cloud, using a variational autoencoder and rectified-flow transformer to generate relightable 3D assets from single images.
Yufan Wu, Yinghui He, +5 more
CritICL improves LLM reasoning at inference time by using structured failure patterns from weaker models as critique-based guidance, reducing generation and token costs.
Jianbo Zhou, Boyuan Zhao, +9 more
TacForcing is a streaming action-generation framework that integrates real-time tactile feedback during execution via a streaming action expert and execution-aware tactile attention, improving contact-rich manipulation.
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.
Zhiyuan Li, Chi-Man Pun, +3 more
EditaLive enables real-time human-centric live-stream video editing by adapting an image animation model to causal streaming generation with distilled two-step sampling and sparse attention.
Yang Xiao, Yusong Sun, +8 more
PILOT enables live self-improvement by allowing a supervisor to steer active workers and distilling execution experience into reusable skills, improving accuracy and efficiency.
Pengfei Zhou, Hexin Wang, +6 more
Game engines provide executable verification and long-horizon trajectories for reinforcement learning post-training of spatial world models, motivating a human-engine verification paradigm.
Youtian Lin, Yikang Yang, +6 more
Procedura is a 3D modeling agent that generates editable, part-structured procedural assemblies with sharp geometry and validated articulation from text prompts.
Chenyang Wu, Fuchen Long, +5 more
An agentic framework combining LLMs and VLMs enables consistent, multi-instruction editing of long multi-shot videos while preserving spatiotemporal structure.
Yuncheng Guo, Zhanqiu Zhang, +2 more
GameWAM is a unified world-action model for native video-game control that jointly predicts future visuals and executable keyboard-mouse actions using block-causal flow matching, mode-specific distributions, and block-cycle replanning.
Test-Time Policy Optimization enables label-free test-time training for mathematical reasoning by asymmetrically distilling agreeing rollouts and penalizing disagreeing ones, matching supervised performance.
Hengyuan Xu, Wei Cheng, +5 more
Aphanta evaluates when image-editing intermediates improve multimodal reasoning by testing direct, editor-generated, and idealized visual states across tasks.
Zhiyuan Li, Linyuan Gao, +4 more
CaSKG calibrates procedural skill relations via counterfactual-causal graph construction to improve compact, executable retrieval for LLM agents.
TaoLive AIGC LLM Team, Yuhan Sun, +8 more
Harness-Aware Training enables compact models to adapt to evolving digital-avatar harness configurations with low latency and high accuracy.
Abhilash Nandy, Rahul Seetharaman, +6 more
CaRGo-T improves multimodal humor understanding by modeling causal relationships as graph-based reasoning structures interpreted by vision-language models.
Yuandong Pu, Le Zhuo, +12 more
The study formalizes probabilistic alignment for world models, introduces PAWBench and PAWEval to evaluate video generators as stochastic samplers, and finds current models fail to match reference behavior distributions.
Tianjie Ju, Zheng Wu, +16 more
UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior.
Jiaming Zhou, Qihang Zhang, +10 more
Zero-WAM enables robotic manipulation of unseen tasks by conditioning a causal video-action model on in-context human video guidance, supported by an automatically generated dataset and a future-chunk prediction objective.
Liyan Tang, Cyrus Rashtchian, +4 more
WikiSkill co-evolves reusable agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models.
Shiyi Zhang, Mushui Liu, +9 more
Self-OPD eliminates task-specific teachers in flow matching by using self-explored stochastic branches and normalized advantages to optimize the velocity field for multi-objective alignment.
Yunpeng Ba, Zhi Zheng, +8 more
Evolution strategies improve reasoning diversity and Pass@K over GRPO through sparse functional updates and population diversity, supporting a hybrid training approach.
Xiaoyu Zhan, Xinyu Wang, +7 more
Magpie is a real-time generative rendering system that separates gameplay logic from visual generation to preserve interactive designability while reducing asset requirements for game prototypes.
Xingshan Zeng, Zishan Xu, +12 more
Agentic data generation is framed as constrained distribution design over factorized experience tuples, emphasizing execution-grounded accuracy, learner-relative complexity, and diversity rather than scale alone.
Md Abrar Jahin, Md Rizwan Parvez
A benchmark of contrastive GUI instructions reveals that vision-language models mostly fail to localize interface elements rather than misunderstand spatial relations, though marking candidates substantially improves selection.
Bobby Cheng, Adam Gaber, +5 more
Multilingual self-play reveals that large language models exhibit significant cross-lingual skill inconsistencies in reasoning and strategy, partly recoverable by altering intermediate reasoning language.