The highest-impact AI research papers trending today on arXiv and Hugging Face. Curated with fast AI summaries, community discussions, and open-source GitHub code.
Siting Li, Zhengyang Wang, +3 more
Using a controlled autoregressive testbed, the study analyzes task-specific validation losses during multimodal pretraining to evaluate how image tokenizer design affects joint text-image modeling and downstream performance.
Parinthapat Pengpun, Simran Khanuja, +1 more
A reasoning-capable vision-language model that iteratively retrieves and reasons over Wikipedia improves multimodal entity linking for rare entities defined by knowledge-graph structure.
Mehrnaz Mofakhami, Ananya Sahu, +6 more
Optimized supervised fine-tuning data composition enables reasoning models to consistently process and respond in diverse non-English languages without requiring reasoning supervision in each target language.
Kaushalraj Puwar, B. Thangaraju
A proxy layer isolates critical ROS 2 subscribers from degraded ones via topic splitting and dynamic rate control to eliminate DDS backpressure.
Vikash Singh, Debargha Ganguly, +6 more
Neurosymbolic reasoning is vulnerable to incorrect but verdict-matching formal translations, which are addressed by a generative verification method that scores reference equivalence without an oracle and improves downstream accuracy.
Yiling Ma, Yilun Zhao, +4 more
ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision.
Yiling Ma, Yilun Zhao, +3 more
The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models.
Sizhe Zhao, Haozhe Xie, +6 more
MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency.
Rongcan Pei, Zhepei Wei, +4 more
Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities.
Vladislav Bargatin, Alexander Yakovenko, +2 more
FreeFlow is a hierarchical transformer for optical flow that eliminates task-specific inductive biases and achieves state-of-the-art accuracy using window, shifted-window, and global attention.