On ScienceQA, Qwen2.5-VL-7B performance drops from 80.76% to 45.48% when reasoning instructions break answer decoding.
SLCA-GRPO resolves cross-segment credit misattribution in tool-calling RL, outperforming GRPO by +2.53 pp on a 7B backbone.
Researchers introduced scale-normalized graph Dirichlet roughness ($D_{\text{IQR}}$) to show semi-empirical baselines improve predictions over complex local descriptors.
Researchers found that $L_0$ ablation reduced object hallucination by 31% in twenty-five VQ-tokenized vision-language models.
Researchers evaluated Large Language Models on two Computer Science exams: a Computer Vision test with 570 dual-graded students and a Machine Learning test with 1,038 dual-graded students. They tested 171 configurations, including closed and open-weight models, using a "strict grader" prompt containing a "never give partial credit" instruction. While the best model achieved a mean absolute error of 1.64/35, the strict preamble caused 14 of 17 open-weight models to exceed an MAE of 8 or stop grading entirely. Light LoRA fine-tuning on approximately 3,900 pooled examples repaired this vulnerability, bringing five small open models to parity or better with human graders and reducing sensitivity to harsh personas to an MAE of 0.32. The authors released an anonymized dataset, ablation grid, and pipelines.
FB-GDM outperforms $\Pi$GDM by up to 14 dB on CelebA-HQ without hyperparameter tuning.
Researchers introduced DSC-Gate to reuse historical credit, reducing mean new tool steps from 472 to 286.
A transformer trained from scratch using only last-token rewards learned iterated non-Abelian group multiplication chains.