DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation Paper • 2606.29961 • Published Jun 29 • 10
DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation Paper • 2606.29961 • Published Jun 29 • 10
MultiHashFormer: Hash-based Generative Language Models Paper • 2606.28057 • Published Jun 26 • 22
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift Paper • 2601.05882 • Published Jan 9 • 21
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks Paper • 2601.03448 • Published Jan 6 • 13
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling Paper • 2510.11602 • Published Oct 13, 2025 • 15
How does the pre-training objective affect what large language models learn about linguistic properties? Paper • 2203.10415 • Published Mar 20, 2022 • 1
Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance? Paper • 2310.17271 • Published Oct 26, 2023 • 2
Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance Paper • 2510.03528 • Published Oct 3, 2025 • 20
Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance Paper • 2510.03528 • Published Oct 3, 2025 • 20
IntrEx: A Dataset for Modeling Engagement in Educational Conversations Paper • 2509.06652 • Published Sep 8, 2025 • 26
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? Paper • 2508.19827 • Published Aug 27, 2025 • 33