Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
NeurIPS (under review) 2026
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
NeurIPS (under review) 2026
CoSe-Co: Text Conditioned Generative CommonSense Contextualizer for Language Models
CSKB AKBC 2021; NAACL 2022
No Need to Know Everything! Efficiently Augmenting Language Models With External Knowledge
CSKB AKBC 2021; NAACL (Findings) 2022