Publication
Spectral Thompson sampling
Video
Matching features, not tokens: Energy-based fine-tuning of language models
Cross-entropy (CE) training provides dense and scalable supervision for language models, but it optimizes next-token prediction under teacher forcing rather than sequence-level behavior under model rollouts. We introduce a feature-matching objective for language-model fine-tuning that…