Why Language Models Love the Em Dash: How Post-Training Recruits a Sparse Writing Circuit
August 12 2026
Mechanistic Interpretability; LLMs Post-Training; Model Behavior
Why do some language models overuse the em dash while others almost never do? We trace this writing habit to a sparse set of late-layer neurons, study how post-training amplifies an internal mechanism already present in the base model, and examine why explicit instructions sometimes fail to suppress it.