Data Science Wire

Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

arXiv cs.LG5d4 min read

arXiv:2607.21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that attention reveals which context the teacher actually uses. We test that assumption on retrieval tasks where the evidence for each answer is known exactly. By masking parts of the context and measuring whether the answer changes, we find that attention and causal dependence often disagree, and distilled selectors inherit the misma

Read the full story at arXiv cs.LG

More in Machine Learning