language models can control their own attention
vlm reliability mechanistic study
softmax free attention gpt2 medium
parallax parameterized local linear attention