moonshot ai kimi k3 2 8t model
tritonsigmoid fast sigmoid attention
how much rank does lora need rank error bounds for transformer attention
semidirect fourier delta attention