Optimizing Transformer Attention: From Multi-Head to Grouped-Query and FlashAttention
LLM FundamentalsInference Optimization
Hi, this is Chengshuo. I'm documenting my learning notes in this blog. Feel free to email me if there is any mistakes in the notes. 😉