[
KongchangAI
]
Feed
Entities
Podcasts
Deep Dives
Developers
About
Get CLI
中文
Product
Flash-Attention 2 / flash-attn
Flash-Attention
由Tri Dao等人提出的IO感知精确注意力算法,通过分块计算减少HBM读写,将注意力显存复杂度从O(N²)降至O(N),是Transformer训练与推理的高性能内核标配
Source Articles
Unsloth 发布 CUDA 13 预编译轮子:告别源码编译烦恼
Sep 27