待验证50% 置信事实精确时间
Mistral和Mixtral等模型采用了交替使用Full Attention和SWA的架构,vLLM的Group分组机制支持这种模式
1
来源数
50%
置信度
长期有效
时效性
2026/9/10
首次发现
来源
涉及实体
相关事实
待验证GQA将注意力头分组共享KV,已被LLaMA 3、Mistral等主流模型采用78% 相似待验证GQA(Grouped-Query Attention)方案已被LLaMA 3、Mistral等主流模型采用以缓解HBM带宽瓶颈72% 相似待验证LlamaFactory supports fine-tuning of models including LLaMA, Qwen, ChatGLM, and Mistral through a unified interface70% 相似已验证Mistral 7B的关键技术创新包括滑动窗口注意力机制(SWA)和分组查询注意力(GQA)68% 相似待验证Mistral模型家族是SWA的典型代表,在每一层使用4096 token的窗口大小68% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/887215API
curl https://kongchang.com/api/v1/knowledge/claims/887215MCP
get_claim(id=887215)