Unverified50% confidenceFactExact time
Mistral和Mixtral等模型采用了交替使用Full Attention和SWA的架构,vLLM的Group分组机制支持这种模式
1
Sources
50%
Confidence
Long-term
Relevance
9/10/2026
First Seen
Sources
Related Entities
Related Claims
UnverifiedGQA将注意力头分组共享KV,已被LLaMA 3、Mistral等主流模型采用78% similarUnverifiedGQA(Grouped-Query Attention)方案已被LLaMA 3、Mistral等主流模型采用以缓解HBM带宽瓶颈72% similarUnverifiedLlamaFactory supports fine-tuning of models including LLaMA, Qwen, ChatGLM, and Mistral through a unified interface70% similarVerifiedMistral 7B的关键技术创新包括滑动窗口注意力机制(SWA)和分组查询注意力(GQA)68% similarUnverifiedMistral模型家族是SWA的典型代表,在每一层使用4096 token的窗口大小68% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/887215API
curl https://kongchang.com/api/v1/knowledge/claims/887215MCP
get_claim(id=887215)