待验证50% 置信事实精确时间
LLaMA-2 70B采用80层、64个注意力头的架构,并使用GQA(8个KV头共享64个Q头)
1
来源数
50%
置信度
长期有效
时效性
2026/7/19
首次发现
来源
相关事实
待验证Llama-3 70B采用80层Transformer、每层64个注意力头、每头128维配置,并采用GQA设计77% 相似待验证LLaMA-2 70B 模型共有 80 个 Transformer 层,每层包含自注意力机制和前馈神经网络两大模块69% 相似待验证LFM2.5 has 24 layers total, with 18 using Liquid AI's proprietary LIV (Liquid) convolution and only 6 retaining traditional attention mechanisms65% 相似待验证GQA被Llama 2 70B和Llama 3系列采用,MQA被Falcon系列采用,通过多个查询头共享同一组KV头减少缓存的KV向量数量63% 相似待验证P106-100基于Pascal架构(GP106芯片),与GTX 1060共享相同的计算核心61% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/561354API
curl https://kongchang.com/api/v1/knowledge/claims/561354MCP
get_claim(id=561354)