[KongchangAI]
Unverified60% confidenceSolutionExact time

GQA(分组查询注意力)是介于多头注意力(MHA)和多查询注意力(MQA)之间的注意力机制,通过将多个 Query 头分组共享 Key-Value 投影来减少 KV Cache 显存占用

2
Sources
60%
Confidence
Long-term
Relevance
9/5/2026
First Seen

Sources

Related Entities

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/860211
API
curl https://kongchang.com/api/v1/knowledge/claims/860211
MCP
get_claim(id=860211)