Verified75% confidenceFactExact time
CoT提示将PaLM 540B在GSM8K数学基准测试中的准确率从17.9%提升至58.1%
5
Sources
75%
Confidence
Medium-term (~90 days)
Relevance
7/2/2026
First Seen
Valid until: 9/30/2026
Sources
提示词工程完全指南:从零基础到实战应用
bilibili大模型教程6/17/2026
Related Claims
Unverified在 GSM8K 基准上,标准思维链准确率约为 56%-63%(GPT-3 规模),PAL 可提升至 72% 以上72% similarUnverifiedAS5048 数据手册中标注的实际精度(考虑积分非线性误差 INL)通常在 ±0.5 度左右65% similarUnverified开发者Ivan测试MTP-LX 0.3.5版本,在数学基准测试中5分30秒内取得了93.3%的正确率64% similarUnverified初代 DeepSeek Math 7B 在 MATH 竞赛基准测试中以70亿参数达到 51.7% 的正确率63% similarUnverifiedLFM2.5在模拟真实Agent场景的电信测试中提升了74个点,指令跟随IFEval达到91.8,数学推理Math500达到88.863% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/44810API
curl https://kongchang.com/api/v1/knowledge/claims/44810MCP
get_claim(id=44810)