Unverified80% confidenceFactTime unknown
开发者Ivan测试MTP-LX 0.3.5版本,在数学基准测试中5分30秒内取得了93.3%的正确率
1
Sources
80%
Confidence
Medium-term (~90 days)
Relevance
5/31/2026
First Seen
Valid until: 8/29/2026
Sources
Mac本地跑Qwen3.6-27B:4种方案实测对比
bilibilikate人不错
Related Entities
Related Claims
Unverified微软的Phi-3仅3.8B参数但在多项基准测试中性能媲美早期GPT-4级别的模型66% similarUnverified在文档摘要基准测试上,Qwen2.5-32B-Instruct和Llama-3.1-70B的量化版本已能达到GPT-4早期版本85%-92%的表现水平66% similarUnverifiedBullet 在 SWE-bench Verified 基准测试中取得 95.8% 的成绩,位列前三,平均每个任务耗时 119 秒65% similarUnverifiedMTP-LX 4bit版本在理财规划推理任务上经GPT-5.5 Thinking评分为55分(满分100),而同题GPT-5.5 Pro得分82分65% similarUnverified盘古2.0 Flash在AIME类数学测试中得分93.3,在HMMT测试中得分91.5,数学能力处于第一梯队65% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/4335API
curl https://kongchang.com/api/v1/knowledge/claims/4335MCP
get_claim(id=4335)