[KongchangAI]
Unverified90% confidenceFactTime unknown

SWE-bench(Software Engineering Benchmark)是目前衡量智能体编码能力的核心基准,从真实的GitHub Issue出发要求模型自主定位代码库中的问题并提交修复补丁

1
Sources
90%
Confidence
Long-term
Relevance
6/1/2026
First Seen

Sources

Related Entities

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/24179
API
curl https://kongchang.com/api/v1/knowledge/claims/24179
MCP
get_claim(id=24179)