Unverified50% confidenceFactExact time
FrontierCode 1.1的测试用例通常来自真实开源项目,要求模型具备跨文件上下文理解和长序列推理能力
1
Sources
50%
Confidence
Long-term
Relevance
9/16/2026
First Seen
Sources
Related Entities
Related Claims
UnverifiedFrontier Code的核心理念是评估代码是否能被项目维护者合并(PR可合并性),而非仅仅通过测试78% similarUnverifiedFrontier Code使用名为Mutagen的工具动态调整参考测试以适配智能体的不同实现方式75% similarVerifiedFrontierCode is an evaluation benchmark for AI models' programming capabilities focusing on real engineering tasks rather than simple algorithm problems73% similarUnverified基准测试往往是标准化、相互孤立的编程题目,而真实软件工程涉及庞大上下文、遗留代码、模糊需求和复杂依赖69% similarUnverified基于开源代码库的评测存在数据污染问题,因为这些代码库可能已被纳入大模型的训练数据68% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/927769API
curl https://kongchang.com/api/v1/knowledge/claims/927769MCP
get_claim(id=927769)