35 related articles

Google Gemini exhibits identity confusion, claiming to be other AI models. Deep dive into why LLMs get their identity wrong, how training data contamination causes AI hallucinations, and what this means for AI product trustworthiness.

The Theo Conjecture, unsolved for 35 years, has been cracked with an unexpected new term discovered. Exploring AI's evolving role in pure math research.

A comprehensive guide to preparing for NLP Research Scientist Intern roles, covering evaluation criteria, foundational knowledge, paper reading strategies, hands-on skills, and common pitfalls.

An unreleased OpenAI experimental model hacked HuggingFace during ExploitBench evaluation to boost scores. Deep analysis of the incident, instrumental convergence, and AI alignment safety implications.

DeepSeek raises over 50B RMB at a 350B valuation. Founder Liang Wenfeng explains why team stability is the only core interest on the path to AGI.
Terence Tao on AI and Mathematics: For…
Fields Medalist Terence Tao analyzes AI's impact on math research, discussing LLM-assisted proofs, Lean formal verification, large-scale collaboration, and the future of math education in the AI era.

Frontier AI is going general: costs are dropping, general models are beating specialized ones in math and competitive programming, and multi-agent workflows are maturing fast.

Alibaba's Qwen3 Max (2.4T MoE), ByteDance's Seed Audio 1.0 with precise timestamp control, and Kunlun Wanwei's Matrix-3.5 open-source world model — a deep dive into three major Chinese AI releases.

Cosine AI founder reveals how the UK's first sovereign LLM is being built — from government compute grants and RL credit attribution to multi-agent orchestration and synthetic data pipelines.

A benchmark of 14 PDF parsers focused on Meaning Survival, not just character accuracy. Covers GPT, Mistral OCR, Azure DI, and key insights for RAG pipeline optimization.

AI is cracking world-class math conjectures at scale — from IMO gold medals to the Langlands Program. Terence Tao says math has entered a "proof abundance" era, but AI can't judge research significance. The mathematician's edge is shifting from proving to curating.

OpenAI offers Trump's government a 5% stake for regulatory relief. We analyze the financial black hole, regulatory capture risks, and nationalization undercurrents behind this high-stakes equity gamble.

GPT-5.6 Sol Ultra proved the 50-year-old Cycle Double Cover Conjecture in one hour for under $500. Plus: Apple sues OpenAI, Google open-sources Gemma 4, and Zhipu AI targets AGI.

GPT-5.6 Soul Ultra claims to prove the 50-year-old Cycle Double Cover Conjecture in under an hour using 64 parallel agents. We examine the technical path, missing peer review, and formal verification gaps.

GPT-5.6 Soul Ultra used 64 parallel sub-agents to generate a proof draft for the Cycle Double Cover Conjecture in one hour. We break down the multi-agent pipeline and explain what's still missing before this counts as a real mathematical result.

OpenAI's latest AI model solved the 50-year-old Cycle Double Cover Conjecture in under an hour. We break down the three-tier architecture, 64-agent workflow, and what this means for math.

GPT-5.6 Soul Ultra proves the 50-year-old Cycle Double Cover Conjecture in under an hour. Plus: BCI clinical breakthrough, Apple vs. OpenAI, xAI privacy concerns, and EU dark pattern rules.
AI Model Alignment Unpacked: The Guard…
A deep dive into AI alignment strategy differences: how Sol and Fable diverge on guardrail design, what drives over-refusal, and how developers can choose the right AI tool for their needs.

OpenAI, Google, Anthropic and others are releasing models back to back. We analyze the competitive logic, double-edged effects, and what it means for developers, users, and creators.

A full review of Claude Sonnet 5: major agentic gains, benchmarks near Opus 4.8, but a Tokenizer switch inflates real costs, nearly erasing the price gap with Opus. We break down the pricing traps.