751 related articles

Just $500 in RL fine-tuning enables a 9B open-source model to outperform frontier LLMs on catalog review tasks. Analysis of when small-model RL works and its enterprise implications.

Qwen launches a massive open-weight model, DeepSeek V4 priced at $0.0028/M tokens, Kimi K3 hits compute limits — Chinese AI labs are aggressively disrupting OpenAI and Anthropic's frontier model dominance.
Frontier AI Models Keep Making Element…
Why do frontier AI models like GPT-5 still make basic errors? This deep dive explores the reliability crisis in advanced LLMs, benchmark gaps, and what developers should do.

How do governments evaluate frontier AI model safety? This deep dive examines opacity in AI safety governance, missing standards, regulatory capacity gaps, and paths toward transparent oversight.
The Diffusion Model Revolution: The Ne…
Former Meta Llama lead Sergey Edunov joins Genesis Molecular AI, betting on diffusion models for drug discovery. PEARL achieves zero-shot top results on OpenBind.
Should Frontier AI Models Like GPT-5.6…
Should frontier AI models be open-sourced? This deep dive explores the key debates around democratization, misuse risks, commercial sustainability, and governance — and the middle paths between open and closed.

Claude Sonnet 5 may launch this week with up to 2M token context; GPT-4.6 Pro arrives with stunning code generation; mysterious Opus 6 exists internally. Full breakdown of this week's frontier AI model updates.

Cursor unveils three major updates: Cursor Mobile, Origin platform challenging GitHub, and a frontier in-house LLM trained from scratch. A deep dive into Cursor's strategy.

Exploring why frontier AI models need mandatory third-party safety testing across cybersecurity, biosecurity, and autonomy risks, and the paradigm shift from voluntary commitments to mandatory oversight.

Sakana AI releases Fugu Ultra, achieving frontier AI performance through autonomous model orchestration. Deep dive into its technology, strategic implications, and impact on global AI competition.
GPT-5.5 Instant's Medical Q&A Capabili…
OpenAI announces GPT-5.5 Instant matches frontier Thinking models in health Q&A, with major improvements in emergency recognition, follow-up questioning, uncertainty expression, and plain-language explanations — free for all users.

Zhipu AI's GLM-5.2 passes the community vibe check, showing capabilities rivaling top closed-source models. Analysis of what this means for open-source AI.

OpenAI launches GPT-Rosalind, its first frontier AI model for biology, drug discovery, and translational medicine. A deep dive into its capabilities, trusted access deployment, and industry impact.

A Google AI PRO subscriber reports Antigravity coding tool subscription issues, hitting quota limits despite paying. We analyze possible causes and offer fixes.

The faithfulness of the Burau representation of braid groups at n=4 has been unresolved for nearly 90 years. A recent proof finally settles this critical case, with implications for knot theory, topological quantum computing, and cryptography.

Google DeepMind announces Gemini 4 pre-training has begun, calling it their most ambitious training yet. A deep dive into its technical direction, compute scale, multimodal breakthroughs, and competitive impact.

An in-depth analysis of the open-weights model debate: public release brings transparency and innovation, but raises safety and misuse risks. Exploring tiered release, red-teaming, and governance challenges.

A viral Reddit post asks: will AI end human history? This article analyzes the blind spots of tech accelerationism, the governance mismatch, and how to rationally navigate AI transformation.

An in-depth analysis of the open-weights model debate: publicly releasing model weights enables transparency and innovation but raises safety risks. Explores tiered release, red-teaming, and the industry dynamics behind open AI governance.

Large models aren't search engines — they're more like super compressors. This article explains how LLMs compress data to learn semantic patterns, and explores the phenomenon of intelligent emergence.