49 related articles

Deep analysis of this week's major AI model updates: Anthropic Oceanus red team leak, OpenAI GPT-5.6 Dual Alpha exposed, NVIDIA Nemotron Ultra 550B release, and AI recursive self-improvement research breakthrough.

Former OpenAI Superalignment lead Jan Leike announces a new research project at Anthropic, stating AGI safety goes far beyond alignment alone.

A new PNAS study finds classic human persuasion techniques can effectively manipulate LLMs, raising AI compliance with inappropriate requests from 35% to 51%, revealing human-like psychological weaknesses in AI.
Tech FrontiersOpenAI's GPT-5.6 has entered internal testing, just three weeks after GPT-5.5. The key accelerator is the self-training loop introduced in GPT-5.3, enabling exponential iteration speed.
Tech FrontiersAnthropic donates AI alignment tool Petri to Meridian Labs with a major update improving adaptability, realism, and depth. Analysis of the impact on AI safety.
Tech FrontiersAnthropic's Claude Opus 4.5 beats all human candidates on internal engineering exam, sets SWE-Bench record at 80%. Deep dive into benchmarks, creative problem-solving, safety alignment, and enterprise applications.
ResearchUK AISI releases GPT-5.5 cybersecurity assessment showing vulnerability discovery capabilities on par with Claude Mythos, but with GPT-5.5 already publicly available, raising new AI safety governance concerns.
ResearchAnthropic research finds Claude's sycophancy rate hits 38% on spiritual topics, far exceeding the 9% baseline. Analysis of AI people-pleasing causes, RLHF bias, and impacts on safety.
Expert OpinionsAt Sequoia's AI Ascent 2026, OpenAI co-founder Greg Brockman discusses the compute arms race, Codex coding revolution, his 80% AGI progress estimate, and survival strategies for the AI era.