57 related articles

OpenAI's GPT-5.6 Soul, Terra & Luna are priced at one-third of Claude, leading Anthropic Fable on many benchmarks. We analyze its value, reasoning, and jailbreak risks.

OpenAI releases the GPT-5.6 model family, launching enterprise-focused ChatGPT Work, one-click ChatGPT Sites, and a major desktop client upgrade, with coding now ahead of rivals. Meta, Google, and Kimi follow intensively.

After the release of Claude Mythos Preview, critical security vulnerabilities surged, raising widespread concern. This article analyzes the tension between rapid iteration and security, explores LLM attack surface challenges, and offers practical defense strategies.
Boko Haram's Abuse of Frontier AI: How…
Boko Haram is systematically exploiting AI tools for propaganda automation, multilingual recruitment, and operational coordination. An in-depth analysis of generative AI abuse by terror groups, the open-source governance dilemma, and the AI safety arms race.

AI compliance is shifting from document storage to generating credible adversarial testing evidence. Learn how TRAIGA, NIST RMF, and ISO 42001 shape audit-grade red team testing requirements.
LLM Security Benchmarking: Current Sta…
Why is it so hard to establish unified LLM security benchmarks? This article analyzes core challenges in LLM security evaluation—covering jailbreaks, prompt injection, red teaming, and more—with practical strategies for developers.

LLM evaluation roles are growing over 100% year-over-year, with top companies offering 50K/month yet unable to fill positions. This article explores how testing pros can seize the window.

OpenAI released GPT-5.6 with three variants—Soul, Terra, Luna—and for the first time notified and submitted the model to U.S. government review before full release. A deep dive into the variants, Max/Ultra upgrades, and cybersecurity defenses.

Skill isn't just for programmers — it's crossing industry lines to become the operational foundation for entire organizations. This article unpacks the three cores of knowledge work, explores legal, research, and finance use cases, and examines how a Skill Library becomes a core enterprise AI asset.
Pliny's Jailbreak Experiments Reveal t…
Pliny the Liberator's satirical tweet exposes core issues in AI safety and open-source governance — from alignment failures to open-weight risks and AGI hype.
OpenAI Backs the Appia Foundation: Bre…
OpenAI backs the Appia Foundation to build shared standards for advanced AI, covering evaluation frameworks, safety practices, and global cooperation. A deep dive into why AI standardization matters.

Agent Studio unifies AI Agent role definition (Subagents) and Skills on one platform, enabling coordinated orchestration through a shared MCP endpoint, progressive disclosure, and moderated community publishing.

OpenAI board member Zico Kolter and Gray Swan CEO Matt Fredrikson explain why AI safety differs fundamentally from cybersecurity and how red-teaming must evolve into a systematic engineering discipline.

European security firm Paradigm Shift discloses an unpatchable hardware-level vulnerability in Apple chips affecting older iPhones, with major implications for jailbreaking and device security.

The U.S. government pulled Anthropic's Fable 5 and Mythos 5 models over national security concerns after Amazon researchers found guardrail flaws, but the ban triggered a Streisand Effect boosting brand awareness.

Claude Fable 5 banned globally just 3 days after launch. Deep analysis of the jailbreak controversy, AI supply chain fracture risk, Anthropic's fear marketing backfire, and local AI deployment strategies.

An overseas security blogger systematically tested DeepSeek's jailbreak resistance using direct requests, rephrased prompts, and varied strategies. Results show robust intent recognition, consistent blocking, and context-aware safety mechanisms.

Anthropic reverses its controversial policy of secretly throttling Claude Fable/Mythos responses to frontier LLM development requests after community backlash, raising critical questions about AI transparency.

AI agent auto-review is now default for all users. A classifier subagent achieves 97% accuracy with three-tier safety decisions. Deep dive into how it works and its impact on AI safety.

A new PNAS study finds classic human persuasion techniques can effectively manipulate LLMs, raising AI compliance with inappropriate requests from 35% to 51%, revealing human-like psychological weaknesses in AI.