66 related articles
Tech FrontiersAnalysis of mythos-research, an open-source replication of Anthropic's Mythos Preview autonomous vulnerability discovery framework using Claude Opus 4.7.
Third-Party Cybersecurity Evaluations …
An in-depth analysis of third-party cybersecurity evaluation methodologies for OpenAI models, covering red teaming, vulnerability discovery assessment, risk classification, and impact on AI governance.

Deep analysis of Nightcrawler, an AI penetration testing agent running entirely on smartphones. Exploring how on-device AI empowers cybersecurity testing, its architecture, use cases, and risks.

VulX Watch is a security audit tool for AI-generated code that connects read-only to GitHub repos, independently reviews vulnerabilities, and provides line-level evidence for every finding.

Analysis of three real cyberattack incidents reveals AI's actual capability boundaries in offensive operations, exposing gaps between lab benchmarks and real-world threats for better security assessment.

Investigating three real cyberattack incidents to analyze AI's true role in offensive operations, examining the gap between lab assessments and real threats for better AI security evaluation.

Cynative is an open-source AI cloud security auditing tool that lets you query AWS, GCP, Azure, and Kubernetes infrastructure using natural language. Its read-only architecture ensures zero risk to production environments.

Cynative is an open-source AI cloud security tool that lets you audit AWS, GCP, Azure, and Kubernetes infrastructure using natural language. Its read-only architecture ensures production environments stay safe.

Google launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—three new models targeting quality-cost balance, extreme affordability, and cybersecurity specialization for AI Agent use cases.

Deep analysis of Hugging Face's frontier lab AI agent intrusion report, covering indirect prompt injection, lateral movement, data exfiltration, and defense-in-depth strategies for AI agent security.

July 24 AI news: Black Forest Labs launches Flux 3 multimodal model, Kimi K3 lags in US-UK gov tests, Alibaba Qwen tops TTS rankings, Etched raises $300M, AMD unveils MI430X.

Moonshot AI releases Kimi K3 open-weight model with 2.8T parameters and 1M token context. Our deep dive covers coding, 3D dev, agent capabilities, and safety concerns.

OpenAI launches the GPT-5.6 family (Sol/Terra/Luna), ChatGPT Work, a new desktop app, and Hosted Sites — marking AI's evolution from Q&A assistant to autonomous task executor.

OpenAI releases the GPT-5.6 model family, launching enterprise-focused ChatGPT Work, one-click ChatGPT Sites, and a major desktop client upgrade, with coding now ahead of rivals. Meta, Google, and Kimi follow intensively.

After the release of Claude Mythos Preview, critical security vulnerabilities surged, raising widespread concern. This article analyzes the tension between rapid iteration and security, explores LLM attack surface challenges, and offers practical defense strategies.

OpenAI releases GPT-5.6 (Sol/Terra/Luna), beating Anthropic on Terminal Bench at ~40% lower cost. But its cybersecurity capabilities hit danger thresholds, limiting access to trusted partners at government request.

OpenInspect's Multi-Repo Automations lets AI coding agents maintain up to 10 repositories on a schedule simultaneously — isolated sessions, independent PRs, and fault-tolerant execution for security sweeps, dependency upgrades, and framework migrations.

GPT-5.6 is officially released, merging ChatGPT and Codex into one app and launching the three-tier Sol, Terra, and Luna models. A detailed breakdown of 16 hands-on tests plus Worker mode and Codex dev upgrades.

AI coding tools are changing development, but Vibe Coding hides risks in code quality and maintenance. This article explores Engineered AI Programming, compares Codex and Claude Code, and reveals real enterprise development paths.

Visa's open-source Agentic security testbed Harness orchestrates threat modeling, vulnerability research, adversarial reproduction, and structured reporting into an auditable pipeline — not just a scan button.