Claude Code Artifacts Now Open: AI Agent Race Shifts Into High Gear

Claude Code Artifacts goes live for Pro/Max users as the AI Agent race accelerates across products, benchmarks, and compute.
Anthropic has opened Claude Code Artifacts to Pro and Max users, enabling AI-generated interactive web pages to be published with a shareable URL. This week also saw Alipay's Agent 'Abao' launch public beta, ByteDance release the EdgeBench long-term learning benchmark, Microsoft form the model-agnostic Frontier Company unit, and reports of Anthropic pursuing custom AI chips with Samsung.
Claude Code Artifacts Now Open to Pro and Max Users
Anthropic recently announced that the Artifacts feature in Claude Code is now officially available to Pro and Max subscribers. This feature allows users to directly ask Claude to generate interactive web pages and publish them in real time to their personal Claude.ai space — closing the loop from generation all the way through to deployment.
Artifacts was a core interaction paradigm Anthropic introduced in mid-2024, initially allowing Claude to generate standalone code snippets, SVG graphics, and HTML pages in the conversation sidebar for preview. Code Artifacts is the evolved version, specifically optimized for runnable web application scenarios — essentially executing JavaScript/HTML/CSS code in a sandboxed environment and presenting the results as a live interactive interface. The newly unlocked "publish to Claude.ai space" capability takes this further by bridging the last mile between generation and hosting: users receive a publicly accessible URL without needing to configure their own server or CDN. This is conceptually similar to Vercel's "one-click deploy" philosophy, but with the key difference that the entire create-to-deploy pipeline is AI-driven and isn't limited to professional developers.
For rapid prototyping, demo showcasing, and lightweight app publishing, this represents a meaningful efficiency leap — developers can complete the entire workflow without repeatedly debugging in a local environment. Notably, Anthropic engineer Zarik responded to community questions by stating that the official goal is to restore these features as standard subscription offerings as soon as compute resources allow. This is a candid acknowledgment that, amid rapid AI service expansion, compute remains the core bottleneck limiting broad feature rollout.
Open Source and Localization: Portuguese LLM Amelia Launches
In the open-source model space, a coalition of Portuguese universities and research institutions jointly released Amelia — the first open-source large language model designed for the Portuguese language. The family includes the 9B-parameter language model Amelia 9B and the vision-language model Amelia VL.

The 9B parameter scale is a strategically significant sweet spot in the current open-source model ecosystem. Models in this range — such as Meta LLaMA 3 and Google Gemma 2 — can run local inference on consumer-grade GPUs (like a single RTX 4090) while retaining enough model capacity to handle complex language tasks. This makes them the ideal balance point between performance and deployability, particularly for resource-constrained academic institutions and SMEs looking to run private deployments or domain-specific fine-tuning. Amelia VL, as a vision-language model, follows a multimodal architecture that typically connects a visual encoder (such as CLIP) to a language model via a projection layer, enabling it to process mixed image-text inputs. For Portuguese-language use cases, this means applications like document OCR and image captioning can be handled natively — without relying on cross-lingual reasoning routed through English.
This release carries clear regional language sovereignty significance. Most mainstream LLMs are built primarily on English and Chinese corpora, leaving smaller languages with notable gaps in comprehension accuracy and cultural context. With Amelia, the Portuguese-speaking world — spanning Brazil, Portugal, and several African nations, totaling roughly 260 million people — now has a freely fine-tunable, privately deployable open-source foundation model. The value of such localized open-source projects lies not just in model performance, but in providing a self-sovereign technical foundation for specific language ecosystems.
AI Agent Race Accelerates: From Code Arena to Chinese Vendors
The AI Agent space is seeing a wave of dense product signals. Arena.ai has introduced FullStack functionality to its Code Arena platform, enabling developers to build complex applications that rely on databases or backend support. A FullStack leaderboard driven by community voting data is also planned — expanding the evaluation dimension for AI-assisted programming from isolated code snippets to full-stack application capability.
Domestic Chinese vendors are moving just as quickly. Kunlun Tech announced an upgrade to Tiangong 3.2 and launched SkyWork Tags, which supports connecting Agents to communication platforms like Slack and Feishu to participate in collaborative group chats. Technically, integrating AI Agents into instant messaging platforms relies on the open Bot APIs and Webhook mechanisms these platforms provide — Agents listen for @mentions or specific keywords to trigger, call a backend LLM to complete inference, and write results back into the group chat. Compared to standalone apps or web-based chat interfaces, the "Agent in the group chat" model's core advantage is eliminating context-switching costs: employees can access AI assistance without leaving their existing workflows, and conversation history is naturally preserved in the team's shared space for easy collaboration and review. This pattern is quickly becoming a new trend in enterprise collaboration — AI is no longer an isolated chat window, but a deeply embedded collaborative participant in existing workflows.
Meanwhile, Alipay's Agent product "Abao" has officially opened to public beta, allowing users to access its improved spoken language understanding and full-scenario task-handling capabilities on iOS or Android without an invite code.

Reports also indicate that Alibaba plans to consolidate three product lines — CodeWork, Wukong, and a third Agent product — into a unified AI product portal, with existing user benefits unaffected. This consolidation sends a clear signal: leading vendors are shifting from a "cast a wide net" approach to concentrating resources on building flagship Agent entry points.
New Evaluation Standards: ByteDance's EdgeBench Targets Long-Term Learning
ByteDance's C-team has released EdgeBench, a benchmark specifically designed to evaluate the long-term learning capabilities of autonomous AI Agents in real-world environments. The benchmark covers 134 tasks across 6 categories, with 51 tasks and a complete evaluation framework currently made public.
EdgeBench's core value lies in what it measures: an Agent's ability to learn continuously, rather than single-task completion rate. The current AI Agent evaluation landscape has developed several main tracks: software engineering task benchmarks like SWE-bench, web navigation benchmarks like WebArena, and general assistant capability benchmarks like GAIA. These benchmarks share a common limitation — task boundaries are well-defined, evaluation cycles are short, and what they're ultimately measuring is "single-task completion rate." EdgeBench differentiates itself by introducing a temporal dimension — Agents must accumulate experience and update their strategies across multiple rounds of interaction, similar to online learning paradigms in reinforcement learning. This more closely mirrors real enterprise deployment scenarios: a truly practical Agent needs to continuously calibrate its behavior from user feedback, rather than starting from scratch every time. While most current Agent evaluations focus on one-time task success rates, EdgeBench emphasizes adaptation and accumulation over long cycles in real environments — which more accurately reflects the core challenges Agents will face as they move toward practical deployment. The evolution of evaluation standards often foreshadows where the competitive focus in a technology space is shifting.
Big Moves: Microsoft, NVIDIA, and OpenAI's New Commercial Plays
On the commercial front, Microsoft announced the formation of a new business unit, Microsoft Frontier Company, backed by $2.5 billion in investment, which will dispatch 6,000 experts to enterprise clients to provide AI deployment and continuous improvement services that are "not locked to a single model."

"Not locked to a single model" is the key phrase in this strategy. Microsoft's deep partnership with OpenAI was once seen as the core moat of its AI strategy. However, as models like Anthropic's Claude, Google's Gemini, and Meta's LLaMA have demonstrated competitive performance on specific tasks, enterprise clients are increasingly inclined to select the best model for each use case. Microsoft Frontier Company enters with a model-agnostic service posture, effectively positioning Microsoft as a systems integrator for enterprise AI transformation rather than a pure model distribution channel — a model highly similar to the enterprise consulting approach of IBM, Accenture, and other traditional IT services giants, with the key difference being Microsoft's full-stack control spanning from cloud infrastructure (Azure) to developer tools (GitHub Copilot). Despite Microsoft's deep ties to OpenAI, its choice to act as a neutral broker in the enterprise services market — helping clients connect to multiple models — reflects the strong demand among enterprise customers for autonomy in model selection, and represents a pragmatic positioning move for the multi-model era.
NVIDIA has launched a revenue-sharing partnership model, teaming up with AI cloud partners like Sharon AI and Firmers to deploy large-scale AI factories and accelerate training and inference workloads for AI-native companies. By deeply binding ecosystem partners through revenue sharing, NVIDIA is transforming from a pure hardware supplier into a deep participant in AI infrastructure.
Most notable, however, is OpenAI. Reports indicate OpenAI has proposed offering the U.S. government a 5% equity stake — worth roughly $42.6 billion at current valuations — with Sam Altman framing it as a way of sharing the benefits of AI development with the public. Separately, Anthropic is reportedly pushing forward on developing its own AI chips and has entered negotiations with Samsung Electronics regarding potential manufacturing partnerships. The strategic motivations behind custom chip development are multi-layered: reducing dependence on NVIDIA's H100/H200 series to mitigate supply chain risks and premium pricing; achieving hardware-level optimization for their own model architectures to improve energy efficiency; and, as compute becomes a competitive moat, controlling cost structures and production capacity by owning the hardware. The choice to partner with Samsung rather than TSMC may reflect a combination of capacity allocation, geopolitical compliance, and cost negotiation considerations — Samsung's HBM (high-bandwidth memory) production capacity is also a critical component for large model inference. The intent of leading AI companies to control their compute destiny and reduce dependence on NVIDIA is becoming increasingly clear.
Summary
From the deployment of Claude Code Artifacts to the full public beta of Agent products, and on to the capital and compute buildouts by major players, the AI industry is showing three clear storylines in recent weeks: features are moving from generation to deployability, Agents are moving from demos to practical utility, and competition is extending from model capability into compute and business models. Agent practicalization and compute autonomy will be the core battlegrounds in the competitive landscape ahead.
Related articles

Oxide Computer Raises $445 Million to Rebuild Server Architecture from the Ground Up
Cloud hardware startup Oxide Computer raises $445M to redefine server architecture with open-source firmware and integrated rack-scale design for on-premises cloud experiences.

Media File Organizer: A Free, Open-Source Tool for Automatically Organizing Your Plex Media Library
Media File Organizer is a free, open-source desktop tool that auto-matches TMDB metadata to batch rename and organize movie and TV files into Plex-compatible formats with preview before changes.

Double Descent Explained: Why Massively Overparameterized Models Don't Overfit
A deep dive into the Double Descent phenomenon in machine learning, explaining why overparameterized models defy the classic bias-variance tradeoff to achieve stronger generalization.