Gemini 3.8 Live Launches; Google Adopts Claude Internally for Coding

Gemini 3.8 Live debuts, OpenAI human review controversy emerges, and Google quietly adopts Claude internally.
Google launched Gemini 3.8 Live with $0.84/hour audio input costs and top Arkham Banking benchmark scores, featuring an Extended Thinking mode that boosts Speech-to-Speech Index from 76% to 82.6%. OpenAI's Project Lily revealed hundreds of contractors reading real ChatGPT conversations, reigniting privacy and transparency debates around RLHF training. Google also made Anthropic's Claude Opus 5 available internally via its AntiGravity platform, signaling engineers' preference for best-in-class tools. In China, ByteDance, Volcano Engine, and Baidu rolled out new AI service tiers, while MediaTek's Dimensity 9600 Pro advanced on-device AI inference with a 2nm dual-NPU design.
Gemini 3.8 Live: Merging Real-Time Conversation and Visual Presentation
Google's Gemini 3.8 Live integrates real-time conversational intelligence, fluid interaction, and visual rendering into a single Live experience — a bid to stake out an early lead in multimodal real-time interaction. Based on publicly available benchmark data, the system's competitive edge lies primarily in two areas: cost efficiency and reasoning capability.
Benchmark charts show Gemini 3.8 Live's input audio cost at $0.84 per hour, placing it among the lower-priced options in its category. For applications requiring extended voice interaction, this cost advantage compounds with scale and directly affects the commercial viability of real-world deployments.
On the Arkham Banking benchmark, Gemini 3.8 Live with Extended Thinking enabled scored 35.1%, ranking first among all models evaluated. On the Speech-to-Speech Index, the standard version achieved 76%, with the Extended Thinking tier pushing that further to 82.6%. These numbers reflect a coherent design philosophy: keep costs low while offering an "Extended Thinking" mode that trades response speed for higher task accuracy, letting users dial between latency and reasoning depth as needed.
Extended Thinking is a reasoning mechanism appearing in recent multimodal large models. Unlike standard inference, this mode allows the model to execute longer internal chain-of-thought steps before producing a final answer — similar to how a person might draft their thinking before responding to a complex question. The tradeoff is increased response latency; the benefit is higher accuracy in scenarios requiring multi-step logical reasoning, mathematical computation, or complex instruction-following. The Arkham Banking benchmark is an evaluation suite designed specifically for financial domain dialogue tasks, testing model accuracy across banking, lending, and compliance scenarios — widely regarded as a key reference for measuring the real-world usability of live voice models in vertical industries. The Speech-to-Speech Index provides a holistic assessment of full-pipeline quality from voice input to voice output, encompassing semantic accuracy, conversational coherence, and related dimensions.
OpenAI's Human Review: Privacy Concerns Behind Model Improvement
According to an investigation by 404 Media, as reported by India Today, OpenAI has hired hundreds of contractors to read and evaluate real ChatGPT conversations in order to improve its models. Under a program called Project Lily, reviewers sometimes have access to complete human-AI exchanges — which may contain sensitive personal information.

Reviewers do more than just read: they also summarize user intent, check whether model responses comply with training guidelines, and score multiple model outputs against each other. This "human-in-the-loop" approach genuinely helps improve model alignment quality, but it resurfaces a persistent question — the private information users share in AI conversations may ultimately be read, word for word, by human beings. Transparency and data boundaries will continue to be issues that large model companies must address head-on.
In contrast to the review controversy, OpenAI has also been active in the open-source ecosystem. Its Codex for Open Source program has introduced a $100 Pro tier and doubled the number of funded slots. Aimed at open-source software maintainers, the program encourages previously funded maintainers to reapply — an effort to channel AI coding tool capabilities back to long-term contributors in the open-source community.
RLHF (Reinforcement Learning from Human Feedback) is one of the core training components in today's mainstream large language models, and human-in-the-loop data annotation is an indispensable part of that process. Annotators compare and score multiple model outputs, helping train a reward model that then guides the language model to optimize toward more human-preferred responses. The workflow described in Project Lily — summarizing user intent, evaluating response compliance, scoring multiple output versions — closely mirrors the standard RLHF annotation pipeline. The fundamental tension here is structural: model training quality correlates positively with annotation scale, and improving annotation quality often requires annotators to engage with real user conversations — which sits in inherent conflict with users' reasonable expectations of data privacy. Major vendors generally include clauses in their privacy policies permitting conversations to be used for model improvement, but the actual transparency of review pipelines varies considerably.
Google Brings Claude In-House: A Nuanced Collaboration Between Giants
Business Insider reports that Google has made Anthropic's coding model available to engineers across the company for internal development. The specific model in use is Opus 5, accessible exclusively through Google's internal development platform, AntiGravity.

This is a telling signal. Google has its own Gemini model family, yet for internal coding workflows it has brought in a competitor's model — Anthropic's Opus 5 — suggesting that when it comes to specific programming tasks, engineers gravitate toward whatever tool delivers the best experience, regardless of corporate allegiance. This pragmatism also reflects a broader reality: differences in model capability are felt directly in day-to-day development workflows, and vendors can't rely on "use our own" mandates to sustain internal adoption.
A Wave of Updates in China's AI Services and Hardware
Domestic players continue to double down on subscription billing and enterprise services. 36Kr reports that ByteDance is set to launch A-Drive, a standalone enterprise intelligent cloud storage product, with applications currently open. Built by Volcano Engine, A-Drive is designed to connect AI applications and keep files and outputs continuously accessible, while unifying file storage and team collaboration.

Volcano Ark's Agent Plan has listed DeepSeek 4.1 Flash, offering credits at half price during a promotion running September 15–28. This personal subscription plan provides access to full-modality models across four tiers, Agent usage tracking, and compatibility with a range of coding agent tools. Baidu Qianfan's personal edition showcases four credit-based subscription tiers — Mini, Lite, Pro, and Max — priced from ¥4.9 to ¥299.9 per month, translating to monthly token quotas of 10 million, 42 million, 230 million, and 700 million tokens respectively. This "tiered subscription + token quota" pricing model is fast becoming the standard approach for AI services in China.
On the hardware front, MediaTek has announced the Dimensity 9600M and 9600 Pro, with the Pro featuring a 2nm process and a 2+3+3 all-large-core architecture, dual NPUs, and emphasized support for running AI models directly on-device. The continued improvement in on-device AI inference capability means more AI tasks could move off the cloud and run locally on smartphones.
On-device AI Inference refers to running AI models directly on the local chip of a terminal device — such as a smartphone or tablet — rather than sending requests to a remote cloud server. Key advantages include reduced network latency, enhanced user data privacy (data never leaves the device), functionality in low-connectivity or offline environments, and lower cloud compute costs. The primary constraints are chip performance and power efficiency — the NPU's integrated specifications directly determine the maximum model parameter count that can run locally. MediaTek's Dimensity 9600 Pro, with its 2nm process and dual NPUs, represents a continued push in this direction. As quantization compression techniques have matured — reducing model precision from FP16 to INT4 and similar formats — language models in the 1B–7B parameter range can now run smoothly on flagship smartphones, opening up new possibilities for offline AI assistants, real-time translation, local code completion, and more.
Industry Observations: Narrative Atmosphere and Workplace Shifts
According to a report cited by MeiJin.com, NVIDIA CEO Jensen Huang said on September 14 that the AI narrative in China is more pragmatic, with no one in those discussions talking about "AI doomsday." He used this observation to capture two starkly different cultural atmospheres — one focused on practical applications, the other fixated on risk and existential fear.
A LinkedIn study, meanwhile, reveals a fracture in workplace relationships in the age of AI: 72% of Gen Z professionals feel disconnected from people who could help advance their careers, with professional influence and social confidence falling out of alignment.

The same study found that 27% of Gen Z find their next job through online professional communities, and candidates with internal connections are hired at nearly seven times the rate of general applicants. Even as AI reshapes how work gets done, the weight of professional networks in career development hasn't diminished — if anything, it's persisting in new forms through online communities.
Related articles

Why Do All AI-Generated Projects Look the Same? The Aesthetic Homogenization Problem in Vibe Coding
Why do vibe coding projects all use purple gradients and dark glassmorphism? We break down the technical roots of AI aesthetic homogenization and how to escape it.

Weave Router 2.0: A Subscription-Aware AI Coding Agent Router with Cross-Service Intelligent Dispatch
Weave Router 2.0 is a subscription-aware AI coding agent router that auto-dispatches requests across Claude, Codex, and GPT subscriptions — claiming half the cost and double the speed of GPT-6 Astra.

LARA: A Lightweight Adaptation Framework for Injecting Composable Behaviors into Frozen LLMs
LARA (Lightweight Additive Residual Adaptation) is an open-source framework that injects pluggable, composable behaviors into frozen LLMs via low-rank residual adapters with soft routing and Mixture of Behaviors support.