Gemini 4 Pro Spotted in Arena Ghost Testing: Major Upgrades in 3D and Physics Capabilities

Gemini 4 Pro appears to be ghost-testing in Arena, with community samples showing major leaps in 3D and interactive capabilities.
According to community leaks, Google may be hiding an early Gemini 4 Pro checkpoint (codenamed Argon) inside Arena's anonymous "Gemini 3.8 Flash" variant for ghost testing. Abnormal signals — including dramatically improved 3D output, a changed checkpoint ID, longer generation times, and a 256K token output limit — suggest a far more powerful model is running under the hood. Community-collected samples showcase interactive 3D modeling, a playable racing game, and scientific visualizations, hinting at major advances in spatial understanding and physics simulation. Rumors also point to 2M token context support and positioning against Astra and Fable 5.1. All information remains unconfirmed community speculation.
Google's next-generation flagship model is quietly undergoing real-world evaluation. According to a Bilibili AI news blogger, Gemini 4.0 Pro has entered "ghost testing" (Ghost Test) on the Chatbot Arena platform, with community tests revealing performance far beyond everyday norms in areas like rendering quality, 3D spatial understanding, and physics simulation. Here's a breakdown of the key signals this gray-scale test has revealed, based on publicly available testing clues.
Ghost Testing and the Internal Codename Argon
Ghost testing is an anonymous gray-scale validation approach used before a model's official release — the model appears in Arena as an anonymous variant for blind testing, eliminating users' brand bias and yielding more objective capability assessments.
The community noticed that a new variant called Gemini 3.8 Flash suddenly appeared in Arena, but its actual performance didn't match its name. Multiple testers speculated that this Flash variant was actually running an early checkpoint of Argon — the internal codename for Gemini 4 Pro — meaning that when users interacted with the Flash variant, they may have actually been calling an unreleased early checkpoint of Gemini 4 Pro.

Several pieces of evidence support this conclusion: output quality and 3D generation effects suddenly far exceeded typical Flash-level performance; the Flash checkpoint ID changed; the model began demonstrating understanding of three-dimensional space and physical laws; single generation time increased noticeably; and the output limit was raised to 256K tokens. This collective leap across all these metrics is hard to explain as a routine optimization of a same-generation smaller model.
Chatbot Arena (created by UC Berkeley's LMSYS team) is currently the most influential crowdsourced model evaluation platform. Users interact with two anonymous models without knowing their identities, then vote for the better response, which the platform uses to generate Elo rankings. Because testers don't know which model they're using, brand halo effects are significantly reduced, making results considered closer to real user preferences than pure benchmark tests. Major labs therefore frequently submit new models as anonymous variants to Arena for "warm-up" before official release — collecting real user feedback while quietly accumulating Elo scores. This is exactly why ghost testing appears so frequently in Arena.
Targeting Astra and Fable 5.1
According to the leaks, Gemini 4.0 Pro is targeting Astra and Fable 5.1 as its competitors (the "Extra" and "Faber" mentioned in the original post are likely voice transcription errors). This positioning signals that Google wants to compete head-on with the industry's most cutting-edge solutions in multimodal interaction and content generation capabilities.
Even more noteworthy is the context window. Rumors suggest the new model may support up to 2M tokens of context length. If true, this would dramatically expand the model's ability to handle long documents, large codebases, and complex multi-turn tasks, further cementing Google's traditional advantage in the ultra-long context niche.

It's worth noting that all of the above parameters currently come from community speculation and third-party observations — Google has not officially confirmed the existence of Gemini 4 Pro or any technical specifications. Before an official release, these numbers should be treated as clues rather than conclusions.
Ultra-long context windows have been a signature technical direction for Google's recent Gemini models. Gemini 1.5 Pro was the first to support 1 million token context, with subsequent versions expanding to 2 million tokens, and the rumored 2M token figure continues this trajectory. The practical significance of context length goes far beyond the number itself: 1 million tokens is roughly equivalent to 750,000 English words, enough to hold an entire mid-sized codebase or dozens of lengthy research papers; 2M tokens means the model can "read" an encyclopedia-sized knowledge base in a single conversation without segmentation. This capability has substantial value for tasks requiring reasoning across large amounts of context (such as full-codebase refactoring or long-running multi-turn research conversations), and is one of Google's core differentiators from competitors like OpenAI.
Suspected Output Samples: 3D and Interactive Capabilities Stand Out
The community has collected numerous output samples suspected to be from Gemini 4 Pro, with the most striking improvements concentrated in 3D modeling, visualization, and interactive content generation.
From the publicly available test lists, the scenarios covered are quite diverse: a PS5 console SVG example shared by Lumina, a road bike visualization interface by UVL, customizable and explorable interactive 3D components, a 3D model of the Airbus H145 helicopter, a fully playable racing game, a small game controlling aircraft flight, and Wolke's pixel-art pagoda modeling, among others. What these cases share in common is that they're no longer static images or plain text — they are interactive, spatially structured, and physically coherent composite outputs.

Several "person riding a bicycle" tests (from 7-Eleven 1121 and Harshice respectively) were repeatedly cited. These scenarios require the model to simultaneously understand human posture, bicycle structure, and motion physics — a classic test of spatial and physical understanding.

Additionally, a 3D cell structure comparison case provided by Hicam demonstrates the model's potential in scientific visualization — converting abstract biological structures into accurate, comparable 3D representations, which is a dual challenge for both the model's knowledge accuracy and spatial modeling capabilities.
Several cases mentioned involve SVG and interactive 3D content generation — tasks that represent a high-difficulty "code-as-visuals" challenge for large models. The model needs to directly translate a user's natural language description into structured vector graphics code (SVG) or 3D rendering code like WebGL/Three.js, which is then executed and rendered by the browser in real time. This process requires the model to simultaneously possess spatial geometry reasoning, object structure understanding, and precise code generation — any failure in one link causes the final output to break down. Past mainstream models commonly suffered from spatial relationship errors, distorted physical proportions, or code logic flaws on such tasks. The appearance of a playable racing game and explorable interactive 3D components in this community showcase suggests the model has made significant progress in the pipeline of translating "spatial concepts" into "runnable code."
How to Interpret This Wave of "Explosive Performance"
Looking at these clues holistically, the core breakthrough of Gemini 4 Pro (if it truly exists) likely lies not in traditional text conversation capabilities, but in an overall leap forward along the path of spatial understanding, physics simulation, and interactive content generation. The 256K output limit and the rumored 2M context window also signal that the model is evolving toward handling more complex and complete tasks.
That said, readers should remain cautious. Performance during ghost testing is subject to sampling bias — the community's curated "explosive" examples tend to be the best of the best and may not represent the model's average performance level. A genuine capability assessment will require systematic benchmark testing after the official release.
Regardless, the direction revealed by this gray-scale test is worth watching: the competitive focus of large models is shifting from "being able to talk and write" toward "understanding the world and building interactive content." Google's positioning along this path may become an important variable in the next round of the model race.
Related articles

What Are APIs and API Keys? A 4-Minute Explainer Using DeepSeek's Docs
What are APIs and API Keys? Using DeepSeek's official docs and a live terminal demo, this guide explains the core concepts in 4 minutes with a simple door-and-key analogy.

AI Agent Beginner's Guide: How Agents Work, What They're Made Of, and Where They're Used
A beginner's guide to AI Agents: learn how agents differ from chatbots, the three core components (brain, memory, tools), and the three types of agents available today.

Beware of 'One-Click GPT-6 Deployment' Scams: Third-Party API Wrapper Traps Explained
Fact-checking viral 'free GPT-6 one-click deployment' tutorials: the GPT-6 X-TRAW model doesn't exist, and third-party token relays pose serious data and security risks.