Minimax H3 Circle-Drawing Positioning Trick: Precisely Control AI Video Scene Placement

Minimax H3 users discover that drawing circles on reference images can precisely anchor AI video scene placement.
Reddit users discovered that drawing freehand circles on reference images in Minimax H3 can specify exactly where AI-generated content appears in the video — a "circle positioning" trick that bypasses text prompts' inherent spatial ambiguity by converting visual annotations into spatial anchors. Circle color can be adjusted freely, but mode selection (match vs. max) affects detail fidelity. While still a community-discovered technique rather than an official feature, it clearly signals AI creative tools' evolution from text-driven toward visual, multimodal interaction.
A Practical Breakthrough in AI Video Scene Control
A Reddit community user recently shared a useful discovery about the Minimax H3 video generation model: drawing a circle on a reference image lets you specify exactly where a scene appears in the AI-generated video. This deceptively simple technique reveals an important advancement in the spatial control capabilities of AI video generation tools.
In traditional text-to-video or image-to-video workflows, creators can only describe visual content through text prompts — but precisely controlling the spatial position of elements or scenes remains frustratingly difficult. This "circle-to-position" method gives creators a more intuitive and controllable way to interact with the model.

How Circle Positioning Works
The core operation is straightforward: use a brush to circle the region in a reference image where you want the scene to appear, and the AI will generate content in that corresponding location.
Key operational details include:
- Accurate spatial correspondence: Draw a circle over a water surface in the image, and the generated scene will appear in the water; circling a background building works the same way. This spatial mapping is genuinely effective
- Flexible color choices: The circle doesn't have to be red — just adjust the prompt accordingly. The visual marker functions more like a "spatial anchor" and needs to work in conjunction with the text prompt
- Effects are still being refined: Close inspection reveals some missing details, which may be related to using "match" mode instead of "max" mode for image reference
The Difference Between Match and Max Modes
"Match" and "max" are two intensity settings within Minimax's image reference feature. These parameters control how closely the model adheres to the reference image — "match" tends to flexibly match the overall style and composition, while "max" more strictly preserves the reference image's details. Users seeking higher fidelity may get better results by switching to "max" mode.
The Technical Significance of This Technique
From a technical perspective, this discovery reflects a broader trend in AI video generation: the shift from "pure text control" to "multimodal hybrid control."
Traditional prompt engineering has an inherent problem: language is naturally ambiguous when describing spatial relationships. When you say "make the explosion happen on the left," it's difficult for AI to accurately interpret the specific location and extent of "the left." By drawing directly on the image, users are providing visual spatial instructions, bypassing the ambiguity of language descriptions.
This "draw-to-guide" interaction paradigm builds on techniques already established in image editing — regional control, inpainting, and similar tools — now being brought into video generation, signaling that video AI tools are becoming increasingly "directable."
Practical Value for Creative Workflows
For content creators, this type of feature delivers an immediate benefit: greater controllability and lower trial-and-error costs. Previously, getting a scene to appear in the right position might require repeatedly rewriting prompts, generating multiple times, and sifting through numerous outputs. A single intuitive circle-drawing operation can now dramatically shorten that iterative process.
This direction of development is closely related to the field of "Spatially Conditioned Generation" in computer vision research. Earlier technologies like ControlNet demonstrated that feeding additional inputs — such as skeleton maps, depth maps, or edge maps — into diffusion models as conditioning signals significantly improves spatial controllability over generated content. Circle positioning is essentially converting a user's freehand annotation into a lightweight spatial condition: during inference, the model jointly interprets this region mask alongside the text prompt, "anchoring" semantic content to the specified coordinate range. Compared to full ControlNet pipelines, this approach has a much lower technical barrier for users — but the trade-off is relatively coarse-grained control, as the model still needs to infer the specific compositional details within the circled area.
Current Limitations and Room to Grow
This technique is still in the "community exploration" stage and is not an officially promoted feature. The original poster maintained an objective tone — excited about the discovery while honestly noting its current limitations: missing details, unclear mechanisms around color influence, and mode selection affecting results.
This also reminds us that many powerful capabilities of AI video tools are often "uncovered" by active user communities in practice, rather than through official documentation alone. This community-driven exploration is a key expression of the vitality of today's AIGC ecosystem.
It's worth noting that the phenomenon of "community-discovered" features is extremely common in the diffusion model tool ecosystem. Because large generative models absorb vast amounts of data with various annotation formats during training, they often possess implicit capabilities that haven't been explicitly documented officially. Users accidentally activate these latent capabilities through unconventional inputs — such as hand-drawn markings on images — and then rapidly share their findings on platforms like Reddit, X (formerly Twitter), and Discord via screenshots and videos. This "discover and spread" community mechanism allows the boundaries of model capabilities to be quickly mapped, objectively accelerating the real-world adoption of AIGC tools — though it also introduces issues of inconsistent results and questionable reproducibility, so maintaining realistic expectations is essential.
Practical Recommendations
Minimax H3's circle positioning trick demonstrates, through a simple operation, the potential of AI video generation for spatial control. While the results aren't perfect yet, it points in a clear direction: future AI creative tools will increasingly rely on intuitive, visual interaction methods — letting creators truly "direct" AI rather than simply "hoping" it understands their words.
Users of Minimax H3 are encouraged to try this technique — remember to adjust your prompts based on the circle color you use, and experiment with "max" mode for more refined results.
Related articles

Vercel AI SDK Releases Vue 3.0.282 Patch Update
Vercel AI SDK releases @ai-sdk/vue@3.0.282 patch update, syncing with core package ai@6.0.282. Learn about the changes, release cadence, and upgrade recommendations.

Vercel AI SDK Sandbox Component Receives Patch Update
Vercel AI SDK releases sandbox-vercel@1.0.109 patch update, syncing the harness dependency to the same version. A look at this maintenance release and what it means for AI app developers.

Vercel AI SDK Vue 4.0.99 Released: Dependency Update Overview
The @ai-sdk/vue 4.0.99 patch release syncs the underlying ai@7.0.99 dependency. Learn what this means for Vue developers building AI apps with Vercel AI SDK.