GPT-6 Astra Auto-Edits 75 Minutes of Footage via MCP: Real-World Results and True Time Costs

GPT-6 Astra natively controls DaVinci via MCP, auto-editing 75 minutes of footage into a publish-ready cut in about an hour.
Creator Mike tested GPT-6 Astra using Blackmagic's native MCP interface to fully automate DaVinci Resolve 21.1: connecting the server, transcribing 75 minutes of footage, rough-cutting to 22 minutes, adding graphics, grading color, hitting -14.69 LUFS, and rendering output. The rough cut length matched the manual version almost exactly. However, the true time cost approached one hour, with slow transcription, heavy token usage, and context compression being unavoidable; micro-edits (two minutes to remove a two-second silence) showed AI still lags humans on instantaneous fine adjustments. Mike's key takeaway: the moat isn't which software you use — it's how you build and orchestrate AI agent workflows.
An Experiment That Could Upend Video Editing
Video editing has long been one of the most time-consuming parts of content creation. From transcription and rough cuts to color grading and loudness normalization, a skilled editor can easily spend hours on a single project. Veteran audio and video creator Mike (Creator Magic channel) recently ran a bold experiment: letting GPT-6 Astra take control of DaVinci Resolve 21.1 via the MCP (Model Context Protocol) to automatically edit 75 minutes of raw footage from scratch.
The experiment was triggered by several converging signals: Higgsfield claimed everything could be edited with GPT-6 Astra on an infinite autonomous canvas; Blackmagic had just released DaVinci Resolve 21.1 with AI assistant integration; and OpenAI officially confirmed that a first-party MCP server had landed in Resolve. This meant an AI agent could, for the first time, "natively" operate professional editing software — no more sluggish, unreliable cursor simulation.

Connecting MCP to Resolve: Near-Zero-Config Handshake
The first step was having Astra find and connect to the MCP server on its own. Mike was explicit: "I want everything done through the MCP link — no computer use, no cursor control."
The connection process went surprisingly smoothly. Astra scanned the local environment, discovered that the DaVinci Resolve 21.1 installer already included Blackmagic's native Resolve MCP, completed the handshake, and registered with Codex. MCP reported that Resolve was running and accessible — all automatic.
The significance of this cannot be overstated. Previous attempts to control editing software with AI agents relied on cursor simulation — slow and unreliable. As a first-party native interface, MCP lets AI directly invoke software functions, moving "AI-controlled professional software" from demo territory into practical territory.
From Transcription to Rough Cut: AI Demonstrates a Complete Editing Workflow
Once connected, Mike gave a moderately complex prompt: place the main video (4K 60fps talking-head footage) and screen recording (5K capture) onto separate tracks, then complete a rough cut using only the DaVinci MCP.
Astra displayed the logic of a seasoned editor. Its first move wasn't to start cutting — it was to transcribe the footage, which is the prerequisite step in any solid AI editing workflow. That said, a first limitation surfaced here: DaVinci's built-in transcription took nearly 6 minutes, while Mike's existing workflow using AssemblyAI takes just 20 seconds. The speed gap is real, though the upside is that DaVinci's transcription is bundled with the Studio subscription at virtually no extra cost.

With the transcript in hand, Astra began making content-based decisions: cutting failed takes, removing long pauses, switching to screen recording footage where it improved clarity, and flagging missing B-roll. It trimmed 75 minutes down to roughly 22 minutes. Mike noted that when he had manually edited the same video using his existing agent-based workflow, the result also came out to around 22 minutes — a compelling validation.
Even more impressive were some proactive touches: Astra spontaneously added audio ducking on the audio tracks, making transitions remarkably smooth. As a self-described perfectionist editor, Mike admitted, "I don't even want to touch it." It also added chapter markers throughout and kept all 96 clips' audio and video links perfectly aligned.
Refinements, Color Grading, and Rendering: A Complete End-to-End Loop
After the rough cut, Mike continued testing finer capabilities.
Targeted Trimming
He asked Astra to fix only a long silence at the 1:20 mark. After checking track alignment, Astra precisely removed 1.67 seconds while keeping the other tracks in sync. The result was correct, but it took nearly two minutes — Mike said bluntly that "spending two minutes on a two-second trim is unacceptable," since a human can do it with Shift+Delete in one second. This clearly defines the human-AI collaboration boundary: rough cuts and high-level decisions suit AI well; instantaneous micro-edits still favor humans.
Lower-Third Graphics and Subtitles
With no style guide provided, Astra designed and inserted a five-second lower-third graphic on its own — white name text with a cyan accent and a fade-in/fade-out animation.

Color Grading and Loudness Normalization
Mike asked Astra to grade the footage like a world-class expert and ensure loudness met YouTube upload standards. Astra enhanced skin tone detail, preserved the studio's purple lighting, and delivered loudness at -14.69 LUFS with a peak of -1.04dB — Mike confirmed "that is absolutely correct." The visual color change was subtle, and this step took over 12 minutes.
Automated Render and Export
Finally, Astra used a preset it created called "Creator Magic YouTube 4K Review," rendered the finished edit via MCP, and saved it to the specified folder — completing the full loop from raw footage import to final output.

The True Time Cost Behind the Demo
Mike was refreshingly honest about this, and it's the most practically valuable part of the entire test. Slick social media demos are often trimmed to seconds or minutes, giving the impression of instant results. The reality: the entire process took Mike nearly an hour. Transcription latency, token consumption, and automatic context compression are all hidden costs.
Video editing is "extremely token-intensive" — midway through the experiment, context hit the auto-compression threshold, triggering a roughly four-minute wait. This is a clear reminder that AI-automated editing is still far from "instant and fully automatic." Speed and cost remain unavoidable real-world constraints.
Where's the Moat? Not the Software — It's Agent Orchestration
Mike's core insight here is worth noting: the real value isn't in "which software you use to edit" or the convenience of visual inspection — it's in how you use AI agents to automate your entire workflow, what skills you build, and how you direct your agents.
He has long been searching for "the Linux of video editing" — a workflow that isn't locked in, one with full freedom and sovereignty. He currently uses Claude Code on Fable 5.1 for automation and is watching open-source alternatives like Diffusion Studio and Chatcut that strip away bloat. By providing a first-party DaVinci MCP, Blackmagic has issued a clear challenge to competitors like Adobe Premiere. Mike explicitly called on Adobe to open up the ability to choose Codex or Claude subscriptions for editing video.
How Far Are We From "Press Stop and You're Done"?
This experiment proved one thing: AI agents can now complete rough cuts, subtitles, color grading, loudness normalization, and rendering end-to-end inside professional editing software, at a quality level that's "ready to publish." For a veteran who has been teaching audio editing for nearly 20 years, that was enough to make him reflect: "Maybe we really are witnessing the early form of AGI."
But a gap remains between "done" and "perfect": it's slow, token-heavy, imprecise on micro-edits, and lacks automatic B-roll insertion. Mike's ideal endgame is a world where creators only press Record and Stop — and AI agents handle everything from editing to publishing. That day may not be far off, but for now, human-AI collaboration remains the most practical answer.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.