98 related articles

Google Gemini Omni Flash is now open via API, supporting multi-turn video editing with text and reference images, audio-video sync, and character consistency. Learn about its capabilities, API usage, pricing, and best practices.

Deep dive into Google Gemini Omni's core capabilities: multimodal input support for images, video, and audio, enabling interactive video generation and editing—a full-modal AI transforming content creation.

Gemini Omni features native multimodal video editing, directly understanding and editing existing videos. See its style transfer and element addition capabilities demonstrated on a classic 1896 film.

How a creator used Google Gemini, Nano Banana, and VEO to produce the dark fantasy samurai short film The Moon Does Not Forget — full workflow and lessons.

Deep dive into Project Rai-chan's tech stack: Ollama+Gemma local LLM, Unity rendering, VOICEVOX speech synthesis, and more — exploring the technical path for local AI companions.

sqzd is an open-source tool that uses Google Gemini's multimodal video understanding to automatically extract high-value playable clips from long videos.

sqzd is an open-source tool that uses Google Gemini's multimodal video understanding to automatically extract high-value playable clips from long videos.

Moonshot AI launches Kimi K3 with 2.8 trillion parameters and 1M token context. Google delays Gemini 3.5 Pro, AI coding tools upgrade collectively as competition shifts to coding and Agent capabilities.

Google Gemini launches Avatar feature — set up your digital likeness once to generate personalized AI images anytime without re-uploading selfies. Powered by Nano Banana for identity consistency.

A Reddit user generated a polished parody movie poster with a single prompt. This article analyzes AI image generation's one-shot breakthroughs and deepfake risks.

LiblibTV's AI Agent feature tested end-to-end: from a one-sentence brief through storyboarding, Seed Audio music, and CapCut editing to a polished brand film in under two hours.

Full Flowva review: complete AI short film pipeline via Agent chat — from script breakdown to asset generation, storyboarding, and editing, all without switching platforms.

DeepSeek seeks $7B for custom AI inference chips; Zhipu AI explores ASIC. Deep dive into China's AI compute independence strategy, multimodal generation, agents, and hardware trends.
OpenCut: Can This Open-Source Video Ed…
OpenCut is a free, open-source video editor with 72K+ GitHub Stars, offering privacy-first, self-hostable editing as an open-source alternative to CapCut.

Hands-on test of ChatCut AI video editor: Codex plugin support, auto filler removal, MG animation & subtitle generation. Full workflow test on a 3min 20sec video.

ComfyUI v0.28.0 adds SeedVR2 native video super-resolution, PixelDiT architecture, 3D Gaussian Splatting export, int4 quantization, and lip-sync integration for a major multimodal AI workflow upgrade.

iOS 27 deep dive: AI photo Extend & Spatial Reframe, a rebuilt Siri with personal data access, 30%+ system-wide speed gains, and long-overdue quality-of-life fixes — all tested and explained.

LibTV's 'Screenshot to Promo Video' Skill lets designers generate promo videos by simply uploading a mockup — no prompts, no MCP setup required.

A tech blogger with zero programming knowledge built a retro DV app in four days using AI tools like Cursor and Codex, and got it published on Huawei App Gallery. A full vibe coding walkthrough.

A user used Gemini to virtually place a desk lamp in a real photo of their home, generated a precise rendering, and placed an order. An in-depth look at AI image editing's practical value in home design.