AI Agent Interior Design End-to-End Test: From Floor Plan to Presentation Deck via Conversation

A conversational AI Agent completes the entire interior design pipeline — from floor plan to PPT — via natural language alone.
This article walks through a real-world test of an AI Agent platform that lets users complete a full interior design workflow — multi-style layout proposals, effect renderings, storyboard animations, and a presentation deck — using only natural language. It examines the underlying technology, the platform's key features, and an honest assessment of both the value and current limitations of this approach.
From Blank Floor Plan to Presentation Deck — All Through Conversation
Traditional interior design workflows are long and laborious: from floor plan analysis and layout design to rendering, material refinement, diagram production, and final client presentations, every step demands specialized software and significant manpower. Now, as conversational AI Agent capabilities mature, this entire pipeline is being fundamentally reimagined.
According to a hands-on demo by a Bilibili creator, using the Agent mode on the "AI Interior Master" platform, users don't need to write a single prompt. By simply issuing natural language instructions — just like a client would — the AI can start from a blank floor plan and progressively complete multi-style commercial interior design, effect rendering, storyboard animation, and even a presentation deck. The entire process is highly automated and accessible even to those with no interior design background.
This article breaks down each stage of this conversational AI interior design workflow and analyzes the underlying product logic and real-world value.
Agent Mode and Model Selection: Understanding Your Tools First
An AI Agent is an AI system capable of perceiving its environment, autonomously planning, and executing multi-step tasks — distinct from single-turn large language models. Its core capabilities lie in Tool Use and contextual memory: within a single conversation, an Agent can invoke multiple tools such as image generation, text analysis, and file export, using the output of one step as the input to the next, forming an automated task chain. This architecture, driven by standards like OpenAI's Function Calling and Anthropic's Tool Use, represents a critical leap from "point-solution AI" to "end-to-end workflows."
To get started on the platform, navigate to "AI Design Agent" in the top navigation bar and switch to Agent mode. This step is essential — Agent mode differs from standard single-shot image generation in that it maintains conversational context, proactively asks clarifying questions, and chains multiple tools into a complete pipeline.
For model selection, the platform offers several options including Image 2, Banana Pro, and Banana 2. Based on the creator's experience:
- Banana series: Superior visual output, ideal for scenarios that prioritize aesthetic impact;
- Image 2: Stronger text comprehension, better suited for complex instructions requiring precise intent execution.
Users can switch between models depending on their needs, and can also configure aspect ratio and resolution. This "visual quality vs. semantic understanding" trade-off reflects a universal tension in AI image generation today. Leading models like Stable Diffusion, DALL-E, and Midjourney all face an architectural trade-off between visual fidelity and text-semantic alignment. Visually strong models tend to prioritize fitting aesthetic data distributions during training, while semantically strong models require more powerful cross-modal alignment (e.g., CLIP scores). This tension stems from diffusion model mechanics: a higher CFG Scale during denoising improves semantic accuracy but can make images feel rigid; a lower CFG Scale produces more natural visuals but risks diverging from the prompt. Achieving both simultaneously remains a genuine challenge.
Generate Multiple Layout Options with a Single Sentence
After uploading a blank floor plan, simply tell the AI to "generate several different commercial interior layout options based on this floor plan," and after a brief moment of reasoning, it will produce multiple options. In the demo, the AI generated three concepts in one pass — an office, a café/restaurant, and a retail showroom — summarizing the functional zoning and applicable scenarios for each, along with a comparison overview for easy selection.

Even more notable is the AI Agent's proactivity. After delivering the initial results, it will proactively ask about optimization needs. When the user expressed interest in "exploring more possibilities for the floor plan," the AI generated two additional concepts, combining them with the original three to produce a final shortlist of five distinct options. This "AI-guided" interaction model dramatically reduces cognitive load — you don't need to know what to ask; the AI will lead you forward.
Scheme Refinement and Style Exploration: Choosing References Like You'd Choose a Designer
Once the "café/restaurant" scheme is selected, users can ask the AI to develop it further. The AI will convert the floor plan into a color-coded zoning diagram with Chinese annotations, and generate effect renderings for the dining area and bar counter.
If the user isn't satisfied with the café rendering style, a single prompt — "show me some café interior reference cases" — will return multiple styles including Japanese minimalist, industrial, and Nordic. If those aren't enough, asking to "generate five more" will produce additional options, automatically categorized alongside the existing styles.

Once a preferred style is chosen, simply tell the AI, and it will apply that style to the café/restaurant floor plan. This "reference → selection → application" workflow essentially simulates the feedback loop between a designer and client, compressing what might take days into a matter of minutes.
Precision Refinement: From Overall Suggestions to Localized Edits
After scheme development, users can reference a specific image and ask the AI "what can be optimized in this interior layout." The AI responds with six professional recommendations covering spatial hierarchy, lighting layers, and material balance — each with a prioritized order of implementation. This kind of design-logic-driven analysis is simply beyond the reach of standard image generation tools.
For specific modifications, the platform supports two approaches:
Conversational Global Edits
Describe your idea in natural language — for example, requesting a material swap — and the AI will adjust the entire composition accordingly, with no need to touch any professional software.
Selection-Based Local Edits
In the image editing interface, draw a selection box around the area you want to modify, then describe the change in text. The AI will make precise local edits without affecting the rest of the image. You can also generate specific lighting scenarios, such as "bar counter with evening lighting."

The platform also includes built-in material replacement and interior lighting tools, accessible via the "More Options" menu on any rendering. This combination of "conversational rough-tuning + tool-based fine-tuning" balances both efficiency and control.
From Static Drawings to Storyboard Animation and Presentation Deck
The workflow doesn't end with floor plans and renderings. Users can also have the AI generate floor tile layouts and ceiling plans. For those unsure which analytical diagrams to produce, the platform offers one-click templates under "Prompt Templates — Interior Analysis Diagrams," lowering the barrier to entry.
Going further, the AI can generate a nine-panel storyboard with auto-written prompts for each scene. The nine-panel storyboard concept originates from the film industry's storyboard tradition — by decomposing a space into multiple fixed camera positions, it helps clients intuitively preview the spatial experience before construction begins. Copy those prompts, open an AI video tool, upload the images, and select a video model to generate interior storyboard animations for the restaurant. The primary technical approach for converting static renderings into interior animation is the Image-to-Video model, with leading examples including Runway Gen-3, Kling, and Pika. These models predict frame-by-frame temporal changes from a static image to simulate camera movements such as dolly-ins, orbits, and pans. For interior scenes, the main challenge lies in maintaining spatial perspective consistency and temporal stability of material details (i.e., avoiding flickering artifacts).

Once all drawings are essentially complete, users can ask the AI to "outline a presentation narrative," and it will draw from the full conversation history to propose an overall structure — or even directly generate several slides for a presentation deck. Finally, uploading the floor plan to the "AI Image-to-CAD" tool yields an editable CAD file. The core technical approaches for AI image-to-CAD conversion include two categories: first, computer vision-based line detection and vectorization, using edge detection algorithms (such as Canny or HED) to identify walls, doors, and window outlines and convert them to vector segments; second, deep learning-based semantic segmentation, which classifies different regions (walls, door swing directions, furniture, etc.) and reconstructs them as structured CAD layers, outputting DXF/DWG formats directly importable into AutoCAD, Revit, or other mainstream engineering software. This means the AI's outputs can seamlessly connect to downstream processes in traditional design software.
Value and Limitations: How AI Is Reshaping the Interior Design Workflow
This complete demonstration makes it clear that conversational AI Agents are shifting interior design from "professional-tool-driven" to "intent-driven." The core value operates on three levels:
First, lowering the barrier to entry. No need to master PS, SketchUp, CAD, or any prompt engineering techniques — anyone can produce a complete design proposal using natural language alone.
Second, connecting the full pipeline. From floor plan analysis to CAD output, the AI Agent integrates what were previously fragmented tools into a single conversational flow, dramatically reducing the cost of switching between applications.
Third, proactive guidance. Rather than passively responding to prompts, the AI acts like a design assistant — proactively suggesting optimizations and asking clarifying questions, effectively bridging the experience gap for non-professional users.
That said, it's important to maintain a realistic perspective on the limitations. AI-generated renderings are best suited for early-stage concept visualization and design communication; precise dimensions, material specifications, and structural details required for construction still need professional oversight. Image-to-CAD accuracy can still fall short in complex scenarios like irregular spaces or hand-drawn sketches, and the temporal coherence of animated storyboards also has room for improvement.
Overall, Agent mode combined with a suite of AI tools is making the early creative and presentation phases of interior design more efficient than ever before. As the demo illustrates, this is just a glimpse of what AI design agents are capable of — and as deep reasoning capabilities continue to advance, the possibilities ahead are far greater still.
Key Takeaways
Related articles

Pinery Prose: Redefining the AI Book-Writing Experience with Diff Review
Pinery Prose is a Mac AI book-writing assistant using code diff review mechanics, letting authors accept or reject each AI edit. Supports Markdown, ePub/PDF export, and covers the full self-publishing workflow.

How Developer Productivity Startups Boost Their Own Efficiency: Practicing What You Preach
How developer productivity startups practice what they preach—from automated toolchains and DORA metrics to engineering culture that shortens feedback loops and reduces cognitive load.

Laxis Review: Bot-Free Meeting Notes & Real-Time Translation AI Tool
In-depth review of Laxis AI meeting tool: bot-free recording, 100+ language real-time translation, voice dictation 4x faster than typing. Features, competitors & value analysis.