Qwen-Image-Edit-2511 vs SenseNova U1.5: A Real-World Test of Multi-Reference Image Fusion

Qwen excels at texture and lighting; SenseNova U1.5 Lite edges ahead on spatial structure in multi-reference image fusion tests.
A Reddit user ran four demanding multi-reference image fusion tasks to compare Alibaba's Qwen-Image-Edit-2511 against SenseTime's SenseNova-U1.5-8B-Lite. Tasks ranged from placing real people in sci-fi scenes to mixing photorealistic and anime characters. Qwen delivered superior texture quality and lighting rendering, while SenseNova showed stronger spatial logic and fewer structural errors — all from a lightweight 8B model. Neither model is a clear winner; the right choice depends on your specific task requirements.
Multi-Reference Image Fusion: The Ultimate Stress Test for Image Editing Models
Image editing models are rapidly pushing their capability boundaries, and "multi-reference image fusion" — integrating visual elements from multiple sources and styles into a single unified scene — is emerging as one of the hardest benchmarks for measuring model capability. A Reddit user recently conducted a head-to-head comparison of two popular open-source image editing models: Alibaba's Qwen-Image-Edit-2511 and SenseTime's open-source SenseNova-U1.5-8B-MoT-Preview (Lite).
The core question was straightforward: SenseNova U1.5 Lite claims to be compact yet capable — but can it actually beat the well-regarded Qwen-Image-Edit-2511 in real-world scenarios? The tester used four demanding tasks to find out.

Test Design: Grounded in Real Creative Workflows
The input images came from mixed sources — some generated by Krea-2, others from real photographs. This mixed-source setup mirrors actual creative workflows, where users often need to combine AI-generated assets with real photos, or even blend elements across different visual styles (photorealistic + anime) within the same composition.
All four prompts fell into the "high-concept, multi-element, multi-constraint" category, involving character placement, prop integration, spatial perspective adjustments, and unified lighting. Tasks like these place extreme demands on a model's instruction-following ability and spatial reasoning.
Where Qwen-Image-Edit-2511 Shines: Texture and Rendering Quality
The tester's first observation was that Qwen-Image-Edit-2511 is "genuinely impressive" in terms of image texture quality — particularly in its handling of lighting and shadow.
This matters a great deal in practice. Take the third task — "orange cat riding a vintage toy scooter" — as an example. The prompt was extremely demanding: it required capturing the cat's distinctive facial features (chubby face, green eyes, a dense white triangular patch on the chest), while also specifying golden afternoon sunlight as the key light, a warm golden rim light along the cat's back and ear edges (backlit effect), metal bodywork reflecting the street scene outside the window, and soft-edged shadows cast by the entire object onto a wooden floor.
For these physically demanding lighting details, Qwen demonstrated solid rendering chops. For social media creators or commercial visual production where final image quality matters, this texture advantage is a real, tangible benefit.
Where SenseNova U1.5 Pulls Ahead: Spatial Understanding
That said, beautiful textures don't automatically mean a coherent scene. The tester found that SenseNova actually held the upper hand when it came to spatial understanding.
The clearest evidence, again, came from the cat-on-scooter comparison:
- Qwen's problem: It generated a mysterious "pillar" beneath a coffee table, and the cat's front paws were placed in a physically implausible position.
- SenseNova's performance: No such artifacts appeared — the spatial relationships between objects and the points of physical contact were far more logical.
This distinction is critical. The real challenge of multi-reference image fusion isn't making elements look attractive in isolation — it's ensuring that elements drawn from different source images are logically consistent within a shared three-dimensional space: perspective alignment, natural occlusion and compression at contact surfaces, consistent shadow direction. Qwen generating a "ghost pillar" reveals structural misjudgments that can arise when the model parses complex spatial constraints.
Why Spatial Understanding Is Harder to Crack Than Texture Rendering
It's worth digging into why. Texture and lighting are fundamentally surface rendering problems, while spatial relationships are structural modeling problems. The former can be addressed through powerful generative priors and style fitting; the latter requires the model to genuinely "understand" the three-dimensional positions of objects and how they relate to one another.
SenseNova-U1.5 employs what it calls a "Native Unified Paradigm" and NEO-unify architecture, where text understanding, image perception, and generation capabilities are jointly trained within a single unified model framework — rather than being handled by separate modules in a sequential pipeline. Its MoT (Mixture of Tokens) mechanism dynamically allocates computational resources across different input modalities, which theoretically helps maintain global spatial consistency when integrating multiple reference images.
By contrast, many image editing models are built by layering instruction fine-tuning on top of pre-trained image generation models (such as Stable Diffusion-based architectures). This "bolt-on" approach may be more prone to strong local perception but weak global structure when tasks require holistic 3D reasoning — and Qwen's generation of a "ghost pillar" that defies the overall spatial logic may well be a manifestation of exactly this architectural limitation.
Four High-Difficulty Fusion Tasks: A Comprehensive Gauntlet
Beyond the cat-on-scooter scenario, the other three tasks were equally demanding:
- Futuristic rooftop scene: Placing a real person on a sci-fi rooftop at sunset, with requirements for depth-of-field blur, golden-hour lighting, and flying drones in the background — testing lighting atmosphere and subject integration.
- Vintage study cross-style fusion: A photorealistic young woman, an anime-style blonde girl, a Roman temple engraving, and a vintage card all sharing the same frame — an extreme test of coexistence between realistic and anime visual styles in one image.
- Creative market panorama: Two realistic human figures and two anime characters naturally sharing a space under unified lighting and consistent spatial perspective — the ultimate stress test for cross-style fusion.
These tasks collectively point to a broader trend: image editing is moving from "localized edits on a single image" toward "scene-level synthesis from multi-source assets." The evaluation criteria are evolving accordingly — it's no longer just about point-level image quality, but about overall logical consistency, stylistic coherence, and physical plausibility.
"Coexistence of photorealistic and anime styles" is widely recognized as one of the most challenging tasks in multi-reference image fusion. The two styles differ fundamentally in line logic, color saturation, and lighting models: photorealistic styles rely on physically accurate illumination, while anime styles use simplified cel-shading. When characters from both styles appear in the same frame, the model must simultaneously maintain two distinct visual grammars without letting them "contaminate" each other — anime characters can't be rendered into an uncanny-valley semi-realistic state, and realistic figures can't be cartoonified. This demands not just style recognition, but the ability to apply different rendering strategies to different regions during generation — an extreme stress test of a model's fine-grained regional control.
Verdict and Selection Guidance
Based on this community benchmark, both open-source image editing models have distinct strengths, and neither is a clear overall winner:
- Qwen-Image-Edit-2511 wins on texture quality and lighting/shadow rendering precision — best suited for use cases where final visual polish is paramount.
- SenseNova-U1.5-Lite wins on spatial understanding and structural plausibility, producing fewer physically implausible artifacts in complex multi-element fusion tasks — and does so with a lightweight 8B parameter footprint.
For developers and creators, the key takeaway is: model selection should be driven by specific task requirements. If your core pain point is multi-image fusion with complex spatial relationships, SenseNova deserves serious evaluation. If you need the highest possible lighting and texture fidelity, Qwen remains a strong contender.
One important caveat: this is a four-task subjective comparison from a single user, with a limited sample size and no coverage of engineering metrics like inference speed or VRAM usage. Any real selection decision still requires more systematic testing within your own workflow. That said, the rapid progress open-source image editing models are making on the hard problem of multi-reference image fusion is itself a signal worth paying attention to.
Related articles

Invalid Source Material: Unable to Generate a Valid AI/Tech Article
This Twitter source material is an irrelevant marketing tweet with no AI or tech content, making it impossible to generate a valid professional article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material was limited to a single broken tweet with no usable content, making it impossible to produce a complete, high-quality article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material provided was a single vacuous social media tweet with a broken link — insufficient to support writing a complete, factual article.