WAN 2.1 Physics Motion LoRA Benchmark: Rankings and Methodology for 11 Models Tested

Optical flow benchmark of 11 WAN 2.1 physics LoRAs: only 3 work, coherence beats raw amplitude.
A Reddit user ran a rigorous controlled benchmark of 11 physics-motion LoRAs on WAN 2.1, fixing seed, prompt, and drive video while using optical flow analysis and inter-frame differential energy as scoring metrics. Only 3 of 11 LoRAs produced meaningful effects, with 4 scoring below the no-LoRA baseline. The top performer, Bouncing_Breasts...HIGH_v1.0, won on both physics and shape — its key advantage being the highest motion coherence and +50% residual motion after body deceleration, demonstrating genuine inertia. The benchmark also revealed that LoRA naming is unreliable, and that motion coherence matters more than raw amplitude when evaluating realism.
In AI video generation, LoRA (Low-Rank Adaptation) models are commonly used to fine-tune specific actions or effects. A Reddit user conducted a controlled benchmark of 11 physics-motion LoRAs on the WAN 2.1 image-to-video model, replacing subjective impressions with optical flow analysis and inter-frame differential energy — producing a ranking with genuine methodological value.
Test Method: Fixed Seed, Fixed Prompt, Fixed Drive Video
What makes this benchmark worth studying is its rigor. The tester locked the random seed, prompt, and drive video — the only variable was the LoRA being used. This ensures that the rankings reflect differences between LoRAs themselves, not random noise from other factors.
Scoring was based on two dimensions:
- Physics: Does the motion behave like real soft tissue?
- Shape: Is the morphology anatomically natural?
Crucially, the tester emphasized that the criterion is the "deformation chain" — not raw displacement magnitude. Moving more does not mean moving realistically. This principle runs throughout the entire benchmark, and it's what sets it apart from ordinary "gut feeling" evaluations.

Optical flow is a computer vision technique for estimating pixel-level motion between video frames. By tracking changes in brightness patterns between adjacent frames, it generates a motion vector field for each pixel. In video quality assessment, it can quantify the magnitude, direction, and coherence of motion — far more objectively than human intuition. Inter-frame differential energy is more straightforward: it computes the sum of squared grayscale differences between corresponding pixels in adjacent frames, with higher values indicating greater frame-to-frame change. Used together, the two metrics measure both how strong the motion is (inter-frame differential) and whether the motion is organized and directional (optical flow coherence) rather than random pixel noise. This is the technical foundation that allows this benchmark to distinguish "real motion" from "frame jitter."
Results: Only 3 of 11 LoRAs Actually Worked
The most surprising finding: only 3 of the 11 LoRAs produced meaningful effects. Worse, 4 of them scored below the no-LoRA control group — meaning they didn't just underperform, they actively reduced motion in the wrong direction.
Top results were as follows:
| LoRA | Physics | Shape | Notes |
|---|---|---|---|
| Bouncing_Breasts...HIGH_v1.0 | ★★★★★ | ★★★★☆ | 🏆 Overall best |
| big_breasts_v2 | ★★★☆☆ | ★★★☆☆ | Solid performer |
| BoobPhysic_HighNoise | ★★★★☆ | ★☆☆☆☆ | Good motion, unnatural shape |
| jiggle_tits-14b | ☆☆☆☆☆ | ★★★★★ | Best shape, zero motion effect |
The tension between the two dimensions is where the real value lies. BoobPhysic had the strongest motion response, but the shape was spherical — overly full on top, lacking the natural downward gradient, moving like a rubber ball rather than a soft suspended mass. Meanwhile, jiggle_tits-14b had the most natural morphology (a clean teardrop shape) but produced zero motion effect — actually a negative gain.
Another counterintuitive finding: the Perky HIGH variant performed worse than Perky LOW within the same series. The "HIGH" in the name should imply stronger effect, but the opposite was true — a reminder that you can't judge intensity by naming alone.
Why the Champion Is "Physically Correct"
The tester offered a well-grounded physics explanation for why the top LoRA won. Soft tissue is not a rigid body — it's a compliant mass suspended from the chest wall, and the realistic behavior of its free end must satisfy three conditions:
- Lag: Inertial delay relative to the chest wall
- Overshoot: Compliance-driven overshoot beyond equilibrium
- Ring-down: Damped oscillation that settles within 2–4 cycles
Key optical flow measurements were as follows:
| LoRA | Cantilever Gain | Propagation Delay | Coherence |
|---|---|---|---|
| Bouncing_Breasts | 2.09 | 62.5 ms | 0.718 |
| BoobPhysic | 2.95 | 31.2 ms | 0.568 |
| big_breasts_v2 | 2.59 | 62.5 ms | 0.668 |
While BoobPhysic had the highest raw gain, it had the lowest coherence — its free end moved chaotically in all directions rather than as an organized deformation chain. The real differentiator: after the body decelerated to a stop, Bouncing_Breasts retained +50% more motion relative to the control group. This "still moving after the body stops" behavior is precisely what inertia looks like, and it's the source of perceived realism.
The "cantilever gain" mentioned here borrows from structural mechanics: soft tissue is modeled as a cantilever beam — fixed at one end, free at the other — and the gain is the ratio of free-end displacement to input motion at the fixed end. Higher gain means the free end is more responsive, but high gain does not necessarily mean realism — coherence is equally critical. Coherence measures how consistently the motion vectors across the free-end region align: coherence near 1 means the region moves as a coordinated whole, while coherence near 0 means each point moves independently, resembling random noise. BoobPhysic's gain was the highest (2.95) but its coherence was the lowest (0.568), indicating localized chaotic pixel disturbance rather than an ordered deformation chain propagating from the fixed end to the free end. This difference is nearly invisible in a single frame — it can only be identified through the optical flow vector field.
Conclusions and Essential Caveats
The final verdict is clear:
- 🏆 Overall best: Bouncing_Breasts_Wan2.2_14B_I2V_HIGH_v1.0 — high scores on both physics and shape.
- Stiffer/faster alternative: BoobPhysic_HighNoise, but with an "implant-like" shape.
- Skip entirely: All models in the lower half of the rankings.
That said, the tester's list of caveats is equally important and reflects the same rigor:
- Overall effect sizes are small — ranging from −3% to +18% relative to baseline motion. These are subtle modifiers, not dramatic transformations.
- Some LoRAs ran at reduced strength in the first pass, so their scores represent floor values.
- RIFE ×2 frame interpolation in the pipeline smooths high-frequency detail, making it impossible to accurately measure the damping ratio.
- Shape was judged from single frames at a single angle; dose-response above 1.0 strength, other poses, identity preservation, and temporal stability were not tested.
Implications for AI Video Fine-Tuning
Setting aside the subject matter, this benchmark's methodology has reference value for anyone working on AI video fine-tuning. It demonstrates several things: LoRA naming and claimed strength are unreliable and require empirical validation; replacing subjective impressions with quantifiable metrics (optical flow, inter-frame energy) reveals differences that are invisible to the naked eye; and when evaluating "realism," motion organization (coherence) is often more important than motion magnitude.
For users looking to tune effects on WAN series models, the "fixed variables + quantitative scoring" approach is far more effective at identifying genuinely useful components than blindly stacking LoRAs.
LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning method whose core idea is: rather than modifying all weights of a pretrained model, attach a small low-rank decomposed matrix alongside the original weight matrix (the product of two tall-and-thin matrices), and only update this small matrix during training. Because the low-rank matrix has far fewer parameters than the original weights, LoRA can adapt a large model to a specific style or motion with minimal compute. In video generation, LoRAs are typically distributed as supplementary weight files that users layer onto the base model at inference time with a specified "strength" (weight). Too low and there's no visible effect; too high and it may degrade image quality or identity consistency — which is why the benchmark's caveats note that dose-response above 1.0 has not yet been tested.
Related articles

Show HN: A New Platform for Sharing AI Workflows and Learning from Others
A new Hacker News Show HN platform lets developers share AI tool setups and workflows. Explore its value, early traction, and what it means for AI users.

Lucid Partners with Bolt to Target European Robotaxi Market
Lucid Motors has signed a letter of intent with European mobility platform Bolt to explore Robotaxi services in Europe, though no vehicle orders have been placed yet.

AI Assistants Enter the "Phone Call" Era: Instinct and Meta Muse Add Voice Task Execution
AI assistants Instinct and Meta Muse now make phone calls on your behalf — booking restaurants, canceling subscriptions — marking a leap from chat tools to real-world agents.