Wan22-I2V-A14B-LOW-rCM1_0_lora_rank_64_bf16.safetensors
UpdateMar 19, 2026 22:34PublishedMar 19, 2026 22:34
0
0
0
Wan22-I2V-A14B-LOW-rCM1_0_lora_rank_64_bf16.safetensors - 1
Avatar
Gugugaga
Type
LoRA
Basic Model
Wan Video
Published time
Mar 19, 2026 22:34
File Signature
1ea30aea1ff08dfe7412785b633a6efd40119785f07a0f6ba8a7ab0c3b9b74d8
Wan22-I2V-A14B-LOW-rCM1_0_lora_rank_64_bf16.safetensors
601.48 MB

rCM Consistency Distillation Diffusion Model (LoRA supports Wan 2.2 / 2.1 / Text-to-Video / Image-to-Video)

rCM can generate high-quality images and videos in just 2 to 4 steps, which significantly speeds up generation compared to traditional diffusion models. By employing score regularization and a joint forward-reverse divergence distillation framework, rCM substantially improves the quality of generated images and videos.

Wan2.2 Model Files (Kijai version):

Wan22-I2V-A14B-HIGH-rCM6_0_lora_rank_64_bf16.safetensors

Wan22-I2V-A14B-LOW-rCM1_0_lora_rank_64_bf16.safetensors

Wan2.1 Model Files (Kijai version):

Wan_2_1_T2V_14B_480p_rCM_lora_average_rank_148_bf16.safetensors

Wan_2_1_T2V_14B_720p_rCM_lora_average_rank_94_bf16.safetensors

Wan_2_1_T2V_14B_480p_rCM_lora_average_rank_83_bf16.safetensors

Wan_2_1_T2V_1_3B_480p_rCM_lora_average_rank_64_bf16.safetensors

  • If you have already loaded the 4-step acceleration LoRA from lightx2v, it is recommended to start testing with a weight of 0.3.

*Note: Regarding the rank parameter in LoRA models—the rank parameter is a key hyperparameter that determines the dimensionality of the low-rank matrices, thereby affecting the model's adaptability and computational efficiency. Choosing an appropriate rank parameter requires a trade-off between adaptability and computational efficiency, which typically needs to be determined through experimentation. In practical applications, the rank parameter can be gradually adjusted according to specific tasks and resource conditions to achieve the best balance between performance and efficiency.

Image Generation Tasks: For image generation tasks, a larger rank parameter can better capture image details and features, but computational costs should be taken into account.

Lightweight Applications: For lightweight applications that need to run on mobile devices, a smaller rank parameter can significantly reduce computational and storage costs, improving user experience.


rCM (Score-Regularized Continuous-Time Consistency Model)

It provides multifaceted practical value for Wan2.1 image/video models, mainly reflected in the following key areas:

  1. Accelerating the Generation Process

Fast Sampling: rCM can generate high-quality images and videos in just 2–4 steps, which greatly speeds up generation compared to traditional diffusion models. For example, traditional diffusion models may require hundreds of steps to generate high-quality samples, whereas rCM accelerates this process by 15x to 50x through an optimized distillation method.

Efficient Resource Utilization: This acceleration not only saves time but also significantly reduces computational resource consumption. For large-scale image and video generation tasks, such as content creation and video editing, this means more data can be processed in a shorter time, improving workflow efficiency.

  1. Enhancing Generation Quality

High-Quality Output: Through score regularization and a joint forward-reverse divergence distillation framework, rCM significantly improves the quality of generated images and videos. It can generate finer and more realistic details while avoiding common blurring and distortion issues found in traditional methods.

Diversity Preservation: rCM not only improves generation quality but also maintains the diversity of generated samples. This means when generating multiple samples, each output retains unique characteristics rather than being repetitive. This is highly valuable for applications requiring diverse content (e.g., advertising design, video special effects).


rCM: Large-Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency (Paper)

Authors: Kaiwen Zheng, Yuji Wang, Qianli Ma, Huayu Chen, Jintao Zhang, Yogesh Ba****, Jianfei Chen, Ming-Yu Liu, Jun Zhu, Qinsheng Zhang

Institutions: NVIDIA, Tsinghua University

Research Highlights

rCM is the first to:

Scale continuous-time consistency distillation (e.g., sCM/MeanFlow) to 10B+ parameter video diffusion models.

Provide an open-source FlashAttention-2 Jacobian-Vector Product (JVP) kernel supporting parallelization techniques such as FSDP/CP.

Identify the quality bottleneck of sCM and overcome it using a joint forward-reverse divergence distillation framework.

Generate high-quality videos with strong diversity in just 2–4 steps.

High-level comparison of diffusion distillation methods. Although forward divergence exists theoretically, practical GANs still suffer from limited diversity and mode collapse.

Abstract: This study scales continuous-time consistency distillation to general application-level image and video diffusion models for the first time. Although continuous-time consistency models (sCM) are theoretically sound and empirically powerful in accelerating academic-scale diffusion, their applicability to large-scale text-to-image and video tasks has remained unclear due to infrastructure challenges in Jacobian-Vector Product (JVP) computation and limitations in standard evaluation benchmarks. We first develop a parallelization-compatible FlashAttention-2 JVP kernel, enabling sCM to be trained on models exceeding 10 billion parameters and high-dimensional video tasks. Our study reveals a fundamental quality limitation of sCM in detail generation, which we attribute to error accumulation and the "mode-covering" nature of its forward divergence objective. To address this, we propose Score-Regularized Continuous-Time Consistency Models (rCM), which incorporate score distillation as a long-horizon regularizer. This integration complements sCM with a "mode-seeking" reverse divergence, effectively boosting visual quality while maintaining high generation diversity. Validated on large-scale models up to 14 billion parameters and 5-second videos (Cosmos-Predict2, Wan2.1), rCM matches or outperforms state-of-the-art distillation methods like DMD2 on quality metrics while offering significant advantages in diversity—all without requiring GAN tuning or extensive hyperparameter searches. Distilled models generate high-fidelity samples in just 1–4 steps, accelerating diffusion sampling by 15x to 50x. These results establish rCM as a practical and theoretically grounded framework for advancing large-scale diffusion distillation.

Results: By distilling Cosmos-Predict2/Wan2.1 image/video models, rCM achieves state-of-the-art results on few-step GenEval and VBench benchmarks.

Gallery

No creation yet