Flux.1-Dev SRPO Large Model FP8 Quantized Version_v1.safetensors
UpdateJan 24, 2026 06:29PublishedJan 24, 2026 06:29
91
0
3
Flux.1-Dev SRPO Large Model FP8 Quantized Version_v1.safetensors - 1
Avatar
User_k6sj24
Type
UNet
Basic Model
FLUX.1 D
Published time
Jan 24, 2026 06:29
File Signature
1db3f42651f1c8f3c3f600fc163663fbe79e25f5803949e71baf4a4ec0b236bc
Flux.1-Dev SRPO 大模型 FP8 量化版_v1.safetensors
11.08 GB

Reposted from author: Tencent - SRPO - wikeeyang SRPO (Semantic Relative Preference Optimization)

SRPO is an optimization method developed by the Tencent Hunyuan team for text-to-image generation tasks.

Performance: Experiments on the FLUX.1 - dev model demonstrate that SRPO significantly improves the realism and aesthetic quality of generated images in human evaluations. The original FLUX model had an excellence rate of only 8.2% for realism, which surged to 38.9% after SRPO training; the aesthetic quality excellence rate increased from 9.8% to 40.5%, and the overall preference reached an excellence rate of 29.4%.


This model is a converted and bf16 / fp8_e4m3fn quantized version of the tencent-SRPO model, adapted for ComfyUI users to load and generate images properly while maintaining the original model's generation quality.


SRPO Workflow Link:

https://www.liblib.art/modelinfo/97c008dcca214e7ca69077b484e740b7


Key Features of SRPO:

Enhance Image Generation Quality: Finely optimizes diffusion models, significantly boosting output images in detail representation, visual realism, and artistic aesthetics.

Support Dynamic Reward Adjustment: Users can adjust reward orientation in real time by inputting positive and negative text prompts, flexibly controlling image style and content preferences without needing to retrain or fine-tune the reward model.

Enhance Model Generalization Capability: Enables the model to quickly adapt to diverse human aesthetics and task requirements, such as different lighting, artistic styles, or level-of-detail generation goals.

Efficient Training Mechanism: Focuses optimization on the early stages of the diffusion process, allowing model fine-tuning to be completed in an extremely short time (e.g., within 10 minutes), greatly improving iteration speed and resource utilization.


Core Technical Principles of SRPO

Direct-Align Technology: By pre-injecting noise and utilizing a preset noise prior to recover the original image from any timestep, it avoids the limitation of optimizing only in the later steps, reduces the "reward hacking" phenomenon, and mitigates the gradient explosion problem of traditional methods during backpropagation in early timesteps.

Semantic Relative Preference Optimization: Models the reward as a differential signal guided by positive and negative text prompts. For the same image, the model calculates rewards using positive and negative prompts respectively, then uses their relative difference as the optimization target to achieve real-time control over the generation process.

Gallery

No creation yet