Ported from https://modelscope.cn/models/MoYouuu/MYHuman-QWen/summary?version=20251019221336
MoYou Human QWen is a fine-tuned and merged model based on qwen image 2509. Utilizing distributed multi-stage training, merging, and differential extraction, it achieves the uncontaminated coexistence of multiple styles, categories, and concepts. It uses qwen2.5vl for captioning to achieve full alignment with CLIP, while maintaining strong compatibility with native qwen image LoRAs.
Model Features
- First Qwen model supporting direct output at 2K resolution.
- Exceptional Portrait Generation Capabilities: Effortlessly responds to prompts for male, female, young, old, single, or multiple subjects. Allows flexible adjustments for strong, shallow, or no depth-of-field (bokeh) effects.
- Broad Applicability: Covers various textures (film, studio, web images, etc.), styles (modern, traditional/ancient Chinese, fantasy, cosplay, etc.), and special compositions (grid layouts, breaking out of the frame, etc.). Capable of stable, batch generation for both Chinese and English posters.
- Easy to Use: Responds well to simple Chinese descriptions, or you can use Qwen interrogators/taggers. Features excellent prompt responsiveness.
Local Usage Recommendations: Since running Qwen image locally requires significant VRAM (around 24GB), it is recommended to use it online. In ComfyUI, select "Browse Templates" from the dropdown menu and use the Qwen image template workflow.
Please use long Chinese prompts to achieve better results.
- Sampler: euler, res_multistep
- Scheduler: simple, sgm_uniform
- Steps: 20, 30, 50
- CFG: 3.5, 4
- Model Sampling Algorithm AuraFlow Shift: 3.1
- Recommended Resolutions: 1152×1536, 992×1776, 928×1984, 1536×2048 (All resolutions work well in both portrait and landscape orientation)