HuMo-1.7B
UpdateMar 16, 2026 08:05PublishedMar 16, 2026 08:05
7
0
0
Avatar
Gugugaga
Type
UNet
Basic Model
Other
Published time
Mar 16, 2026 08:05
File Signature
e2d5d8f8ffa6a97dbc4142f41ffbf6aa4c069ce047dcf934e525141207481b48
humo_1.7B_fp16.safetensors
3.24 GB

HuMo is a unified, human-centric video generation framework designed to produce high-quality, fine-grained, and controllable human videos from multimodal inputs—including text, images, and audio. It supports strong text prompt following, consistent subject preservation, synchronized audio-driven motion.

  • VideoGen from Text-Image - Customize character appearance, clothing, makeup, props, and scenes using text prompts combined with reference images.
  • VideoGen from Text-Audio - Generate audio-synchronized videos solely from text and audio inputs, removing the need for image references and enabling greater creative freedom.
  • VideoGen from Text-Image-Audio - Achieve the higher level of customization and control by combining text, image, and audio guidance.

GitHub Reposted from Huggingface

Gallery

No creation yet