HuMo_17B_fp8_e4m3fn
更新2026-03-16 07:56发布时间2026-03-16 07:56
6
0
0
头像
Gugugaga
类型
UNet
基础模型
Other
发布时间
2026-03-16 07:56
文件签名
1e20716f122a59d77a6e0c26f5c20ff0610563b6d78c7f3b33681eb0284f592d
humo_17B_fp8_e4m3fn.safetensors
15.89 GB

HuMo is a unified, human-centric video generation framework designed to produce high-quality, fine-grained, and controllable human videos from multimodal inputs—including text, images, and audio. It supports strong text prompt following, consistent subject preservation, synchronized audio-driven motion.

  • VideoGen from Text-Image - Customize character appearance, clothing, makeup, props, and scenes using text prompts combined with reference images.
  • VideoGen from Text-Audio - Generate audio-synchronized videos solely from text and audio inputs, removing the need for image references and enabling greater creative freedom.
  • VideoGen from Text-Image-Audio - Achieve the higher level of customization and control by combining text, image, and audio guidance.

GitHub Reposted from Huggingface

作品

暂无作品