

Main H3-Omni-Transformer of the MiniMax-H3 Base FL2VA checkpoint. Handles text-to-audio-video (t2va) and first/last-frame-to-audio-video (fl2va), jointly predicting video and audio latents. With INT8 quantization and pruning to cut memory and speed up inference while preserving generation quality.
暂无作品
