MiniMax H3 Ref2VA+FlashVSR

Update:Sep 20, 2026 15:23|Published:Sep 20, 2026 15:23
1
0
0
MiniMax H3 Ref2VA+FlashVSR. MiniMax H3 Ref2VA → FlashVSR Sharp Finish This workflow does two jobs in one run: it generates a finished video with sound from a single reference image…
Avatar
DBclown(LH)
Type
Workflow
Basic Model
MiniMax H3
Published time
Sep 20, 2026 15:23
File Signature
8255a44f84eb5388d7c0501762fb5eb353e802d197adeee4b19754751d8d44a3

MiniMax H3 Ref2VA → FlashVSR Sharp Finish This workflow does two jobs in one run: it generates a finished video with sound from a single reference image, then runs the whole thing through a fast upscaler before saving — so what you download is already a sharp, ready-to-use MP4. The first stage is built on MiniMax H3's Ref2VA model. You drop in one still image to lock the look, write a prompt, and the model renders picture and audio together — footsteps, water, wind, even a bit of score, all placed where the action happens. The sample prompt in the template shows how much detail the model happily works with: shot-by-shot timing, a soundscape line, a music line. The more you describe, the better it lands. Then, instead of stopping there, the finished clip gets unpacked, enlarged 2× with FlashVSR, and re-packed into a clean H.264 file with the original audio untouched. You get the sharpness of a bigger render without actually paying for one. What you control: One image, one look — the reference image keeps your character and style steady across the entire clip. Duration in plain seconds — type 5 to 15; the workflow converts it into the exact frame count the model prefers, so you never have to think about it. Auto-fitted resolution — your reference image is resized to fit the model's pixel budget, letterboxed if needed. Sharper output, same cost — the FlashVSR pass doubles resolution at the end, with tiled processing so it runs comfortably on modest GPUs. Lighter model build — uses the compressed INT8 version of H3, so it loads faster and needs less VRAM. You bring an image, a prompt, and a length. The frame math, resizing, upscaling, and audio muxing are already wired for you.

Gallery

No creation yet