

language:
An experimental character-replacement LoRA trained by Akatz Labs for 1,000 updates. It aims to replace a selected character using an image reference while preserving the source scene. In our local comparisons it often preserved the background more closely than the base model, but motion timing, facial expressions, and hard cuts remain unreliable.
Download: final 1,000-step LoRA. Only the final 1,000-step checkpoint is published. Start with the final checkpoint at strength 1.0. This is an adapter, not a standalone model or a Turbo distillation LoRA.
Dataset: H3 Character Swap v1.
.safetensors file in ComfyUI/models/loras/ (or your configured shared LoRA directory), and apply it to the H3 model at strength 1.0 using a compatible model-only LoRA loader.<Video 1> and replacement image or character sheet as <Picture 1>.Example:
Replace only the man in the purple shirt in <Video 1> with the character in <Picture 1>. Keep the replacement character's identity, outfit, and art style from <Picture 1>. Preserve the source video's camera, background, lighting, objects, and all other people. Match the target person's position, scale, pose, and movement. Do not show the reference sheet or its background.
The training captions were shorter, for example Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>. Preservation instructions helped some local evaluations, but stronger expression instructions sometimes suppressed the swap entirely. Prompt wording is not a guarantee of strict source alignment.
Short, continuous shots of roughly 4–5 seconds were more promising than our full 14-second tests. A precise maximum duration has not been established. Use 24 fps and your runtime's supported H3 frame grid. The character-swap LoRA does not require a Turbo LoRA, Spectrum, or Sol attention.
Training used minimax_h3_ref2va_pruned_int8_convrot.safetensors from Comfy-Org/MiniMax-H3, plus the frozen Ostris Ref2VA training assistant. The assistant and base weights are not merged into this adapter and are not distributed here. Model revisions and hashes are in training/base-model-files.json.
We evaluated on the Ref2VA base and a local FL2VA/Ref2VA hybrid (blocks 25–49, INT8). A later experimental configuration combined this character-swap adapter with a separate 768p Turbo 8-step LoRA, res_multistep / simple, and native Sol attention. Those are evaluation choices, not the training base or universal compatibility claims. The local hybrid is not bundled. The text-to-video-only FastH3 experiment did not demonstrate a useful replacement workflow.
| Setting | Recorded value |
|---|---|
| Updates / saves | 1,000 / every 250 updates |
| Hardware | RunPod RTX PRO 4500 Blackwell, 32 GB |
| LoRA rank / alpha | 16 / 16, excluding adaln_proj |
| Optimizer / learning rate | AdamW8bit / 5e-5 |
| Batch / accumulation | 1 / 1 |
| Precision | BF16, convrot8 transformer, NVFP4 text encoder |
| Edit target resolution | Area budget 1024; 1344×768 buckets |
| Video regularization resolution | Reduced area budget 384 |
| Regularization duration | 73 frames at 24 fps, approximately 3.04 seconds |
| Memory measures | Gradient checkpointing, layer offload, cached latents/text, chunked MLP |
| Sampling during training | Disabled |
The dataset has 94 synthetic image-edit triplets and 40 unchanged video/audio examples. Optimization used 76 edits and 32 regularization clips; 18 edits and 8 clips were held out. The edit targets are single still images, with five-frame static source-video controls. This is not training on long moving character-swap targets. AI Toolkit interleaves regularization, so file-count ratios are not update ratios. The data include cross-style swaps and varied character sheets, but only one-character replacement targets.
See the full description on the original page
暂无作品
