h3_character_swap_pro4500_1000.safetensors

更新:2026-09-26 13:38|发布时间:2026-09-26 13:38
0
0
0
h3_character_swap_pro4500_1000.safetensors. --- language: - en license: other license name: minimax-h3-community-license-agreement license link: LICENSE base model: Comfy-Org/MiniM…
头像
Gugugaga
类型
LoRA
基础模型
MiniMax H3
发布时间
2026-09-26 13:38
文件签名
664d82ab9689fdcc4f601f6008b96021b7bc8cd938e7f843e6500c860805ff3d
h3_character_swap_pro4500_1000.safetensors
147.92 MB

language:

  • en license: other license_name: minimax-h3-community-license-agreement license_link: LICENSE base_model: Comfy-Org/MiniMax-H3 base_model_relation: adapter pipeline_tag: video-to-video library_name: diffusion-single-file datasets:
  • akatz-ai/H3-Character-Swap-v1 tags:
  • minimax-h3
  • lora
  • character-swap
  • video-editing
  • ref2va
  • comfyui
  • ai-toolkit

MiniMax H3 Character Swap LoRA v1

An experimental character-replacement LoRA trained by Akatz Labs for 1,000 updates. It aims to replace a selected character using an image reference while preserving the source scene. In our local comparisons it often preserved the background more closely than the base model, but motion timing, facial expressions, and hard cuts remain unreliable.

Download: final 1,000-step LoRA. Only the final 1,000-step checkpoint is published. Start with the final checkpoint at strength 1.0. This is an adapter, not a standalone model or a Turbo distillation LoRA.

Dataset: H3 Character Swap v1.

Use

  1. Use an H3 Ref2VA-capable runtime and obtain the base model and VAEs separately.
  2. Place the final .safetensors file in ComfyUI/models/loras/ (or your configured shared LoRA directory), and apply it to the H3 model at strength 1.0 using a compatible model-only LoRA loader.
  3. Supply the source video as <Video 1> and replacement image or character sheet as <Picture 1>.
  4. Specify the target person in the prompt. No additional trigger word was trained.

Example:

Replace only the man in the purple shirt in <Video 1> with the character in <Picture 1>. Keep the replacement character's identity, outfit, and art style from <Picture 1>. Preserve the source video's camera, background, lighting, objects, and all other people. Match the target person's position, scale, pose, and movement. Do not show the reference sheet or its background.

The training captions were shorter, for example Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>. Preservation instructions helped some local evaluations, but stronger expression instructions sometimes suppressed the swap entirely. Prompt wording is not a guarantee of strict source alignment.

Short, continuous shots of roughly 4–5 seconds were more promising than our full 14-second tests. A precise maximum duration has not been established. Use 24 fps and your runtime's supported H3 frame grid. The character-swap LoRA does not require a Turbo LoRA, Spectrum, or Sol attention.

Base model and tested configurations

Training used minimax_h3_ref2va_pruned_int8_convrot.safetensors from Comfy-Org/MiniMax-H3, plus the frozen Ostris Ref2VA training assistant. The assistant and base weights are not merged into this adapter and are not distributed here. Model revisions and hashes are in training/base-model-files.json.

We evaluated on the Ref2VA base and a local FL2VA/Ref2VA hybrid (blocks 25–49, INT8). A later experimental configuration combined this character-swap adapter with a separate 768p Turbo 8-step LoRA, res_multistep / simple, and native Sol attention. Those are evaluation choices, not the training base or universal compatibility claims. The local hybrid is not bundled. The text-to-video-only FastH3 experiment did not demonstrate a useful replacement workflow.

Training record

SettingRecorded value
Updates / saves1,000 / every 250 updates
HardwareRunPod RTX PRO 4500 Blackwell, 32 GB
LoRA rank / alpha16 / 16, excluding adaln_proj
Optimizer / learning rateAdamW8bit / 5e-5
Batch / accumulation1 / 1
PrecisionBF16, convrot8 transformer, NVFP4 text encoder
Edit target resolutionArea budget 1024; 1344×768 buckets
Video regularization resolutionReduced area budget 384
Regularization duration73 frames at 24 fps, approximately 3.04 seconds
Memory measuresGradient checkpointing, layer offload, cached latents/text, chunked MLP
Sampling during trainingDisabled

The dataset has 94 synthetic image-edit triplets and 40 unchanged video/audio examples. Optimization used 76 edits and 32 regularization clips; 18 edits and 8 clips were held out. The edit targets are single still images, with five-frame static source-video controls. This is not training on long moving character-swap targets. AI Toolkit interleaves regularization, so file-count ratios are not update ratios. The data include cross-style swaps and varied character sheets, but only one-character replacement targets.

See the full description on the original page

作品

暂无作品