
This is ICEdit, an image editing framework technology based on the Black Forest Flux-Fill inpainting model and ICEdit-MoE-LoRA—an efficient instruction-based image editing framework.
Compared to previous methods, ICEdit requires only 1% of trainable parameters (200M) and 0.1% of training data (50k) to demonstrate strong generalization capabilities across a wide range of editing tasks.
Compared to commercial models like Gemini and GPT-4o, ICEdit is more open-source, lower cost, faster (taking approximately 9 seconds to process an image), and delivers powerful performance.
ICEdit leverages the enhanced generative capabilities and native in-context awareness of large-scale Diffusion Transformers (DiT) to tackle this challenge. The ICEdit solution introduces three key contributions: (1) An in-context editing framework that uses in-context prompts to achieve zero-shot instruction compliance while preventing structural changes; (2) A LoRA-MoE hybrid tuning strategy that enhances flexibility through efficient adaptation and dynamic expert routing, eliminating the need for extensive retraining; (3) An early-filter inference-time scaling method that utilizes Vision-Language Models (VLMs) to select superior initial noise at an early stage, thereby improving editing quality.
• Project Homepage: https://river-zhang.github.io/ICEdit-gh-pages/ • GitHub: https://github.com/River-Zhang/ICEdit • Hugging Face: https://huggingface.co/sanaka87/ICEdit-MoE-LoRA

No creation yet
