GPT Image 2.5 Flare is OpenAI's fast, balanced image tier supporting text-to-image generation and multi-reference editing, with five quality tiers, transparent background, and accurate in-image text. Billed per output image by quality and image size, plus per additional input image.
GPT Image 2.5 Flare (channel edition) is OpenAI's fast, balanced image tier for text-to-image and multi-reference editing via a low-price channel, with flat per-call pricing and flexible resolutions and aspect ratios. Priced lower than the official edition.
GPT Image 2.5 Sunburst (channel edition) is OpenAI's image generation tier for text-to-image and multi-reference editing via a low-price channel, with flat per-call pricing and flexible resolutions and aspect ratios. Priced lower than the official edition.
GPT Image 2.5 Sunburst is OpenAI's precision-focused image tier supporting text-to-image generation and multi-reference editing, with five quality tiers, transparent background, and accurate in-image text. Billed per output image by quality and image size, plus per additional input image.
Qwen Image 3.0 is Alibaba's image generation and editing model with native multilingual prompting and best-in-class text rendering, capable of multi-line layouts and paragraph-level typography in images. It supports text-to-image and editing with 1-3 reference images, billed per output image.
Seedream 5.0 Pro is ByteDance's flagship image generation and editing model, supporting text-to-image, single-reference image editing, and layer decomposition with output up to 2K resolution. Output is billed per image (or per generated layer) by pixel tier, suited to posters, brand visuals, and professional design workflows.
Grok Imagine Image 2.0 is xAI's image generation and editing model, creating detailed images from text prompts or reference images, with quality tiers (low/medium) at 1K/2K resolution. Billed per image.
Qwen Image 3.0 Pro is the premium edition of Alibaba's image model, tuned for higher-quality editing and transformation with up to 4.5k-token prompts, 12 languages, and 20+ fonts for native text rendering in complex multi-line layouts. It supports text-to-image and editing with 1-3 reference images; output is billed per image by size tier, plus $0.003 per input image in edit mode.
ByteDance Seedream 5.0 is a new-generation text-to-image model focused on knowledge reasoning and intelligent editing, and the first Seedream release with real-time web retrieval. It offers deeper world-knowledge understanding, improved semantic comprehension and spatial-logical reasoning for complex prompts, with native 4K output for infographics, knowledge-reasoning illustrations, and e-commerce design. Edit mode supports multi-turn editing via natural language without regeneration.
Google Nano Banana 2 is a new-generation image generation model. Inheriting the advanced intelligence of the Pro edition while delivering faster generation, it supports output from 512px up to 4K. It achieves significant gains in text rendering, world-knowledge integration, character consistency, and instruction adherence, and can directly produce finished-grade layouts for infographics, multilingual menus, and dynamic illustrations.
Google Nano Banana 2 is a high-performance multimodal image-generation base model. It supports direct text-to-image generation of high-fidelity images, or precise controllable generation with reference images, with outstanding visual expression and detail fidelity. The channel edition is more economical than the official, suitable for social-media assets, marketing collateral, and e-commerce product images. Its edit mode performs stable editing at high speed, enhanced by real-time search.
Google Nano Banana Pro is a professional image generation model supporting up to 4K output, leveraging advanced physical light-shadow rendering and text understanding to generate industrial-grade visuals from precise instructions. The channel edition is priced lower than the official, suitable for daily creative scenarios. Its edit mode supports up to 14 reference images with control over lens angle, focal length, color grading, and scene lighting.
Google Nano Banana Pro is a professional image generation model supporting text-to-image generation and multi-image reference fusion, with ultra-clear rendering and up to 4K output. It features advanced semantic understanding and structured reasoning, and excels at complex instruction understanding, detail presentation, and precise multi-language text rendering. Suitable for brand advertising, product posters, and other visual content with strict text-presentation requirements.
Alibaba Wan 2.7 Image is the standard-tier text-to-image model with an upgraded core rendering architecture and precise text-semantic understanding. It offers coordinated composition, soft, layered color rendering, and compatibility with realistic, anime, and minimalist styles, with custom size control, built-in thinking mode, and broad aspect-ratio support. Edit mode preserves original composition and subject identity, supporting style transfer, quality enhancement, and fine-tuning.
OpenAI GPT Image 2 turns natural-language prompts into high-quality images. It supports text-to-image generation and multi-image reference editing, excelling at complex instruction understanding, detailed rendering, and precise multi-language text layout. The official edition adds a reasoning-verification mode and can output up to 8 stylistically consistent images per prompt.
OpenAI GPT Image 2 is an image-generation interface built on a frontier large-model architecture. It supports deep text-semantic understanding, accurately captures complex prompt details, and transforms them into high-quality, well-composed visuals. The channel edition is priced lower than the official, suitable for daily visual and marketing collateral scenarios. Its edit mode maintains character appearance, lighting, and subject identity without manual masking for pixel-level edits.
Alibaba Wan 2.7 Pro Image is the professional tier of the Wan 2.7 text-to-image model, supporting up to 4K (4096x4096) output with built-in thinking mode and custom size control. It delivers higher-fidelity compositions with precise proportions, rich scene detail, and high-end light-shadow, compatible with realistic, Chinese-style, and anime styles for commercial visual design and concept art. Edit mode preserves composition via intelligent style remodeling.
Seedance 2.5 is ByteDance's latest-generation video model supporting up to 4K output with native audio, cinematic camera control, and multi-character consistency. It covers generation from text, first-last frames, or multimodal references as well as prompt-driven video editing and extension, billed per second of output video by resolution tier.
MiniMax H3 is MiniMax's multimodal video model that generates video from text prompts, first and last frames, or reference images, videos, and audios, at up to 4K resolution with strong prompt understanding and cinematic quality, billed per second of output video by resolution tier.
MiniMax H3 (channel edition) is the economy edition of the MiniMax H3 video model, generating video from text prompts, first and last frames, or reference images, videos, and audios at 2K or 768P resolution. Billed per second of output video by resolution tier, suited for high-frequency, cost-sensitive production scenarios.
ByteDance Seedance 2.0 is a new-generation native multimodal video-generation model supporting joint input of text, images, video clips, and audio. With powerful audio-visual sync and motion perception, it locks onto character visuals, camera movement, and audio rhythm via @-tags, outputting film-grade high-fidelity video. The channel edition supports text-to-video, first-last-frame, and reference-to-video generation with up to 9 images / 3 videos / 3 audios as reference.
Happy Horse 1.1 is Alibaba's video generation model (v1.1), generating cinematic 480p/720p/1080p video from text, reference images, or a first-frame image, with smooth camera movement and expressive motion. Billed per second of output video by resolution tier.
Wan 3.0 is Alibaba's video generation model producing high-quality video from text prompts, reference images or videos, or first and last frames, at 480P/720P/1080P with flexible duration and aspect ratio options. Billed per second of output video by resolution tier; uploaded reference videos are also billed per input second.
ByteDance Seedance 2.0 Fast is the official accelerated text-to-video variant of the Seedance 2.0 multimodal audio-video joint generation model on Seed's unified architecture. It supports mixed input of text, image, audio, and video, delivers strong physical realism and fine instruction control, and generates high-quality multi-shot audio-video content with native binaural audio at low latency for commercial teams needing stable service and standard quality.
ByteDance Seedance 2.0 is ByteDance's flagship text-to-video model, deeply tuned for full-scenario general generation quality. It generates high-fluency, physically consistent video from text or images, with native binaural audio, excellent multi-shot continuity and detail rendering. It autonomously plans storyboards and maintains strong subject consistency across multi-character, multi-plot narratives, making it a key tool for film studios and ad teams delivering high-quality finished clips.
ByteDance Seedance 2.0 Fast is a film-grade video model optimized for faster, lower-cost generation on Seed's unified multimodal architecture. With a powerful physics engine and aesthetic understanding, it recreates shot design, motion changes, and camera rhythm from text, images, or audio, suiting advertising, creative shorts, and film pre-visualization. The channel edition compresses multimodal inference overhead for batch production while keeping brand-visual alignment and audio-video sync.
Alibaba Happy Horse 1.0 generates cinematic 720p/1080p videos from text or a reference image, with smooth camera movement, expressive motion, and strong prompt fidelity. The official edition is tuned for ad spots, short-drama segments, and scenarios demanding higher visual quality, with up to 15-second multi-shot narration and native audio-visual sync.
Built on the Omni One architecture, Kuaishou Kling O3 Pro supports text-to-video generation and precise modification of existing videos via natural-language instructions — covering scene atmosphere, element substitution, and lighting style — while preserving original motion trajectories and character continuity without manual masking. The channel edition is priced lower than the official version, suiting creative short videos, ad-script visualization, and film-concept dynamic previews.
Alibaba Wan 2.7 Video is an advanced video generation model supporting five modes: text-to-video, image-to-video, reference-to-video, video editing, and video extension. It relies on high-precision semantic understanding to produce high-definition dynamic footage with natural character motion, smooth camera movement, and rich detail, handling complex scenes and multi-character narratives for commercial, film-short, and artistic-video workflows.
GPT Image 2.5 Sunburst Text to Image, OpenAI's precision-focused GPT Image 2.5 tier for everyday text-to-image generation. Five quality tiers (low to max), strong prompt fidelity and accurate in-image text. Billed per output image by quality and image size.
GPT Image 2.5 Sunburst Image to Image, OpenAI's precision-focused GPT Image 2.5 tier for natural-language image editing. Accepts up to 16 reference images, supports mask for scoped edits, transparent background, five quality tiers. Billed per output image by quality and image size, plus per additional input image.
GPT Image 2.5 Sunburst Text to Image, OpenAI's GPT Image 2.5 Sunburst tier for text-to-image generation, served via a low-price with flat per-call pricing. Exposes resolution and aspect ratios. The edition is priced lower than the official version.
GPT Image 2.5 Sunburst Image to Image, OpenAI's GPT Image 2.5 Sunburst tier for natural-language image editing, served via a low-price with flat per-call pricing. Accepts up to 10 reference images. The edition is priced lower than the official version.
GPT Image 2.5 Flare Text to Image, OpenAI's fast, balanced GPT Image 2.5 tier for everyday text-to-image generation. Five quality tiers (low to max), strong prompt fidelity and accurate in-image text. Billed per output image by quality and image size.
GPT Image 2.5 Flare Image to Image, OpenAI's fast, balanced GPT Image 2.5 tier for natural-language image editing. Accepts up to 16 reference images, supports mask for scoped edits, transparent background, five quality tiers. Billed per output image by quality and image size, plus per additional input image.
GPT Image 2.5 Flare Text to Image, OpenAI's GPT Image 2.5 Flare tier for text-to-image generation, served via a low-price with flat per-call pricing. Exposes resolution and aspect ratios. The edition is priced lower than the official version.
GPT Image 2.5 Flare Image to Image, OpenAI's GPT Image 2.5 Flare tier for natural-language image editing, served via a low-price with flat per-call pricing. Accepts up to 10 reference images. The edition is priced lower than the official version.
Alibaba Happy Horse 1.1 Image to Video is Alibaba's Image to video model (v1.1), generating cinematic video based on the first frame at 480p/720p/1080p with smooth camera movement and expressive motion. Billed per second of output video by resolution tier.
Alibaba Wan 3.0 FLF to Video is Alibaba's FLF to video model generating high-quality video from first and last frames at 480p/720p/1080p with flexible duration and aspect ratio options. Billed per second of output video by resolution tier.
ByteDance Seedream 5.0 Pro Image to Image Layer Decomposition, ByteDance Seedream 5.0 Pro Layer Decomposition is ByteDance's image layer separation model that decomposes a generated image into multiple layers for professional design and compositing workflows. Billed per generated layer by pixel tier.
ByteDance Seedream 5.0 Pro Image to Image Edit is ByteDance's premium image editing model that transforms existing images via natural language instructions, supporting multi-reference image input with consistent subject identity. The first input image is free; subsequent images are $0.003 each. Output is billed per image by pixel tier.
ByteDance Seedance 2.5 Video Extend is ByteDance's video extend model that extends an existing video with a prompt-driven cinematic continuation generated seamlessly from its last frame. Billing is based on the combined duration of the reference video and the new segment by resolution tier; the reference duration is clamped to 2-30 seconds and the new segment duration is set by the duration parameter.
ByteDance Seedance 2.5 Video Edit is ByteDance's video editing model that applies prompt-driven edits to an existing video while maintaining motion and temporal consistency. Billing is based on the combined duration of input and output video by resolution tier; the input video is handled within the 4-30s range before pricing.
ByteDance Seedance 2.5 Text to Video is ByteDance's latest video generation model supporting up to 4K resolution with native audio, cinematic camera control, and multi-character consistency. Billed per second of output video by resolution tier.

BizyAirPlus
Run ComfyUI on Cloud GPUS
No local GPU needed. Install the plugin and run any workflow in the cloud.
Get the PluginAfter installation, start ComfyUI and choose BizyAirPlus mode from the run button to begin.
