

Z-image-Image Refinement & Upscaling-Ultimate Aesthetics
Base Model
🚀 Z-Image-Turbo—a refined version of Z-Image, matching or exceeding leading competitors in just 8 NFEs (Number of Function Evaluations). It delivers ⚡️sub-second inference latency⚡️ on enterprise-grade H800 GPUs and fits comfortably on 16GB VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and strict instruction following.
📸 Realistic Image Quality: Z-Image-Turbo delivers powerful photorealistic image generation while maintaining exceptional aesthetic quality.
📖 Accurate Bilingual Text Rendering: Z-Image-Turbo excels at accurately rendering complex Chinese and English text.

💡 Prompt Enhancement & Reasoning: The prompt enhancer empowers the model with reasoning capabilities, enabling it to go beyond surface descriptions and draw upon underlying world knowledge.

🧠 Creative Image Editing: Z-Image-Edit demonstrates a deep understanding of bilingual editing instructions, achieving imaginative and flexible image transformations.

We adopt a Scalable Single-Stream DiT (S3-DiT) architecture. In this setup, text, visual semantic tokens, and image VAE tokens are concatenated at the sequence level as a unified input stream, maximizing parameter efficiency compared to dual-stream approaches.

According to Elo-based human preference evaluations (AI Arena), Z-Image-Turbo demonstrates high competitiveness against other leading models while achieving state-of-the-art results among open-source models.

Reposted from: Z-image-Ultimate Aesthetics Portrait Photography
No creation yet
