Introducing Illustrious Z: A Shift Toward DiT-Based Scene Understanding

userImage
Onoma TechApr 15, 2026

Introduction

Illustrious Z is a newly developed model designed to extend complex prompt understanding and scene-level illustration generation to a new level.


The Illustrious XL series has steadily improved image quality, prompt controllability, and real-world usability on top of the SDXL architecture. In particular, versions v3.5 and v3.6 showed clear advancements in handling complex prompts, multi-character compositions, tag control, and pose expression. These improvements made the models significantly more practical for production workflows.


However, as prompts became longer and more descriptive—and as users increasingly required multi-character interactions and structured scene composition—the limitations of the existing architecture became more apparent. Handling relationships between characters, maintaining spatial consistency, and interpreting narrative-style prompts at scale required a fundamentally different approach.


Illustrious-Z was developed in response to this shift. Rather than extending the SDXL lineage, it is built on a new S3-DiT-based architecture, enabling deeper natural language understanding and more accurate interpretation of structured visual relationships. This transition allows the model to move beyond element-level generation and toward scene-level reasoning and composition.



What's New in Illustrious Z

Illustrious-Z introduces several key advancements:

  1. Fine-tuned on Z-image-turbo
  2. Training data updated through January 2026
  3. Enhanced natural language prompt understanding
  4. Improved multi-character composition and spatial reasoning
  5. Stronger text rendering accuracy

While previous models focused on improving aesthetics, control, and prompt fidelity step-by-step, Illustrious-Z shifts toward a more holistic goal:

Understanding and constructing the entire scene—not just individual elements.



Image Style

Illustrious-Z supports both tag-based prompting and natural language prompting, but the output characteristics differ noticeably depending on the approach.

Natural Language Prompting

Natural language prompts excel at constructing full scenes with contextual depth.

  1. Richer detail in faces, hair, and clothing
  2. Stronger lighting contrast and depth perception
  3. Better integration of background, composition, and atmosphere

When using descriptive prompts, the model interprets not only character attributes but also mood, spatial relationships, and cinematic composition, resulting in more visually complete outputs.


Tag-based Prompting

Tag-based prompts are optimized for structured and direct control.

  1. Faster and more intuitive input
  2. Precise control over character attributes
  3. Simpler and. more predictable outputs

However, compared to natural language prompting:

  1. Lines appear softer and less defined
  2. Details (e.g., fabric folds, facial features) are simplified
  3. Shading is lighter, resulting in a flatter visual style

In short:

Natural language → richer, scene-driven outputs
Tags → faster, controlled, and simplified outputs



Key Improvements

Illustrious-Z shows clear improvements over previous models across multiple dimensions, including complex prompt understanding, multi-character composition, text rendering, and character interaction.


To better illustrate these improvements, we first compare outputs across different models under the same prompt conditions.

Prompt: A girl, wearing a blue blouse with no exposed breasts and no cleavage. She is holding an ice-cream cup with the text "onoma" printed on it.


As you can see from the example image above, Illustrious-Z demonstrates stronger prompt fidelity, cleaner structure, and more accurate text rendering.


1. Complex Prompt Understanding

Illustrious-Z improves how long, structured prompts are interpreted by maintaining consistency across multiple attributes and instructions. Rather than selectively reflecting parts of a prompt, the model preserves the overall context and translates it into a more coherent visual result.


The following examples highlight how Illustrious-Z maintains consistency across complex poses and detailed instructions.

Prompt: She is performing a dancer pose:balancing on one leg, with her other leg lifted and bent backward, one hand reaching back to hold her foot behind her, while her other arm is extended straight forward for balance. Her body leans slightly forward with a graceful arch, emphasizing balance and flexibility.


In previous models, complex pose descriptions like this often resulted in missing elements or unstable body structure. Illustrious-Z, however, is able to reflect multiple pose constraints simultaneously while maintaining natural proportions and balance.


This difference becomes more apparent when multiple visual attributes are described simultaneously.

Prompt: She is wearing a dark navy off-shoulder frill dress with a layered skirt structure. The chest area features semi-transparent mesh fabric, along with detached puff sleeves, lace decorations, and ribbon details…


In previous models, complex pose descriptions like this often resulted in missing elements or unstable body structure. Illustrious-Z, however, is able to reflect multiple pose constraints simultaneously while maintaining natural proportions and balance.


Illustrious-Z integrates these layered visual attributes more coherently into a single composition. In contrast, previous models tend to overemphasize certain elements or omit finer details. Illustrious-Z maintains the relationships between these attributes, resulting in a more complete and balanced output.


While v3.5 and v3.6 significantly improved prompt fidelity, longer prompts with multiple simultaneous attributes could still lead to partial omissions or uneven emphasis. Illustrious-Z addresses this by not only reflecting more information, but by organizing it into a structurally consistent scene.


2. Multi-Character Composition

Multi-character composition and spatial relationship understanding are also key areas of improvement in Illustrious-Z. Multi-character scenes further demonstrate how positional relationships are interpreted.

Prompt: An illustration of narita top road from umamusume and lumine from genshin impact. On the left is Narita Top Road: a horse girl with horse ears and horse tail, purple eyes, short blonde hair, and a soft blush, with an open-mouth smile showing upper teeth. Body details include medium breasts. She is wearing a purple dress and gloves, with bare shoulders, collarbone, bow accents, and light confetti detail. On the right is Lumine: yellow eyes, short slightly wavy blonde hair with hair flower, hair ornament, and jewelry, smiling with a closed mouth. Body details include medium breasts. She is wearing a white sleeveless dress with detached sleeves, scarf, gold trim, and brown thigh-high boots. Both characters are standing side by side, looking at the viewer, in a clean simple white background. Absurdly high resolution.


Previous models made significant progress in handling left-right positioning, character differentiation, and role separation. However, in more complex multi-character prompts, attributes could still become mixed, and spatial relationships could become unstable.


Illustrious-Z improves this by more accurately preserving positional cues, character separation, and distinct attributes. As a result, each character maintains a clear identity, and the overall composition remains stable and structured.


The difference becomes more pronounced as the number of characters increases.

Prompt: 3girls, absurdres, angels_of_delusion, animal_ear_hairband, animal_ears, aria_(zenless_zone_zero), bare_shoulders, bell, black_garter_straps, black_hair, black_thighhighs, blue_eyes, bow, breasts, cat_ear_hairband, cat_ears, commentary_request, dress, fake_animal_ears, garter_straps, green_eyes, green_hair, hair_bow, hair_ornament, hairband, hairclip, halo, heart, heart_arms, heart_arms_duo, heart_hands, highres, interlocked_fingers, maid, maid_headdress, multicolored_hair, multiple_girls, nangong_yu, neck_bell, official_alternate_costume, pink_hair, red_eyes, red_nails, simple_background, streaked_hair, sunna_(afternoon_tea_break)(zenless_zone_zero), sunna(zenless_zone_zero), thighhighs, white_dress, zenless_zone_zero


In scenarios involving three or more characters, previous models often struggled with attribute mixing or loss of structural clarity. Illustrious-Z maintains clearer separation between characters while preserving interaction and visual balance across the scene. This leads to more coherent multi-character outputs, where each subject retains their role and visual identity without compromising the overall composition.


3. Text Rendering

Text rendering is another area where Illustrious-Z shows noticeable improvement. Text rendering improvements can be observed in scenarios where readable text is required within the image.

Prompt: She is holding a cup with the text "IL-Z" printed on it.


In previous models, generating readable text within images was often inconsistent, with distortions or incorrect characters appearing frequently. Illustrious-Z improves the preservation of character shapes, resulting in more legible and stable text outputs.


Additionally, text is more naturally integrated into the overall composition, rather than appearing as a separate or distorted element. This makes it more suitable for use cases involving signs, labels, or decorative text within illustrations. While text rendering still remains an area for further improvement, Illustrious-Z demonstrates clear progress in terms of readability, structure, and placement stability compared to earlier models.



Trade-offs and Current Strengths of ILXL v3.6

Despite the advancements introduced in Illustrious-Z, Illustrious XL v3.6 continues to maintain strong advantages in several areas. The tag-based workflow remains one of the most stable aspects of v3.6. Over time, the model has been extensively optimized for tag-driven prompting, making it highly reliable for users familiar with structured inputs such as Danbooru-style tags. This accumulated optimization results in predictable and consistent outputs.


In addition, character recognition and consistency are still areas where v3.6 can outperform Illustrious-Z in certain cases. Due to longer training cycles and repeated tuning, v3.6 has developed stronger stability when reproducing specific characters or maintaining consistent visual identities across generations.


From an aesthetic perspective, v3.6 also retains a high level of refinement. Its outputs are often more consistent in style, particularly for character-focused illustrations where visual stability is prioritized over scene complexity.


Taken together, the distinction between the two models can be understood as a difference in focus. Illustrious-Z emphasizes interpretation and composition, while v3.6 continues to excel in stability, control, and stylistic consistency.



Closing

Illustrious-Z is an actively evolving model, and further improvements are planned in areas such as character consistency and overall output stability. At present, the most reliable results are achieved at a resolution of 1024×1024.


Building on the foundation established by the Illustrious XL series, Illustrious-Z represents a new stage in development—one that prioritizes deeper understanding of prompts and more structured scene generation. The transition from an SDXL-based architecture to a DiT-based approach enables meaningful progress in how the model interprets language, organizes visual elements, and constructs interactions within a scene.


Moving forward, we will continue to refine Illustrious-Z with the goal of delivering a more stable, expressive, and controllable generation experience.

Read More

Illustrious XL v3.6: Enhanced Capabilities for Cutting-Edge Illustrations

Illustrious XL v3.6: Enhanced Capabilities for Cutting-Edge Illustrations

Introducing Illustrious XL v3.6, our newest model release packed with significant advancements to elevate your illustration generation experience.

userImage
Onoma TechJul 14, 2025

by Onoma AI

Address : 201, 2F, D-dong, 47, Maeheon-ro 8-gil, Seocho-gu, Seoul, 06770, Rep. of KOREA

Business Registration Certificate : 450-86-02454

CEO : Min Song

Contact : illustrious@onomaai.com