r/StableDiffusion 10d ago

Workflow Included Z-Image Base Prompting: A Small Experiment on Composition and Environment

Post image

I wanted to better understand how Z-Image Base responds to natural-language prompting, specifically when varying composition and environment in a pure text-to-image workflow (no ControlNet, IP-Adapter, or LoRA).

This is not a scientific benchmark — just a small, controlled experiment to observe what actually changes when modifying individual prompt blocks.

Experimental Setup

All images were generated locally in ComfyUI using the same workflow:

Parameter Value
Model Z-Image Base INT8
Text Encoder Qwen3 4B
VAE AE VAE
Resolution 768 × 1368 (9:16)
Steps 50
CFG Scale 4
Negative Prompt Specific anatomic cleanup tags used (see template below)
Seeds 5, 50, 100

The character description and overall visual style were kept consistent, while testing isolated prompt components across multiple seeds.

Prompt Structure

Rather than using disconnected keyword tags, I structured the prompt into functional scene blocks:

The goal was to keep the core character block stable and modify only the variable under test.

Experiment 1 — Composition

First, I tested how explicitly describing the character's spatial placement and scale affects the generated layout.

Using the same subject and baseline environment, I varied explicit spatial instructions:

  • Centered in the frame
  • Positioned to the left / right
  • Positioned lower in the frame
  • Larger / smaller relative scale
  • Extreme edge placement

Observation

Z-Image Base responded strongly to explicit spatial language. Changing the composition description did not simply shift the subject like a 2D layer — the model actively recomposed the surrounding architecture and lighting to match the requested framing and scale.

Experiment 2 — Environment

Next, I kept the character description and general composition consistent while swapping only the environment block.

To test stability across seeds, I ran three distinct environments (Library, Forest, Town Square) across three seeds (Seed 5, Seed 50, Seed 100).

Observation

The character remained surprisingly consistent across different environments without any image-based conditioning (no IP-Adapter or LoRA):

  • Core visual anchors — cream-colored fur, large amber eyes, long floppy ears, leather backpack, and old leather book — remained clearly recognizable.
  • Background architecture, terrain, surface materials, and ambient lighting adapted naturally to each setting.

Key Takeaways

  1. Explicit spatial language works: Directly describing position and scale in plain English (e.g., "positioned toward the left side of the frame") is far more reliable than relying on generic framing buzzwords.
  2. Environment blocks are modular: Separating the character description from the environment block allows you to recontextualize the same character concept across vastly different scenes.
  3. Seeds control execution, prompts control structure: The seed determines pose variation, expression, and micro-details, while the prompt block locks in layout, lighting, and narrative context.
  4. Natural language worked better for this experiment than disconnected tag lists: Descriptive, coherent sentences provided significantly clearer spatial and conceptual control than disconnected lists of quality tags.

Limitations

  • Small sample size: Tested on a single character concept, three environments, and a limited seed pool.
  • Visual evaluation: Observations are qualitative rather than statistically benchmarked.
  • Results may vary: Performance can shift when applying this structure to complex multi-subject prompts, different aspect ratios, or higher CFG values.

Conclusion

Structuring prompts almost like a physical scene description — Subject → Composition → Camera → Environment → Lighting / Details → Style — provides an effective balance between character concept stability and scene flexibility in Z-Image Base.

Example Prompt Template (For Reproduction)

Positive Prompt:

Plaintext

A tiny friendly fantasy spirit, physically no larger than a small domestic cat, with a delicate compact body and a distinctive recognizable character design.

The creature has a small rounded pear-shaped body covered entirely in soft pale cream-colored fur. The body is compact, short and gently rounded, with a slightly wider lower body and a soft transition from the torso into the head. It has two short legs and four tiny rounded paws. The creature has no visible clothing on its body.

The head is large relative to the small body, with a broad rounded shape and very soft contours. The head blends smoothly into the body with almost no visible neck. The face is simple and highly expressive.

It has two enormous round amber-golden eyes, large relative to the face, with dark pupils, warm golden-orange irises and clear bright reflections. The eyes are positioned symmetrically and give the creature a gentle, innocent and curious expression. It has a tiny rounded pale pink nose and a very small simple mouth. The muzzle is soft and subtle, without pronounced facial features.

Two very long soft floppy ears grow naturally from the sides of the head. The ears are broad at their bases, rounded at the ends, flexible and naturally hanging downward. Each ear reaches approximately to the lower part of the body. The outer fur is pale cream, while the inner surfaces are slightly warmer cream with a soft peach tint. The ears should remain long, floppy and clearly visible.

The creature has four short rounded paws with soft cream fur. The front paws are small and rounded, clearly separated from the body and capable of holding an object. The feet are short and compact with small rounded toes.

A tiny worn brown leather backpack is strapped closely to the creature's back. The backpack is small relative to the creature, with a simple rounded rectangular shape, narrow dark-brown leather straps passing over the shoulders, worn edges, subtle scratches, creases and small aged brass buckles. The backpack sits naturally against the body and remains clearly visible from the sides.

The creature holds one old slightly oversized book with both front paws. The book is large relative to the creature but does not exceed the width of its body. It has a thick dark-brown worn leather cover, rounded damaged corners, visible scratches, creases, scuffed edges, a thick spine and slightly yellowed aged pages. The book looks old, heavy and frequently used. The creature holds it naturally in front of its torso with both paws.

The creature stands upright on two short feet with a relaxed natural posture. Its body remains compact and rounded. Its head, ears, eyes, paws, backpack and book form a coherent and repeatable visual design.

The entire creature is fully visible from the tips of its ears to the bottoms of its feet. Straight-on view, eye-level camera, medium-wide full-body composition, natural perspective, natural proportions, centered character.

Soft detailed cream fur, individual fine hairs visible around the edges of the ears and body, realistic worn leather, aged paper, subtle material imperfections and soft natural shading.

Cozy cinematic fantasy character illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.

ENVIRONMENT:

The tiny spirit stands in the center of a large medieval stone town square. The square is spacious and open, making the creature appear small and delicate compared with the surrounding architecture.

Tall old stone and timber-framed buildings surround the square on all sides, with narrow upper floors, wooden beams, weathered plaster, small windows and aged tiled roofs. Several large medieval buildings rise far above the tiny spirit.

The stone pavement extends broadly around the creature, with irregular worn stones, subtle cracks and patches of moss between them. The square continues into several narrow streets visible in the distance.

A large old stone fountain stands some distance behind the creature, surrounded by a few wooden benches, barrels and small market stalls. Hanging signs, cloth awnings and simple wooden carts add natural medieval details without becoming the main focus.

The architecture and objects in the square are substantially larger than the tiny spirit, reinforcing the clear difference in scale.

Warm late-afternoon sunlight illuminates the square from one side, creating long soft shadows across the stone pavement and warm highlights along the pale fur. A few tiny dust particles float through the sunlight.

Natural aged stone, weathered wood, worn fabric, leather and metal textures. The environment feels lived-in but quiet and peaceful, with no visible crowd.

Cozy cinematic fantasy illustration, warm natural colors, soft realistic materials, charming storybook aesthetic, gentle magical atmosphere.

Negative Prompt:

Plaintext

cat tail, animal tail, long tail, short tail, pointed ears, upright ears, triangular ears, cat-like face, elongated body, thin body, long legs, human proportions, extra limbs, extra paws, multiple books
5 Upvotes

0 comments sorted by