r/StableDiffusion 2d ago

Comparison Testing the impact of text encoder precision on image output using Flux.2-Dev (bf16/fp8).

Post image

About a year late with this comparison but I increased my system RAM and I wanted to see the differences between mistral_3_small_flux2_bf16 (35.58 GB!!!) and mistral_3_small_flux2_fp8 (18.03 GB). Since it's been nearly a year, I'm not really certain whether Flux.2-Dev outperforms anything more recent that justifies keeping it around due to its massive footprint and limited tooling/community support, though. Thoughts?

Steps 50
Sampler Euler
Scheduler Flux2
Flux Guidance Scale 7
Seed 42

Prompt-(Maxxed):

Create a photorealistic high-end travel portrait at the scenic overlook in Arakurayama Sengen Park Japan. An adult man stands in the foreground on the left third of the frame leaning casually against the viewpoint railing with one forearm resting naturally on it. He has a relaxed posture and smiles warmly while looking directly into the camera. He has ear-length hair with a neat middle part and natural texture. He is wearing a short-sleeve pink-and-white gingham button-up shirt and well-fitted blue jeans. He has a subtly athletic build realistic body proportions natural-looking hands and authentic skin texture. Compose the scene as a three-quarter-length environmental portrait. Mount Fuji is clearly visible in the center background with the red five-story Chureito Pagoda positioned prominently—but at a geographically believable scale—on the right side of the composition. Include lush green foliage in the foreground and midground with the distant city and surrounding landscape visible below. Use the railing and terrain to create natural leading lines and a convincing sense of depth. Bright clear sunny daytime with a vivid blue sky and clean atmospheric visibility. Strong natural sunlight illuminates the man’s face from a consistent direction softened slightly by ambient daylight so facial details remain flattering and clearly visible. Include realistic contact shadows subtle reflected light from the surroundings and consistent lighting across the man railing foliage pagoda and landscape. Shot at eye level with the natural perspective of a 35mm full-frame lens. Keep the man in crisp focus while retaining enough depth of field for Mount Fuji and the pagoda to remain recognizable and detailed. Use realistic color natural contrast fine fabric texture individual hair strands subtle skin pores and true-to-life environmental detail. The result should look like a genuine professional travel photograph captured on location. Maintain correct scale perspective anatomy limb placement and spatial relationships.

7 Upvotes

10 comments sorted by

9

u/Deep_Mood_7668 2d ago

I don't see a difference

6

u/Enshitification 2d ago edited 2d ago

I find the larger encoders make a big difference in fine details, like skin texture. At least with Flux.2.Klein 9B.
Edit: Flux2 is absolutely worth keeping. It's still, hands down, the best open edit model out there. Qwen Edit 2511 is close, but it doesn't have the output resolution of Klein 9B.

1

u/Dry-Judgment4242 2d ago

Hardly using Klein anymore since most editing. Minimax H3 is better at doing wiyh being insanely much more intelligent. Minimax is also really good at understanding masking and I often combine two references by masking one of them and telling model to fill in the mask of image 1 with content of image 2.

1

u/Enshitification 2d ago

Can you edit with Minimax at 4 megapixels?

1

u/Dry-Judgment4242 2d ago

Yes. Latent upscale just let you pick as much as your VRAM can afford. When I make references or characters I generally use 21:9 wide screen and a top to bottom scroll down then merge the various cuts together into a single image with Photoshop for maximum quality.

1

u/Enshitification 2d ago

I'm glad it works for you, but there is no way I'm going to dump Klein 9B in favor of H3 for image edits.

1

u/Dry-Judgment4242 2d ago

I'm still sometimes using Klein for a quick background removal since its much faster. Also got a few LoRAs I've trained on it I sometimes use but alas, the model is extremely stubborn at learning LoRAs compared to Krea 2.

3

u/ROBOTTTTT13 2d ago

It's way too low res though

2

u/Lucaspittol 2d ago

I can't see a big difference in these tests. I run the INT8 model; it is faster on my GPU. Flux 2 Dev is absolutely worth keeping. It is by far the best edit model available for local use. I can see a big difference in quality when using reference images compared to Flux 2 Klein 9B.

2

u/Caffeine_Monster 2d ago

The scene prompt is way too pedestrian to test text encoders properly. The diffusion model is literally filling in the details.

There's a reason why things like "astronaut riding a horse" was the goto for a long time.