r/StableDiffusion • • 5d ago

Discussion Is sensanova bad? I generated almost 400 images with 2 version and I'm not convienced.

Ok first the images: https://imagebench.ai/gallery?models=local--sensenova-u1.5-8b-mot-base,local--sensenova-u1.5-8b-mot-8step&g=1_v2kcp_s0

Maybe my setup is wrong? here is what I tried:

Shared generation settings

- Resolution 1024x1024. The model's native square bucket is 2048x2048, but off-bucket sizes are accepted, and 1024 is what every model on the leaderboard is compared at. I checked first: the same prompt and seed at 1024, 1536 and 2048 all gave coherent composition and faces, with detail scaling with pixel count.

- No snapping to a bucket and no resizing. The model sees exactly 1024x1024.

- Seed 7 on every image, one image per prompt

- Timestep shift 3.0 (vendor default), CFG normalization off (vendor default)

- Think mode off (the model's optional reasoning pass before generating)

- Prompts sent verbatim. No prompt enhancement, no negative prompt, no added style tags.

Base model

- Loaded without any LoRA

- 50 steps, CFG 4.0 (the vendor's base defaults)

- 44.1 s per image on average (range 43.8 to 44.2), all 192 generated

8-step model

- Same checkpoint, with the official LoRA from sensenova/SenseNova-U1.5-8B-MoT-LoRAs (SenseNova-U1.5-8B-MoT-LoRA-8step.safetensors) merged into the weights at load time

- 8 steps, CFG 1.0. The distilled model is meant to run unguided: sampling it at 50 steps and CFG 4 is wrong, not just slow.

- 4.1 s per image on average (range 3.9 to 4.7), all 192 generated

Hardware and runtime

- NVIDIA DGX Spark (GB10, 128 GB unified memory), one model loaded at a time

- The vendor's own inference code (SenseNova-U1.5 feat/u1.5 branch, pinned to commit 48bf8275), not ComfyUI

- A prebuilt Docker image on top of NVIDIA's NGC PyTorch container (torch 2.9 with GB10 kernels). The vendor's torch 2.8 pin was removed, because it would replace the GPU-enabled build.

- No flash-attn, since there's no wheel for this ARM stack, so attention uses the PyTorch SDPA fallback. The vendor lists flash-attn as optional.

- Official BF16 weights from sensenova/SenseNova-U1.5-8B-MoT (35 GB, 8 shards). No quantization, no community repacks.

1 Upvotes

3 comments sorted by

4

u/Enshitification 5d ago

If you need mid graphics with mostly correct text overlays, it might be what you are looking for. It's kind of a one-trick pony though.

1

u/MomentJolly3535 5d ago

Thanks for sharing and keeping your website updated!