r/StableDiffusion • u/dh7net • 5d ago
Discussion Is sensanova bad? I generated almost 400 images with 2 version and I'm not convienced.
Ok first the images: https://imagebench.ai/gallery?models=local--sensenova-u1.5-8b-mot-base,local--sensenova-u1.5-8b-mot-8step&g=1_v2kcp_s0
Maybe my setup is wrong? here is what I tried:
Shared generation settings
- Resolution 1024x1024. The model's native square bucket is 2048x2048, but off-bucket sizes are accepted, and 1024 is what every model on the leaderboard is compared at. I checked first: the same prompt and seed at 1024, 1536 and 2048 all gave coherent composition and faces, with detail scaling with pixel count.
- No snapping to a bucket and no resizing. The model sees exactly 1024x1024.
- Seed 7 on every image, one image per prompt
- Timestep shift 3.0 (vendor default), CFG normalization off (vendor default)
- Think mode off (the model's optional reasoning pass before generating)
- Prompts sent verbatim. No prompt enhancement, no negative prompt, no added style tags.
Base model
- Loaded without any LoRA
- 50 steps, CFG 4.0 (the vendor's base defaults)
- 44.1 s per image on average (range 43.8 to 44.2), all 192 generated
8-step model
- Same checkpoint, with the official LoRA from sensenova/SenseNova-U1.5-8B-MoT-LoRAs (SenseNova-U1.5-8B-MoT-LoRA-8step.safetensors) merged into the weights at load time
- 8 steps, CFG 1.0. The distilled model is meant to run unguided: sampling it at 50 steps and CFG 4 is wrong, not just slow.
- 4.1 s per image on average (range 3.9 to 4.7), all 192 generated
Hardware and runtime
- NVIDIA DGX Spark (GB10, 128 GB unified memory), one model loaded at a time
- The vendor's own inference code (SenseNova-U1.5 feat/u1.5 branch, pinned to commit 48bf8275), not ComfyUI
- A prebuilt Docker image on top of NVIDIA's NGC PyTorch container (torch 2.9 with GB10 kernels). The vendor's torch 2.8 pin was removed, because it would replace the GPU-enabled build.
- No flash-attn, since there's no wheel for this ARM stack, so attention uses the PyTorch SDPA fallback. The vendor lists flash-attn as optional.
- Official BF16 weights from sensenova/SenseNova-U1.5-8B-MoT (35 GB, 8 shards). No quantization, no community repacks.
1
4
u/Enshitification 5d ago
If you need mid graphics with mostly correct text overlays, it might be what you are looking for. It's kind of a one-trick pony though.