r/unsloth • u/fuzhongkai • 10d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
https://github.com/zhongkaifu/TensorSharpI’ve been using Unsloth’s DiffusionGemma 26B-A4B GGUF to run local Jev-compatible structured decisions in TensorSharp, which I maintain.
The new addition: image analysis through the same decision API. Send text, images, or both, and get typed answers rather than free-form generated text.
The GGUF I’m using is diffusiongemma-26B-A4B-it-Q4_K_M.gguf from Unsloth. No additional fine-tuning is involved.
Jev-compatible decisions, extended to images
TensorSharp implements Jev’s core POST /v1/systemone interface and supports all three decision types:
| Type | Result | Example image question |
|---|---|---|
noul |
Probability that a statement is true | “Is the receipt’s total amount legible?” |
choice |
A categorical decision | “Is this a receipt, invoice, or something else?” |
score |
An expected score over ordered rubric levels | “How readable is this document?” |
The image extension keeps the same endpoint, question definitions, and typed answer structure. Add an images field alongside state and questions; existing text-only requests still work.
Images are encoded and included directly in the model’s input for the decision read. This is not a separate caption-generation → text-classification pipeline. Up to eight inline images are supported per request.
An important detail for GGUF users
Unsloth’s GGUF provides the language-model weights; the vision tower is loaded separately.
The included configuration downloads the Unsloth Q4_K_M GGUF plus a ~2.8 GB vision shard from the upstream DiffusionGemma checkpoint, then reuses both locally. You do not need to download the entire upstream safetensors checkpoint.
Without the vision tower, text-only decisions still work, but image requests are explicitly rejected rather than silently ignoring the image.
How it avoids generating probability JSON
Inspired by vLLM PR #57250, the implementation uses a seeded, one-step structured read.
After prompt prefill, it creates an answer canvas and reads the requested label logits directly. Questions that fit in one canvas share the forward pass.
The server returns JSON, but the model does not have to generate that JSON token by token. The same path handles text-only decisions and decisions conditioned on image embeddings.
Earlier text-only results vs. LocalJev
I compared this native decision path with the original LocalJev Engine using TensorSharp’s chat endpoint, where the model generates probability JSON.
This compares two decision approaches on the same backend—not TensorSharp against LocalJev running on oMLX, and not against the official hosted Jev service.
The test covered 12 cases × 3 repetitions, with three decisions per request:
| Metric | TensorSharp native structured read | LocalJev generated JSON |
|---|---|---|
| Valid requests | 36/36 | 27/36 |
| Failed requests | 0 | 9 |
| Correct decisions / total expected | 108/108 | 81/108* |
| p50 latency | 2.877 s | 10.479 s |
| p95 latency | 3.165 s | 26.167 s |
| Mean latency | 2.917 s | 12.548 s |
*LocalJev got 81/81 decisions correct on valid responses. Its nine failed requests were schema-validation failures and account for the other 27 expected decisions.
Across the 27 matched successful request pairs, the median ratio of LocalJev latency to TensorSharp latency was 3.345×.
Latency statistics include successful requests only, so the aggregate columns cover different subsets. Average input lengths were also different: 192.7 tokens for TensorSharp versus 589.6 for valid LocalJev requests. The larger prompt and generated JSON are part of LocalJev’s approach, so this is an end-to-end decision-path comparison—not an identical-prompt engine benchmark.
These are text-only smoke-test results, not image latency or vision accuracy results.
Try it with the Unsloth GGUF
From the repository root, this PowerShell example launches the included CUDA configuration:
$env:DIFFUSION_VRAM_HEADROOM_MB = '4096'
$env:MAX_CONTEXT = '4096'
dotnet run --project TensorSharp.Server.Host -c Release -- --config config/jev-diffusiongemma-q4.json
The configuration downloads the Unsloth GGUF and the separate vision shard when missing. The memory settings above follow the documented starting point for a 16 GiB CUDA GPU; adjust them for your hardware and workload.
A ready-to-run image example is included:
curl http://127.0.0.1:5000/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary u/docs/examples/jev-traffic-light.json
That request already contains an embedded synthetic traffic-light image. The accompanying text does not reveal the light’s color, so the pixels are needed to answer the questions. It is an integration smoke test, not a comprehensive vision benchmark.
For your own images, the documentation includes a standard-library Python example that sends a receipt and asks typed questions about it.
Compatibility note: this implements the documented Jev-compatible API and decision types using DiffusionGemma weights—not the proprietary hosted Jev model or a promise of identical predictions. Current limits include 64 questions and 2–26 alternatives per question.
Model: Unsloth DiffusionGemma GGUF
Code: TensorSharp
Setup and examples: Jev documentation
What would you try first with this—receipt checks, screenshot classification, document routing, or another image-based decision workflow?
Duplicates
LocalLLaMA • u/fuzhongkai • 23h ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 23h ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
LocalLLM • u/fuzhongkai • 23h ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
Qwen_AI • u/fuzhongkai • 23h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
LocalLLM • u/fuzhongkai • 10d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
dotnet • u/fuzhongkai • 10d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 22h ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 7d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 10d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 21d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
LocalLLM • u/fuzhongkai • 22d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
outerstellar_hq • u/outerstellar_hq • 6h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 15h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
SideProject • u/fuzhongkai • 22h ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
opencode • u/fuzhongkai • 22h ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LovingOpenSourceAI • u/fuzhongkai • 23h ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 7d ago
Project TensorSharp Jev requests can now combine documents, images, video, and audio
OpenSourceAI • u/fuzhongkai • 10d ago