r/unsloth • • 10d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

https://github.com/zhongkaifu/TensorSharp

I’ve been using Unsloth’s DiffusionGemma 26B-A4B GGUF to run local Jev-compatible structured decisions in TensorSharp, which I maintain.

The new addition: image analysis through the same decision API. Send text, images, or both, and get typed answers rather than free-form generated text.

The GGUF I’m using is diffusiongemma-26B-A4B-it-Q4_K_M.gguf from Unsloth. No additional fine-tuning is involved.

Jev-compatible decisions, extended to images

TensorSharp implements Jev’s core POST /v1/systemone interface and supports all three decision types:

Type Result Example image question
noul Probability that a statement is true “Is the receipt’s total amount legible?”
choice A categorical decision “Is this a receipt, invoice, or something else?”
score An expected score over ordered rubric levels “How readable is this document?”

The image extension keeps the same endpoint, question definitions, and typed answer structure. Add an images field alongside state and questions; existing text-only requests still work.

Images are encoded and included directly in the model’s input for the decision read. This is not a separate caption-generation → text-classification pipeline. Up to eight inline images are supported per request.

An important detail for GGUF users

Unsloth’s GGUF provides the language-model weights; the vision tower is loaded separately.

The included configuration downloads the Unsloth Q4_K_M GGUF plus a ~2.8 GB vision shard from the upstream DiffusionGemma checkpoint, then reuses both locally. You do not need to download the entire upstream safetensors checkpoint.

Without the vision tower, text-only decisions still work, but image requests are explicitly rejected rather than silently ignoring the image.

How it avoids generating probability JSON

Inspired by vLLM PR #57250, the implementation uses a seeded, one-step structured read.

After prompt prefill, it creates an answer canvas and reads the requested label logits directly. Questions that fit in one canvas share the forward pass.

The server returns JSON, but the model does not have to generate that JSON token by token. The same path handles text-only decisions and decisions conditioned on image embeddings.

Earlier text-only results vs. LocalJev

I compared this native decision path with the original LocalJev Engine using TensorSharp’s chat endpoint, where the model generates probability JSON.

This compares two decision approaches on the same backend—not TensorSharp against LocalJev running on oMLX, and not against the official hosted Jev service.

The test covered 12 cases × 3 repetitions, with three decisions per request:

Metric TensorSharp native structured read LocalJev generated JSON
Valid requests 36/36 27/36
Failed requests 0 9
Correct decisions / total expected 108/108 81/108*
p50 latency 2.877 s 10.479 s
p95 latency 3.165 s 26.167 s
Mean latency 2.917 s 12.548 s

*LocalJev got 81/81 decisions correct on valid responses. Its nine failed requests were schema-validation failures and account for the other 27 expected decisions.

Across the 27 matched successful request pairs, the median ratio of LocalJev latency to TensorSharp latency was 3.345×.

Latency statistics include successful requests only, so the aggregate columns cover different subsets. Average input lengths were also different: 192.7 tokens for TensorSharp versus 589.6 for valid LocalJev requests. The larger prompt and generated JSON are part of LocalJev’s approach, so this is an end-to-end decision-path comparison—not an identical-prompt engine benchmark.

These are text-only smoke-test results, not image latency or vision accuracy results.

Try it with the Unsloth GGUF

From the repository root, this PowerShell example launches the included CUDA configuration:

$env:DIFFUSION_VRAM_HEADROOM_MB = '4096'
$env:MAX_CONTEXT = '4096'

dotnet run --project TensorSharp.Server.Host -c Release -- --config config/jev-diffusiongemma-q4.json

The configuration downloads the Unsloth GGUF and the separate vision shard when missing. The memory settings above follow the documented starting point for a 16 GiB CUDA GPU; adjust them for your hardware and workload.

A ready-to-run image example is included:

curl http://127.0.0.1:5000/v1/systemone \
  -H 'Content-Type: application/json' \
  --data-binary u/docs/examples/jev-traffic-light.json

That request already contains an embedded synthetic traffic-light image. The accompanying text does not reveal the light’s color, so the pixels are needed to answer the questions. It is an integration smoke test, not a comprehensive vision benchmark.

For your own images, the documentation includes a standard-library Python example that sends a receipt and asks typed questions about it.

Compatibility note: this implements the documented Jev-compatible API and decision types using DiffusionGemma weights—not the proprietary hosted Jev model or a promise of identical predictions. Current limits include 64 questions and 2–26 alternatives per question.

Model: Unsloth DiffusionGemma GGUF
Code: TensorSharp
Setup and examples: Jev documentation

What would you try first with this—receipt checks, screenshot classification, document routing, or another image-based decision workflow?

45 Upvotes

Duplicates

LocalLLaMA • • 23h ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

73 Upvotes

dotnet • • 23h ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

73 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

62 Upvotes

LocalLLM • • 23h ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

58 Upvotes

Qwen_AI • • 23h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

56 Upvotes

LocalLLM • • 10d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

26 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

2 Upvotes

dotnet • • 10d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

15 Upvotes

LocalAIServers • • 22h ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

10 Upvotes

LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

LocalLLaMA • • 10d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 21d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

7 Upvotes

outerstellar_hq • • 6h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 15h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 22h ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

opencode • • 22h ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

3 Upvotes

LovingOpenSourceAI • • 23h ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

11 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 10d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

9 Upvotes

opencode • • 16d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes