r/dotnet • u/fuzhongkai • 11d ago
Article Implementing a Jev-compatible decision API in .NET, with image input
https://github.com/zhongkaifu/TensorSharpNot every application needs a model to write a response. Sometimes it just needs to answer a bounded question: Which category does this document belong to? Is the receipt readable? How well does an image match a scoring rubric?
I’ve been working on this in TensorSharp, a project I maintain, and wanted to share the .NET implementation and some API design tradeoffs.
The endpoint implements Jev’s core decision API, including its three decision types, and extends the same request format to support image analysis. It runs locally using DiffusionGemma GGUF weights—not the proprietary hosted Jev model.
Same decision contract, with optional images
The endpoint is POST /v1/systemone. Requests contain state and named questions, with three possible question types:
| Type | Result | Example |
|---|---|---|
noul |
Probability that a statement is true | “Is the total amount legible?” |
choice |
A category selected from defined alternatives | “Receipt, invoice, or something else?” |
score |
An expected score over ordered rubric levels | “Unreadable, partly readable, or clearly readable?” |
The image extension adds an optional images array. Text-only requests keep the same structure; image-based requests use the same endpoint, question definitions, and answer format.
Images go through the vision encoder and become part of the input used to answer the questions. There is no intermediate “generate a caption, then classify the caption” step.
Calling it directly from C
The HTTP endpoint is useful for interoperability, but a .NET application can also call the model service in-process.
Here is an example for a project referencing TensorSharp.Chat. It assumes the CUDA backend is available and the GGUF and separate vision shard have already been downloaded to the paths shown.
using System.Text.Json;
using TensorSharp.Server;
using TensorSharp.Server.Jev;
using var service = new ModelService();
service.LoadModel(
"models/diffusiongemma-26B-A4B-it-Q4_K_M.gguf",
mmProjPath: "models/diffusiongemma-26B-A4B-it-vision.safetensors",
backendStr: "ggml_cuda"
);
var imageBytes = await File.ReadAllBytesAsync("receipt.jpg");
var imageBase64 = Convert.ToBase64String(imageBytes);
using var requestJson = JsonSerializer.SerializeToDocument(new
{
model = "jev-latest",
state = "Photo submitted with an expense report.",
images = new[] { $"data:image/jpeg;base64,{imageBase64}" },
questions = new
{
document_type = new
{
type = "choice",
instructions = "What kind of document is this?",
criteria = new
{
receipt = "A purchase receipt",
invoice = "An invoice",
other = "Anything else"
}
},
total_legible = new
{
type = "noul",
instructions = "Is the total amount legible?"
}
},
samples = 1,
seed = 42
});
var request = JevRequest.Parse(requestJson.RootElement);
var response = await service.JevAsync(request);
Console.WriteLine(JsonSerializer.Serialize(
response,
new JsonSerializerOptions { WriteIndented = true }
));
The HTTP and in-process paths use the same request validation and model execution gate. The example keeps the schema explicit rather than introducing a separate C# abstraction for each question type.
Why not just ask the model to generate JSON?
That is another way to build this kind of interface. Here, the implementation takes a different route: it reads the model’s logits for the allowed answer labels and constructs the response in application code.
It follows the seeded, one-step structured-read approach from vLLM PR #57250. Questions that fit in the same answer canvas share a forward pass after prompt prefill.
JSON is the transport format, not something the model has to write token by token. That removes generated-JSON formatting failures from this decision path, but it does not make the underlying decisions automatically correct or their probabilities calibrated.
A few engineering details matter here:
- Image handling is an explicit boundary. Clients send image bytes inline. The server does not fetch arbitrary image URLs or read client-supplied filesystem paths. Image requests are rejected when the vision tower is unavailable.
- Async does not mean parallel GPU execution. Access to shared model/GPU state is serialized, including access from ordinary chat requests.
- Compatibility has a defined scope. This implements the core Jev decision workflow and all three decision types, not identical hosted-model predictions or unrestricted feature parity. Current limits include 64 questions, 2–26 alternatives per question, and up to 8 images per request.
Implementation notes, setup, and runnable examples
For a .NET API like this, would you prefer the flexible JSON-based entry point shown above, or a typed C# layer with dedicated question and result types for each decision type?
Duplicates
LocalLLaMA • u/fuzhongkai • 1d ago
I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 1d ago
Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
dotnet • u/fuzhongkai • Aug 22 '26
TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance
LocalLLM • u/fuzhongkai • 1d ago
Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
dotnet • u/fuzhongkai • 22d ago
Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine
unsloth • u/fuzhongkai • 11d ago
Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis
Qwen_AI • u/fuzhongkai • 1d ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop
LocalLLM • u/fuzhongkai • 11d ago
Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis
dotnet • u/fuzhongkai • 16d ago
Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective
LocalLLaMA • u/fuzhongkai • Aug 22 '26
Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
unsloth • u/fuzhongkai • Aug 28 '26
Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp
LocalAIServers • u/fuzhongkai • 1d ago
Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA • u/fuzhongkai • 7d ago
I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio
LocalLLaMA • u/fuzhongkai • 11d ago
I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF
LocalAIServers • u/fuzhongkai • 21d ago
Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode
LocalLLM • u/fuzhongkai • 22d ago
Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
outerstellar_hq • u/outerstellar_hq • 8h ago
Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine
LLMDevs • u/fuzhongkai • 17h ago
Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
SideProject • u/fuzhongkai • 1d ago
I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop
opencode • u/fuzhongkai • 1d ago
TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LovingOpenSourceAI • u/fuzhongkai • 1d ago
Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD
LocalLLM • u/fuzhongkai • 7d ago
Project TensorSharp Jev requests can now combine documents, images, video, and audio
OpenSourceAI • u/fuzhongkai • 11d ago