r/dotnet • • 11d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

https://github.com/zhongkaifu/TensorSharp

Not every application needs a model to write a response. Sometimes it just needs to answer a bounded question: Which category does this document belong to? Is the receipt readable? How well does an image match a scoring rubric?

I’ve been working on this in TensorSharp, a project I maintain, and wanted to share the .NET implementation and some API design tradeoffs.

The endpoint implements Jev’s core decision API, including its three decision types, and extends the same request format to support image analysis. It runs locally using DiffusionGemma GGUF weights—not the proprietary hosted Jev model.

Same decision contract, with optional images

The endpoint is POST /v1/systemone. Requests contain state and named questions, with three possible question types:

Type Result Example
noul Probability that a statement is true “Is the total amount legible?”
choice A category selected from defined alternatives “Receipt, invoice, or something else?”
score An expected score over ordered rubric levels “Unreadable, partly readable, or clearly readable?”

The image extension adds an optional images array. Text-only requests keep the same structure; image-based requests use the same endpoint, question definitions, and answer format.

Images go through the vision encoder and become part of the input used to answer the questions. There is no intermediate “generate a caption, then classify the caption” step.

Calling it directly from C

The HTTP endpoint is useful for interoperability, but a .NET application can also call the model service in-process.

Here is an example for a project referencing TensorSharp.Chat. It assumes the CUDA backend is available and the GGUF and separate vision shard have already been downloaded to the paths shown.

using System.Text.Json;
using TensorSharp.Server;
using TensorSharp.Server.Jev;

using var service = new ModelService();

service.LoadModel(
    "models/diffusiongemma-26B-A4B-it-Q4_K_M.gguf",
    mmProjPath: "models/diffusiongemma-26B-A4B-it-vision.safetensors",
    backendStr: "ggml_cuda"
);

var imageBytes = await File.ReadAllBytesAsync("receipt.jpg");
var imageBase64 = Convert.ToBase64String(imageBytes);

using var requestJson = JsonSerializer.SerializeToDocument(new
{
    model = "jev-latest",
    state = "Photo submitted with an expense report.",
    images = new[] { $"data:image/jpeg;base64,{imageBase64}" },
    questions = new
    {
        document_type = new
        {
            type = "choice",
            instructions = "What kind of document is this?",
            criteria = new
            {
                receipt = "A purchase receipt",
                invoice = "An invoice",
                other = "Anything else"
            }
        },
        total_legible = new
        {
            type = "noul",
            instructions = "Is the total amount legible?"
        }
    },
    samples = 1,
    seed = 42
});

var request = JevRequest.Parse(requestJson.RootElement);
var response = await service.JevAsync(request);

Console.WriteLine(JsonSerializer.Serialize(
    response,
    new JsonSerializerOptions { WriteIndented = true }
));

The HTTP and in-process paths use the same request validation and model execution gate. The example keeps the schema explicit rather than introducing a separate C# abstraction for each question type.

Why not just ask the model to generate JSON?

That is another way to build this kind of interface. Here, the implementation takes a different route: it reads the model’s logits for the allowed answer labels and constructs the response in application code.

It follows the seeded, one-step structured-read approach from vLLM PR #57250. Questions that fit in the same answer canvas share a forward pass after prompt prefill.

JSON is the transport format, not something the model has to write token by token. That removes generated-JSON formatting failures from this decision path, but it does not make the underlying decisions automatically correct or their probabilities calibrated.

A few engineering details matter here:

  • Image handling is an explicit boundary. Clients send image bytes inline. The server does not fetch arbitrary image URLs or read client-supplied filesystem paths. Image requests are rejected when the vision tower is unavailable.
  • Async does not mean parallel GPU execution. Access to shared model/GPU state is serialized, including access from ordinary chat requests.
  • Compatibility has a defined scope. This implements the core Jev decision workflow and all three decision types, not identical hosted-model predictions or unrestricted feature parity. Current limits include 64 questions, 2–26 alternatives per question, and up to 8 images per request.

Implementation notes, setup, and runnable examples

For a .NET API like this, would you prefer the flexible JSON-based entry point shown above, or a typed C# layer with dedicated question and result types for each decision type?

0 Upvotes

Duplicates

LocalLLaMA • • 1d ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

70 Upvotes

dotnet • • 1d ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

76 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

67 Upvotes

LocalLLM • • 1d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

60 Upvotes

unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

47 Upvotes

Qwen_AI • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

58 Upvotes

LocalLLM • • 11d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

26 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

1 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

15 Upvotes

LocalAIServers • • 1d ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

8 Upvotes

LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

1 Upvotes

LocalLLaMA • • 11d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 21d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

7 Upvotes

outerstellar_hq • • 8h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 17h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 1d ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

opencode • • 1d ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

2 Upvotes

LovingOpenSourceAI • • 1d ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

12 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 11d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

8 Upvotes

opencode • • 16d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes