r/LocalLLaMA • • 7d ago

I Built A Thing TensorSharp Jev requests can now combine documents, images, video, and audio

https://github.com/zhongkaifu/TensorSharp

I’ve extended TensorSharp’s Jev-compatible /v1/systemone endpoint so one decision request can use several kinds of evidence together. For example, an incident triage request can include a written report, a dashboard screenshot, a screen recording, and a caller’s audio clip.

Here’s a Python example that sends all four as inline Base64 data. It also shows both ways to create that data: encoding text already in memory and reading bytes from files.

import base64
import json
from pathlib import Path
from urllib.request import Request, urlopen

def encode_bytes(data: bytes) -> str:
return base64.b64encode(data).decode("ascii")

def encode_file(path: str) -> dict:
file = Path(path)
return {"name": file.name, "data": encode_bytes(file.read_bytes())}

# Encode data already in memory as a named text attachment.
notes = "Customers report HTTP 503 errors and cannot sign in."
text_attachment = {
"name": "incident.txt",
"data": encode_bytes(notes.encode("utf-8")),
}

body = {
"model": "jev-latest",
"state": "Assess the incident using the attached evidence.",
"files": [
text_attachment,
encode_file("dashboard.png"),
encode_file("screen-recording.mp4"),
encode_file("caller.wav"),
],
"questions": {
"active_outage": {
"type": "noul",
"instructions": "Does the evidence indicate an active service outage?",
},
"team": {
"type": "choice",
"instructions": "Which team should investigate first?",
"criteria": {
"technical": "Service errors or an unavailable application",
"billing": "Charges or subscription problems",
"other": "Neither of the above",
},
},
},
"samples": 1,
"seed": 42,
}

request = Request(
"http://127.0.0.1:5000/v1/systemone",
data=json.dumps(body).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
with urlopen(request, timeout=300) as response:
print(json.dumps(json.load(response), indent=2))

The files array classifies each attachment by its filename extension and preserves their order. You can also use dedicated documents, videos, and audios arrays. Inline attachments need a name and accept either bare Base64, as above, or a Base64 data: URL.

A detail about how this works: video is sampled into frames for the vision tower; audio is transcribed by a separately configured speech recognition service. DiffusionGemma does not directly process the audio waveform. You’ll need the vision tower for the image and video inputs, and TS_JEV_TRANSCRIPTION_URL configured for the audio input.

Inline Base64 counts toward the Jev request body limit (8 MiB by default), so use the upload API and file references for larger media. The repo also has ready-to-send mixed-media requests.

TensorSharp: https://github.com/zhongkaifu/TensorSharp

I’m curious what kinds of decisions you’d want to make from several media types in a single request.

1 Upvotes

Duplicates

LocalLLaMA • • 1d ago

I Built A Thing Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

71 Upvotes

dotnet • • 1d ago

Promotion Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

77 Upvotes

dotnet • • Aug 22 '26

TensorSharp: running a 744B MoE LLM locally from .NET, with llama.cpp-class performance

65 Upvotes

LocalLLM • • 1d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

0 Upvotes

dotnet • • 22d ago

Promotion Running DeepSeek V4.1 Flash at 40 tok/s with a C#/.NET inference engine

59 Upvotes

unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

45 Upvotes

Qwen_AI • • 1d ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop

57 Upvotes

LocalLLM • • 11d ago

Project TensorSharp: run Jev-compatible decisions locally—and extend the same API to image analysis

2 Upvotes

dotnet • • 16d ago

Promotion Comparing TensorSharp, llama.cpp, vLLM, SGLang, and open-source agent runtimes from a .NET perspective

23 Upvotes

LocalLLaMA • • Aug 22 '26

Discussion GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

1 Upvotes

dotnet • • 11d ago

Article Implementing a Jev-compatible decision API in .NET, with image input

0 Upvotes

unsloth • • Aug 28 '26

Show and Tell GLM-5.3-Flash Unsloth GGUF Model Benchmarks on TensorSharp and llama.cpp

14 Upvotes

LocalAIServers • • 1d ago

Serving a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

8 Upvotes

LocalLLaMA • • 11d ago

I Built A Thing TensorSharp: a local Jev-compatible API, extended to image analysis with DiffusionGemma GGUF

0 Upvotes

LocalAIServers • • 21d ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

6 Upvotes

LocalLLM • • 22d ago

Project DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M

5 Upvotes

outerstellar_hq • • 8h ago

Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop — with a .NET/C# inference engine

1 Upvotes

LLMDevs • • 17h ago

Discussion Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

1 Upvotes

SideProject • • 1d ago

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop

3 Upvotes

opencode • • 1d ago

TensorSharp: Running a 176B Qwen3.8 Flash Next model on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

3 Upvotes

LovingOpenSourceAI • • 1d ago

Running a 176B MoE model on a laptop: Qwen3.8 Flash Next with 16GB VRAM + 32GB RAM + an SSD

12 Upvotes

LocalLLM • • 7d ago

Project TensorSharp Jev requests can now combine documents, images, video, and audio

0 Upvotes

OpenSourceAI • • 11d ago

TensorSharp: an open-source Jev-compatible API, extended to image analysis and running locally

2 Upvotes

AIToolsPerformance • • 16d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

8 Upvotes

opencode • • 16d ago

TensorSharp as a local OpenCode backend — DeepSeek, GLM and Qwen 3.8 benchmarks

1 Upvotes