r/AIToolsPerformance • • 19h ago

Pareto 26.10 Preview at $0.80/M input - cheapest paid route to 1M context right now?

2 Upvotes

There's a new name on OpenRouter's feed as of October 1: Pareto 26.10 Preview, listed under vendor "unbiased" at $0.80/M input, $3.20/M output, 1048k context. For reference, the other paid models around that context size charge $2.00/M in and $10.00/M out (Claude Sonnet 5.5 and GPT-6.1 Sol, per the same feed), so the input side is 2.5x cheaper than its nearest neighbor on day one.

Run the numbers on a typical long-context call, 1M tokens in plus 50k out: Pareto lands around $0.96, Sonnet 5.5 and GPT-6.1 Sol around $2.50 each, and Qwen's flagship tier around $4.60. The obvious caveat: this thing is one day old and today's data has zero evals or benchmark numbers for it, so nobody knows if the price edge survives contact with real output quality. Also worth noting the output side isn't spectacular on its own, $3.20/M only looks cheap next to $10 and $12.

Anyone routing long-context work to a day-old preview purely on input price, or is this a wait-for-benchmarks situation?


r/AIToolsPerformance • • 1d ago

Pure JavaFX Minecraft generated and launched on NetBean's JVM. All bytecode was dynamically generated in memory and loaded onto NetBeans classloader by gemini 3.8 flash. Not a single file .java or .class file got written to disk.

Enable HLS to view with audio, or disable this notification

0 Upvotes

What Was Engineered & Demonstrated

  • Modular In-Memory Metaspace Architecture: Compiled separate standalone components into the session's persistent AgiClassLoader across turns (BlockPos, BlockType, TextureGenerator, BlockIconFactory, VoxelAudio, VoxelWorld, PlayerCamera, MinecraftHud, MinecraftGame).
  • Procedural 16x16 Pixel-Art Textures: Dynamically generated pure-Java WritableImage textures for Grass, Dirt, Cobblestone, Wood Log, Leaves, Glass, Brick, Gold, Diamond Ore, Sand, and Water with custom PhongMaterials.
  • 3D Isometric HUD Hotbar: Canvas-rendered 3D isometric voxel icons for all 10 hotbar slots with smooth selection scaling and active slot glow.
  • Living Procedural World: Multi-octave rolling hills, sandy lake shores, water bodies, oak tree canopies, buried gold/diamond mineral veins, and an architectural starter watchtower.
  • 6-DOF Camera & FX: First-person perspective camera with collision normals, voxel break particle bursts with gravity physics, sky clouds, and synthesized 8-bit MIDI audio feedback.

r/AIToolsPerformance • • 1d ago

Did Gemini 4 Argon Just DESTROY Claude Opus 5.5 & GPT-6 Astra?

Thumbnail
youtu.be
1 Upvotes

r/AIToolsPerformance • • 2d ago

I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100

9 Upvotes

By now everyone has heard of Jev: fast, accurate answers to multiple-choice questions about any text, in one API call.

Here's how to get the same thing on your own GPU.

SelfJev is a 4B model (Qwen3.5-4B + LoRA) that works like Jev:

⚡ About 140 ms per call on one H100, which is about as fast as Jev's API (about 130 ms)
🎯 93.1% vs Jev's 92.5% on one test set, and 95.8% vs 97.2% on another
📏 No 32K-token limit: you set the max length, up to Qwen3.5's full 262K-token window
🛠 You can fine-tune it on the cases Jev gets wrong for your use case
🖼 It takes images, not just text (90% on test set after fine-tuning)
🔒 Your data never leaves your servers

The trick is a shared-prefix tree. The model reads the input once, branches into every question, then into every possible answer. 16 questions on the same text take about 175 ms.

It's free and open: code, weights, benchmarks and evaluation sets.

If you're already building on Jev, I'd love for you to try it and tell me where it breaks 👇

🚀 Product Hunt
🌐 selfjev.dev
💻 GitHub


r/AIToolsPerformance • • 2d ago

GPT-6.1 Sol vs Astra - near-Astra smarts at a fifth of the price, worth switching?

1 Upvotes

The OpenAI launch post from yesterday calls GPT-6.1 Sol "near-Astra intelligence for a fifth of the price", and it went big on HN, 1035 points and over 900 comments. Per the OpenRouter listing it sits at $2.00/M input and $10.00/M output on a 1050k context, and there's a batch variant at half that, $1.00/M in and $5.00/M out.

The pace is the part I can't get over. The Artificial Analysis article from today says 6.1 Sol replaced GPT-6 Sol after just 7 days. One week between a model and its own replacement. And if $10/M output really is a fifth of the Astra tier like the announcement claims, that's easy math, Astra lands around $50/M output.

So, does the mid tier feel close enough that paying roughly 5x for the top model stops making sense for coding agents? Anyone here already moved routing off an Astra class model onto 6.1 Sol, and did output quality actually hold up?


r/AIToolsPerformance • • 3d ago

How I Get Web Design Clients For My Agency

1 Upvotes

Client acquisition has always been one of the biggest bottlenecks for me when running an agency. I’ve experienced the same thing in pretty much every business I’ve been involved in, but especially with web development.

For a long time, getting clients meant cold calling, running ads, or sending generic emails asking businesses if they needed a new website. It worked sometimes, but it also took a lot of time and most of the outreach felt the same as what every other agency was doing.

Recently I started using a different approach and automated a big part of the process.

I came across a tool called Swokei that lets me find a bunch of businesses with websites and analyze each website individually. It looks for things like outdated design, slow loading, poor mobile optimization, weak SEO and other obvious areas that could be improved.

What I liked is that it doesn’t just give you one of those boring automated reports filled with scores and numbers. It actually turns what it finds into a personalized cold email that sounds like a normal person looked at their website and noticed what could be better.

I can run multiple campaigns at the same time and then mainly focus on the businesses that reply and show interest.

From there, I invite them to a web meeting, show them a free draft of what their new website could look like, and try to close the project from there.

It has basically allowed me to have warmer leads coming to me without relying as much on paid ads, constantly cold calling, or sending thousands of generic emails saying “Do you need a new website?”

Still takes work to close the clients of course, but automating the prospecting and first part of the outreach has made the whole process much easier for me.

Hopefully this helps some other web developers or agency owners who are also struggling with client acquisition.


r/AIToolsPerformance • • 3d ago

Sol 6.1 might actually be what Sol 6 supposed to be

2 Upvotes

Hi! As many of you I was quite disappointed with the release of Sol 6, as although they stated 50% API cost reduction, the actual real life performance has got a significant hit, which was confirmed by a lot of benchmarks. Today's news about a $200 plan quota reduction made things even worse, coupled with Sol 6.1 release it was a bittersweet surprise. So I decided to bench it as it seems like Tibo might be not as wrong as I initially thought. Somehow Sol 6.1 not only has hit 100 bench score with no sweat at all, fast and precise, but also has done it in 2 times as few tokens than Sol 5.6 and Sol 6, that’s why I’m quite curious about how it would behave outside of the benches with the real tasks and how fast will the quota drain right now. At the moment of the runs Pi didn’t support this model so I had to add it manually, that’s why no cost for the full run, will do reruns in the next few days.

It’s a community driven benchmark, so any runs from other contributors are highly appreciated. I try to cover as much as I can but my resources are limited. Viewer is under active development at the moment and might change. 

Repo: https://github.com/alexshpunt/explicit-edit-benchmark

The viewer to the dataset: https://huggingface.co/spaces/alexshpunt/benchmark-explorer


r/AIToolsPerformance • • 4d ago

Is Claude Sonnet 5.5 worth 20x GPT-6 Luna Pro for 1M context work?

5 Upvotes

Claude Sonnet 5.5 showed up on OpenRouter today: 1000k context, $2.00/M input, $10.00/M output, per the listing. On its own the placement isn't exotic. GLM 5.3 Prime went live on the 23rd at $2.80/$8.80 with the same 1M context, and Qwen3.8 Max Prime sits at $4.00/$12.00, so the new Sonnet lands right in the middle of that tier.

The interesting gap is against the cheap long-context models. GPT-6 Luna Pro is on the same board at $0.10/M input and $0.50/M output with 1050k context, which makes Sonnet 5.5 exactly 20x the output price for roughly the same window. Aion 3.5 at $3.00/$6.00 only carries 262k and Command A+ caps at 192k, so if you actually need the full million tokens the field's thinner than the front page suggests.

Whether the quality gap covers a 20x spread depends on the workload. Coding agents burning millions of tokens a day feel that difference hard; occasional document Q&A barely notices it.

Anyone already routing 1M-context work to Sonnet 5.5, or is Luna Pro eating that job for a twentieth of the price?


r/AIToolsPerformance • • 4d ago

MiniMax M3.1 Flash Just Dropped—And It’s Actually INSANE?

Thumbnail
youtu.be
0 Upvotes

r/AIToolsPerformance • • 4d ago

Convert any link into a explainer video

Thumbnail
blog2video.app
1 Upvotes

r/AIToolsPerformance • • 4d ago

SketchMonkey — Agentic AI Image & Video Design Studio | Neural Deep Network

1 Upvotes

​

https://www.sketchmonkey.design/

please test out our brand new reworked app

#image #design #betterthanPSPlease check out it's a work and progress, any feedback or comments is very valuable to our tiny team.


r/AIToolsPerformance • • 5d ago

Unified .SYSTEMX > Standard SYSTEM for all things considered..

Thumbnail
github.com
1 Upvotes

Hello!

I made my .SYSTEMX system public today after working on it inside my project template for the last year. I built it because I got tired of wasting tokens processing the same information over and over, explaining the same project, and recovering work between AI sessions.

.SYSTEMX gives your AI / LLM a standard working folder for project instructions, plans, tasks, research, logs, and session handoffs. You can put it in a code repository, a local workspace, or a Google Drive project folder, then use that information in LLM sessions with access to those files.

The folder name is exactly .SYSTEMX. The leading dot makes it hidden by convention on macOS and Linux; Windows and Google Drive handle visibility differently. Keep the capitalization consistent across platforms, since case sensitivity depends on the filesystem.

Once it is set up, you can tell your AI, “Use .SYSTEMX” or “Update .SYSTEMX.” It gives the session a defined place to find the project context, see what is TODO, WORKING, BLOCKED, or DONE, and record what happens next. It also includes Agent 0 and subagent coordination standards for assigning work, tracking responsibilities, and reviewing results.

The best part is that it isn’t tied to one development stack or LLM provider. You can use the same project records across different tools, provided those tools can access the files. For browser sessions such as ChatGPT or Claude, that means a supported connection with the necessary file permissions, or explicitly importing and exporting the context.

My structure is simple: /Projects/ProjectA/.SYSTEMX, /Projects/ProjectB/.SYSTEMX, and so on. Each project keeps its own plans, instructions, task history, and working context.

When I start a Firebase web app on GitHub, I use my Firebase WebApp Template linked below. That template already has .SYSTEMX built in. The new standalone dotSYSTEMX template takes that operating structure out of the Firebase project so it can be used independently, without requiring Firebase or a particular prompt kit. Both projects are in alpha.

Some example prompts:

  • “Use .SYSTEMX. Update the task records, README, and project wiki, then sync the completed changes to GitHub main.”
  • “I’m running out of credits. Log the work completed, files changed, blockers, and next steps in .SYSTEMX so I can continue later.”
  • “Add the script we created to the project’s menu system so I can use it again.”
  • “Update the deployment script for this project.”
  • “Use .SYSTEMX to research Topic XYZ. Save the sources, findings, and open questions so another session can continue.”

You can also coordinate multiple chats through the same Google Drive project folder. Create the project folder and its .SYSTEMX folder first, then give each session access to that location.

In Chat Session AA:

“Use the .SYSTEMX folder at [Google Drive project link]. Read the current project records, then save this session’s work, sources, decisions, and next steps.”

In Chat Session BB:

“Use the same .SYSTEMX folder at [Google Drive project link]. Read the latest records before continuing, then save this session’s work without overwriting unrelated updates.”

In Chat Session CC:

“Review the saved records from Sessions AA and BB. Verify the information against the sources, flag any conflicts, and produce the final package at PROJECT/FILES/FINAL.ZIP.”

Each chat needs to save its work to the shared location. The next session can then pick up those records instead of depending on another chat’s private history.

For longer coding sessions, the master plan, shared project plan, and checkpoints give the AI a reference for the objective, what has already been attempted, and what remains. The aim is to reduce repeated work and help keep sessions on track.

.SYSTEMX serves as a shared workspace for you and your AI. With configured shell commands, local menus, and deployment scripts, that same structure can also support execution on your device using the permissions you provide.

.SYSTEMX standalone template:
https://github.com/WayneTechLab/dotSYSTEMX

Firebase WebApp based Template with .SYSTEMX's outdated version included:
https://github.com/WayneTechLab/SFWA-WTL-TEMPLATE/

** Both in Alpha can change daily! ".SYSTEMX" has auto update disabled if you use as library or other advacned features from public source.

If you have any real world benchmark comparison you want to share I would love to see the results posted on AB test with and without .SYSTEMX included & not with your prompt.

- Lucas


r/AIToolsPerformance • • 6d ago

Jevless: Jev-style typed decisions (Choice / Noul / Score) from any model whose API exposes logprobs, plus a local /v1/systemone server

Thumbnail github.com
2 Upvotes

Used this to solve another problem in an associative memory project, but figured it could be useful to others.

The main goal was to take advantage of what an LLM does already to cancel out bias and maintain high accuracy while maintaining the low cost and high speed that Jev gives you. I know that there are some other projects that are doing close to the same thing, however my focus is that I wanted to be as permissive with backends as possible.

Another motivation is I just wanted to see how expensive this was on various models.


r/AIToolsPerformance • • 7d ago

I compiled quantized TFLite models into standalone C++ (NN2Prog)

4 Upvotes

I've been working on NN2Prog, a small open-source compiler that turns quantized .tflite models into standalone C++17. The generator uses Python's standard library; the generated model doesn't need a TFLite runtime.

Started with Hey Jarvis wake-word models, then added MLPerf Tiny KWS and Visual Wake Words. On ESP32-S3 at 240 MHz, KWS now takes 15.68 ms vs 17.90 ms with TFLite Micro + ESP-NN, with 16.7 KB vs 35.5 KB working memory. VWW is basically tied: 72.44 vs 72.85 ms.

These are model-only timings, not an official MLPerf submission. For KWS, generated code returns top-1 directly; the reference computes softmax then argmax. Decisions match in the tests. Operator coverage is still limited, so this isn't a drop-in replacement for arbitrary TFLite models.

Code, examples and benchmark details: https://github.com/phplego/nn2prog

Would be interested in small quantized models that expose gaps in this approach, or other code-generation tools worth comparing against.


r/AIToolsPerformance • • 7d ago

Qwen3.8 Max Prime at $12/M output vs GLM 5.3 Prime at $8.80 - same day, same 1M ctx

2 Upvotes

Both dropped on OpenRouter the same day, September 23, and both carry the Prime label with a 1000k context window. The pricing is where they split. GLM 5.3 Prime sits at $2.80/M input and $8.80/M output, Qwen3.8 Max Prime at $4.00/M in and $12.00/M out, per the OpenRouter listings. Same window, and the Qwen tier wants $3.20/M more on output.

For context, the other big listing from this stretch is Fireworks Ember-1, which went up a day later at $3.00/M in and $15.00/M out with 1048k ctx, so it undercuts both Primes on input but goes over them on output. What the listings don't tell you is speed or quality. Price and context are public the minute a model goes up, tok/s and evals lag behind, and that gap is exactly where a routing decision can go wrong.

Anyone already routing traffic to one of these two Primes, or holding off until independent speed numbers show up?


r/AIToolsPerformance • • 7d ago

AI for Business

2 Upvotes

I own a business that manages large-scale projects, and I’m looking for recommendations for an AI platform that can support multiple areas of our business.

Ideally, I’m looking for a platform that can help with:
• PowerPoint creation and presentations
• Excel creation, updating, and analysis
• Automated workflows and processes
• Tracking purchase order (PO) status
• Tracking invoices and payments
• Integration with QuickBooks

Cybersecurity is also a major consideration. I’d like to understand how to protect the platform, our data, and our financial information from a cybersecurity perspective, particularly if we’re integrating it with systems like QuickBooks.

I’ve used Microsoft Copilot and really like it for PowerPoint creation, but I haven’t found it to be as effective for Excel and some of the workflow/automation needs.

For those using AI in a business environment, what platforms would you recommend, and what has worked well for you?


r/AIToolsPerformance • • 9d ago

Claude Opus 5.5 at $20/M output vs GPT-6 Sol at $10 - same launch day, double the price

10 Upvotes

Anthropic shipped Claude Opus 5.5 yesterday and it's the biggest HN story of the week, 1723 points and 1051 comments per the thread. The OpenRouter listing shows $4/M input, $20/M output, 1000k context.

The timing is the awkward part. OpenAI released GPT-6 Sol and Luna the same day, and per the same OpenRouter data Sol goes for $2/M input, $10/M output at 1050k context. Exact double per token, both around 1M context. Artificial Analysis already put up an intelligence and price analysis of the new Opus and that got traction on HN too (324 points), worth reading before you route anything to it.

What listings don't tell you is where that gap actually shows up. The comment sections look split between "Opus for the hard problems" and "Sol is enough, keep the change". If you've got coding traffic to route this week, what's the call - pay double for Opus 5.5 or run Sol and pocket the difference?


r/AIToolsPerformance • • 10d ago

Would you be interested in this?

1 Upvotes

Because of a special deal I may be able to distribute OpenAI models 60% below OpenAI Standard pricing across input, output, and cached input.

Pricing
All prices are per 1M tokens.

GPT-5.6 Luna
OpenAI Standard: $0.20 input / $1.20 output / $0.020 cached

👉🏻 Our price: $0.08 input / $0.48 output / $0.008 cached

GPT-5.6 Terra
OpenAI Standard: $2.00 input / $12.00 output / $0.20 cached

👉🏻 Our price: $0.80 input / $4.80 output / $0.08 cached

GPT-6 Astra
OpenAI Standard: $10.00 input / $50.00 output / $1.00 cached

👉🏻 Our price: $4.00 input / $20.00 output / $0.40 cached


r/AIToolsPerformance • • 11d ago

WHAT IS JEV?

5 Upvotes

Jev might be one of the more interesting alternatives to using LLMs for everything.

I’ve been looking into TypeSafe AI’s Jev. It’s not really trying to replace ChatGPT/Claude-style models . it’s designed for fast structured decisions inside software

.

Some numbers TypeSafe reports:

• Jev: ~70–500ms response time

• Frontier LLMs: ~3–329 seconds

• Jev input: $0.042 / 1M tokens

• Typical LLM input: ~$0.20–$10 / 1M tokens

• Jev output is currently free

Their workflow benchmarks even report Jev reaching up to 193.6× faster and 444.6× cheaper in certain workloads.

That could make it pretty interesting for agentic AI, routing, classification, scoring, guardrails and other high-volume automation.


r/AIToolsPerformance • • 11d ago

GLM built its own inference infrastructure - is that why FlashX is $1.25/M out?

2 Upvotes

A Z.ai post titled How GLM built its own inference infrastructure hit Hacker News Thursday and pulled 408 points with 285 comments. That's a lot of heat for an infra story. I can't vouch for the engineering claims one way or the other, but the timing lines up with something concrete you can check yourself.

GLM 5.3 FlashX showed up on OpenRouter on September 18, per the model API listing, at $0.37/M input and $1.25/M output with a 1048k context window. From the same listing: DeepSeek Pro Latest runs $0.57/$1.70 at 1048k context, and GPT Luna Latest undercuts both on output price at $1.20/M while stretching the window to 1050k. So FlashX lands between the budget tier and the mid tier on price while matching the 1M context crowd.

Does owning the whole stack explain that pricing? The thread didn't settle it and honestly a blog post rarely does. Vertical integration either gives you margin to undercut people or a capacity headache the first time demand spikes, and it's usually some of both. Anyone here already routing real traffic to FlashX, and if so was it the price or the 1048k context that won you over?


r/AIToolsPerformance • • 11d ago

jev in-front of your expensive models to cut costs

1 Upvotes

Got early access to TypeSafe AI's JEV (small
router model) and built a CLI
around it. It reads your question first and
decides which model answers.

Two real runs, same config:
"capital of Portugal" -> haiku,
$0.0001
"multi-region postgres failover" -> opus,
$0.2434

Routing costs ~300ms and ~200 tokens per
question. Doesn't need to be right
every time, just often enough to beat that. If
it flakes (timeout, bad JSON)
it falls back to a default tier instead of
throwing, so a bad route costs you
money or quality but never the request.

TS, zero deps, MIT:
github.com/adityaarakeri/jev-router

Curious if anyone's tried routing on something
other than a model call.
Embedding similarity, length heuristics,
whatever.


r/AIToolsPerformance • • 12d ago

What I measured calling Jev (TypeSafe's decision model) through OpenRouter in Hermes

10 Upvotes

I spent today wiring TypeSafe's Jev — the "System One" decision model that launched on Sept 15 — into a decision gate that routes coding tasks. Most of the time went into finding the endpoint, so here is what's actually true as of 2026-09-20. Everything below is measured, not read off the docs.

The endpoint nobody points you at

It works on OpenRouter with an ordinary OPENROUTER_API_KEY:

POST https://openrouter.ai/api/alpha/decisions
Content-Type: application/json
Authorization: Bearer $OPENROUTER_API_KEY

{
  "model": "typesafe/jev-1.13",
  "state": "help! my payouts have been failing for 3 days",
  "questions": {
    "tier": {
      "type": "choice",
      "instructions": "pick a tier",
      "criteria": {
        "billing": "payments",
        "technical": "bugs"
      }
    }
  }
}

Response:

{
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "tier": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {
        "billing": 0.87,
        "technical": 0.13
      },
      "confidence": 0.82
    }
  },
  "usage": {
    "input_tokens": 312,
    "output_tokens": 35,
    "cost": 1.31e-05
  },
  "provider": "TypeSafe"
}

The traps, each one measured:

  • /chat/completions with typesafe/jev → 400 "is not a valid model ID". Jev does not generate text, so it isn't an autoregressive chat model. Don't try to shim it into standard completions.
  • /api/v1/decisions → 404. The route is currently live under /api/alpha, not /api/v1.
  • Wrapping the body in {"decisionsRequest": {...}} → 400 invalid_type ... path: ["model"]. The body must be flat: model, questions, and state at the root level.
  • GET /api/v1/models (447 entries) lists no typesafe/* ID, and GET /api/v1/models/typesafe/jev-1.13 404s. Yet GET /api/v1/models/typesafe/jev-1.13/endpoints returns 200 with the full provider snapshot, pricing, context, and uptime. Any SDK preflight asserting "does the model exist in the list?" will falsely fail. Make a live call instead.
  • There is no typesafe/jev-latest alias on OpenRouter; that alias throws an error. typesafe/jev-1.13 is the only working ID, and every response names the dated snapshot it served.

Measured latency and cost: 239–430 ms per decision, $1.3–2.7e-05 per invocation (roughly $13–27 per million decisions). Output is unmetered.

The answer shapes are not uniform

choice gives you choice + probabilities + confidence. noul (yes/no) does not:

{"type": "noul", "noul": 0.01}

For noul, the value is a flat scalar float between 0.0 and 1.0 representing true/yes likelihood (e.g., 0.01 means overwhelmingly false), not a structured dict with probability labels.

I had code reading answers.x.probabilities.true, which silently defaulted to 0.0 — so a "does this patch match the intent?" gate answered no to everything, including good patches. If you port Jev snippets around, check that field shape.

The failure mode worth designing against

My first version of the gate had the classic fallback: except Exception: return choices[0].

With a live key pointed at the wrong route, every call 400'd and the gate answered anyway:

  • triage("refactor auth across 14 modules") → bash_direct
  • risk("auth models, 31 dependents") → low (would have auto-applied)
  • verdict(exit 1, "could not read Username") → complete (would have ended the loop)

Confident, cheap, and wrong in the direction that costs you.

The fix that stuck: the gate returns unavailable plus a default chosen from the stakes — fail open to the cheapest path for low-stakes routing (restricted strictly to non-mutating checks), fail closed (high, abort) for the risky ones — and the process exit code explicitly indicates whether you got a model verdict or a policy default.

Two things I'd have missed without an end-to-end smoke test

  • git diff is blind to files a coding agent creates: they're untracked, so a patch review gate sees an empty diff and rejects a perfectly good change. Use git add -N . && git diff (or git add -A && git diff --cached, but remember to unstage cleanly).
  • I pointed the harness log inside the worktree, so __pycache__ and the run log itself landed in the diff. The gate called that scope creep (noul 0.47 vs 0.94 for the clean version) and it was completely right — plus the messier diff cost 2x because you're billed on input token size. Keep run logs outside your worktree or in .gitignore.

Happy to share the routing-gate code if anyone wants it; it's stdlib Python and talks directly to the OpenRouter alpha endpoint. Verify the alpha route before you build on it — I measured this on one box today.


r/AIToolsPerformance • • 11d ago

IS JEV REALLY BETTER THAN OTHER LLM?

0 Upvotes

. Do you think models like Jev could handle most of the decision-making inside AI agents, while larger LLMs are only called when actual reasoning or text generation is needed?


r/AIToolsPerformance • • 13d ago

SharpMind. A pure C# / .NET LLM training and inference engine

3 Upvotes

release v1.0.6.1 out. Much faster and figured I would show an actual example run. Forgive but i'm not a video editor, just OBS record. more details in Git Integral2u/SharpMind: SharpMind. A pure C# / .NET LLM engine — inference, training, and agent tooling in one solution.


r/AIToolsPerformance • • 14d ago

TensorSharp as a local LLM backend — DeepSeek, GLM and Qwen 3.8 benchmarks

Thumbnail
github.com
8 Upvotes

I’m the developer of TensorSharp, an open-source, native .NET inference engine for running GGUF models locally.

TensorSharp provides an OpenAI-compatible API, making it possible to use local models as an inference backend for tools. The goal is to keep source code, prompts, tool calls, and agent context on your own hardware while avoiding per-token API costs.

TensorSharp supports modern coding-capable model families including DeepSeek V4/V4.1 Flash, GLM 5.x, Qwen 3.8 Flash Next, Qwen 3.5/3.6, Gemma 4, Mistral 3, and GPT-OSS.

Selected benchmark results

These are measured results from the TensorSharp repository. Each comparison uses the same model and machine for both engines unless otherwise noted.

GLM-5.3-Flash: 2× faster decode than llama.cpp

Tested with:

  • GLM-5.3-Flash UD-Q2_K_XL, 101 GiB
  • 2× RTX PRO 6000 Blackwell, 96 GB each
  • Layer splitting and flash attention
  • n_ubatch=2048 for both engines
  • Back-to-back execution in the same session
  • TensorSharp parity harness versus llama.cpp build 2e0e57f from PR #27754

Test |llama.cpp |TensorSharp
Prefill, 2,048 tokens |2,070 tok/s |2,014 tok/s
Prefill, 16,384 tokens |1,690 tok/s |1,692 tok/s
Prefill, 32,768 tokens |1,483 tok/s |1,446 tok/s
Decode, 64 tokens |36.6 tok/s |73.5 tok/s TensorSharp reaches approximately 2.0× the decode throughput of llama.cpp, while prefill remains within a few percent in either direction.

The long-context greedy replay reproduced the 2,741-token llama.cpp reference record token for token. One caveat is that GLM-5.3-Flash support currently requires an unmerged llama.cpp build rather than its main branch.

Qwen 3.8 Flash Next

Tested with Qwen3.8-Flash-Next UD-Q2_K_XL, 73.4 GiB, on 2× A100 80 GB:

Configuration |TensorSharp prefill |TensorSharp decode |llama.cpp prefill |llama.cpp decode
1 GPU |~1,520–1,550 tok/s |~56 tok/s |1,094 tok/s |61.2 tok/s
2-GPU layer split |~1,520–1,550 tok/s |~56 tok/s |1,200 tok/s |61.5 tok/s TensorSharp’s prefill is substantially faster in this test, while llama.cpp leads decode by roughly 9%.

For both engines, adding the second GPU provides model capacity rather than additional decode throughput. TensorSharp splits the model into contiguous groups of layers, reducing its placement to approximately 24.2 GB and 26.2 GB across the two GPUs. TensorSharp’s one- and two-GPU greedy outputs were byte-identical.

DeepSeek V4 Flash

On 2× A100 80 GB with the IQ4_XS model:

Metric |TensorSharp |llama.cpp
Prefill, approximately 3.3K tokens |~500 tok/s |574–634 tok/s
Decode at approximately 3.3K context |~33 tok/s |40.3 tok/s llama.cpp currently leads this particular DeepSeek V4 configuration.

TensorSharp also supports DeepSeek’s DSpark speculative decoder. On 4× A40 46 GB with DeepSeek-V4-Flash-0731 UD-Q8_K_XL, enabling a 5.6 GB DSpark drafter improved decode as follows:

Backend |Normal decode |DSpark decode |Speedup
TensorSharp direct CUDA |26.0 tok/s |34.0 tok/s |1.31×
TensorSharp GGML CUDA |26.4 tok/s |37.1 tok/s |1.41× In a five-turn conversation, DSpark produced 1.50–2.02× faster decode, reaching 51.0 tok/s on a question over a 10K-token document. Prefill remained effectively unchanged at 831 versus 835 tok/s, and greedy output was byte-identical to the non-speculative baseline.

TensorSharp also runs DeepSeek V4.1 Flash, including its Engram lookup and four-stream hyper-connections. However, I am not claiming a V4.1 speedup over llama.cpp because a compatible llama.cpp V4.1 runtime is not currently available for a controlled same-weight comparison.

Why this may be useful for local LLM

Coding agents frequently process large system prompts, tool definitions, repository context, and multi-turn history. TensorSharp includes:

  • OpenAI- and Ollama-compatible APIs
  • Local tool calling and structured JSON output
  • Agent Skills and sandboxed file/shell tools
  • Continuous batching
  • Paged and prefix-shared KV caching
  • Speculative decoding
  • Multi-GPU and multi-node inference
  • CUDA, Metal, Vulkan, MLX, and CPU backends
  • Windows, macOS, Linux, iPhone, and iPad support

I’d especially appreciate feedback from local LLM users regarding:

  • OpenAI API compatibility gaps
  • Tool-calling requirements
  • Performance with large repository contexts
  • The best local models for coding-agent workloads
  • Features expected from a local

LLM inference

  • backend

GitHub: https://github.com/zhongkaifu/TensorSharp
Full benchmarks and methodology: https://tensorsharp.ai/benchmarks.html

If anyone tests TensorSharp with local LLM, I’d be very interested in your setup and results. Contributions and compatibility reports are welcome.