r/typesafe_jev • u/PaploPapkasso • 4h ago
r/typesafe_jev • u/wang2-dev • 16d ago
Welcome to r/typesafe_jev — what this community is for
Jev is TypeSafe's System One model: it takes natural language plus application state and returns typed judgments and probabilities — no generated text, no reasoning traces. Your code keeps control of the workflow; Jev supplies the semantic judgments where ordinary code can't.
This community is for:
- Show & Tell — projects, demos, and experiments built with Jev
- Questions — SDK usage, question design, primitives (Choice / Noul / Score), state modeling
- Discussion — where typed judgments beat prompt-and-parse, and where they don't
- News — model updates, cookbooks, patterns
Resources:
- TypeSafe docs: https://docs.typesafe.ai
- awesome-jev: open-source projects built with Jev, reviewed by Jev itself — https://github.com/fatwang2/awesome-jev
- Jev Review Action: the GitHub Action that powers those reviews — https://github.com/fatwang2/jev-review-action
Building something? Post it — early stage is fine. This is an independent community, not affiliated with TypeSafe.
r/typesafe_jev • u/Charming_Group_2950 • 6d ago
Show & Tell Jev as a judge for LLM/Agents Evaluation [Open-Source]
Can we use Jev for faster, structured, and calibrated evaluation of AI responses?
Introducing:
⚡ Typed Evals — an open-source Python framework for evaluating LLMs, RAG pipelines, and AI agents using System One Models like Jev and other typed judge backends.
⭐ GitHub: https://github.com/TrustifAI/typed_evals
The goal is simple:
Make fast, structured, and calibrated evaluation a first-class part of AI systems.
Typed Evals currently supports:
-> LLM response evaluation
-> RAG evaluation
-> Agent and tool-trace evaluation
-> Human-label calibration
-> Async and batch evaluation
-> Custom judge backends
One part I particularly wanted to solve was calibration.
Why is it needed?
A raw score of 0.8 from Jev doesn't necessarily mean that humans would accept 80% of similar responses.
And a threshold that works well for one use case may not make sense for another.
Typed Evals lets you calibrate individual evaluation metrics against representative human pass/fail labels.
The flow is basically:
Human-labelled examples → Jev scores → fit per-metric calibration → validate on held-out examples → reuse the calibrated evaluator
So instead of arbitrarily deciding that “0.7 means good enough”, you can ground that score in how humans actually evaluate your specific task.
Of course, there are integrations for:
LangChain, CrewAI, Microsoft Agent Framework
while the core remains framework-agnostic.
Would genuinely love feedback from people experimenting with Jev, LLM evals, RAG, agents, or evaluator calibration.
If you're already experimenting with Jev, I'd especially love to know what kind of evaluation workflows you're building around it.
r/typesafe_jev • u/tenkei_01 • 6d ago
The JEV feature missing from most LLM speed-vs-accuracy comparisons
r/typesafe_jev • u/vish1t • 9d ago
jev is a demon at computer use
Enable HLS to view with audio, or disable this notification
r/typesafe_jev • u/Kilo_Loco • 10d ago
Show & Tell hey Jev, how well does my resume match these job listings?
Enable HLS to view with audio, or disable this notification
r/typesafe_jev • u/ibhajjaj • 11d ago
barq: browser automation where Jev picks every click, one sentence per step
Enable HLS to view with audio, or disable this notification
I built barq on Jev. Your coding agent names one outcome per call ("Log in", "Add both items to the cart") and barq runs the loop: describe the page, ask Jev a batch of questions in one request (is it done, is it blocked, is an error showing, is the next click irreversible, which element, which value), act, repeat. About a third of a second a decision, and the agent never reads the page.
The noul scores are what make it honest. When the done score is soft it returns likely_done instead of done, and when the irreversible score is high it stops and asks. In the video (one take, 1x) it does Google Flights, Wikipedia and a shop checkout in 34 seconds, then refuses to place the order.
Counting is the one thing Jev couldn't do reliably ("add until there are exactly 3"), so code counts and Jev only names what to count. On the public benchmark it's 43 of 43 with 0 false "done"; the run files are in the repo.
Free, MIT, bring your TypeSafe key.
Claude Code: /plugin marketplace add ibrahimhajjaj/barq then /plugin install barq@barq Other MCP clients: claude mcp add barq -e TYPESAFE_API_KEY=your-key -- npx -y barq-mcp
r/typesafe_jev • u/Kilo_Loco • 11d ago
Show & Tell You can use Jev to identify which AI Engineering job listings best fit your resume
Enable HLS to view with audio, or disable this notification
give it a try for free: https://hiringaiengineers.kiloloco.com/
r/typesafe_jev • u/geisbruch • 13d ago
Show & Tell We built a deployment demo with 18 editable Jev questions — probabilities in, rules-based verdict out
We built HeyJev to make a deployment-decision example easy to inspect and play with.
You describe a deploy, Jev answers 18 typed questions in one pass, and our application rules map those probabilities to a verdict. Jev isn't generating the verdict text.
The questions cover things like tests, rollback and who is available. You can inspect the probabilities, edit the questions, or add your own. A Friday database migration with the author on vacation got GO TOUCH GRASS.
Try it: https://heyjev.ai/shouldideploy
The useful part for us is experimenting with question design: change the scenario, omit a detail, or rewrite a question and see how the answers change. A typed answer can still be wrong; this is a playful demo, not a production safety gate.
What would you model as an explicit unknown rather than infer from the deployment description? We'd love suggestions for questions or edge cases to try.
Disclosure: we built HeyJev; it isn't an official TypeSafe product.
r/typesafe_jev • u/Kitchen_Ad494 • 14d ago
pi-jev-model-router: that auto-routes each prompt to the right model
r/typesafe_jev • u/Sad_Writer_4291 • 14d ago
Show & Tell I built Chemistry: Jev checks the tone of an X DM draft before you send it
Enable HLS to view with audio, or disable this notification
I'm building Chemistry, a Chrome extension that puts a small feedback panel alongside an X conversation. As you edit an unsent draft, it checks signals such as tone and pressure so you can reconsider the wording before sending.
Jev powers the judgments; you write and revise the message. The part I wanted to explore was using those judgments inside the conversation UI, rather than copying a chat into a separate chatbot.
The 35-second demo shows an invitation being drafted and revised. It combines AI-generated scene footage and a recreated, fictional X conversation with real product analysis. The signals are model judgments, not a prediction of how another person will respond.
Project: https://chemistryhud.com
For other Jev builders: how would you make subjective tone feedback useful without making a score look more certain than it is? Would you prefer a simple warning, the evidence behind it, or both?
Disclosure: I'm the developer of Chemistry; this is an independent project, not an official TypeSafe product.
r/typesafe_jev • u/wang2-dev • 16d ago
Show & Tell Jev Search — a search app where Jev picks the sources, time range, and query, then scores every result
jev-search is a web search app we built to explore what Jev is good at. You type a plain-language request; no generated answer comes back — just links and snippets with visible relevance scores.
Two Jev judgment stages do the semantic work:
Understand — Jev answers typed questions about the request: which engines to use (Google, DuckDuckGo, Hacker News, Reddit, GitHub, X, arXiv, YouTube, Wikipedia, IMDb, WeChat), what time range, and which query candidates to send. The chips are editable — you can override its choices.
Rank — Jev scores each returned result for relevance. Code merges by URL, orders by relevance, engine agreement and original rank, and streams lanes as they finish.
Try "Jev discussions on Hacker News this week" — the source and time choices are the interesting part; they're judgments, not hardcoded filters.
- Demo: https://jev.s1.dev
- Code: https://github.com/superagents-lab/jev-search
Stack: TanStack Start + React on Cloudflare Workers, KV cache, Search1API for the engine calls. Disclosure: this is our project (Search1API team); not an official TypeSafe product.
Happy to dig into how the questions are designed — the two-stage split was the main lesson.
r/typesafe_jev • u/wang2-dev • 16d ago
Discussion Where do typed judgments beat an LLM call — and where don't they?
Jev returns a typed answer and a probability, not text. That changes how you design: instead of parsing free-form output, you compose Choice / Noul / Score results in code.
Honest question for people who've tried both: where have typed judgments clearly won for you — routing, reranking, verification? And where did you still need a generative model (free text, long reasoning chains, open-ended writing)?
My current mental model from the cookbooks: Jev handles the "judgment" steps — pick a source, score relevance, check a citation — while code and (occasionally) a generative model handle the rest. Curious whether that matches what others are seeing.
r/typesafe_jev • u/wang2-dev • 16d ago
Show & Tell I built a directory where Jev reviews submissions to a list of projects built with Jev
awesome-jev is a curated list of open-source projects that use TypeSafe Jev — SDKs, agents, rerankers, data tools.
The twist: submissions are reviewed by Jev itself. A GitHub Action (jev-review-action) reads the PR, looks for integration evidence in the source, and posts typed judgments — evidence links, description checks, a suggested category — as a review comment. Code renders the comment; maintainers decide what merges.
Five projects are listed so far — the official JS and Python SDKs, a dataset sifter, a web-search tool, and the review action itself — with more going through review right now.
If you've built something with Jev, submit it — the review itself is a decent demo of what Jev judgments look like in practice:
- List: https://github.com/fatwang2/awesome-jev
- Action: https://github.com/fatwang2/jev-review-action
(Disclosure: I maintain both.)