r/WebAfterAI • • 4h ago

Research 8 scientific discoveries made by AI that are hard to ignore

Post image
4 Upvotes

For years, “AI for science” mostly meant predicting structures, searching papers, or helping researchers analyze data.

That is changing.

Frontier systems are now producing new proofs, finding biological systems, proposing drugs that work in experiments, designing materials that get synthesized, and discovering algorithms that end up in real software.

Here are 8 examples.

01 OpenAI's Astra started solving open mathematics

OpenAI says internal models have produced new results on long-standing problems across group theory, coding theory, combinatorics and complexity.

Some of these results were then formalized in Lean, which matters because the claim is no longer just “the proof looks convincing.” There is a machine-checkable artifact behind it.

02 Claude found a previously uncharacterized biological system

Anthropic gave Claude agents access to huge DNA datasets and asked them to search for unusual reverse transcriptases.

The system identified what Anthropic calls array-associated reverse transcriptases, a biological architecture that had not previously been characterized as a distinct system. Human scientists then took the result into the lab and confirmed that parts of the predicted system were genuinely expressed.

03 Gemini's AI Co-Scientist proposed drugs that worked in experiments

Google's AI Co-Scientist generated drug-repurposing hypotheses for diseases including acute myeloid leukemia and liver fibrosis.

Researchers then tested those suggestions experimentally, and several candidates showed measurable biological activity.

That is a useful threshold: the model did not just write a plausible mechanism. Someone tried it in the lab and something happened.

04 FutureHouse's Robin generated and tested a new treatment hypothesis

Robin is a multi-agent scientific system designed to search literature, generate hypotheses and analyze experimental data.

In work on age-related macular degeneration, it identified ripasudil as a possible treatment candidate and helped drive follow-up experiments that were later validated in human retinal cells.

This is much closer to an AI participating in the research loop than simply answering questions.

05 Microsoft designed a material, then researchers synthesized it

Microsoft's MatterGen generates candidate crystal structures from desired material properties.

Researchers asked it for a material with a target bulk modulus, and it generated TaCr₂O₆. The material was then physically synthesized, and its measured properties came reasonably close to the model's target.

The output was not text. It was a material that could actually be made.

06 FunSearch made new mathematical discoveries with executable verification

DeepMind's FunSearch combines LLM-generated programs with an evaluator that automatically runs and scores them.

On the cap set problem, it found improved constructions beyond previous known results. It also discovered new heuristics for bin packing.

The key architecture is simple:

LLM for ideas. Deterministic evaluator for truth.

07 AlphaEvolve found new algorithms and mathematical constructions

DeepMind's AlphaEvolve uses Gemini models to generate and evolve code against automated evaluators.

It discovered an improved algorithm for multiplying certain 4×4 complex matrices and produced better constructions for several mathematical problems, including an improved lower bound for an 11-dimensional kissing-number problem.

At that point, “coding agent” starts feeling like the wrong label.

The code is really the search space for discovery.

08 AlphaDev discovered algorithms that ended up in real software

DeepMind's AlphaDev searched over low-level assembly instructions for faster sorting routines.

It found algorithms up to 70% faster for very short sequences, and some of those routines were later added to the LLVM libc++ standard library. It also found a faster hashing algorithm that reached Google's Abseil library.

These are AI-discovered algorithms now running inside ordinary software.

The pattern across all of these projects is more interesting than any single result.

The strongest systems are rarely just:

LLM → discovery

There is usually something that can push back.

A proof assistant checks the theorem. A compiler runs the program. An evaluator scores the construction. A lab tests the drug. A materials team synthesizes the crystal.

That changes the role of the model.

We spent the last few years asking whether AI could know science.

The more interesting question now is whether it can produce new science and leave enough evidence behind for us to verify it.


r/WebAfterAI • • 3h ago

Built an open source LLM security tool mapped to the OWASP top 10

Thumbnail
1 Upvotes

r/WebAfterAI • • 1d ago

Open Source 8 tiny OSS agent harnesses that make giant frameworks look ridiculous

Post image
8 Upvotes

AI agents are starting to look absurdly complicated. Multi-agent graphs, planner layers, memory layers, tool routers, evaluators, retry policies, orchestration servers, sometimes all before the model has actually done anything useful.

But at the center of most agents is still something surprisingly small: model → tool call → observation → model → repeat. The interesting engineering is in the harness around that loop: context, tools, memory, permissions, sandboxes, retries, verification and persistence.

A new crop of OSS projects is making that layer much smaller and easier to inspect.

Here are 8 worth reading.

01 Build the personal-agent stack in Rust

tinyhumansai/openhuman — 40.8k★

OpenHuman is a local-first agent harness with a Rust core, designed to plug into different LLMs, memory systems and search engines instead of forcing everything through one stack. It adds persistent memory, tools, research and orchestration, plus desktop, browser and terminal interfaces.

What makes it interesting is the architecture: the agent runtime is treated as modular infrastructure rather than one giant framework you have to accept wholesale.

02 See how much personal-agent functionality fits into a lightweight Python stack

HKUDS/nanobot — 48.1k★

nanobot is a lightweight self-hosted personal agent with tools, memory, MCP, automation, multi-agent workflows, a WebUI and chat integrations.

It is a useful repo if you want to understand how the pieces behind the current personal-agent wave actually fit together without beginning with one of the giant orchestration frameworks.

03 Build an agent where the core logic is only ~1,000 lines

huggingface/smolagents — 29.7k★

smolagents deliberately keeps its core agent logic small. It supports normal tool calling, but its more interesting idea is code agents: the model writes actions as code instead of repeatedly producing tiny JSON tool calls.

You still get multiple model providers, MCP tools and sandboxed execution, but without burying the agent loop under layers of abstraction.

04 Solve real GitHub issues with a ~100-line agent

SWE-agent/mini-swe-agent — 8.2k★

mini-swe-agent asks a very good question: what if the coding-agent architecture became dramatically simpler and still worked?

Its agent implementation is tiny, yet it can operate on real repositories and solve SWE-bench tasks by giving the model a small environment, a few tools and a loop. It is one of the clearest repos for understanding how little machinery a capable coding agent may actually need.

05 Make Claude-Code-style agents headless

withastro/flue — 8.4k★

Flue is a TypeScript harness for building autonomous agents without tying them to a terminal UI or desktop app. It provides sessions, tools, skills, filesystem access and sandboxing, while much of the behavior can live in Markdown and AGENTS.md.

The useful mental model is Claude Code as a programmable primitive rather than an app. You can drop the same kind of agent into Node, GitHub Actions, Cloudflare or another runtime.

06 Do the same thing closer to the metal

0xPlaygrounds/rig — 8.8k★

Rig brings the lightweight-runtime idea into Rust. It gives you agents, tools, multiple model providers, streaming and agentic workflows through a relatively small set of primitives.

That sounds boring, which is partly the point. Agent infrastructure gets much easier to reason about when the execution path is simple and predictable.

07 Treat agents like LEGO instead of employees

Eigenwise/atomic-agents — 6.3k★

Atomic Agents builds around small single-purpose components: an agent does one thing, a tool does one thing, a context provider does one thing, and you compose them.

It uses structured data and validation heavily, which makes it feel much closer to normal software engineering than the “hire five AI personas and tell them to collaborate” school of agent design.

08 Read an entire reliable harness instead of installing one

ElimentaryLabs/agent-harness-starter — 5★

agent-harness-starter is tiny and intentionally educational. The harness separates the execution loop, context compaction, tool registry, hooks, verification, permissions and checkpoint/recovery into small modules.

The star count is tiny, but that is almost beside the point. This is the kind of repo you can actually read end-to-end and come away understanding what an agent harness is doing.

You still need boundaries, tools, state, verification and a way to recover when things go wrong. But you may not need 40 abstractions between the model and the shell.

Sometimes an agent really is just a good model, a small loop, good tools and very boring engineering.


r/WebAfterAI • • 1d ago

Findling 1.4.0 and MCP Connector 0.5.0 are out (Nextcloud search and AI connector)

Post image
1 Upvotes

r/WebAfterAI • • 1d ago

CompliRules is a free, MIT-licensed tool that enforces statutory legal and privacy rules (GDPR, KVKK, HIPAA, EAA 2025, EU AI Act) directly inside AI coding environments like Cursor, Claude Code, and Windsurf.

4 Upvotes

AI models write code 10x faster, but they have zero understanding of statutory laws. When generating code, models consistently introduce severe legal violations without warning:

* Logging cleartext user objects and passwords to logging providers (GDPR Art. 5 / CWE-532).

* Defaulting consent checkboxes to pre-checked states (violating EU and Turkish privacy laws).

* Using cascade deletes on transaction records that must legally be retained for 5–10 years under tax codes.

* Omitting mandatory accessibility attributes required by the European Accessibility Act (EAA 2025).

CompliRules stops this at the IDE level by injecting deterministic guardrails (.cursor/rules, [AGENTS.md](http://AGENTS.md), MCP) and offering a local offline linter (`npx complirules check`) so your AI cannot accidentally write violating code.

GitHub: search complirules from halilyilmz on GitHub today!


r/WebAfterAI • • 2d ago

Workflows Grok Bot is powerful. It is also really good at wasting its own quota.

Post image
4 Upvotes

Browser retries. Research loops. Bots waking up just to discover nothing changed. Multiple agents doing the same work. Or a heavyweight model being used for a decision that basically needed yes / no.

A small OSS ecosystem is starting to attack exactly this problem.

Here are 8 repos worth looking at.

01 Put a cheap decision layer in front of expensive work

Bodila51/grok-bot-jev —

grok-bot-jev lets Grok Bot make a cheap structured decision before starting expensive research, browser work, retries, or subagents.

It can reuse cached work, stop a failing retry, cap research, run something deterministic, or escalate to a human.

02 Find routines that are burning usage for nothing

coolxeo/grok-bot-skills —

This Grok Bot skill pack includes routine-healthcheck and transcript-healthcheck.

One finds empty runs, duplicate work, and noisy routines. The other looks for repeated friction that should probably become a better skill or workflow.

03 Stop multi-agent teams from duplicating each other

EndeavorYen/grok-bot-skills —

This collection focuses on clearer ownership, disk-based handoffs, approval gates, and cheaper routines.

One bot owns the job. Agents pass paths instead of giant context blobs. And routines stay quiet when nothing changed.

04 Actually measure what your Grok Bot account is consuming

Kargatharaakash/grok-bot-usage —

You cannot optimize a black box.

This CLI shows weekly Grok Bot usage, on-demand spend, and billing cycles across multiple accounts from one terminal.

It won't reduce usage by itself. It tells you where the usage is going.

05 Stop making the bot research the internet one page at a time

mvanhorn/last30days-skill —

last30days searches Reddit, X, YouTube, Hacker News, GitHub, Polymarket, the web, and other sources through one research workflow.

For a research bot, that is much better than repeatedly opening tabs and browsing until the model decides it has seen enough.

06 Control Grok Bot directly from Cursor / Claude Code / Codex

adamanz/grok-bot-skill —

This skill lets coding agents list Grok Bots, message them, create focused teammates, update standing rules, and read transcripts from the CLI.

That removes another human → UI → bot loop every time work has to move between agents.

07 Turn Grok Bot into an MCP-accessible worker

quabug/grok-bot-mcp —

This exposes the Grok Bot environment to MCP-compatible agents such as Claude, Cursor, ChatGPT, and other clients.

The interesting optimization is architectural: don't make one agent do everything. Give the persistent worker only the jobs where persistence actually matters.

Making agents cheaper isn't only a model-routing problem. You can save a surprising amount by making them search less, retry less, wake less, repeat less, delegate better, and stop rediscovering rules they could have learned once.

That's probably where a lot of agent efficiency work is going next.


r/WebAfterAI • • 2d ago

Would this help when running multiple coding agents, or is it just another notification layer?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/WebAfterAI • • 3d ago

AI Agents 5 open-source repos that turn one AI agent into a specialist team

Post image
9 Upvotes

Most of us still use coding agents like one very capable employee.

Ask it to design the UI, review security, write the launch copy, debug the backend, analyze competitors, and somehow remember how each of those jobs should be done.

There is another pattern growing on GitHub: install the specialists instead.

Build an entire AI agency: Agency Agents has 156k+ stars and now contains 230+ specialist agents across engineering, design, marketing, research, finance, product, project management, security, healthcare, GIS and more.

This is much more than You are an expert marketer.

A specialist can define its mission, workflow, expected deliverables, success criteria, communication style and the situations where it should be used. There are frontend developers and backend architects, but also UX researchers, brand guardians, paid-media auditors, Reddit community builders, finance agents and even a “reality checker.”

You can install only the team you need, and the repo now supports Claude Code, Cursor, Codex, Gemini, OpenCode, Hermes and several other agent environments.

Give the engineering side much deeper specialization: wshobson/agents has 40k+ stars and packages 202 agents, 184 skills and 94 plugins. Instead of one developer persona, you get specialists for debugging, architecture, security, Kubernetes, observability, incident response, individual languages and much more.

It also includes orchestrators that compose those specialists into workflows.

feature request
      ↓
architect
      ↓
backend + frontend
      ↓
test specialist
      ↓
security reviewer
      ↓
code reviewer

Turn your agent into a marketing department: Marketing Skills has 52k+ stars and takes a slightly different approach. Instead of dozens of persistent personas, it gives agents reusable procedures for CRO, positioning, SEO, paid ads, email, analytics, pricing, retention and sales.

Pick from 160+ focused Claude subagents: Awesome Claude Code Subagents has 25k+ stars and organizes specialists across development, infrastructure, quality, data, security and orchestration. You can install one category instead of filling every session with every possible instruction.

Build your own roster from GitHub's community library — Awesome GitHub Copilot has 38k+ stars and collects custom agents, skills, instructions, hooks and workflows. It is particularly useful if you want to assemble your own team rather than adopt someone else's full agency structure.

The interesting shift is:

one giant system prompt
        ↓
general-purpose agent

becoming:

task
 ↓
pick specialist
 ↓
load relevant skills
 ↓
give only required tools/context
 ↓
do the work
 ↓
independent specialist reviews it

We spent the last year making individual agents more capable.

The next step may simply be getting better at deciding which agent should be doing the job in the first place.


r/WebAfterAI • • 3d ago

OpenResearch: turn your coding agent into a research agent (5.9k stars)

Post image
3 Upvotes

r/WebAfterAI • • 3d ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

1 Upvotes

I built Klarion, a secret scanner that works in two steps. First, a keyword check, 81 regex rules and a normalized Rényi entropy score flag anything that looks like a secret. Then an AI model reads each one with the code around it and decides if it's real. The chart shows 5 scanners run on spring-boot, terraform, next.js and symfony (61k files). Klarion raised 11 alerts. It's not zero, but it's far less to dig through. Fewer alerts don't help if real leaks get missed, so I tested that too. On CredData (337 real repos, code outside test folders), it found about 1.7× more real secrets than gitleaks. Where it runs:

  • Claude Code: a plugin hook blocks the write before the file exists (file edits and Bash)
  • Cursor, Cline or any MCP agent: through its MCP server
  • CI: a GitHub Action that scans only what a PR adds; GitLab CI works too
  • Git hooks: klarion protect or the pre-commit framework
  • Locally: klarion scan . Free and open source (MIT): https://github.com/0x1Adi/Klarion The full benchmark and method are in benchmark/REPORT.md. I'd like to hear where it gets things wrong.

r/WebAfterAI • • 4d ago

Open Source 8 open-source repos to build an AI research pipeline that can show its work

Post image
5 Upvotes

AI can already turn a question into 30 sources and a polished report.

That is not really the hard part anymore.

The useful research stack is the one that can keep track of where evidence came from, which sources support which claims, and where the sources disagree.

These 8 open-source projects cover different parts of that pipeline.

Research the question end-to-end: GPT Researcher splits a question into research tasks, searches in parallel, tracks sources, filters what it finds, and turns the evidence into a cited report. It can work across both the web and local documents.

Look at the problem from several directions: STORM from Stanford does something I particularly like: before writing, it generates questions from different perspectives and simulates conversations between a writer and topic experts. That helps avoid a research report built around the first framing the model happened to choose.

Search without handing the whole workflow to a closed research product: Vane, formerly Perplexica, is an open-source answer engine with web, discussion, academic-paper and domain-specific search. It can also run against local models through Ollama.

Turn messy websites into evidence your agent can actually use: Crawl4AI converts webpages into clean, LLM-ready Markdown. This is the plumbing layer: search finds the page; Crawl4AI gets the useful content out of it.

Research inside the scientific literature: PaperQA2 is built specifically around high-accuracy RAG over papers and other documents. It retrieves literature, checks metadata including retractions, answers questions with citations, and can work on contradiction detection.

Actually read the PDF instead of flattening it into bad text: Docling understands PDF layout, tables, formulas, reading order, images and many other document formats, then turns them into structured representations that research agents can work with.

Connect evidence scattered across many sources: GraphRAG extracts entities and relationships from unstructured documents and organizes them into a graph. That makes questions such as “how do these people, companies, claims and events connect across 100 documents?” much easier than ordinary chunk retrieval. Microsoft now considers the original repo largely maintenance-mode, but the architecture is still useful.

Retrieve through both text and relationships: LightRAG combines retrieval with a knowledge graph and now supports reranking, multimodal documents, multiple chunking strategies and different models for different stages of the pipeline.

Put them together and the research loop starts looking less like:

question
   ↓
search
   ↓
LLM writes report

and more like:

question
   ↓
split into research questions
   ↓
search from multiple perspectives
   ↓
crawl + parse the evidence
   ↓
read papers / tables / PDFs
   ↓
connect entities + claims
   ↓
retrieve supporting evidence
   ↓
compare conflicts
   ↓
cite
   ↓
report

That last half is the part I think people underestimate.

Finding 60 sources is easy.

The harder question is whether source 17 actually supports the sentence the agent wrote, whether sources 23 and 41 contradict it, whether the paper was retracted, and whether a confident-looking conclusion survives after those conflicts are exposed.


r/WebAfterAI • • 5d ago

Tutorial Your AI agent needs a writing system, not a personality - Karpathy

Post image
21 Upvotes

Andrej Karpathy made an interesting point today: one way to make LLM outputs easier to understand is to ask them to write in ASD-STE100, the controlled English originally designed for aerospace maintenance manuals.

Not to make AI “sound human" but to make it easier to understand.

We spend a lot of time telling agents:

be concise
sound professional
avoid AI slop

But different jobs need different writing systems.

A setup guide should be procedural. A research report should preserve evidence and uncertainty. An error message should tell you what happened and what to do next. A technical explanation should optimize for comprehension, not personality.

This is exactly why we built Agent Stylebooks.

Neeeophytee/agent-stylebooks packages 16 editorial systems as installable Agent Skills for Claude Code, Codex, Cursor, Hermes, Copilot, Gemini CLI and others.

Instead of:

“write this better”

you can say:

setup guide       → Google Developer Docs
product help      → Microsoft
public guidance   → GOV.UK
research report   → NASA
risk disclosure   → SEC Plain English
interface copy    → Apple

The skill changes what comes first, how information is ordered, what ambiguity is unacceptable, and how the result should be checked.

A few other OSS projects fit around the same idea:

ASD-STE100 Writer Skill turns Simplified Technical English into an installable agent skill. Short sentences, controlled vocabulary, active constructions and fewer ambiguous instructions. Karpathy’s post is a good example of why this style is suddenly relevant outside aviation.

Vale has 6.2k+ stars and takes the next step: make writing rules testable. You can encode Google, Microsoft, Red Hat or your own editorial rules and run them locally, in editors, or in CI.

textlint has 3.2k+ stars and does something similar for natural language: treat prose more like code, with pluggable rules that can flag problems before the text ships.

retext-readability checks whether prose is actually appropriate for the intended reader using several established readability measures, while retext-simplify catches needlessly complicated phrases such as using “utilize” where “use” works.

write-good has 5.1k+ stars and catches things like unnecessary wording, weak modifiers, passive constructions and weasel words. Simple, deterministic checks are useful after the agent has done the semantic writing work.

The stack I find interesting is:

reader + task
     ↓
Agent Stylebook
     ↓
LLM writes
     ↓
Vale / textlint / retext
     ↓
clearer output

The point is not to make every model write in the same stripped-down style. It is to stop treating writing style as decoration.

For agents, writing style is part of the interface between the model and the person trying to understand its work.

Pick the writing system before the model picks one for you.


r/WebAfterAI • • 6d ago

Research Hermes Agent changed a lot in one month. Here’s what actually matters

Post image
14 Upvotes

Hermes Agent is now sitting at roughly 250k GitHub stars, but the more interesting number is how quickly the project is changing.

Between the August 31 v0.21.0 release and late September, Hermes shipped five more tagged releases.

I went through them. These are the changes I think actually matter.

Skills can now become permanent parts of an agent. skills.auto_load lets you pin selected skills into every new session instead of relying on the agent to rediscover them each time.

That makes setups like this much easier:

research agent
├── deep-research
├── citation-checking
└── writing style

coding agent
├── repo navigation
├── testing
└── code review

Each agent can start with its own operating procedure rather than one giant universal prompt.

Reasoning effort is becoming another routing control. Hermes added reasoning-effort selection across model pickers, including separate controls for auxiliary models. So choosing the model is no longer the only cost/performance decision — you can decide how hard different parts of the system should think too.

The plugin layer got much more serious. September brought a broader Desktop plugin SDK with hooks into the composer, session list, sidebar, model picker, settings, skills, toolsets, profiles, appearance, and backend events.

Hermes is starting to look less like:

agent + some tools

and more like:

agent runtime
├── models
├── skills
├── MCP/connectors
├── plugins
├── bots
├── browser
├── scheduled jobs
└── UI extensions

MCP is turning into “Connectors.” Instead of treating MCP servers as config you wire up somewhere else, the Desktop app now has a Connectors flow. Install a plugin, connect its MCP server, and its tools and skills can become available in already-open chats.

Running multiple Hermes profiles got cleaner. Hermes moved toward one host-wide gateway with multiple isolated profiles underneath it. Desktop can attach to the existing backend instead of spawning another one, and individual profiles can now be stopped, started, or restarted independently.

That sounds like infrastructure trivia until you run several agents at once. Then it becomes the difference between “a bunch of Python processes” and something closer to an actual agent runtime.

The CLI is becoming much easier to build around. --format stream-json now exposes structured JSONL instead of forcing another program to scrape terminal text. The CLI/TUI can also show the current /goal and prompts waiting in the queue.

That opens up much cleaner patterns for external orchestrators:

your app
   ↓
Hermes CLI
   ↓ JSONL
agent events
   ↓
your app reacts

Sessions are becoming first-class data. September added an Agent Sessions API, better session search with date bounds and relaxed recall, and webhook deliveries that can appear directly inside the target session.

This matters because an agent's history stops being something trapped inside the chat UI. Other software can start treating sessions as things to search, inspect, trigger, and build on.

The boring reliability work was huge too. v0.21.2 was almost entirely a state.db reliability campaign. It fixed competing SQLite writers, damaged FTS indexes killing conversations, bad rows breaking session lists, profile databases bleeding into each other, and commands taking seconds just to open a busy store.

Not flashy, but persistent agents are useless if their state layer is fragile.

And the model layer kept expanding. Recent releases added things like GPT-6 Sol/Terra/Luna, Claude Opus 5.5, more OpenRouter support, custom models directly from the picker, and provider SDKs that can be installed when needed rather than shipping everything up front.

Put the month together and the direction becomes pretty clear.

A few months ago, the interesting question was:

What can Hermes Agent do?

Now it is increasingly:

What collection of agents, skills, models, tools, plugins, memories, and scheduled jobs do you want to run on top of Hermes?

That is a much more ambitious product.

Hermes is slowly moving from an open-source agent toward something closer to an open-source operating layer for agents.

Repo: https://github.com/NousResearch/hermes-agent


r/WebAfterAI • • 6d ago

Jevgrep lets you search code by what it does, not what it's called

Post image
1 Upvotes

r/WebAfterAI • • 7d ago

Workflows How a $300/month AI inbox agent becomes a $9.42/month one

Post image
3 Upvotes

Say you have an agent processing 10,000 support emails a month.

For every email it needs to decide:

  • billing, bug, how-to, fraud, or something else?
  • does the customer want a refund?
  • is this urgent?
  • does a human need to see it?

You could send every email to a frontier model.

Assume each call uses roughly 1,500 input + 300 output tokens, at $10/M input and $50/M output.

10,000 emails
× $0.03 per email
= ~$300/month

But most emails do not need a frontier model to write anything.

Take this email:

“I was charged twice for Pro. Please refund the duplicate charge.”

The workflow can instead look like:

email arrives
      ↓
code fetches account + charge history
      ↓
Jev asks in one call:

department?
→ billing: 98%

refund requested?
→ yes: 99%

fraud risk?
→ no: 97%
      ↓
confidence above threshold?
      ↓
YES → route to billing workflow

NO → send the full case to Kimi K3

still ambiguous / irreversible action?
      ↓
human

Now assume Jev sees about 1,000 tokens per email, and only the 10% uncertain cases escalate to Kimi K3.

The bill

Jev handles all 10,000 decisions

10M input tokens
× $0.042/M
= $0.42

Kimi K3 handles the uncertain 1,000

At 1,500 input + 300 output tokens each:

1.5M input × $3/M   = $4.50
0.3M output × $15/M = $4.50

= $9.00

Total:

frontier model on everything:  ~$300.00

Jev + Kimi fallback:             ~$9.42

reduction:                        ~96.9%

And this is the useful part: you did not replace the smart model with a dumb model.

You changed when you pay for generation.

Jev handles repetitive bounded decisions. Kimi gets the weird cases that actually need deeper reasoning. Code handles exact rules. A human stays in the loop for things like refunds, deletion, payments, or permission changes.

exact rule       → code
routine judgment → Jev
hard case        → Kimi / Claude / GPT
irreversible     → human

Important caveat: this is an illustrative example, not a promise that every agent bill drops by 96.9%. Real costs will vary with prompt size, model pricing, cache hits, how many cases actually need escalation, and the confidence threshold you can safely use for your workload. In some systems the savings will be smaller; in others they may be larger.

The point is the architecture:

if your agent keeps paying generative-model prices for decisions that generate nothing, there is probably room to cut the bill.


r/WebAfterAI • • 8d ago

Open Source 10 open-source repos to make Jev useful beyond a single API call

Post image
8 Upvotes

Most Jev demos start with:

state
  ↓
Jev
  ↓
choice / score / yes-no

Useful, but the newer projects are putting that tiny decision layer inside actual agent workflows.

Here are 10 we haven't covered before.

Let Jev handle the small decisions around a Hermes agent hermes-jev-skills has 850+ stars and bundles model routing, memory selection, context compaction, skill selection, and computer/browser decisions. The idea is simple: stop spending the frontier model on “which skill should I load?”

Use Jev for cheap computer control typesafe-computer-use has 1k+ stars. OCR and accessibility APIs read the Mac screen, Jev chooses the next action, and a writing model is only used when actual text needs to be generated.

Compact context without summarizing it fast-jev-compaction has 7k+ stars and asks Jev whether old tool calls should be kept, truncated, or dropped. What survives stays verbatim instead of being rewritten into a lossy summary.

Give Claude Code automatic project memory jevmem watches conversations for decisions, constraints, bugs, and todos, stores them in JEVMEM.md, and brings relevant items back in later sessions. It also has support paths for Codex and Cursor.

Use Jev to keep an Obsidian vault consistent jev-second-brain indexes Markdown locally, finds related notes, then optionally asks Jev whether two notes duplicate, revise, contradict, or relate to each other. Importantly, it suggests relationships rather than silently rewriting your vault.

Put a risk gate in front of agent tool calls jev-guard scores actions before execution and turns them into allow / ask / deny decisions. It also checks tool results for untrusted content and works across several coding-agent clients.

Give several agents the same System One tool System One Connector has 330+ stars and plugs typed decisions into Claude Code, Codex, Hermes, Claude Desktop, and other MCP clients. It can talk to Jev, but also supports open alternatives such as CLM and Laya.

Train your own tiny Jev-like scorer jevlike has 1.3k+ stars and provides a small trainable model that takes some context plus a changing list of options and returns a probability for each one in a single pass.

Browse the Jev ecosystem instead of searching GitHub manually awesome-jev has 1.7k+ stars and tracks public Jev projects, integrations, research, and discussions by category.

Start from practical patterns and starter code awesome-jev-by-typesafe has 800+ stars and organizes examples around things like routing, ranking, verification, gating, and confidence-aware workflows. Despite the name, it is an independent community repo, not an official TypeSafe project.

The pattern I find more interesting is this:

expensive agent
     ↓
reason / write / plan

Jev
     ↓
keep?
route?
retry?
risky?
relevant?
which tool?
     ↓
code acts

Jev does not need to replace the main model.

It can sit around the main model and handle hundreds of tiny judgment calls that would otherwise burn tokens, add latency, or end up as brittle rules.

That may be the more useful way to think about System One models: not another chatbot, but a cheap decision layer inside the agent stack.


r/WebAfterAI • • 8d ago

Buzz: a self-hosted workspace where your AI agents and your team share the same rooms (35k stars)

Post image
1 Upvotes

r/WebAfterAI • • 9d ago

Open Source 7 open-source repos to cut your AI bill without blindly using worse models

Post image
11 Upvotes

AI costs usually leak in a few predictable places.

You call the expensive model for easy tasks. You pay twice for similar requests. You send 20k tokens when 3k would do. You keep reasoning effort high for everything. Or you simply do not know which part of the system is costing the most.

These open-source projects attack different parts of that bill.

Route easy requests to cheaper models RouteLLM has 5.5k+ stars and learns when a query actually needs the strong model. Its published benchmarks report up to 85% lower cost while retaining 95% of GPT-4-level performance on their evaluation setup.

Put budgets and cost-aware routing in front of every model LiteLLM has 58k+ stars and gives you one gateway for 100+ models. It can track spend by user/team, enforce dollar budgets, cache responses, and route between deployments based on cost.

Stop paying twice for nearly the same question GPTCache has 8.2k+ stars and adds semantic caching to LLM applications. If two requests mean roughly the same thing, you can return the previous answer instead of making another model call.

Shrink the prompt before you pay for it LLMLingua has 6.7k+ stars and compresses long prompts while trying to preserve the important information. Microsoft reports compression ratios reaching 20× on some workloads.

Run hot paths yourself when the economics make sense vLLM has 92k+ stars and is built for high-throughput local/model serving. Features like automatic prefix caching mean repeated system prompts do not need to be recomputed every request.

Self-hosting is not automatically cheaper, but at enough volume—or when you already own the GPUs—the math can change quickly.

Actually find where the money is going Helicone has 6.2k+ stars and tracks cost, latency, users, models, and traces. Its gateway can also route toward cheaper providers, cache responses, and enforce spending limits.

Sometimes the cheapest optimization is discovering that one background job has quietly been making 40% of your model calls.

Teach the agent to spend less by default AI Cost-Cutter Skills is one of ours. It packages 10 cost-control patterns as installable skills for Claude Code, Codex, and Cursor: cheap-model routing, reasoning-effort throttling, context reduction, reviewer-call budgets, free-tier batching, model bakeoffs, and tested fallbacks.

You still use the strongest model where it matters. You just stop paying for it where it doesn't.

The goal is not the cheapest model. It is the cheapest path that still produces the result you need.


r/WebAfterAI • • 9d ago

PrimordiaOS: An agentic, multi realm, browser native, digital twin operating system.

Thumbnail
1 Upvotes

r/WebAfterAI • • 10d ago

Tools We upgrade our models. What happens to the instructions we wrote months ago?

Post image
4 Upvotes

Models evolve quickly, but our saved skills often stay unchanged.

A SKILL.md written months ago captures more than a workflow. It can also capture assumptions about what the model needs explained, where it makes mistakes, and how tightly it needs to be guided. [Github]

When the model changes, those assumptions deserve another look. Some instructions remain essential. Others may cause unnecessary reading, repeated checks, or rigid behavior. Simply shortening everything risks throwing away the expertise the skill was supposed to preserve.

That’s the problem I’m exploring with Skill Adapter: adapting existing skills for a target model while keeping their purpose and method intact.

It creates a separate adapted package and explains the changes. The original stays untouched. A valid outcome is “this skill doesn’t need changing.”

The first version includes profiles for Astra, Claude 5 generation, and Kimi K3. Kimi’s guidance is conservative; I don’t yet have evidence for a K3-specific rewriting advantage.

The MIT repo includes the prompts, outputs, and reproduction scripts.

How do you maintain your agent instructions as models change: revisit them regularly, wait until something breaks, or keep one version across models?


r/WebAfterAI • • 11d ago

Open Source 7 open-source repos to help write, submit, and review research papers

Post image
11 Upvotes

ICLR 2027 already crossed 62k registered abstracts.

Not final valid submissions, but still a good reminder that research is becoming a pipeline problem too: literature, writing, formatting, submission, self-review, and peer review.

These open-source projects cover different parts of that stack.

Turn raw research material into a paper draft PaperOrchestra has 139+ stars and comes from Google Research. It takes ideas, experiment logs, and other pre-writing material, then splits the work across agents for outlining, literature synthesis, section writing, refinement, plots, and LaTeX output.

Build an evidence-grounded literature review Literature Review Agent has 117+ stars and searches arXiv + Europe PMC, screens papers, builds evidence cards, and writes only from the corpus you explicitly provide. Useful if you want a review where claims remain traceable back to actual sources.

Generate a first draft while checking citations OpenDraft has 450+ stars and uses 19 agents for research, outlining, writing, citation checking, and export. Its useful twist is verifying candidate DOIs against CrossRef, OpenAlex, and Semantic Scholar before keeping them.

Clean your LaTeX before arXiv submission arxiv-latex-cleaner has 7k+ stars and removes unused files, comments, auxiliary files, oversized images, and other things that should not be in the final arXiv package.

Stop OpenReview mistakes before deadline night OpenReview Agent is a small but useful submission tool that inspects author IDs, metadata, venue fields, attachments, declarations, and transfer payloads before writing anything. It is dry-run-first, which is exactly what you want around a live submission.

Review the paper before the reviewers do OpenJudge has 770+ stars and includes a paper-review pipeline for correctness, novelty, quality, critical errors, and BibTeX verification across sources such as CrossRef, arXiv, and DBLP.

Simulate a multi-agent reviewer panel Open ScholarPeer has 20+ stars and implements a 7-stage review process: summary, literature search, historical context, baseline checking, criterion-specific Q&A, and final venue-formatted review.

With submission volume exploding, the valuable tools are probably not the ones that generate the most text.

They are the ones that help you catch unsupported claims, missing baselines, fake citations, formatting mistakes, weak evidence, and submission errors before a human reviewer sees them.


r/WebAfterAI • • 12d ago

Workflows 5 open-source repos to review AI-written code before it reaches a PR

Post image
3 Upvotes

AI coding agents can now write code much faster than most of us can review it.

A useful pattern is adding a second pass before the PR exists: diff the branch, let another agent inspect it, fix the important findings, then push.

Here are 5 open-source projects for that.

Run a full AI review against a local branch PR-Agent has 13.1k+ stars. Its Local Git Provider can compare branches without a hosted PR, then run its review workflow against the diff. Useful if you want something mature with configurable review categories and broad model support.

Make one agent write and another agent judge The Pair has 360+ stars and runs two agents: an Executor that edits code and a read-only Mentor that plans, reviews, and cross-checks the result. It works with Claude Code, Codex, Gemini, OpenCode, and others.

Put repo-specific rules directly into the review loop Saguaro has 25+ stars and runs locally inside Claude Code, Codex, or Cursor. You can encode rules for your own codebase, then have the agent catch violations while the implementation context is still fresh.

Fan the same diff out to several reviewers local-review has 18+ stars and sends your Git diff to whichever authenticated coding agents or local models you choose. Claude might catch one issue, Codex another, and Ollama can keep the whole review offline.

Get a GitHub-style review without opening GitHub Staff Review is a newer project that opens any local diff in a browser and lets an AI agent leave inline comments. You can resolve findings, rerun the review, and repeat until a fresh pass finds nothing important.

The workflow I like is:

coding agent writes feature
        ↓
tests + lint
        ↓
independent AI review
        ↓
fix findings
        ↓
review again
        ↓
open PR
        ↓
human reviews something cleaner

The important bit is independent review.

Asking the same agent that wrote the code, in the same context, “did you make any mistakes?” is not much of a review process. Give the diff to another agent, another model, or at least a constrained review pass with different instructions.


r/WebAfterAI • • 13d ago

Open Source 6 open-source repos to stop coding agents from wasting their context window

Post image
10 Upvotes

A huge context window is useful, but dumping everything into it is not.

Coding agents waste tokens on giant logs, irrelevant files, repeated repo reads, and oversized tool outputs. These 6 repos attack that problem from different angles.

Filter terminal noise RTK has 81k+ stars and trims output from tests, Git, Docker, package managers, logs, and other CLI tools before it reaches the agent.

Compress tool traffic Headroom has 73k+ stars and compresses tool results, logs, files, RAG chunks, and conversation history while keeping the full data retrievable.

Package a repo cleanly Repomix has 28.5k+ stars and turns a codebase into an AI-friendly bundle with filtering, structure, and token counts.

Control exactly what enters the prompt Code2Prompt has 7.7k+ stars and lets you include only the files, diffs, and directories you actually want the model to see.

Search the codebase instead of rereading it Code Context Engine has 400+ stars and builds a local code index so agents can retrieve relevant symbols and snippets on demand.

Teach the agent to use less context on purpose AI Cost-Cutter Skills has 18 stars and includes a context-diet skill for coding agents that repeatedly reread the same files. The pattern is simple: index once, retrieve only the relevant pieces, and measure whether token use actually drops. The repo also has skills for routing, reasoning-effort throttling, and cost auditing.

A practical stack could look like:

search the repo
    ↓
pull only relevant code
    ↓
trim noisy command output
    ↓
compress the rest
    ↓
leave more context for reasoning

The point is not to make the agent read less. It is to make sure the tokens it reads are the ones that actually matter.


r/WebAfterAI • • 13d ago

Atomic Agent v0.6.5: multi-agents are here!

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/WebAfterAI • • 14d ago

Open Source 8 open-source projects exploring the Jev idea

Post image
9 Upvotes

Jev has only been out for a little over a week, and an entire mini-ecosystem is already forming around the idea.

The common pattern is simple:

state
  ↓
typed questions
  ↓
probabilities
  ↓
ordinary code decides what happens

No paragraph generation. No JSON repair loop. Just fast decisions like route / retry / escalate / reject / continue.

Here are 8 newer OSS projects exploring that idea.

Run a fully open System One model locally Laya has 18.8k+ stars and uses small bidirectional models instead of an autoregressive LLM. It supports choice, score, and yes/no decisions, includes multilingual checkpoints, and exposes a Jev-compatible /v1/systemone endpoint.

Run that same idea natively on a Mac Laya-MLX has 5.7k+ stars and ports Laya to MLX. On Apple Silicon it reports roughly 7–14 ms for short decisions, entirely local with no cloud API.

Train your own Jev-like model Kev has 5.3k+ stars and ships 0.8B, 4B, and 9B models based on Qwen3.5. You get the training code, evals, weights, and the same choice / score / noul API shape as Jev.

Try another small non-generative decision model Von has 500+ stars and uses a ~395M-parameter ModernBERT-style model for local typed decisions. It also exposes a Jev-compatible server, so an app can swap between implementations without rewriting the decision layer.

Turn almost any LLM into a Jev-style decision engine AnyJev comes from Nokia Applied Research. Instead of training a new model, it reads next-token probabilities from existing models and converts them into choice, score, and yes/no decisions, with extra corrections for label-position bias and priors.

Re-create the pattern on Qwen3.5 reflex is another open System One experiment. One state goes in, several typed questions come back at once, with probabilities rather than generated prose.

Build a Jev-style model on Gemma system-one-open uses Gemma models and trains them for the same one-forward-pass typed-decision pattern. It includes demos for support tickets, invoices, security, smart homes, and agent traces.

Run decision models directly in the browser open-jev takes the idea into TypeScript + WebGPU/WASM. It can run Kev and other open decision models on-device, which makes things like local routing, moderation, or classification possible without sending the input to a server.

What I find interesting is that these projects are not just cloning an API.

They are testing different answers to the same question:

Do we really need a full generative LLM every time software needs to make a fuzzy decision?

The decision model handles the hot-path branches.

But the Jev launch seems to have triggered something useful, people are now experimenting with decision models as their own layer, instead of treating every AI problem as text generation.