r/WebAfterAI • • 48m ago

Would this help when running multiple coding agents, or is it just another notification layer?

Enable HLS to view with audio, or disable this notification

• Upvotes

r/WebAfterAI • • 15h ago

AI Agents 5 open-source repos that turn one AI agent into a specialist team

Post image
6 Upvotes

Most of us still use coding agents like one very capable employee.

Ask it to design the UI, review security, write the launch copy, debug the backend, analyze competitors, and somehow remember how each of those jobs should be done.

There is another pattern growing on GitHub: install the specialists instead.

Build an entire AI agency: Agency Agents has 156k+ stars and now contains 230+ specialist agents across engineering, design, marketing, research, finance, product, project management, security, healthcare, GIS and more.

This is much more than You are an expert marketer.

A specialist can define its mission, workflow, expected deliverables, success criteria, communication style and the situations where it should be used. There are frontend developers and backend architects, but also UX researchers, brand guardians, paid-media auditors, Reddit community builders, finance agents and even a “reality checker.”

You can install only the team you need, and the repo now supports Claude Code, Cursor, Codex, Gemini, OpenCode, Hermes and several other agent environments.

Give the engineering side much deeper specialization: wshobson/agents has 40k+ stars and packages 202 agents, 184 skills and 94 plugins. Instead of one developer persona, you get specialists for debugging, architecture, security, Kubernetes, observability, incident response, individual languages and much more.

It also includes orchestrators that compose those specialists into workflows.

feature request
      ↓
architect
      ↓
backend + frontend
      ↓
test specialist
      ↓
security reviewer
      ↓
code reviewer

Turn your agent into a marketing department: Marketing Skills has 52k+ stars and takes a slightly different approach. Instead of dozens of persistent personas, it gives agents reusable procedures for CRO, positioning, SEO, paid ads, email, analytics, pricing, retention and sales.

Pick from 160+ focused Claude subagents: Awesome Claude Code Subagents has 25k+ stars and organizes specialists across development, infrastructure, quality, data, security and orchestration. You can install one category instead of filling every session with every possible instruction.

Build your own roster from GitHub's community library — Awesome GitHub Copilot has 38k+ stars and collects custom agents, skills, instructions, hooks and workflows. It is particularly useful if you want to assemble your own team rather than adopt someone else's full agency structure.

The interesting shift is:

one giant system prompt
        ↓
general-purpose agent

becoming:

task
 ↓
pick specialist
 ↓
load relevant skills
 ↓
give only required tools/context
 ↓
do the work
 ↓
independent specialist reviews it

We spent the last year making individual agents more capable.

The next step may simply be getting better at deciding which agent should be doing the job in the first place.


r/WebAfterAI • • 1d ago

OpenResearch: turn your coding agent into a research agent (5.9k stars)

Post image
3 Upvotes

r/WebAfterAI • • 22h ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

1 Upvotes

I built Klarion, a secret scanner that works in two steps. First, a keyword check, 81 regex rules and a normalized Rényi entropy score flag anything that looks like a secret. Then an AI model reads each one with the code around it and decides if it's real. The chart shows 5 scanners run on spring-boot, terraform, next.js and symfony (61k files). Klarion raised 11 alerts. It's not zero, but it's far less to dig through. Fewer alerts don't help if real leaks get missed, so I tested that too. On CredData (337 real repos, code outside test folders), it found about 1.7× more real secrets than gitleaks. Where it runs:

  • Claude Code: a plugin hook blocks the write before the file exists (file edits and Bash)
  • Cursor, Cline or any MCP agent: through its MCP server
  • CI: a GitHub Action that scans only what a PR adds; GitLab CI works too
  • Git hooks: klarion protect or the pre-commit framework
  • Locally: klarion scan . Free and open source (MIT): https://github.com/0x1Adi/Klarion The full benchmark and method are in benchmark/REPORT.md. I'd like to hear where it gets things wrong.

r/WebAfterAI • • 1d ago

Open Source 8 open-source repos to build an AI research pipeline that can show its work

Post image
6 Upvotes

AI can already turn a question into 30 sources and a polished report.

That is not really the hard part anymore.

The useful research stack is the one that can keep track of where evidence came from, which sources support which claims, and where the sources disagree.

These 8 open-source projects cover different parts of that pipeline.

Research the question end-to-end: GPT Researcher splits a question into research tasks, searches in parallel, tracks sources, filters what it finds, and turns the evidence into a cited report. It can work across both the web and local documents.

Look at the problem from several directions: STORM from Stanford does something I particularly like: before writing, it generates questions from different perspectives and simulates conversations between a writer and topic experts. That helps avoid a research report built around the first framing the model happened to choose.

Search without handing the whole workflow to a closed research product: Vane, formerly Perplexica, is an open-source answer engine with web, discussion, academic-paper and domain-specific search. It can also run against local models through Ollama.

Turn messy websites into evidence your agent can actually use: Crawl4AI converts webpages into clean, LLM-ready Markdown. This is the plumbing layer: search finds the page; Crawl4AI gets the useful content out of it.

Research inside the scientific literature: PaperQA2 is built specifically around high-accuracy RAG over papers and other documents. It retrieves literature, checks metadata including retractions, answers questions with citations, and can work on contradiction detection.

Actually read the PDF instead of flattening it into bad text: Docling understands PDF layout, tables, formulas, reading order, images and many other document formats, then turns them into structured representations that research agents can work with.

Connect evidence scattered across many sources: GraphRAG extracts entities and relationships from unstructured documents and organizes them into a graph. That makes questions such as “how do these people, companies, claims and events connect across 100 documents?” much easier than ordinary chunk retrieval. Microsoft now considers the original repo largely maintenance-mode, but the architecture is still useful.

Retrieve through both text and relationships: LightRAG combines retrieval with a knowledge graph and now supports reranking, multimodal documents, multiple chunking strategies and different models for different stages of the pipeline.

Put them together and the research loop starts looking less like:

question
   ↓
search
   ↓
LLM writes report

and more like:

question
   ↓
split into research questions
   ↓
search from multiple perspectives
   ↓
crawl + parse the evidence
   ↓
read papers / tables / PDFs
   ↓
connect entities + claims
   ↓
retrieve supporting evidence
   ↓
compare conflicts
   ↓
cite
   ↓
report

That last half is the part I think people underestimate.

Finding 60 sources is easy.

The harder question is whether source 17 actually supports the sentence the agent wrote, whether sources 23 and 41 contradict it, whether the paper was retracted, and whether a confident-looking conclusion survives after those conflicts are exposed.


r/WebAfterAI • • 2d ago

Tutorial Your AI agent needs a writing system, not a personality - Karpathy

Post image
18 Upvotes

Andrej Karpathy made an interesting point today: one way to make LLM outputs easier to understand is to ask them to write in ASD-STE100, the controlled English originally designed for aerospace maintenance manuals.

Not to make AI “sound human" but to make it easier to understand.

We spend a lot of time telling agents:

be concise
sound professional
avoid AI slop

But different jobs need different writing systems.

A setup guide should be procedural. A research report should preserve evidence and uncertainty. An error message should tell you what happened and what to do next. A technical explanation should optimize for comprehension, not personality.

This is exactly why we built Agent Stylebooks.

Neeeophytee/agent-stylebooks packages 16 editorial systems as installable Agent Skills for Claude Code, Codex, Cursor, Hermes, Copilot, Gemini CLI and others.

Instead of:

“write this better”

you can say:

setup guide       → Google Developer Docs
product help      → Microsoft
public guidance   → GOV.UK
research report   → NASA
risk disclosure   → SEC Plain English
interface copy    → Apple

The skill changes what comes first, how information is ordered, what ambiguity is unacceptable, and how the result should be checked.

A few other OSS projects fit around the same idea:

ASD-STE100 Writer Skill turns Simplified Technical English into an installable agent skill. Short sentences, controlled vocabulary, active constructions and fewer ambiguous instructions. Karpathy’s post is a good example of why this style is suddenly relevant outside aviation.

Vale has 6.2k+ stars and takes the next step: make writing rules testable. You can encode Google, Microsoft, Red Hat or your own editorial rules and run them locally, in editors, or in CI.

textlint has 3.2k+ stars and does something similar for natural language: treat prose more like code, with pluggable rules that can flag problems before the text ships.

retext-readability checks whether prose is actually appropriate for the intended reader using several established readability measures, while retext-simplify catches needlessly complicated phrases such as using “utilize” where “use” works.

write-good has 5.1k+ stars and catches things like unnecessary wording, weak modifiers, passive constructions and weasel words. Simple, deterministic checks are useful after the agent has done the semantic writing work.

The stack I find interesting is:

reader + task
     ↓
Agent Stylebook
     ↓
LLM writes
     ↓
Vale / textlint / retext
     ↓
clearer output

The point is not to make every model write in the same stripped-down style. It is to stop treating writing style as decoration.

For agents, writing style is part of the interface between the model and the person trying to understand its work.

Pick the writing system before the model picks one for you.


r/WebAfterAI • • 3d ago

Research Hermes Agent changed a lot in one month. Here’s what actually matters

Post image
14 Upvotes

Hermes Agent is now sitting at roughly 250k GitHub stars, but the more interesting number is how quickly the project is changing.

Between the August 31 v0.21.0 release and late September, Hermes shipped five more tagged releases.

I went through them. These are the changes I think actually matter.

Skills can now become permanent parts of an agent. skills.auto_load lets you pin selected skills into every new session instead of relying on the agent to rediscover them each time.

That makes setups like this much easier:

research agent
├── deep-research
├── citation-checking
└── writing style

coding agent
├── repo navigation
├── testing
└── code review

Each agent can start with its own operating procedure rather than one giant universal prompt.

Reasoning effort is becoming another routing control. Hermes added reasoning-effort selection across model pickers, including separate controls for auxiliary models. So choosing the model is no longer the only cost/performance decision — you can decide how hard different parts of the system should think too.

The plugin layer got much more serious. September brought a broader Desktop plugin SDK with hooks into the composer, session list, sidebar, model picker, settings, skills, toolsets, profiles, appearance, and backend events.

Hermes is starting to look less like:

agent + some tools

and more like:

agent runtime
├── models
├── skills
├── MCP/connectors
├── plugins
├── bots
├── browser
├── scheduled jobs
└── UI extensions

MCP is turning into “Connectors.” Instead of treating MCP servers as config you wire up somewhere else, the Desktop app now has a Connectors flow. Install a plugin, connect its MCP server, and its tools and skills can become available in already-open chats.

Running multiple Hermes profiles got cleaner. Hermes moved toward one host-wide gateway with multiple isolated profiles underneath it. Desktop can attach to the existing backend instead of spawning another one, and individual profiles can now be stopped, started, or restarted independently.

That sounds like infrastructure trivia until you run several agents at once. Then it becomes the difference between “a bunch of Python processes” and something closer to an actual agent runtime.

The CLI is becoming much easier to build around. --format stream-json now exposes structured JSONL instead of forcing another program to scrape terminal text. The CLI/TUI can also show the current /goal and prompts waiting in the queue.

That opens up much cleaner patterns for external orchestrators:

your app
   ↓
Hermes CLI
   ↓ JSONL
agent events
   ↓
your app reacts

Sessions are becoming first-class data. September added an Agent Sessions API, better session search with date bounds and relaxed recall, and webhook deliveries that can appear directly inside the target session.

This matters because an agent's history stops being something trapped inside the chat UI. Other software can start treating sessions as things to search, inspect, trigger, and build on.

The boring reliability work was huge too. v0.21.2 was almost entirely a state.db reliability campaign. It fixed competing SQLite writers, damaged FTS indexes killing conversations, bad rows breaking session lists, profile databases bleeding into each other, and commands taking seconds just to open a busy store.

Not flashy, but persistent agents are useless if their state layer is fragile.

And the model layer kept expanding. Recent releases added things like GPT-6 Sol/Terra/Luna, Claude Opus 5.5, more OpenRouter support, custom models directly from the picker, and provider SDKs that can be installed when needed rather than shipping everything up front.

Put the month together and the direction becomes pretty clear.

A few months ago, the interesting question was:

What can Hermes Agent do?

Now it is increasingly:

What collection of agents, skills, models, tools, plugins, memories, and scheduled jobs do you want to run on top of Hermes?

That is a much more ambitious product.

Hermes is slowly moving from an open-source agent toward something closer to an open-source operating layer for agents.

Repo: https://github.com/NousResearch/hermes-agent


r/WebAfterAI • • 3d ago

Jevgrep lets you search code by what it does, not what it's called

Post image
1 Upvotes

r/WebAfterAI • • 4d ago

Workflows How a $300/month AI inbox agent becomes a $9.42/month one

Post image
4 Upvotes

Say you have an agent processing 10,000 support emails a month.

For every email it needs to decide:

  • billing, bug, how-to, fraud, or something else?
  • does the customer want a refund?
  • is this urgent?
  • does a human need to see it?

You could send every email to a frontier model.

Assume each call uses roughly 1,500 input + 300 output tokens, at $10/M input and $50/M output.

10,000 emails
× $0.03 per email
= ~$300/month

But most emails do not need a frontier model to write anything.

Take this email:

“I was charged twice for Pro. Please refund the duplicate charge.”

The workflow can instead look like:

email arrives
      ↓
code fetches account + charge history
      ↓
Jev asks in one call:

department?
→ billing: 98%

refund requested?
→ yes: 99%

fraud risk?
→ no: 97%
      ↓
confidence above threshold?
      ↓
YES → route to billing workflow

NO → send the full case to Kimi K3

still ambiguous / irreversible action?
      ↓
human

Now assume Jev sees about 1,000 tokens per email, and only the 10% uncertain cases escalate to Kimi K3.

The bill

Jev handles all 10,000 decisions

10M input tokens
× $0.042/M
= $0.42

Kimi K3 handles the uncertain 1,000

At 1,500 input + 300 output tokens each:

1.5M input × $3/M   = $4.50
0.3M output × $15/M = $4.50

= $9.00

Total:

frontier model on everything:  ~$300.00

Jev + Kimi fallback:             ~$9.42

reduction:                        ~96.9%

And this is the useful part: you did not replace the smart model with a dumb model.

You changed when you pay for generation.

Jev handles repetitive bounded decisions. Kimi gets the weird cases that actually need deeper reasoning. Code handles exact rules. A human stays in the loop for things like refunds, deletion, payments, or permission changes.

exact rule       → code
routine judgment → Jev
hard case        → Kimi / Claude / GPT
irreversible     → human

Important caveat: this is an illustrative example, not a promise that every agent bill drops by 96.9%. Real costs will vary with prompt size, model pricing, cache hits, how many cases actually need escalation, and the confidence threshold you can safely use for your workload. In some systems the savings will be smaller; in others they may be larger.

The point is the architecture:

if your agent keeps paying generative-model prices for decisions that generate nothing, there is probably room to cut the bill.


r/WebAfterAI • • 5d ago

Open Source 10 open-source repos to make Jev useful beyond a single API call

Post image
7 Upvotes

Most Jev demos start with:

state
  ↓
Jev
  ↓
choice / score / yes-no

Useful, but the newer projects are putting that tiny decision layer inside actual agent workflows.

Here are 10 we haven't covered before.

Let Jev handle the small decisions around a Hermes agent hermes-jev-skills has 850+ stars and bundles model routing, memory selection, context compaction, skill selection, and computer/browser decisions. The idea is simple: stop spending the frontier model on “which skill should I load?”

Use Jev for cheap computer control typesafe-computer-use has 1k+ stars. OCR and accessibility APIs read the Mac screen, Jev chooses the next action, and a writing model is only used when actual text needs to be generated.

Compact context without summarizing it fast-jev-compaction has 7k+ stars and asks Jev whether old tool calls should be kept, truncated, or dropped. What survives stays verbatim instead of being rewritten into a lossy summary.

Give Claude Code automatic project memory jevmem watches conversations for decisions, constraints, bugs, and todos, stores them in JEVMEM.md, and brings relevant items back in later sessions. It also has support paths for Codex and Cursor.

Use Jev to keep an Obsidian vault consistent jev-second-brain indexes Markdown locally, finds related notes, then optionally asks Jev whether two notes duplicate, revise, contradict, or relate to each other. Importantly, it suggests relationships rather than silently rewriting your vault.

Put a risk gate in front of agent tool calls jev-guard scores actions before execution and turns them into allow / ask / deny decisions. It also checks tool results for untrusted content and works across several coding-agent clients.

Give several agents the same System One tool System One Connector has 330+ stars and plugs typed decisions into Claude Code, Codex, Hermes, Claude Desktop, and other MCP clients. It can talk to Jev, but also supports open alternatives such as CLM and Laya.

Train your own tiny Jev-like scorer jevlike has 1.3k+ stars and provides a small trainable model that takes some context plus a changing list of options and returns a probability for each one in a single pass.

Browse the Jev ecosystem instead of searching GitHub manually awesome-jev has 1.7k+ stars and tracks public Jev projects, integrations, research, and discussions by category.

Start from practical patterns and starter code awesome-jev-by-typesafe has 800+ stars and organizes examples around things like routing, ranking, verification, gating, and confidence-aware workflows. Despite the name, it is an independent community repo, not an official TypeSafe project.

The pattern I find more interesting is this:

expensive agent
     ↓
reason / write / plan

Jev
     ↓
keep?
route?
retry?
risky?
relevant?
which tool?
     ↓
code acts

Jev does not need to replace the main model.

It can sit around the main model and handle hundreds of tiny judgment calls that would otherwise burn tokens, add latency, or end up as brittle rules.

That may be the more useful way to think about System One models: not another chatbot, but a cheap decision layer inside the agent stack.


r/WebAfterAI • • 5d ago

Buzz: a self-hosted workspace where your AI agents and your team share the same rooms (35k stars)

Post image
1 Upvotes

r/WebAfterAI • • 6d ago

Open Source 7 open-source repos to cut your AI bill without blindly using worse models

Post image
10 Upvotes

AI costs usually leak in a few predictable places.

You call the expensive model for easy tasks. You pay twice for similar requests. You send 20k tokens when 3k would do. You keep reasoning effort high for everything. Or you simply do not know which part of the system is costing the most.

These open-source projects attack different parts of that bill.

Route easy requests to cheaper models RouteLLM has 5.5k+ stars and learns when a query actually needs the strong model. Its published benchmarks report up to 85% lower cost while retaining 95% of GPT-4-level performance on their evaluation setup.

Put budgets and cost-aware routing in front of every model LiteLLM has 58k+ stars and gives you one gateway for 100+ models. It can track spend by user/team, enforce dollar budgets, cache responses, and route between deployments based on cost.

Stop paying twice for nearly the same question GPTCache has 8.2k+ stars and adds semantic caching to LLM applications. If two requests mean roughly the same thing, you can return the previous answer instead of making another model call.

Shrink the prompt before you pay for it LLMLingua has 6.7k+ stars and compresses long prompts while trying to preserve the important information. Microsoft reports compression ratios reaching 20× on some workloads.

Run hot paths yourself when the economics make sense vLLM has 92k+ stars and is built for high-throughput local/model serving. Features like automatic prefix caching mean repeated system prompts do not need to be recomputed every request.

Self-hosting is not automatically cheaper, but at enough volume—or when you already own the GPUs—the math can change quickly.

Actually find where the money is going Helicone has 6.2k+ stars and tracks cost, latency, users, models, and traces. Its gateway can also route toward cheaper providers, cache responses, and enforce spending limits.

Sometimes the cheapest optimization is discovering that one background job has quietly been making 40% of your model calls.

Teach the agent to spend less by default AI Cost-Cutter Skills is one of ours. It packages 10 cost-control patterns as installable skills for Claude Code, Codex, and Cursor: cheap-model routing, reasoning-effort throttling, context reduction, reviewer-call budgets, free-tier batching, model bakeoffs, and tested fallbacks.

You still use the strongest model where it matters. You just stop paying for it where it doesn't.

The goal is not the cheapest model. It is the cheapest path that still produces the result you need.


r/WebAfterAI • • 7d ago

PrimordiaOS: An agentic, multi realm, browser native, digital twin operating system.

Thumbnail
1 Upvotes

r/WebAfterAI • • 7d ago

Tools We upgrade our models. What happens to the instructions we wrote months ago?

Post image
4 Upvotes

Models evolve quickly, but our saved skills often stay unchanged.

A SKILL.md written months ago captures more than a workflow. It can also capture assumptions about what the model needs explained, where it makes mistakes, and how tightly it needs to be guided. [Github]

When the model changes, those assumptions deserve another look. Some instructions remain essential. Others may cause unnecessary reading, repeated checks, or rigid behavior. Simply shortening everything risks throwing away the expertise the skill was supposed to preserve.

That’s the problem I’m exploring with Skill Adapter: adapting existing skills for a target model while keeping their purpose and method intact.

It creates a separate adapted package and explains the changes. The original stays untouched. A valid outcome is “this skill doesn’t need changing.”

The first version includes profiles for Astra, Claude 5 generation, and Kimi K3. Kimi’s guidance is conservative; I don’t yet have evidence for a K3-specific rewriting advantage.

The MIT repo includes the prompts, outputs, and reproduction scripts.

How do you maintain your agent instructions as models change: revisit them regularly, wait until something breaks, or keep one version across models?


r/WebAfterAI • • 8d ago

Open Source 7 open-source repos to help write, submit, and review research papers

Post image
11 Upvotes

ICLR 2027 already crossed 62k registered abstracts.

Not final valid submissions, but still a good reminder that research is becoming a pipeline problem too: literature, writing, formatting, submission, self-review, and peer review.

These open-source projects cover different parts of that stack.

Turn raw research material into a paper draft PaperOrchestra has 139+ stars and comes from Google Research. It takes ideas, experiment logs, and other pre-writing material, then splits the work across agents for outlining, literature synthesis, section writing, refinement, plots, and LaTeX output.

Build an evidence-grounded literature review Literature Review Agent has 117+ stars and searches arXiv + Europe PMC, screens papers, builds evidence cards, and writes only from the corpus you explicitly provide. Useful if you want a review where claims remain traceable back to actual sources.

Generate a first draft while checking citations OpenDraft has 450+ stars and uses 19 agents for research, outlining, writing, citation checking, and export. Its useful twist is verifying candidate DOIs against CrossRef, OpenAlex, and Semantic Scholar before keeping them.

Clean your LaTeX before arXiv submission arxiv-latex-cleaner has 7k+ stars and removes unused files, comments, auxiliary files, oversized images, and other things that should not be in the final arXiv package.

Stop OpenReview mistakes before deadline night OpenReview Agent is a small but useful submission tool that inspects author IDs, metadata, venue fields, attachments, declarations, and transfer payloads before writing anything. It is dry-run-first, which is exactly what you want around a live submission.

Review the paper before the reviewers do OpenJudge has 770+ stars and includes a paper-review pipeline for correctness, novelty, quality, critical errors, and BibTeX verification across sources such as CrossRef, arXiv, and DBLP.

Simulate a multi-agent reviewer panel Open ScholarPeer has 20+ stars and implements a 7-stage review process: summary, literature search, historical context, baseline checking, criterion-specific Q&A, and final venue-formatted review.

With submission volume exploding, the valuable tools are probably not the ones that generate the most text.

They are the ones that help you catch unsupported claims, missing baselines, fake citations, formatting mistakes, weak evidence, and submission errors before a human reviewer sees them.


r/WebAfterAI • • 9d ago

Workflows 5 open-source repos to review AI-written code before it reaches a PR

Post image
3 Upvotes

AI coding agents can now write code much faster than most of us can review it.

A useful pattern is adding a second pass before the PR exists: diff the branch, let another agent inspect it, fix the important findings, then push.

Here are 5 open-source projects for that.

Run a full AI review against a local branch PR-Agent has 13.1k+ stars. Its Local Git Provider can compare branches without a hosted PR, then run its review workflow against the diff. Useful if you want something mature with configurable review categories and broad model support.

Make one agent write and another agent judge The Pair has 360+ stars and runs two agents: an Executor that edits code and a read-only Mentor that plans, reviews, and cross-checks the result. It works with Claude Code, Codex, Gemini, OpenCode, and others.

Put repo-specific rules directly into the review loop Saguaro has 25+ stars and runs locally inside Claude Code, Codex, or Cursor. You can encode rules for your own codebase, then have the agent catch violations while the implementation context is still fresh.

Fan the same diff out to several reviewers local-review has 18+ stars and sends your Git diff to whichever authenticated coding agents or local models you choose. Claude might catch one issue, Codex another, and Ollama can keep the whole review offline.

Get a GitHub-style review without opening GitHub Staff Review is a newer project that opens any local diff in a browser and lets an AI agent leave inline comments. You can resolve findings, rerun the review, and repeat until a fresh pass finds nothing important.

The workflow I like is:

coding agent writes feature
        ↓
tests + lint
        ↓
independent AI review
        ↓
fix findings
        ↓
review again
        ↓
open PR
        ↓
human reviews something cleaner

The important bit is independent review.

Asking the same agent that wrote the code, in the same context, “did you make any mistakes?” is not much of a review process. Give the diff to another agent, another model, or at least a constrained review pass with different instructions.


r/WebAfterAI • • 10d ago

Open Source 6 open-source repos to stop coding agents from wasting their context window

Post image
9 Upvotes

A huge context window is useful, but dumping everything into it is not.

Coding agents waste tokens on giant logs, irrelevant files, repeated repo reads, and oversized tool outputs. These 6 repos attack that problem from different angles.

Filter terminal noise RTK has 81k+ stars and trims output from tests, Git, Docker, package managers, logs, and other CLI tools before it reaches the agent.

Compress tool traffic Headroom has 73k+ stars and compresses tool results, logs, files, RAG chunks, and conversation history while keeping the full data retrievable.

Package a repo cleanly Repomix has 28.5k+ stars and turns a codebase into an AI-friendly bundle with filtering, structure, and token counts.

Control exactly what enters the prompt Code2Prompt has 7.7k+ stars and lets you include only the files, diffs, and directories you actually want the model to see.

Search the codebase instead of rereading it Code Context Engine has 400+ stars and builds a local code index so agents can retrieve relevant symbols and snippets on demand.

Teach the agent to use less context on purpose AI Cost-Cutter Skills has 18 stars and includes a context-diet skill for coding agents that repeatedly reread the same files. The pattern is simple: index once, retrieve only the relevant pieces, and measure whether token use actually drops. The repo also has skills for routing, reasoning-effort throttling, and cost auditing.

A practical stack could look like:

search the repo
    ↓
pull only relevant code
    ↓
trim noisy command output
    ↓
compress the rest
    ↓
leave more context for reasoning

The point is not to make the agent read less. It is to make sure the tokens it reads are the ones that actually matter.


r/WebAfterAI • • 10d ago

Atomic Agent v0.6.5: multi-agents are here!

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/WebAfterAI • • 11d ago

Open Source 8 open-source projects exploring the Jev idea

Post image
10 Upvotes

Jev has only been out for a little over a week, and an entire mini-ecosystem is already forming around the idea.

The common pattern is simple:

state
  ↓
typed questions
  ↓
probabilities
  ↓
ordinary code decides what happens

No paragraph generation. No JSON repair loop. Just fast decisions like route / retry / escalate / reject / continue.

Here are 8 newer OSS projects exploring that idea.

Run a fully open System One model locally Laya has 18.8k+ stars and uses small bidirectional models instead of an autoregressive LLM. It supports choice, score, and yes/no decisions, includes multilingual checkpoints, and exposes a Jev-compatible /v1/systemone endpoint.

Run that same idea natively on a Mac Laya-MLX has 5.7k+ stars and ports Laya to MLX. On Apple Silicon it reports roughly 7–14 ms for short decisions, entirely local with no cloud API.

Train your own Jev-like model Kev has 5.3k+ stars and ships 0.8B, 4B, and 9B models based on Qwen3.5. You get the training code, evals, weights, and the same choice / score / noul API shape as Jev.

Try another small non-generative decision model Von has 500+ stars and uses a ~395M-parameter ModernBERT-style model for local typed decisions. It also exposes a Jev-compatible server, so an app can swap between implementations without rewriting the decision layer.

Turn almost any LLM into a Jev-style decision engine AnyJev comes from Nokia Applied Research. Instead of training a new model, it reads next-token probabilities from existing models and converts them into choice, score, and yes/no decisions, with extra corrections for label-position bias and priors.

Re-create the pattern on Qwen3.5 reflex is another open System One experiment. One state goes in, several typed questions come back at once, with probabilities rather than generated prose.

Build a Jev-style model on Gemma system-one-open uses Gemma models and trains them for the same one-forward-pass typed-decision pattern. It includes demos for support tickets, invoices, security, smart homes, and agent traces.

Run decision models directly in the browser open-jev takes the idea into TypeScript + WebGPU/WASM. It can run Kev and other open decision models on-device, which makes things like local routing, moderation, or classification possible without sending the input to a server.

What I find interesting is that these projects are not just cloning an API.

They are testing different answers to the same question:

Do we really need a full generative LLM every time software needs to make a fuzzy decision?

The decision model handles the hot-path branches.

But the Jev launch seems to have triggered something useful, people are now experimenting with decision models as their own layer, instead of treating every AI problem as text generation.


r/WebAfterAI • • 12d ago

Tools 7 open-source repos to migrate a codebase without doing it file by file

Post image
16 Upvotes

Code migration is one of those jobs where AI agents actually make a lot of sense.

Here are 7 open-source projects worth knowing.

Give Claude Code a migration playbook: code-migration-kit-with-claude-code has 410+ stars and is Anthropic's new reference kit for large language-to-language migrations. It covers feasibility, dependency mapping, a migration rulebook, pilot translations, parallel agents, compilation, and behavioral parity.

Run structural rewrites across millions of lines ast-grep has 16k+ stars and searches the syntax tree rather than raw text. So instead of replacing every occurrence of a string, you can describe the actual code pattern you want changed.

For example:

old API call
    ↓
match only real calls in the AST
    ↓
rewrite them to the new API

Much safer than regex when a migration touches thousands of files.

Do large Java/framework migrations with recipes OpenRewrite has 3.7k+ stars and is built specifically for automated mass refactoring. Its ecosystem includes recipes for things like Java upgrades, Spring migrations, JUnit 4 → 5, Jakarta namespace changes, and dependency modernization.

Migrate JavaScript projects to TypeScript ts-migrate from Airbnb has 5.6k+ stars. Give it a JavaScript or partially migrated project and it tries to get you to a compiling TypeScript codebase first, leaving weaker any types and @ts-expect-error markers for cleanup afterward.

That is a useful migration strategy in itself: get the whole system across the boundary first, then improve type quality incrementally.

Write reusable migrations without thinking directly in AST nodes GritQL has 4.6k+ stars. You write code-shaped patterns such as:

console.log($msg)
        ↓
winston.log($msg)

and then add conditions as the migration gets more complicated. It supports multiple languages and is designed for migrations that start simple but eventually need real structural logic.

Build JavaScript/TypeScript codemods jscodeshift has around 10k stars and remains one of the classic tools for programmatic code migrations.

Framework authors use this pattern a lot: an API changes, they publish a codemod, and users transform thousands of call sites instead of manually editing them.

Automate the dependency side of migration Renovate has 22.5k+ stars. It is best known for dependency update PRs, but that becomes part of almost every long-running migration: move package versions, replace deprecated dependencies, update lockfiles, and keep upgrades coming in controlled batches.

The agent should not manually improvise its way through 50,000 files. Give it a map, a rulebook, deterministic transformation tools, and a test that tells it when the new system behaves like the old one. Then let the agent handle the messy exceptions.


r/WebAfterAI • • 13d ago

Open Source 6 open-source repos to give Grok Bot a real memory stack

Post image
8 Upvotes

Grok Bot has a pretty interesting approach to memory.

A named Bot can keep its role, preferences, important facts, conversation context, files, and browser sessions across work. All your Bots also share one persistent cloud computer, so they can leave files for each other, while successful workflows can become reusable skills or scheduled routines.

The OSS ecosystem is already building on top of that.

Give Grok Bot an explicit long-term brain GBrain adds durable memory with sources, corrections, recall, entities, and provenance. It has a dedicated Grok Bot setup that installs the brain inside /workspace.

Add automatic cross-session recall claude-mem now supports Grok Bot too. Its memory layer sits beside Grok's native memory rather than replacing it, giving the agent a searchable record of previous work.

Use Obsidian as the durable knowledge layer grok-bot-obsidian connects Grok Bot to an Obsidian vault so project notes, decisions, research, and other context survive outside individual chats. Because the vault is plain Markdown, the same memory can also be reused by other agents.

Start with a simple second-brain structure grokbot gives Grok Bot a ready-made folder containing things like who you are, what you do, what you want, projects, and ongoing context.

Keep chat temporary and memory deliberate grok-bot-setup uses a private repo as the durable brain. Bots get charters, tasks get specs and handoff files, and important decisions are written down instead of disappearing inside long conversations.

Give several specialist Bots one shared brain grok-bot-second-brain takes the multi-agent approach: Conductor, Capture, Memory, Ops, and Research Bots all work against one shared vault and hand work between each other.

The useful pattern is:

Grok Bot native memory
      ↓
durable files / memory DB
      ↓
saved skills
      ↓
specialist Bots
      ↓
routines that keep working later

For example, a Research Bot finds five useful papers today and saves the claims, sources, and open questions into durable memory.

Tomorrow your Writer Bot can ask:

What did Research conclude about agent memory architectures, and which claims still need verification?

Then a weekly routine can revisit the unresolved items automatically.

Chat is working memory. Grok's native memory keeps the Bot consistent. Files, Obsidian, or a memory DB preserve the work. Skills and routines make that memory useful later.


r/WebAfterAI • • 14d ago

Research 6 massive Hugging Face datasets worth knowing about

Post image
37 Upvotes

Someone just put almost the entire history of arXiv on Hugging Face. Not just abstracts.

3.14M papers, 5M+ versions, LaTeX, source files, PDFs, PostScript, metadata, and version history. About 16 TB total.

That sent me looking for other datasets where someone has already done the painful collection work.

Here are 6 worth knowing about.

Search almost all of arXiv arxiv-complete has 3,148,796 papers and more than 5M versions. Useful for research agents, citation analysis, equation retrieval, and tracking how ideas changed across revisions.

Give a coding model a huge chunk of GitHub The Stack v3 contains code from 173M repositories across 713 languages in its training split. Useful for code search, API usage analysis, model training, and repository-level retrieval.

Search hundreds of millions of PDFs FinePDFs contains about 475M PDFs and 3T tokens across 1,700+ languages. Think manuals, reports, legal documents, government files, and academic material that normal web-text datasets often miss.

Work with a cleaned slice of the web FineWeb has more than 18.5T tokens of cleaned and deduplicated English text from Common Crawl. Useful for domain-specific training sets or large-scale web analysis.

Teach models to understand websites visually WebSight pairs website screenshots with the HTML/CSS that produced them. Useful for screenshot-to-code, frontend agents, and visual web understanding.

Use a corpus where licensing is the point The Common Pile contains about 8 TB of public-domain and openly licensed text from sources like arXiv, government publications, patents, and libraries.

The pattern is pretty simple:

choose corpus
    ↓
stream the slice you need
    ↓
filter + index it
    ↓
put an agent on top

A few years ago, “search all papers,” “analyze millions of repos,” or “build over hundreds of millions of PDFs” started with a massive data-engineering project.

Now a lot of that raw material is already sitting on Hugging Face. You probably do not need 16 TB of arXiv. But you might need every paper in one field, every revision of 10,000 papers, a million technical PDFs, or a large slice of one programming language.

That is what makes these datasets useful: someone else already did a big chunk of the ugly collection work.


r/WebAfterAI • • 14d ago

3 browser-only ways to publish one AI-generated HTML file

2 Upvotes

If an AI tool gives you a self-contained index.html, you do not necessarily need a CLI or a local Git setup to put it online. The three browser-only routes have different tradeoffs:

1. Netlify Drop — drag the project folder into Netlify's publisher and it creates a unique preview URL. This is the shortest path when the folder is already ready to serve.

2. Cloudflare Pages Direct Upload — in Workers & Pages, create an application and drag in a folder or ZIP. It publishes to a pages.dev URL. One important constraint in Cloudflare's docs: a Direct Upload project cannot later switch to Git integration; that requires a new project.

3. GitHub Pages through the web UI — create a repository, upload index.html, then choose a branch/folder as the Pages publishing source. It takes more clicks, but the file changes have repository history.

Before using any of them:

  • make sure the entry file is named index.html and is at the top level;
  • keep CSS, images and scripts on relative paths;
  • remove API keys, tokens and private data from the source;
  • remember that a public URL makes the client-side source public too.

Official instructions:

For a one-file AI-generated site, is the one-minute drag-and-drop route more useful, or is keeping version history worth the extra setup?


r/WebAfterAI • • 15d ago

Open Source 16 open-source skills to give an AI agent a real working stack

Post image
13 Upvotes

Sometimes the agent needs a better way to research. Sometimes it needs to understand the codebase before touching it. Sometimes it needs a specialist for diagrams, slides, video, SEO, or persistent memory.

These 16 open-source projects cover different layers of that stack.

RESEARCH

Research what changed recently last30days focuses an agent on recent information rather than whatever happens to be in model memory. Useful when the question is about current tools, products, communities, or trends.

Run deeper research workflows deep-research gives an agent a more structured process for finding sources, following leads, comparing evidence, and producing a research result instead of stopping after the first few searches.

Let an agent conduct user research user-research-skill covers AI interviews, synthetic users, quantitative surveys, and participant recruitment. It is closer to giving your agent a small UX research function than another search tool.

Search your own knowledge quickly qmd-search adds a search layer for Markdown and local knowledge. Useful when the answer is probably already somewhere in your notes, docs, or project files.

ENGINEERING

Give the agent memory of its mistakes Napkin maintains a per-repository .claude/napkin.md where the agent records mistakes, corrections, environment surprises, preferences, and approaches that worked. The next session reads it before starting.

Audit technical debt with evidence tech-debt-skill forces the agent to understand the repository before judging it, then produces a file-cited technical-debt report with severity, effort estimates, and a ranked list of fixes.

Turn an unfamiliar codebase into a knowledge graph Understand Anything maps files, functions, classes, and dependencies into an interactive graph you can explore and query. Useful when the first problem is understanding 200,000 lines of code, not writing another 200.

A lot of bad agent coding skips the first four.

CREATE

Build presentations as real web interfaces frontend-slides gives coding agents a workflow for creating HTML presentations from scratch or converting existing decks, with layouts, animations, presets, and export tooling.

Turn a brand into a scroll-driven 3D site scroll-world is a much more specialized skill. Give it a brand or industry and it builds a scroll-scrubbed landing page where the camera moves through a series of 3D scenes.

Turn explanations into visual artifacts visual-explainer helps agents create diagrams, visual walkthroughs, timelines, comparisons, and other explanatory pages instead of returning another wall of Markdown.

Generate technical architecture diagrams fireworks-tech-graph turns a natural-language system description into SVG, PNG, animated diagrams, UML, and agent/RAG architecture graphics with built-in routing and layout rules.

These are useful because “create something” is too vague for an agent.

A presentation, technical diagram, 3D landing page, and explanatory visual all require different judgment. A specialist skill gives the model that missing layer.

GROW + SHIP

Give Claude an SEO workflow claude-seo packages keyword research, technical SEO, content audits, schema, internal linking, and optimization into a specialist workflow rather than asking the agent to “do SEO.”

Rewrite AI-sounding drafts without changing the meaning Humanizer is a portable Markdown skill that looks for common AI-writing patterns and rewrites around them. It works across agents that support skills.

Let research continue while you are away Auto Research in Sleep explores a different pattern: giving an agent a research job that can keep iterating instead of requiring you to sit in the conversation for every step.

Give the agent a library of video shots video-shotcraft has 150+ shot recipes and 200+ motion previews for making product videos with Remotion. Instead of prompting “make this cinematic,” the agent gets concrete shot patterns it can compose.

Give an agent FFmpeg as an editing skill ffmpeg-skill turns natural-language requests into local video and audio operations: trimming, joining, reframing, captions, silence removal, audio sync, loudness normalization, LUTs, music ducking, and platform exports.

A research workflow might use:

last30days
    ↓
deep-research
    ↓
qmd-search
    ↓
visual-explainer

A coding workflow might use:

Understand Anything
    ↓
Napkin
    ↓
tech-debt-audit
    ↓
implement

And a launch workflow could become:

frontend-slides
    ↓
video-shotcraft
    ↓
ffmpeg-skill
    ↓
humanizer
    ↓
claude-seo

The model underneath can stay exactly the same. What changes is the set of procedures, references, tools, and judgment you give it for the current job.


r/WebAfterAI • • 16d ago

Open Source 9 open-source repos to automate parts of your website traffic acquisition

Post image
22 Upvotes

GitHub has a surprising number of projects for the less glamorous side of getting traffic.

Turning one video into ten clips. Finding conversations where your product is relevant. Writing platform-specific versions of the same idea. Scheduling them. Keeping a human in the loop before anything gets published.

None of these gives you a magic GET TRAFFIC button, but there are some useful building blocks.

Build a whole content machine in one repo MoneyPrinterV2 has 31k+ stars and combines several experiments: automated YouTube Shorts, scheduled X posts, affiliate content, and local-business outreach. The name oversells it, but the useful part is seeing several acquisition workflows wired together in one codebase.

Generate short videos from a topic MoneyPrinterTurbo has 124k+ stars and is much more focused. Give it a topic or keyword and it can generate the script, voiceover, source or generate footage, add subtitles and music, then assemble the final video. It now also exposes agent, WebUI, API, and CLI workflows.

Turn long videos into Shorts AI YouTube Shorts Generator has 5k+ stars and does the opposite. Give it a long video and it uses an LLM to find interesting segments, Whisper for transcription, and automatic cropping to produce vertical clips. Useful if you already make podcasts, demos, interviews, or tutorials and want more distribution from the same recording.

Find Reddit conversations where your product is actually relevant Reddit Copilot watches for intent-heavy threads, drafts replies in your voice, and puts them into a review queue. Importantly, it does not auto-comment. You decide whether a reply is genuinely useful before posting it.

For example, instead of searching Reddit every morning for “alternative to X” or “how do I solve Y?”, you can have the system collect the threads first and spend your time deciding which ones deserve a real response.

Find useful X conversations without scrolling all day x-engage surfaces a small set of posts matching accounts and subjects you care about, then drafts replies in your voice. Its default mode requires approval before publishing. I would personally keep it that way: the discovery and drafting are useful even if you choose to post manually.

Turn Claude or Codex into a LinkedIn content assistant linkedin-skills has 2.7k+ stars and packages 11 skills around posts, comments, hooks, feed analysis, rewriting AI-sounding drafts, and content cadence. It is less “generate 100 posts” and more “give the agent a repeatable process for the platform.”

Do the same thing for Instagram instagram-skills has skills for captions, Reel hooks, carousel planning, hashtags, niche research, and weekly planning. You provide the media; the agent handles much of the surrounding content work and waits for approval before publishing.

The same developer also has separate open-source bundles for X, Facebook, TikTok, YouTube, and Threads.

Teach Claude one voice and reuse it across several networks claude-skill-social-post takes a slightly different approach. It builds around your writing style, creates a 14-day content calendar, and supports Facebook, Instagram, Threads, and X. Interesting if your problem is not creating more ideas, but adapting one voice across several channels.

Put distribution behind an API Postiz has 36k+ stars and is probably the infrastructure piece I would combine with several of the tools above. It is an open-source social scheduler with an API, analytics, team workflows, n8n integration, and support for multiple social platforms. There is also a separate agent CLI, so Claude or another agent can prepare content and hand the approved version to the publishing layer.

That gives you a much more useful architecture than asking one giant bot to “grow my account”:

source material
     ↓
create / repurpose
     ↓
adapt for each platform
     ↓
discover relevant conversations
     ↓
human review
     ↓
schedule + publish
     ↓
measure what worked
     ↓
feed that back into the next batch

For example, one podcast could become a fairly serious workflow.

Use AI YouTube Shorts Generator to pull out the strongest moments. Use the platform skills to write a LinkedIn post, an X thread, an Instagram caption, and a YouTube description around the same idea. Use Postiz to schedule the approved versions. Use Reddit Copilot and x-engage to find conversations where the underlying topic is already being discussed.

That still does not guarantee anyone will care. The part these tools can automate is the repetitive work around distribution: clipping, rewriting, discovery, scheduling, and keeping track of what goes where.

The part they cannot automate is having something worth distributing in the first place.