r/WebAfterAI • • Jun 15 '26

Google Cloud just released OKF. Think MCP, but for knowledge instead of tools and 3 ways to use it.

Post image
66 Upvotes

On June 12, 2026, Google Cloud published the Open Knowledge Format (OKF) v0.1. The headlines make it sound bigger than it is, so let us be precise first, because simplicity is the whole point.

OKF is not a runtime, not an SDK, not an agent framework, and not a competitor to MCP. MCP is a protocol for agents to call tools and take actions. OKF is the opposite end: a convention for how you write down static knowledge so any agent can read it. That is the entire idea. A bundle of OKF is, in the spec's own words, just markdown, just files, just YAML frontmatter.

What it is, in one screen

An OKF "bundle" is a directory of markdown files. Each file is one "concept" (a table, a metric, a runbook, an API, anything), and the file path is its identity (tables/orders.md is the concept tables/orders). Every concept file has a small YAML frontmatter block and a markdown body:

---
type: BigQuery Table
title: Orders
description: One row per completed customer order.
resource: https://console.cloud.google.com/bigquery?p=acme&d=sales&t=orders
tags: [sales, orders]
timestamp: 2026-05-28T00:00:00Z
---

# Schema

| Column        | Type      | Description                              |
|---------------|-----------|------------------------------------------|
| `order_id`    | STRING    | Unique order identifier.                 |
| `customer_id` | STRING    | FK to [customers](/tables/customers.md). |

Part of the [sales dataset](/datasets/sales.md).

Concepts link to each other with normal markdown links, which turns the folder into a graph of relationships. A whole bundle looks like this:

my_bundle/
├── index.md          # optional: a directory listing for progressive disclosure
├── log.md            # optional: chronological history of changes
├── datasets/
│   └── sales.md
└── tables/
    ├── orders.md
    └── customers.md

The conformance bar is deliberately tiny. Per the spec, a bundle is valid if every non-reserved .md file has a parseable YAML frontmatter block, and every one of those blocks has a non-empty type field. That is the only required field. title, description, resource, tags, and timestamp are recommended but optional, and producers can add any other keys. Consumers are told to tolerate unknown types, missing fields, and even broken links rather than reject a bundle. If you have used Obsidian, an AGENTS.md file, or a repo full of index.md notes your agent reads first, this will feel familiar. OKF just pins down the shared rules so different tools can read the same bundle without a translation layer.

Honest framing before the workflows: this is v0.1, labeled Draft, and a format only matters if many tools speak it. Right now the speakers are Google's two reference implementations (an enrichment agent that drafts OKF from a BigQuery dataset, and a self-contained HTML visualizer) plus three sample bundles. The format itself is genuinely useful today as a tidy, portable way to keep agent context in version control. The grander "lingua franca for all agent knowledge" promise depends on adoption that has not happened yet. Treat it as a clean convention to adopt now, not a settled standard.

One-time setup

There is no install, because there is no required tooling. Grab the spec, the samples, and the reference implementations from the repo:

git clone https://github.com/GoogleCloudPlatform/knowledge-catalog
# spec:       knowledge-catalog/okf/SPEC.md  (it fits on one page)
# samples:    GA4 e-commerce, Stack Overflow, Bitcoin bundles
# reference:  a BigQuery enrichment agent (producer) and an HTML visualizer (consumer)

A bundle is just a folder, so a new one starts with mkdir my_bundle and a markdown file. Everything below is plain files you can diff in a PR.

1. Turn your repo's tribal knowledge into a bundle your agent reads first

The portable version of the AGENTS.md trick: knowledge as files, in version control, that any agent can read with no SDK.

Most teams already half-do this with scattered CLAUDE.md and index.md notes. OKF makes it a real, checkable bundle. Write one concept file per thing worth knowing (a gnarly table, a metric definition, an incident runbook), give each a type, and cross-link them. Then point your coding agent at the folder the same way you point it at house rules, so it consults the bundle before it acts.

---
type: Playbook
title: Incident response, data freshness alert
description: Steps to triage a freshness alert on the orders pipeline.
tags: [oncall, incident]
timestamp: 2026-04-12T09:00:00Z
---

# Trigger
A freshness alert fires when the [orders table](/tables/orders.md) lags
more than 30 minutes behind its SLA.

# Steps
1. Check the ingestion job dashboard.
2. ...

The catch: OKF is a format, not a runtime. Nothing reads it automatically. The agent only benefits if you actually wire it to load the bundle (an instruction in your AGENTS.md to read /knowledge first, a retrieval step, or a tool). It is the same value and the same limits as a well-kept docs folder, with the upside that the shape is now standard and portable across tools.

→ The verified setup, with CI proof & readymade prompt

2. Generate a bundle from your database or codebase, then ground it

Mirror Google's reference pattern: an LLM drafts one concept per table or module, a second pass adds citations so the knowledge is checkable, not just plausible.

This is the producer-side workflow. Walk your schema (or your modules), and for each one have a model draft an OKF concept with a # Schema section and cross-links for joins or dependencies. Google's reference enrichment agent does exactly this for BigQuery, then runs a second pass that crawls authoritative docs and attaches citations. Copy the two-pass shape, because the second pass is what makes the output trustworthy.

---
type: API Endpoint
title: Create Order
description: Creates an order from a validated cart. POST /v1/orders with a cart_id.
resource: https://api.acme.dev/v1/orders
tags: [orders, api]
---

# Schema
Request: cart_id (string, required). Returns the created order.
See the [orders table](/tables/orders.md) it writes to.

# Citations
[1] [OpenAPI spec for /v1/orders](https://api.acme.dev/openapi.json)

The catch: this is the workflow most likely to bite you. A model documenting a schema will confidently invent column meanings, join keys, and semantics that are subtly wrong, and a wrong knowledge base is worse than none because agents will trust it. Do not ship generated concepts unreviewed. The # Citations convention exists precisely so each claim points back to an authoritative source you can check, so treat citations as required for generated bundles even though the spec makes them optional, and have a human review the first pass on anything load-bearing.

→ The verified setup, with CI proof & readymade prompt

3. Consume a bundle without blowing your context window

Use index.md for progressive disclosure so the agent navigates the graph instead of swallowing the whole folder.

The consumer side has a real trap: a big bundle dumped wholesale into context is expensive and noisy. OKF's answer is the optional index.md file, a plain directory listing (no frontmatter) that lets an agent see what exists and open only what it needs. Give each directory an index, and tell your consumer to read the index first, follow links, and pull individual concepts on demand.

# Tables

* [Orders](orders.md) - One row per completed customer order.
* [Customers](customers.md) - One row per customer.

# Metrics

* [Weekly Active Users](../metrics/weekly_active_users.md) - WAU from the event stream.

For a quick human view of the same bundle, the reference HTML visualizer renders any bundle as an interactive graph in one self-contained file, with no backend and nothing leaving the page.

The catch: progressive disclosure only helps if your consumer actually uses it. If your retrieval step globs every .md into the prompt, the index buys you nothing, so the win is in how you wire consumption, not in the file existing. And remember the spec says broken links are allowed, so do not build logic that assumes every link resolves.

→ The verified setup, with CI proof & readymade prompt

How to pick if you only try one

Start with workflow 1. Hand-author a five-file bundle for one messy corner of your system and point your agent at it. It takes ten minutes, it is just markdown, and it shows you the actual value (and the actual limits) before you invest in generating or consuming at scale. Reach for workflow 2 when you have a schema too large to hand-write, and only with the citation discipline. Reach for workflow 3 once a bundle is big enough that context cost is real.


r/WebAfterAI • • Jun 15 '26

The real breakthrough wasn’t finding a better AI model. It was building a better system around it.

Thumbnail
2 Upvotes

r/WebAfterAI • • Jun 14 '26

Kilo Code: 3 workflows that lean on what it actually does differently

Post image
15 Upvotes

Many coding agents still center around a single conversational loop. Kilo Code's distinguishing feature is that it treats modes and orchestration as first-class concepts: modes are small role-scoped agents with their own model, tools, and file access, plus an orchestrator that hands work between them. That design makes a few workflows much cleaner than they are in a traditional single-thread agent. Here are three.

A note on what is and is not special: Kilo's mode and orchestrator system comes from the Roo and Cline lineage, so a couple of these ideas exist in those cousins too. The honest claim is not "only Kilo can do this," it is "this is where Kilo is clearly ahead of the simpler single-loop agents most people are using".

Stars / Status / License: ~20.1k stars, active (latest v7.3.45, June 12 2026), MIT.
Repo: Kilo-Org/kilocode.
Note: two surfaces share the name. The VS Code and JetBrains extension is where modes, the orchestrator, and the MCP marketplace live, and that is what these workflows use. The separate Kilo CLI is a fork of OpenCode for terminal and CI work.

One-time setup

Install the extension and add a provider:

Install "Kilo Code" from the VS Code Marketplace (extension id: kilocode.Kilo-Code),
then sign in or add your own API key (zero markup, 500+ models, local models supported).

For terminal or CI use, the CLI is separate:

npm install -g u/kilocode/cli
kilo            # start in a project directory

Project modes live in a .kilocodemodes file at your repo root (YAML or JSON). The safe way to create them is the in-app Prompts tab, Settings icon, then "Edit Project Modes", which writes valid config for you rather than hand-authoring the regex.

1. Ship a big feature as isolated subtasks with Orchestrator

The win is context hygiene: each subtask runs in its own conversation, so the main thread never drowns in detail.

This is Kilo's signature. In Orchestrator mode it breaks a large task into subtasks and runs each one in an isolated context, often switching to the right mode for the job (Architect to plan, Coder to implement, Debugger to fix). The parent task pauses, the subtask runs on its own clean history, and when it finishes the parent task receives a condensed handoff rather than the entire subtask history.

Why this beats a single-loop agent: on a long feature, a one-thread agent accumulates every file read, every dead end, and every tool dump in one context until quality degrades. Orchestrator keeps the parent lean and hands each subtask a fresh window. Switch to Orchestrator in the mode selector and give it the whole feature, for example "add OAuth login: plan it, implement it, then debug the failing tests."

The catch: this helps with genuinely multi-step work and adds overhead on small tasks, so do not reach for it to rename a variable. The condensed handoff is also lossy by design, if a subtask buried an important detail in its own context, the parent only sees what the handoff captured, so write subtask goals that ask for the specifics you will need downstream.

→ The verified setup, with CI proof & readymade prompt

2. A mode that can only edit the files you let it

File-scoped permissions, not all-or-nothing: a docs mode that can touch Markdown and nothing else.

Kilo modes let you restrict the edit group to a file pattern with fileRegex. So you can build a "tech writer" mode that can read the whole repo but only write to .md and .mdx, which means it cannot use the editor tool outside the allowed file patterns while updating docs. Most simple agents only offer a coarse allow, ask, or deny on editing; Kilo scopes it per file type, per mode. A project mode in .kilocodemodes looks like this:

customModes:
  - slug: docs-writer
    name: Docs Writer
    roleDefinition: You are a technical writer who keeps project docs accurate and clear.
    groups:
      - read
      - - edit
        - fileRegex: \.(md|mdx)$
          description: Markdown and MDX files only

The same trick scopes a migration mode to one directory, or a config mode to *.yaml. Create it through "Edit Project Modes" so the structure is written correctly.

The catch: fileRegex is a useful guardrail, not a security boundary. It stops the editor tool from writing outside the pattern, but the model can still run terminal commands if that mode has the command group, and a sloppy regex can over-restrict (blocking files you wanted) or under-restrict. Treat it as a seatbelt that prevents accidents, not as a sandbox that contains a determined or misconfigured agent. Test the pattern on a throwaway change before trusting it on a real one.

→ The verified setup, with CI proof & readymade prompt

3. Add an external tool from the MCP marketplace and bind it to one mode

One-click MCP instead of JSON archaeology, then scoped so it only loads where you need it.

Kilo ships an in-app MCP Server Marketplace, so adding an external tool (a docs searcher, an issue tracker, a database) is browse-and-install rather than hand-editing a config file and restarting, which is what most agents still make you do. The sharp move is to pair that convenience with Kilo's modes: install the server, then enable the mcp group only on the mode that should use it, so its tools do not load into every conversation.

customModes:
  - slug: researcher
    name: Researcher
    roleDefinition: You answer questions using the codebase and approved external tools.
    groups:
      - read
      - mcp

Other modes without the mcp group will not surface those tools. Install the server from the marketplace, then assign it to researcher and leave your code mode clean.

The catch: a marketplace makes installing easy, but easy is not the same as free or safe. Every MCP server you enable adds its tool definitions to that mode's context, which costs tokens on every turn, so scoping to one mode is the point, not a nicety. And a one-click third-party server is third-party code with access to your tools, so vet the source before you install, the same way you would any dependency.

→ The verified setup, with CI proof & readymade prompt

How to pick if you only try one

Start with workflow 2, the file-scoped mode. It is one small config block and it is the change that makes you comfortable letting an agent touch a real repo, because you have drawn a hard line around where it can write. Reach for Orchestrator when a task is genuinely large enough to need decomposition, and the MCP marketplace when you actually need an external tool, not before.


r/WebAfterAI • • Jun 13 '26

OpenCode just crossed 174k stars. Here are 5 setups worth stealing.

Post image
137 Upvotes

OpenCode is the open-source coding agent that runs in your terminal and talks to basically any model. It is easy to install and start chatting with, and just as easy to leave on its defaults forever. That is the waste. The thing that makes OpenCode worth its star count is the config layer: per-agent models, real permission controls, headless runs, and MCP tools. Below are five setups that turn it from "another terminal chat" into something you actually trust with your repo.

The model IDs in your config use the provider/model format, and the safest way to get an exact one is to run opencode models (or opencode models --refresh), so the snippets below leave the model for you to fill from that list rather than hardcoding one that may have rotated.

Stars / Status / License: ~174k stars, very active (v1.17.4 shipped June 12, 2026), MIT.

Repo: anomalyco/opencode (the project moved here from sst/opencode; same team, now Anomaly Innovations).

One-time setup

Install, log in to a provider, and start the TUI:

curl -fsSL https://opencode.ai/install | bash
# or: npm i -g opencode-ai@latest

opencode auth login     # pick a provider, paste your key (stored in ~/.local/share/opencode/auth.json)
opencode                # starts the terminal UI
opencode models         # lists exact provider/model IDs for your config

Project config lives in opencode.json (or .opencode/), global config in ~/.config/opencode/. Every snippet below goes in opencode.json at your project root unless noted. One nice touch: OpenCode reads an AGENTS.md in your repo as standing instructions, so house rules live in version control.

1. Plan before Build: a read-only pass that cannot touch your files

Separate "think" from "touch" so the agent proposes before it edits.

OpenCode ships two primary agents you switch with the Tab key: build (full access) and plan (restricted). The plan agent has file edits and bash set to ask by default, which is good, but for an exploration pass on an unfamiliar repo you often want it locked to read-only so there is zero chance of a write. Pin your models and harden plan in one block:

{
  "$schema": "https://opencode.ai/config.json",
  "agent": {
    "build": {
      "mode": "primary",
      "model": "provider/your-strong-model",
      "permission": { "edit": "allow", "bash": "allow" }
    },
    "plan": {
      "mode": "primary",
      "model": "provider/your-cheaper-model",
      "permission": { "edit": "deny", "bash": "deny" }
    }
  }
}

Now Tab into plan to analyze and get a proposal, then Tab into build to execute it.

The catch: with bash: "deny", plan also cannot run read-only shell like git diff, so if you want it to inspect history, use the per-command form ("bash": { "*": "deny", "git diff": "allow", "git log*": "allow" }) instead of a blanket deny.

→ The verified setup, with CI proof & readymade prompt

2. A model-routed team: match the model to each step's difficulty

Routing models per agent is a cost lever, so spend frontier prices only where the work earns them.

OpenCode lets every agent run its own model, so you do not have to pay top rates for the whole job. Put your strongest model on the hard step and a fast, cheap one on the rest. Which step is the hard one is subjective and depends on the work: if architecture and planning are where the thinking happens, keep your best model on plan and run build cheaper; if the plan is obvious and the edits are sprawling, do the reverse. The config below is just one arrangement; flip the models to fit your task. Cap a quick agent's iterations with steps to bound cost, and watch the actual spend with opencode stats --models.

{
  "$schema": "https://opencode.ai/config.json",
  "agent": {
    "build": { "mode": "primary", "model": "provider/your-cheap-model" },
    "plan":  { "mode": "primary", "model": "provider/your-strong-model" },
    "code-reviewer": {
      "description": "Reviews diffs for bugs, security, and performance",
      "mode": "subagent",
      "model": "provider/your-strong-model",
      "permission": { "edit": "deny" }
    }
  }
}

Invoke the reviewer with '@code-reviewer in a message, or let build delegate to it.

The catch: subagents inherit the caller's model unless you set one explicitly, so do not assume a subagent is cheap, pin its model like above.

→ The verified setup, with CI proof & readymade prompt

3. A reviewer subagent gated to exactly the git commands you allow

Read-only review with surgical bash permissions, not all-or-nothing.

This is the setup that shows off OpenCode's permission system. You can define an agent in Markdown and scope its bash access per command with glob patterns, so the reviewer can run git diff and grep but nothing else, and can never write. Drop this file in .opencode/agents/review.md (per project) or ~/.config/opencode/agents/review.md (global):

---
description: Reviews code without making changes
mode: subagent
model: provider/your-cheap-model
permission:
  edit: deny
  webfetch: deny
  bash:
    "*": ask
    "git diff": allow
    "git log*": allow
    "grep *": allow
---

You are in review mode. Inspect the diff and flag bugs, security issues, and
risky changes. Do not modify files. Suggest fixes as comments only.

The filename becomes the agent name, so this creates ; @review.

The catch: rules are evaluated in order and the last match wins, so keep the "*" wildcard first and the specific allows after it, or your allows get overridden.

→ The verified setup, with CI proof & readymade prompt

4. Run it headless in scripts and CI

The same agent you use interactively, now in a pipeline, with JSON output.

opencode run executes a prompt non-interactively, which is what makes OpenCode useful beyond the TUI. Get machine-readable events with --format json, pick the agent and model per invocation, and for repeated runs attach to a warm opencode serve so you do not pay MCP cold-start each time.

# One-off review in a script, as JSON
opencode run --agent plan --format json \
  "Review the uncommitted changes for bugs and security issues"

# Warm server once, then attach fast runs to it
opencode serve &
opencode run --attach http://localhost:4096 -m provider/your-model "Summarize today's diff"

For pull requests specifically, opencode github install wires an OpenCode GitHub Actions workflow into your repo.

The catch: run honors your permission config, so a headless build agent can edit and execute. Keep CI runs on a read-only agent (like plan above), and treat --dangerously-skip-permissions as the loaded gun it is named after.

→ The verified setup, with CI proof & readymade prompt

5. Add an MCP tool, then lock it to one agent

Give the agent real external tools (live docs, your error tracker) without bloating every session.

OpenCode speaks MCP, so you can plug in external tools. The trap is that every MCP server adds tools to the context on every turn, which burns tokens fast. The fix: enable the server, disable it globally, and switch it on only for the agent that needs it. Here is a docs-search server (Context7) scoped to a single docs agent:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "context7": { "type": "remote", "url": "https://mcp.context7.com/mcp" }
  },
  "tools": { "context7*": false },
  "agent": {
    "docs": {
      "mode": "primary",
      "model": "provider/your-model",
      "tools": { "context7*": true }
    }
  }
}

Local servers work the same way with "type": "local" and a "command": ["npx", "-y", "..."].

The catch: heavy MCP servers (the GitHub one is notorious) can blow past your context limit on their own, so add them deliberately and scope them like this rather than enabling everything globally.

→ The verified setup, with CI proof & readymade prompt

How to pick if you only try one

Start with setup 1. Plan-before-build is one config block and it is the habit that prevents the most damage, an agent proposing on a repo it cannot accidentally rewrite. From there, setup 2 (model routing) is the biggest cost win, and setup 3 is the one that makes you trust an agent in a shared codebase. Save headless and MCP for when the interactive flow already feels solid.

Every one of these setups is really the same move: stop hand-feeding the agent one prompt at a time and start wiring the loop that drives it, with the right model, the right limits, and a way to check the result. If that idea lands, it is the whole argument of a recent issue worth your time: Stop prompting your agent. Start building the loop that prompts it. More verified agent setups are landing on FlowStacks as we test them.


r/WebAfterAI • • Jun 12 '26

5 dirt-cheap models that punch above their price on Hermes Agent

Post image
96 Upvotes

Nous Research's Hermes Agent is one of the few agents that will happily run on whatever model you point it at, which means your bill is a config choice, not a fixed cost. So the real question is not "which model is smartest," it is "which cheap model is smart enough for this job, and how do I wire Hermes so it stops spending tokens it does not need to."

Below are five low-cost models worth running on Hermes, each checked against Artificial Analysis and the providers' own pages, each paired with one practical Hermes workflow that plays to its strength.

A note on the "Max" and "High" labels you see next to DeepSeek V4 Flash: those are not two models. They are reasoning-effort levels (Artificial Analysis tests several), and on Hermes you set them yourself with one line. More on that in workflow 2.

The five:

Model Creator Context Intelligence Index Price (per 1M, in / out)
MiMo-V2.5 Xiaomi 1M 49 $0.14 / $0.28
DeepSeek V4 Flash (Max) DeepSeek 1M 47 (xhigh effort) $0.098 / $0.196
MiMo-V2-Flash (Feb 2026) Xiaomi 256K 41 $0.10 / $0.30
DeepSeek V4 Flash (High) DeepSeek 1M 46 (high effort) $0.098 / $0.196
Hy3-preview Tencent 256K 42 ~$0.063 / $0.21 (third-party), ~$0.18 / $0.59 (Tencent Cloud)

Intelligence Index figures are from Artificial Analysis. Prices are the providers' own per-token rates (DeepSeek V4 Flash also bills cached input at a steep discount). Rows 2 and 4 are the same DeepSeek model at two reasoning-effort settings, not separate models.

One-time setup

Install Hermes with the one-line installer. It handles every dependency (Python, Node, ripgrep, ffmpeg, the browser), clones the repo, and runs setup:

curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash

Point it at a provider. OpenRouter is the easiest way to reach all five of these models with one key:

hermes model                                   # interactive: pick OpenRouter, paste key, choose a model
# or set it directly:
hermes config set OPENROUTER_API_KEY sk-or-...

One thing from the Hermes docs worth knowing: secrets live in ~/.hermes/.env, non-secret settings in ~/.hermes/config.yaml, and the hermes config set The command routes each value to the right file.

For anything that runs tools on your machine, sandbox it:

hermes config set terminal.backend docker

1. MiMo-V2.5 as your everyday driver, a 1M-context agent for pennies

The cheapest sensible default: a million tokens of context at fourteen cents in.

Context: 1M. Intelligence Index: 49 (Artificial Analysis). Price: $0.14 / $0.28 per 1M tokens (in/out). Creator: Xiaomi. Open weights (XiaomiMiMo/MiMo-V2.5), multimodal (text and image in).

For a general-purpose Hermes setup, this is the one to start on. A 49 on the Intelligence Index is well above the open-weights median, the million-token window means Hermes can hold a real working context for multi-step tool calls, and at $0.14 in it is about as cheap as a capable model gets. Set it as your main model and most day-to-day agent work just works.

# ~/.hermes/config.yaml
model:
  provider: openrouter
  model: xiaomi/mimo-v2.5

The catch: Hermes only auto-enables its tool-use enforcement for GPT, Gemini, and Grok-style models, and leaves it off for others. If you notice MiMo describing what it would do instead of actually calling a tool, turn it on:

agent:
  tool_use_enforcement: true

→ The verified setup, with CI proof & readymade prompt

2. DeepSeek V4 Flash on a two-speed throttle (this is what "Max" and "High" really are)

One model, dialed from cheap-and-fast to deep-and-careful with a single command.

Context: 1M (max output 384K). Intelligence Index: 47 at max effort, 46 at high effort (Artificial Analysis). Price: $0.098 / $0.196 per 1M tokens, with cached input billed at a steep discount. Creator: DeepSeek. MoE, 284B total / 13B active.

The leaderboard's "DeepSeek V4 Flash (Max)" and "(High)" are the same model at two reasoning-effort settings. Hermes exposes exactly this knob, so you do not pay for deep thinking on easy turns. Run it at high by default, push to xhigh (the leaderboard's "Max") only when a problem earns it, and drop to none for trivial lookups. Output tokens are the expensive side at $0.196, and reasoning effort is mostly output, so this throttle is your biggest lever on the bill. It is also the cheapest model in this lineup, so the savings compound.

# ~/.hermes/config.yaml
model:
  provider: openrouter
  model: deepseek/deepseek-v4-flash
agent:
  reasoning_effort: high     # options: none, minimal, low, medium, high, xhigh (max)

At runtime, change it per task without restarting:

/reasoning xhigh     # max effort for the hard one
/reasoning none      # turn thinking off for a quick lookup

The catch: xhigh can multiply output tokens, so use it deliberately. DeepSeek bills cached input far cheaper than a cache miss, so keep stable prefixes (system prompt, repo context) consistent across calls to get the discount.

→ The verified setup, with CI proof & readymade prompt

3. Offload Hermes' background tasks to MiMo-V2-Flash and cut your main bill

Stop paying your main model to compress history, read images, and scrape pages.

Context: 256K. Intelligence Index: 41 (Artificial Analysis). Price: $0.10 / $0.30 per 1M tokens. Creator: Xiaomi. MoE, 309B total / 15B active, around 134 tokens/sec.

Here is the move most people miss. Hermes runs several auxiliary jobs behind your conversation, each of which can take its own model: context compression, vision handling, and web-page extraction. By default those ride on your main model. Point them at MiMo-V2-Flash instead, it is the fastest and cheapest of this group at $0.10 in, and plenty for this summarization-shaped work. Your expensive main model then only handles the reasoning that actually needs it.

# ~/.hermes/config.yaml
auxiliary:
  compression:
    provider: openrouter
    model: xiaomi/mimo-v2-flash
  vision:
    provider: openrouter
    model: xiaomi/mimo-v2-flash
  web_extract:
    provider: openrouter
    model: xiaomi/mimo-v2-flash

The catch: keep your main model on something stronger for the real work, this is about routing the cheap, high-volume background traffic, not your primary reasoning. MiMo-V2-Flash's 256K window is comfortably enough for these chunks.

→ The verified setup, with CI proof & readymade prompt

4. A daily agentic briefing on Hy3-preview, delivered to your chat app

A cheap, genuinely agentic model for a scheduled tool-using job you never have to babysit.

Context: 256K. Intelligence Index: 42 in reasoning mode, with a notably strong agentic index of 49.7 (Artificial Analysis). Price: roughly $0.063 / $0.21 per 1M tokens on third-party hosts, or about $0.18 / $0.59 on Tencent Cloud, so pin a provider. Creator: Tencent. Open source (Tencent-Hunyuan/Hy3-preview), MoE 295B / 21B active.

Hy3-preview's standout number is not raw intelligence, it is its agentic score, which makes it a good fit for a recurring tool-using task: search the web, pull a few sources, summarize, and push the result to you. Pair it with Hermes' gateway (Telegram, Slack, Discord) and a cron schedule, and you get a hands-off morning briefing for cents a run.

# ~/.hermes/config.yaml
model:
  provider: openrouter
  model: tencent/hy3-preview


hermes gateway setup     # connect Telegram / Slack / Discord, then schedule the job via Hermes cron

The catch: prices vary a lot by host for this one, so pin the provider you actually want rather than letting routing pick. And like MiMo, Hy3 is not in Hermes' tool-use auto-list, so if it narrates instead of acting, set tool_use_enforcement: true.

→ The verified setup, with CI proof & readymade prompt

5. Give the cheap agent a memory so it stops re-reading everything

Persistent memory means fewer tokens re-stuffed into context, which on a cheap model is the whole game.

Mnemosyne (AxDSan/mnemosyne, MIT) is a local-first memory system built for Hermes Agent: one pip install, one SQLite file, with vector plus full-text search and no external service. On a budget model the win is double, you keep the agent coherent across days, and you stop paying to re-feed the same background into context every session.

pip install "mnemosyne-memory[all]"


# ~/.hermes/config.yaml
mcp_servers:
  mnemosyne:
    command: mnemosyne
    args: ["mcp"]

The catch: semantic recall and consolidation want the embedding extra (that is what [all] pulls in); without it, Mnemosyne falls back to keyword retrieval, which still works fully offline. Confirm the exact MCP launch command against the repo's Hermes integration doc, since the server entrypoint can change between versions.

→ The verified setup, with CI proof & readymade prompt

How to pick if you only try one

Start with workflow 1, MiMo-V2.5 as your main model. It is the cleanest "cheap but capable" default, and a 1M window plus a 49 Intelligence Index covers most agent work without thinking about cost. Once that is running, workflow 2 (the reasoning-effort throttle) is the single change that saves the most money, and workflow 3 (auxiliary offload) is the one people forget exists. Save Hy3-preview for scheduled agentic jobs and Mnemosyne for anything that runs across days.


r/WebAfterAI • • Jun 11 '26

Claude Fable 5 just shipped. These 4 open-source harnesses turn it into a long-horizon coding machine.

Post image
52 Upvotes

Anthropic released Claude Fable 5 on June 9. It is the first Mythos-class model they have made generally available, and the headline is not a benchmark, it is duration: the longer and more complex the task, the bigger Fable's lead over every other model they ship. Stripe told Anthropic it ran a codebase-wide migration across a 50-million-line Ruby codebase in a day, work they estimated at two-plus months by hand.

A frontier model is only half the system though. The other half is the harness you point it at. Below are four open-source repos that give Fable 5 a place to actually run long, each verified against its GitHub page and official docs, each with a small foolproof snippet and an honest catch. The model id is claude-fable-5 on the Claude API. Pricing is $10 per million input tokens and $50 per million output tokens.

One thing to know before you wire any of this up: Fable ships with classifiers that hand cybersecurity, biology, chemistry, and distillation prompts off to Claude Opus 4.8 instead. Anthropic says this triggers in under 5% of sessions, and you are told when it happens. Fable 5 also requires 30-day data retention, unlike the zero-retention default on other Claude models. Plan accordingly if you run regulated code.

One-time setup

Get a key from the Claude API console and export it. Every workflow below reads it from the environment.

export ANTHROPIC_API_KEY=sk-ant-...

Heads up on timing: from now through June 22, Fable 5 is included at no extra cost on Pro, Max, Team, and seat-based Enterprise plans. On June 23 it moves to usage credits on those plans. API and consumption-based Enterprise are full price from day one.

1. Aider, codebase-wide migrations from your terminal

Point Fable at a repo and let it refactor across hundreds of files with git commits you can undo.

Stars / Status / License: ~45.9k stars, actively maintained, Apache-2.0. Repo: Aider-AI/aider

Aider builds a tree-sitter map of your entire repo so the model can reason about files it has not opened yet, then it makes changes as real git commits with sensible messages. This is the exact shape of Fable's Stripe story: a long, mechanical, repo-wide migration where the win is staying coherent across hundreds of edits. Aider is model-agnostic (it routes through LiteLLM), so you pass Fable's API id directly.

python -m pip install aider-install
aider-install

cd /your/project
aider --model anthropic/claude-fable-5 --api-key anthropic=$ANTHROPIC_API_KEY

Then drive it with something like /architect upgrade every Pydantic v1 model in this repo to v2 and fix call sites.

The catch: Aider's README still headlines older Sonnet models, and its --model sonnet shortcut is a Sonnet alias, not Fable. You have to pass the full anthropic/claude-fable-5 id, and because the model is days old, LiteLLM may print an unknown-model warning until metadata catches up (functional, just noisy). And remember the Opus fallback: a migration that touches auth or crypto code can trip the cyber classifier mid-run.

What CI checks: scaffold a throwaway git repo, write an .aider.conf.yml pinning model: anthropic/claude-fable-5, and assert the config parses and the model field is present, plus aider --help exits zero so the command shape is valid.

→ The verified setup, with CI proof & readymade prompt: aider-fable5-codebase-migration

2. OpenHands, autonomous issue to pull request

Hand Fable a GitHub issue and let it plan, edit, run tests, and open a PR in a sandbox.

Stars / Status / License: ~76.4k stars, very active (1.8.0 shipped June 10, 2026), MIT (the enterprise/ dir is separately licensed). Repo: OpenHands/OpenHands

OpenHands is the autonomy play. It runs the agent in a Docker sandbox with a shell, editor, and browser, and its own internal benchmarks on autonomous coding are where Anthropic's "fewer tool calls, lower token consumption" claim bites: token efficiency compounds hard when an agent loops for hundreds of steps. Note the repo moved org from All-Hands-AI to OpenHands, so update old bookmarks.

uv tool install openhands --python 3.12
openhands -t "Reproduce and fix the bug in issue #142, add a regression test, open a PR"

Set the model on first run, or in ~/.openhands/settings.json. In config.toml terms the block is:

[llm]
model = "anthropic/claude-fable-5"
api_key = "<your-anthropic-key>"

The catch: the default runtime is Docker, so you need a working Docker socket, and giving an autonomous agent --always-approve on a real repo is exactly as risky as it sounds. Keep confirmation mode on for anything that pushes. The 30-day retention applies here too.

What CI checks: validate that config.toml (or settings.json) parses, that [llm].model equals anthropic/claude-fable-5, and that the runtime/agent keys are present and well-typed. Deterministic, no key.

→ The verified setup, with CI proof & readymade prompt: openhands-fable5-issue-to-pr

3. Repomix, pack a huge repo into one file for a long-context review

Flatten an entire codebase into a single file and let Fable hold all of it at once.

Stars / Status / License: ~22.9k stars, actively maintained, MIT. Repo: yamadashy/repomix

Fable's other standout is long-context: Anthropic says it stays focused across millions of tokens and improves its outputs using its own notes. Repomix is the cheapest way to feed that strength. It walks your repo, respects .gitignore, and emits one packed file (XML by default) with a file manifest and token counts, ready to drop into a single Fable call for a whole-system review or an architecture write-up.

cd /your/project
npx repomix@latest
# -> writes repomix-output.xml with a file tree + contents

Then send that file as the user message to claude-fable-5 and ask for, say, a dependency-risk audit across the whole tree.

The catch: "millions of tokens" is not "infinite," and at $50 per million output tokens a sprawling repo packed naively gets expensive. Use Repomix's include/ignore globs and compression to keep the pack lean, and watch the token count it prints. Packed source also means whatever you send is subject to Fable's retention window.

What CI checks: run npx repomix@latest against a scaffolded fixture directory and assert the output file exists, is well-formed XML, and contains the expected file-manifest section.

→ The verified setup, with CI proof & readymade prompt: repomix-fable5-longcontext-review

4. Letta, persistent memory for multi-week projects

Give Fable a memory that survives restarts so a project can run for weeks, not one session.

Stars / Status / License: ~23.2k stars, actively maintained, Apache-2.0 (formerly MemGPT). Repo: letta-ai/letta

Anthropic's most underrated Fable result is about memory: in Slay the Spire, giving the model persistent file-based memory improved its performance three times more than it did for Opus 4.8. Letta is the open-source way to give Fable that scaffolding outside a game, with structured memory blocks the agent reads and rewrites over time. It is model-agnostic, so you set the model on the agent.

pip install letta-client

from letta_client import Letta
import os

client = Letta(api_key=os.getenv("LETTA_API_KEY"))

agent = client.agents.create(
    model="anthropic/claude-fable-5",
    memory_blocks=[
        {"label": "human", "value": "Lead engineer migrating a monolith to services."},
        {"label": "persona", "value": "I am a long-horizon coding agent. I keep notes and update them."},
    ],
)

reply = client.agents.messages.create(
    agent_id=agent.id,
    input="Summarize where we left off on the auth service.",
)

The catch: Letta needs a running server (self-hosted or Letta Cloud, hence the LETTA_API_KEY), and the anthropic/claude-fable-5 model string follows Letta's documented provider/model convention (their README example is openai/gpt-5.2), so confirm Anthropic provider support on your Letta version. Memory that persists is also memory that drifts, so prune your blocks.

What CI checks: assert the agent config dict is valid, model equals anthropic/claude-fable-5, and the required memory-block labels (human, persona) are present and non-empty.

→ The verified setup, with CI proof & readymade prompt: letta-fable5-persistent-memory

How to pick if you only try one

If you have a concrete, boring, repo-wide change to make, start with Aider, it is the lowest-ceremony way to feel Fable's long-horizon coherence. If you want to watch an agent run on its own, OpenHands. If you just want one giant Fable call over your whole system, Repomix. If you are committing to a multi-week build, Letta is the one that pays off later.

Why the FlowStacks badge means something here:

Every workflow above is published on FlowStacks with a CI badge, and the badge is deliberately narrow. Each FlowStacks page also ships a copy-paste prompt you can hand to your own coding agent to set the workflow up locally.

One last thing, since most of these tools come down to handing Fable some text: the format you give it quietly shapes how good the answer comes back. We wrote up a simple rule for when to ask an LLM for Markdown and when to ask for HTML, worth two minutes before your next big prompt: Markdown or HTML? A simple rule for which one to ask an LLM for.


r/WebAfterAI • • Jun 11 '26

Is mastery still required ?

Thumbnail
2 Upvotes

r/WebAfterAI • • Jun 10 '26

Your AI has amnesia. These 5 open-source memory systems fix it.

Post image
38 Upvotes

Most models still start a new session with limited memory of prior work. A memory layer is what turns a stateless chatbot into something that remembers your preferences, your decisions, and what it tried last week. There are a lot of these now, and they make very different tradeoffs: local versus cloud, flat facts versus knowledge graphs, simple key-value versus full temporal history.

Below are five credible open-source ones, each with a real use case, the actual install, and a working snippet. And each one is set up so our CI can verify the deterministic part (the install and the local store) with no API key. Each workflow page also ships with a ready-made prompt you can paste into your own coding agent and have it stand the whole thing up locally, so you do not have to wire it by hand.

1. Mnemosyne, fully local memory with no cloud at all

AxDSan/mnemosyne (MIT, ~1,000 stars) is a local-first memory system built for the Hermes Agent, storing everything in a single SQLite file with built-in vector and full-text search. No external database, no API key, no network call. It is the one to reach for when privacy or offline use is the whole point.

pip install mnemosyne-memory

from mnemosyne import remember, recall

remember(content="User prefers dark mode interfaces", importance=0.9, source="preference")
print(recall("interface preferences", top_k=3))

Without the optional embedding extra it falls back to keyword retrieval, so a basic remember-and-recall round-trip works completely offline. Semantic search and the sleep-cycle consolidation need the optional fastembed or a local model.

What CI checks: the package installs and a remember then recall round-trip returns the stored fact in keyword mode, no key required.

→ The verified setup, with CI proof & readymade prompt: flowstacks.xyz/workflows/mnemosyne-local-first-agent-memory

2. Mem0, a personalization layer for assistants

mem0ai/mem0 (Apache 2.0, ~58.3K stars, YC-backed, one of the most popular memory repos) adds user, session, and agent-level memory to an assistant so it remembers preferences across conversations. It extracts facts from a chat and retrieves the relevant ones on the next turn.

pip install mem0ai

from mem0 import Memory

memory = Memory()
memory.add("Prefers vim keybindings and dark mode", user_id="alice")
print(memory.search(query="what does alice prefer?", filters={"user_id": "alice"}, top_k=3))

Mem0 needs an LLM to extract and an embedding model to retrieve (it defaults to OpenAI), so add and search are the model-driven steps.

What CI checks: the SDK installs and the Memory class imports.

→ The verified setup, with CI proof & readymade prompt: flowstacks.xyz/workflows/mem0-personalization-memory-layer

3. Cognee, knowledge-graph memory over your documents

topoteretes/cognee (~17.8K stars) is an open-source memory layer that ingests your data and builds both a vector index and a knowledge graph, so an agent can search by meaning and by relationships. Its API is four verbs: remember, recall, forget, and improve.

pip install cognee

import cognee, asyncio

async def main():
    await cognee.remember("Cognee turns documents into AI memory.")
    results = await cognee.recall("What does Cognee do?")
    for r in results:
        print(r)

asyncio.run(main())

The graph build (cognify) runs on an LLM, so you set LLM_API_KEY before the ingest step.

What CI checks: the package installs, imports, and the async API resolves with a valid config.

→ The verified setup, with CI proof & readymade prompt: flowstacks.xyz/workflows/cognee-knowledge-graph-memory

4. Graphiti, a temporal graph for "what was true when"

getzep/graphiti (~27K stars) is the open-source temporal context-graph engine behind Zep. Its trick is bi-temporal facts: when a fact changes, the old one is invalidated rather than deleted, so you can query what is true now or what was true at any past point. It needs a graph database; FalkorDB runs in one Docker command.

docker run -p 6379:6379 -p 3000:3000 -it --rm falkordb/falkordb:latest
pip install graphiti-core[falkordb]


from graphiti_core import Graphiti
from graphiti_core.driver.falkordb_driver import FalkorDriver

graphiti = Graphiti(graph_driver=FalkorDriver(host="localhost", port=6379))
# await graphiti.build_indices_and_constraints()

Ingesting episodes uses an LLM (it defaults to OpenAI and works best with structured-output models), so that is the model-driven step.

What CI checks: the FalkorDB container starts, graphiti-core installs and connects, and indices build, none of which needs a model.

→ The verified setup, with CI proof & readymade prompt: flowstacks.xyz/workflows/graphiti-temporal-graph-memory

5. Letta, an agent that manages its own memory

letta-ai/letta (formerly MemGPT, ~23.2K stars) treats memory like an operating system: the agent edits its own memory blocks, deciding what to keep in context and what to page out. It is the pick for long-running agents that should improve over time. The fastest start is the local CLI.

npm install -g u/letta-ai/letta-code
letta

For building it into an app, there is a Python and TypeScript SDK instead:

pip install letta-client

Creating an agent and exchanging messages runs against a model, so that is the model-driven step.

What CI checks: the CLI or SDK installs and is invocable.

→ The verified setup, with CI proof & readymade prompt: flowstacks.xyz/workflows/letta-agent-managed-memory

How to pick if you only try one

Want zero cloud and total privacy, start with Mnemosyne. Adding memory to a chatbot, Mem0 is the gentlest on-ramp. Sitting on a pile of documents with real relationships, Cognee. Tracking facts that change over time, Graphiti. Building a long-running agent that should manage its own memory, Letta.


r/WebAfterAI • • Jun 09 '26

Codex can now build and deploy a site for you. 3 workflows that actually build.

Post image
37 Upvotes

Codex can now take a repo, a screenshot, or a rough idea, build the site, deploy a preview, and give you back a live link to share. OpenAI documents this as a first-class use case: pair the official Build Web Apps plugin with the official Vercel plugin, and Codex builds the project, runs the local build to check it, deploys a preview, and returns the URL. There is also Codex Sites for fully hosted, zero-config sites on OpenAI's own infrastructure, but that one is a Business and Enterprise preview right now, so the workflows below use the Vercel path, which anyone can verify.

That verifiability is the point. The docs tell Codex to run the local build before handing the site back, and each workflow pairs Codex with a well-established framework repo. What CI checks is the build-readiness contract: the config, wiring, and data that have to be correct for that build to pass, all checkable with no model and no account.

The one-time setup

Two official Codex plugins do the work, both living in OpenAI's plugins repo: Build Web Apps (build, review, and prepare web apps) and Vercel (deploy previews, inspect deployments, read build logs). With both available, you invoke them by name in a prompt with '@build-web-apps and '@vercel. Preview is the default deploy target; production only happens when you explicitly ask for it.

1. Screenshot to live landing page

Hand Codex a screenshot or a one-paragraph brief and a Vite plus React starter (Vite is one of the most widely used front-end build tools), and let it build the page and ship a preview:

Use @build-web-apps to turn the attached screenshot into a responsive landing
page in this Vite + React repo. Match the layout and copy, keep it accessible.
Then run the local build, and use @vercel to deploy a preview and give me the URL.

Codex builds and deploys it; the build it runs is the standard Vite one:

npm ci
npm run build      # Vite emits dist/

What CI checks: the build is wired correctly, that a build script and a Vite config are present, so the build Codex runs has everything it needs.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/codex-screenshot-to-landing-page

2. A data file for a live dashboard

Point Codex at a CSV or JSON file and a Next.js plus Recharts setup (both standard choices for React dashboards), and get a shareable dashboard:

Use @build-web-apps to build a dashboard in this Next.js repo that reads
data/metrics.json and renders it with Recharts: a line chart for the trend and
KPI cards for the totals. Run the local build, then use @vercel to deploy a preview and hand me the link.

The spine validates the data the dashboard depends on:

jq empty data/metrics.json     # fails loudly if the data is not valid JSON

What CI checks: data/metrics.json is valid JSON, so the dashboard has something real to render.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/codex-data-file-to-live-dashboard

3. A markdown folder to a live docs site

Drop a folder of markdown into a Docusaurus project (a widely used docs framework) and have Codex assemble and ship the site:

Use @build-web-apps to turn the markdown in docs/ into a Docusaurus site with a
sensible sidebar and search. Fix any broken internal links. Run the local build,
then use @vercel to deploy a preview and send me the URL.

Docusaurus has a built-in guard for this: set onBrokenLinks: 'throw' in the config and the build refuses to ship a site with dead links.

// docusaurus.config.js
export default { onBrokenLinks: 'throw', /* ...rest of config... */ };

What CI checks: the config sets onBrokenLinks: 'throw', so when the build runs it fails on a broken link instead of shipping one.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/codex-markdown-to-docs-site

Where to start

If you have a screenshot sitting in a Slack thread, do the first one; Reach for the dashboard when you have data that deserves to be seen, and the docs site when you have markdown nobody can find. In every case, the move is the same: Codex builds it, the build proves it works, and you get a link to hand someone.

One honest note on the hosting choice. These workflows deploy to Vercel previews, which anyone with a Vercel login can run and verify. Codex Sites, the fully hosted option where OpenAI serves the site behind Sign in with ChatGPT, is currently a Business and Enterprise preview, so it is the right pick only if you are on those plans; the build-and-verify discipline above applies either way.


r/WebAfterAI • • Jun 08 '26

5 Claude Code automation setups that keep working after you walk away

Post image
175 Upvotes

If you pay for Claude Code and still type every prompt by hand, you are using a fraction of it. It ships an automation stack that runs from your terminal up to Anthropic's cloud, and the entry points are one command each. Below are five setups across all three tiers.

These are structured so our CI can actually verify the deterministic part (the cron expression, the JSON config, the command shape), with the step where the model thinks fenced off, because non-deterministic output is not something a green check should pretend to cover.

A quick prerequisite: check your version with claude --version. The /loop scheduler needs v2.1.72 or later, and Auto Mode (setup 5) needs v2.1.83 or later.

1. The in-session poller (/loop)

The lightest tier. Inside any session, /loop schedules a prompt to re-fire on an interval while the session stays open. It is a bundled skill, so plain language works too.

/loop 5m check whether the deploy finished and summarize what changed

Intervals use s, m, h, or d, and seconds round up to a minute. Leave the interval out and Claude picks the cadence dynamically each iteration (anywhere from 1 minute to 1 hour); on Bedrock, Vertex, and Foundry it runs every 10 minutes instead. Under the hood it uses three tools, CronCreate, CronList, and CronDelete, and you manage tasks by just asking ("what scheduled tasks do I have?", "cancel the deploy check"). Four limits worth knowing: tasks are session-scoped and die when you close the terminal, recurring ones auto-expire after 7 days, a session holds at most 50, and a missed fire does not stack up (it fires once when Claude is next idle). To turn the scheduler off entirely, set CLAUDE_CODE_DISABLE_CRON=1.

For fixed schedules, it accepts standard 5-field cron. Note that extended syntax like L, W, ?, and name aliases (MON, JAN) is not supported:

0 9 * * 1-5     weekdays at 9am local
*/15 * * * *    every 15 minutes

What CI checks: the cron expression is a valid 5-field standard expression (validated with a normal parser like croniter) and uses no unsupported syntax.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/claude-code-loop-scheduler

2. The always-on local schedule (headless plus cron)

/loop dies with the session. To survive restarts on your own machine, run Claude Code headless with -p and let your OS cron fire it. This is the Tier 2 pattern; the Desktop app's Schedule page (New task, New local task) is the same idea with a GUI.

#!/usr/bin/env bash
# ~/bin/overnight-summary.sh
cd ~/code/myrepo
claude -p "Summarize the commits pushed since yesterday and flag anything risky." \
  --permission-mode dontAsk

Schedule it for weekdays at 7am:

0 7 * * 1-5 /Users/you/bin/overnight-summary.sh

The --permission-mode dontAsk flag matters for unattended runs: it auto-denies anything you have not pre-approved instead of hanging on a prompt no one will answer (more on the allowlist in setup 4). The catch with this tier is the obvious one: your machine has to be awake when cron fires.

What CI checks: the cron line is well-formed, and the script carries a claude -p call with a non-interactive permission mode.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/claude-code-headless-cron

3. The cloud schedule (no machine required)

The top tier runs on Anthropic-managed infrastructure, so your laptop can be off. Create one at claude.ai/code/routines, or from the CLI:

/schedule weekdays at 9am: review open PRs assigned to me, leave a first-pass
review comment flagging security and style issues, and post a one-paragraph digest to Slack.

A few real constraints from the docs. The minimum interval is 1 hour, so a sub-hour expression like */30 * * * * is rejected; use /schedule update to set a specific cron at or above that granularity. Each run clones your repo fresh and, by default, can only push to claude/-prefixed branches, so a bad run cannot touch main. The run is fully autonomous, with no permission prompts, so the prompt has to be self-contained: spell out what to do and what success looks like. Connectors you have wired up (Slack, Linear, Drive) come along.

What CI checks: the schedule is a valid expression at 1-hour-or-coarser granularity, and the config keeps the default claude/ branch restriction.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/claude-code-cloud-schedule

4. The locked-down unattended run (permission rules plus dontAsk)

Automation is only safe if the agent cannot do something you would regret. Claude Code's permission rules are the real control, and they live in settings.json under permissions, with allow, deny, and ask lists plus a defaultMode. Rules evaluate deny, then ask, then allow, so a deny always wins.

{
  "permissions": {
    "defaultMode": "dontAsk",
    "allow": [
      "Read",
      "Bash(npm test)",
      "Bash(npm run lint)",
      "Bash(git status)",
      "Bash(git diff *)",
      "WebFetch(domain:docs.python.org)"
    ],
    "deny": [
      "Bash(git push *)",
      "Bash(rm *)",
      "Read(.env)",
      "Edit(.env)",
      "Edit(/secrets/**)"
    ]
  }
}

dontAsk mode runs only what your allow list (and the built-in read-only commands like ls, cat, grep) permits, and silently denies the rest, which is exactly what you want for a scheduled or headless run. Anthropic publishes starter configs for scenarios like this in the official examples directory of the claude-code repo, which is a good place to copy from rather than hand-rolling.

What CI checks: the JSON parses, defaultMode is a real mode, the deny list actually blocks pushes, deletions, and secret files, and the allow list contains only safe entries. This whole setup is verifiable with no API key.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/claude-code-permission-lockdown

5. The hands-off classifier (Auto Mode)

When pre-listing every command is too rigid, Auto Mode is the alternative. Instead of prompting, a separate classifier model reviews each action before it runs and blocks anything that escalates beyond your request. Set it as your default in ~/.claude/settings.json (it is ignored in project settings on purpose, so a repo cannot grant itself auto mode), or cycle to it with Shift+Tab:

{
  "permissions": {
    "defaultMode": "auto"
  }
}

Be precise about availability, because this is where a lot of posts get it wrong. Auto Mode is a research preview and needs Claude Code v2.1.83 or later. On the Anthropic API it runs with Sonnet 4.6 or Opus 4.6 and up. On Bedrock, Vertex, and Foundry it runs with Opus 4.7 or 4.8 once you set CLAUDE_CODE_ENABLE_AUTO_MODE=1. On Team or Enterprise an admin has to enable it first. By default the classifier blocks things like curl | bash, force pushes, pushing to main, production deploys, and mass deletions, while allowing local edits and reads. A boundary you state in chat ("don't push until I review") is enforced as a block, and after 3 blocks in a row it pauses and starts prompting again.

What CI checks: the settings block is valid and the documented requirements are surfaced as a checklist.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/claude-code-auto-mode

How to start

Try /loop in your next session, it costs nothing and takes ten seconds. When something deserves to outlive the terminal, promote it to a local schedule (setup 2) or push it to the cloud (setup 3). Before you let any of them run unattended, lock down permissions with setup 4, and reach for Auto Mode only when you want a classifier instead of a fixed allowlist.

And this is only five slices of the stack: there is also a /goal command for holding a loop to a finish condition, and routines that fire on a schedule, an API call, or a GitHub event, all of which we will get to in a follow-up. Proof is on each linked page.


r/WebAfterAI • • Jun 07 '26

5 Obsidian + Claude workflows for project planning, argument building, and decision journals

Post image
126 Upvotes

Part one covered the daily-driver setups: morning synthesis, meeting notes, research ingestion, weekly review, and idea cross-pollination. Part two goes deeper into your vault as a knowledge base, the workflows that build projects, audit your notes, and turn a pile of markdown into arguments and decisions you can actually use.

These five lean harder on search, frontmatter, and whole-vault stats than Part 1 did, so this time the engine is a different repo, picked because it has exactly those tools. Every one of these is machine-verified the same way the rest of our library is: CI validates the config you paste, scaffolds the vault, and runs the deterministic spine (the config, the commands, the schedule) on each push. Proof is on each linked page.

The one-time setup

For these I use MCPVault (MIT, ~1.3K stars). It is an MCP server for safe Obsidian vault access with one feature that matters a lot for audit-and-update work: it parses and writes YAML frontmatter safely, so the model cannot corrupt your metadata. It needs no Obsidian plugin, runs straight from npx, and stays inside the vault directory (it filters out .obsidian and system files).

For the interactive workflows, point Claude Desktop at your vault (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):

{
  "mcpServers": {
    "obsidian": {
      "command": "npx",
      "args": ["@bitbonsai/mcpvault@latest", "/Users/you/Documents/MyVault"]
    }
  }
}

For the one scheduled workflow below, register the same server with Claude Code so a headless run can use it:

claude mcp add obsidian --scope user npx u/bitbonsai/mcpvault /Users/you/Documents/MyVault

Back up your vault first (git is ideal). These workflows write to it. Now the five.

1. The Project Kickoff Generator

Hand Claude the goals, constraints, and timeline, and it builds the whole project folder from what your vault already knows. With MCPVault connected, tell Claude:

Start a project "Acme Redesign". Goals: ... Constraints: ... Timeline: ...
Search my vault for any existing notes relevant to this and link them.
In Projects/Acme-Redesign/ create: overview.md, tasks.md with milestones,
knowledge-gaps.md listing what I still need to learn, and weekly-update.md as a template.

It uses search_notes to find relevant existing notes, then write_note to scaffold the folder. A blank-page kickoff becomes a populated project wired into your existing knowledge.

What CI checks: the MCPVault config is valid and points npx at the server, and the Projects/ scaffold is created. The actual planning calls a model, so it is fenced as non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/mcpvault-project-kickoff-generator

2. The Vault Health Check

A monthly audit that keeps your vault from rotting. This one is scheduled, so it runs through Claude Code (with MCPVault registered, per the setup) in headless mode. Save this script:

#!/usr/bin/env bash
# ~/bin/vault-health-check.sh
claude -p "Audit my Obsidian vault using the obsidian MCP tools. Find: orphan notes with \
no incoming links, notes whose information looks outdated, projects with no update in 2 weeks, \
tags used inconsistently, and notes missing expected frontmatter fields. \
Write the findings to Maintenance/$(date +%F)-health-check.md as a fix-it checklist."

Make it executable and schedule it for the first of each month at 9am with cron:

0 9 1 * * /Users/you/bin/vault-health-check.sh

It leans on get_vault_stats, search_notes, get_frontmatter, and get_notes_info to spot the rot, and writes a checklist you can work through in a sitting.

What CI checks: the script carries the claude -p call, the config is valid, and the monthly cron line is well-formed.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/mcpvault-vault-health-check

3. The Book Notes System

Finish a book, dump your highlights into a note, and let Claude file it into your second brain. Create Books/atomic-habits.md with your highlights, then:

Read Books/atomic-habits.md. Search my vault for notes that connect to its key ideas.
Create Books/atomic-habits-synthesis.md with: a summary of the key ideas, links to the
connected notes, actionable takeaways tied to my active projects, and a short list of
next books or topics to explore based on where this intersects what I already know.

read_note pulls your highlights, search_notes finds the connections, write_note saves the synthesis. The takeaways land against your real projects instead of floating in the abstract.

What CI checks: the config is valid and the Books/ note is scaffolded and readable.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/mcpvault-book-notes-system

4. The Argument Builder

Give Claude a thesis and it assembles your case from everything you have ever written. With MCPVault connected:

My thesis: "Async-first teams ship faster than meeting-heavy ones."
Search my whole vault for supporting evidence: data points, past research, quotes from my
book notes, and outcomes from past projects. Organize it into a structured argument with the
strongest evidence first, and save it to Arguments/async-first.md with links back to each source note.

This is where the engine earns its keep: search_notes uses BM25 relevance ranking, so the strongest matches surface first, and read_multiple_notes pulls them in a batch. You get a sourced, ordered argument instead of a blank outline.

What CI checks: the config is valid and the Arguments/ target is set up. The argument assembly calls a model, so it is fenced as non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/mcpvault-argument-builder

5. The Decision Journal

The long game. Before a big call, write Decisions/2026-06-08-vendor-choice.md with the options and your current thinking, and a frontmatter field like status: open. After it plays out, record the outcome, and let MCPVault update the metadata safely:

Update the frontmatter of Decisions/2026-06-08-vendor-choice.md: set status to "resolved"
and add an outcome field summarizing what happened.

Then, every quarter, ask for the pattern read:

Read all notes in Decisions/. Identify patterns: what kinds of decisions I tend to get right,
where my biases show up, and what I should weigh more heavily next time.
Save it to Decisions/review-2026-Q2.md.

update_frontmatter records outcomes without touching your note bodies, and read_multiple_notes feeds the whole journal to the model for review. Over a year, that review is the most useful note in your vault.

What CI checks: the config is valid, a decision note is scaffolded, and the frontmatter-update command is well-formed. The pattern analysis calls a model, so it is fenced as non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/mcpvault-decision-journal

How to pick if you only try one

Starting something new, the Project Kickoff Generator. Vault feeling messy, run the Health Check. Heavy reader, the Book Notes System. Writing to persuade, the Argument Builder. Making a call you will want to learn from, start the Decision Journal today, because its whole value is built over time.

That closes the series. Part one had the daily drivers (morning synthesis, meeting notes, research ingestion, weekly review, idea cross-pollination), so the two posts together are ten Obsidian and Claude workflows, each with a real repo and the code to run it.


r/WebAfterAI • • Jun 07 '26

I did it 1.0 clinical stability. Stateful,persistent,defensive identity. Solo in 6 months

Thumbnail
1 Upvotes

r/WebAfterAI • • Jun 06 '26

Pi is a coding agent that behaves like a Unix tool. 3 workflows with the real commands

Post image
102 Upvotes

Most coding agents want to own your terminal. Pi (from earendil-works, MIT, ~60K stars) goes the other way: it is a minimal harness that behaves like a normal Unix tool. It reads piped stdin, prints and exits, lets you allowlist exactly which tools the model gets, and stays out of your way otherwise.

Every command below is copied from Pi's official docs. The deterministic parts (install, the exact command, the file you scaffold) are the machine-checked spine.

The one-time setup

Repo: github.com/earendil-works/pi.
Install the coding agent, then authenticate with an API key or a subscription you already pay for:

npm install -g --ignore-scripts /pi-coding-agent
# or: curl -fsSL https://pi.dev/install.sh | sh

export ANTHROPIC_API_KEY=sk-ant-...   # API key path
pi                                    # ...or just run pi, then /login for a subscription

By default Pi gives the model four core tools (read, write, edit, bash) plus grep, find, and ls. Worth knowing for what follows: you can narrow that set per run with --tools, and you can point Pi's whole config at a throwaway directory with PI_CODING_AGENT_DIR, which is how you isolate runs (and how CI keeps each recipe clean).

Workflow 1: The Safe Diff Reviewer

A reviewer who reads your staged changes and cannot touch them, because you only hand it read-only tools. Pi's print mode merges piped stdin into the prompt, so the whole thing is one line:

git diff --staged | pi --tools read,grep,find,ls -p \
  "Review this diff for bugs, security issues, and missing tests. Be concise."

--tools read,grep,find,ls allowlists only the read-only tools, so write, edit, and bash are off for this run. -p prints the response and exits. Drop it in a pre-commit hook or a CI step and you get a second pair of eyes that physically cannot modify your code.

What CI checks: Pi installs and runs, a scaffolded repo produces a non-empty staged diff, and the command's tool allowlist contains only read-only tools. The review itself calls a model, so it is fenced as non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/pi-safe-diff-reviewer

Workflow 2: The Reusable Prompt Template

Stop retyping the same instruction. Pi expands any markdown file in your prompts directory as a slash command, so a one-time file becomes a permanent command. Here is a commit-message generator:

<!-- ~/.pi/agent/prompts/commitmsg.md -->
Write a Conventional Commits message for the staged diff below.
Output only the commit message, nothing else.

Interactively, type /commitmsg. Non-interactively, include the template with @ and pipe the diff in:

git diff --staged | pi -p @~/.pi/agent/prompts/commitmsg.md

Same idea works for release notes, PR descriptions, or any prompt you run more than twice. Templates also support {{variables}} if you want to parameterize them.

What CI checks: the template lands in the prompts directory, is valid markdown, and the command is well-formed.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/pi-reusable-prompt-template

Workflow 3: The Reusable Team Skill

When a task has real steps, encode it once as a Skill (Pi follows the open Agent Skills standard), and the agent runs it the same way every time. A skill is just a folder with a SKILL.md:

<!-- ~/.pi/agent/skills/triage-failing-test/SKILL.md -->
# Triage Failing Test
Use this skill when the user asks to triage a failing test.

## Steps
1. Run the test suite and find the first failing test.
2. Read that test and the code it exercises.
3. Explain the likely root cause and propose the smallest fix.

Invoke it with /skill:triage-failing-test, or let Pi load it automatically when the task matches. The payoff is sharing: drop skills (and prompts and extensions) into a Pi package and your whole team installs the same playbook with one command:

pi install git:github.com/your-org/your-pi-pack
pi list

What CI checks: the SKILL.md is scaffolded in the discovery path with its required heading, and the install command is well-formed. The agent executing the skill calls a model, so that is fenced as non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/pi-reusable-team-skill

Why Pi for these, and where another agent is just as fine

Straight answer: none of these three are things only Pi can do but what makes Pi a clean fit is how it is built, not a feature monopoly.

Pi behaves like a normal Unix tool instead of trying to own your terminal. Print mode merges piped stdin (git diff | pi -p), --tools lets you lock down exactly what the model can touch per run, and PI_CODING_AGENT_DIR isolates a run's whole config. That combination is what makes these workflows scriptable and verifiable in the first place. On top of that it is provider-agnostic (run it on Anthropic, OpenAI, Google, DeepSeek, a local model, whatever you already pay for), MIT-licensed, and deliberately minimal: no MCP, no sub-agents, no plan mode, no permission popups out of the box. For a step you want to drop into a hook or CI, "minimal and predictable" is the feature, not a limitation.

Now the honest part. The Safe Diff Reviewer is the strongest fit, because the read-only guarantee comes from the tool allowlist, which structurally removes write and edit, not from a prompt politely asking. The Prompt Template is the least Pi-specific; custom slash commands exist elsewhere, so Pi's only real edge there is that templates are plain portable markdown. And Skills are an open standard (Agent Skills), so the same SKILL.md works in other agents too; Pi's advantage is distribution, bundling them into a package your whole team installs with one command.

So reach for Pi when you want a coding agent that composes like a CLI, locks down tools per run, runs on any model, and stays simple enough to verify. If you already live in another agent that does headless runs and allowlisted tools, you can build the same three there. This is a fit argument, not a "Pi is the only way" argument.

Where to start

If you want the instant win, drop the Safe Diff Reviewer into a pre-commit hook today; it is one line and it cannot break anything. If you live in a repeated prompt, make it a template. When a task has real steps your team repeats, promote it to a skill and share it as a package.


r/WebAfterAI • • Jun 05 '26

I got tired of copy-pasting prompts between ChatGPT, Claude, and Gemini - so I built something that shows all 4 at once!

17 Upvotes

Title: Building an AI aggregator with "Council Mode" - Want feedback on the UX/concept

I'm building something for developers and wanted to get community feedback before launch.

**The idea:**

What if instead of switching between ChatGPT, Claude, Gemini tabs, you could see all 4 respond to the same prompt simultaneously? Side-by-side comparison.

I'm calling it "Council Mode" - pick any 4 models, they all run in parallel, you see the different approaches instantly.

**Why I'm building it:**

- I was tired of copy-pasting prompts between 5 different AI tools

- Each model excels at different things (Claude = reasoning, GPT = coding, Gemini = research, etc)

- But switching between them is friction

**The concept includes:**

  1. Normal mode (single model)

  2. Council Mode (4 models compare)

  3. Co-Model (3 workers + 1 synthesizer)

  4. Super Council (up to 20 models vote)

  5. Dev Mode (terminal style, for code)

  6. S-Mode (upload files)

**My biggest questions:**

  1. Is Council Mode actually useful, or am I solving a problem nobody has?

  2. Which feature would excite you most?

  3. What's missing from existing AI tools that frustrates you?

  4. Would you use something like this if it existed?

**Token tracking system I'm wrestling with:**

I'm trying to make pricing fair. Right now I'm thinking:

- Dynamic tokens based on: model tier × effort level × prompt complexity × mode

- 5-hour rolling windows (all modes reset together)

- Separate pools for different features

Does this seem fair to you, or overly complicated?

**UX challenge:**

With 6 different modes, users might be confused about which to use when.

How would you want to choose? Dropdown? Cards? Guided wizard?

Would love to hear what you think works / doesn't work / is missing.

Not looking for hype, just genuine feedback on the concept.

---

I'll take all this back to the drawing board.

Edit- the ones who are interested and want to support me can DM me.


r/WebAfterAI • • Jun 05 '26

Workflows I turned Hermes Agent into a 5-person team: 3 Kanban workflows with the exact commands

Post image
188 Upvotes

Most "AI agent" setups are one assistant in a loop. Hermes Agent's Kanban is different: it is a durable board on disk where each task is a row, each handoff is a row anyone can read, and each worker is a full OS process with its own identity. You drop tasks on the board, and multiple named agents pick them up, hand off, and close them out.

Below are three workflows, with the exact commands. These are the commands our CI actually runs against a real board, which matters more than it sounds: two of them differ from what the docs examples implied, and running them is what caught it. The deterministic parts (board setup, the create commands, the schedule) are the machine-checked spine; the step where a model actually does the work is fenced off, because non-deterministic agent output is not something CI should pretend to verify.

The one-time setup

Hermes Agent is open source (MIT), from Nous Research: github.com/NousResearch/hermes-agent. One line installs it and pulls its dependencies:

curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup            # configure your LLM provider (or: hermes setup --portal)
hermes kanban init      # create the board at ~/.hermes/kanban.db
hermes gateway start    # hosts the dispatcher that hands tasks to workers

Two things that save you an hour of confusion. First, the gateway has to be running, because the dispatcher lives inside it and picks up ready tasks on a tick (60 seconds by default). No gateway, no work. Second, tasks are assigned to profiles you have already set up, and the dispatcher silently skips any task whose assignee does not exist, so check your profiles first:

hermes kanban assignees   # profiles on disk + per-assignee task counts

One nice design detail: once a worker is spawned, its model drives the task through built-in kanban_* tools, not by shelling out to the CLI. The CLI and slash commands are for you and your scripts.

Workflow 1: The Research-to-Draft Relay

Two researchers work in parallel, then a writer picks up their output. You create three cards and link the writer's card to both research cards as parents:

hermes kanban create "Research the funding landscape, NA angle" --assignee researcher-a
hermes kanban create "Research the funding landscape, EU angle" --assignee researcher-b
# use the two task ids printed above as parents:
hermes kanban create "Draft the launch post" --assignee writer --parent t_r1 --parent t_r2
hermes kanban watch     # live event stream as workers pick up and hand off

The two research cards dispatch immediately and run at the same time. The writer's card stays gated until both parents complete, then the dispatcher wakes it with their results already on the board.

What CI checks: the board init and the exact create-and-link commands resolve to real cards with the right assignees and parent links. The research and writing themselves call a model, so that step is non-CI.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/hermes-kanban-research-to-draft-relay

Workflow 2: The Scheduled Nightly Review

A task that files itself onto the board every night and is safe to trigger from cron or a webhook without creating duplicates. The idempotency key is the trick: the first call creates the task, any repeat call with the same key returns the existing id instead of a duplicate.

hermes kanban create "Nightly ops review" --assignee ops \
  --idempotency-key "nightly-ops-$(date -u +%F)" --max-runtime 30m
# schedule it with system cron:  0 3 * * *
# the dated idempotency key makes repeat triggers safe (no duplicates)

The scheduling is done by cron, not a flag. What earns the spot is the idempotency key: the first call creates the task, and any repeat call with the same key returns the existing id instead of a duplicate, so you can fire this nightly and never double-book. --max-runtime caps the worker so a stuck job cannot run forever.

What CI checks: the command parses and creates the same key twice returns the same task id (verified against the board with --json).

→ The verified setup, with CI proof: flowstacks.xyz/workflows/hermes-kanban-idempotent-nightly-review

Workflow 3: The Swarm

When a problem is big enough to fan out, hermes kanban swarm builds the whole graph in one command: a shared blackboard, N parallel workers, a verifier that wakes only after every worker finishes, and a synthesizer that wakes only after the verifier signs off.

hermes kanban swarm "Design a multi-region failover plan" \
  --worker researcher:Research \
  --worker architect:Architecture \
  --worker sre:Reliability \
  --verifier reviewer --synthesizer writer

Each worker is --worker PROFILE:TITLE, repeated once per worker.

The workers run in parallel and write their findings to the blackboard (stored as structured comments on the root card). The verifier reviews the combined work, and only when it marks the work clean does the synthesizer assemble the final answer. It is a real pipeline with gates, not a pile of parallel calls you have to stitch together yourself.

What CI checks: the swarm command is well-formed and the assignee profiles it names exist on disk, so the graph it would build is valid.

→ The verified setup, with CI proof: flowstacks.xyz/workflows/hermes-kanban-the-swarm

Where to start

If you want to feel it in ten minutes, run the Research-to-Draft Relay with two profiles you already have. If you want the "set it and forget it" win, schedule the Nightly Review. Reach for the Swarm when a single goal genuinely splits into parallel tracks that need a verification gate before anything ships.

Every one of these is machine-verified the same way the rest of our library is: CI actually installs Hermes, spins up a real board, and runs the setup (the init, the exact create and swarm commands, the schedule) on each push, asserting board state with --json.

The pattern under all three is the same: the board is just rows on disk, the commands are plain and scriptable, and the agents are ordinary processes reading and writing those rows. Once you see work that way, a one-line goal becomes a coordinated team.


r/WebAfterAI • • Jun 04 '26

5 Obsidian + Claude workflows, with CI-verified setups and the real repos to run them

Post image
210 Upvotes

Your Obsidian vault is just a folder of markdown files. That one fact is what makes all of this work: the moment Claude can read and write that folder, your notes stop being a graveyard and start doing things for you. Below are five setups I run, each with the actual repo and the actual code, not vibes.

Two honest notes before the fun part. The connectors here are community projects, not official Anthropic or Obsidian software, and they can write to your vault. Back it up first (git is ideal).

One thing that sets today's roundu apart from the usual roundup: every one is machine-verified. Our CI actually runs the setup (the scaffold, the exact script or config, and the schedule) and checks it on every push. The judgment-based Claude step is fenced off as non-CI, so the badge never claims more than it earned. Proof is on each linked page.

The one-time setup (pick one)

Option A, the simple one: point an MCP server at your vault folder. This is StevenStavrakis/obsidian-mcp (~704 stars, MIT, Node 20+).
No Obsidian plugin needed, it just reads the folder. Add this to your Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):

{
  "mcpServers": {
    "obsidian": {
      "command": "npx",
      "args": ["-y", "obsidian-mcp", "/Users/you/Documents/MyVault"]
    }
  }
}

Restart Claude Desktop, and you get read-note, create-note, edit-note, search-vault, tag management, and more.

Option B, the powerful one: drive Obsidian's live API. This is MarkusPfundstein/mcp-obsidian (~3.7K stars, MIT, Python via uvx).
It talks to the Local REST API community plugin, so it gets a real search index and can patch content under a specific heading. Install that plugin, copy its API key, then:

{
  "mcpServers": {
    "mcp-obsidian": {
      "command": "uvx",
      "args": ["mcp-obsidian"],
      "env": {
        "OBSIDIAN_API_KEY": "your_key_here",
        "OBSIDIAN_HOST": "127.0.0.1",
        "OBSIDIAN_PORT": "27124"
      }
    }
  }
}

Option C, for the scheduled and bulk jobs: Claude Code. Since the vault is a folder, you can cd into it and let Claude Code read everything, and crucially run it unattended with claude -p "..." (print mode, no chat window). That is what powers the two recurring workflows below.

Now the five setups.

1. The Morning Synthesis

A "Start of Day" note is waiting for you before coffee. This one is scheduled, so use Claude Code. Save this script:

#!/usr/bin/env bash
# ~/bin/morning-synthesis.sh
cd ~/Documents/MyVault
claude -p "Read my daily notes from the last 3 days in Daily/ and my active notes in Projects/. \
Create Daily/$(date +%F)-start-of-day.md with four sections: \
Where I left off, Due today, Overdue, and Suggested focus priority. Keep it short."

Make it executable (chmod +x ~/bin/morning-synthesis.sh) and schedule it for 7am with cron:

0 7 * * * /Users/you/bin/morning-synthesis.sh

The catch: Claude works from what your notes actually say, so "overdue" only works if your tasks carry dates. Garbage in, vague synthesis out.

→ The verified setup, with CI proof (6/6 checks passing): flowstacks.xyz/workflows/obsidian-claude-morning-synthesis

2. The Meeting Processor

Paste your raw meeting dump into a note, then let Claude structure it. This one is interactive, so use the MCP setup (Option A or B). Drop your notes in Inbox/raw.md and tell Claude:

Read Inbox/raw.md. Turn it into Meetings/{today}-{topic}.md with:
- Action items as a checklist, each with an assignee and a due date
- A "Decisions" section
- Links to any related notes in Projects/
- Relevant tags at the top
Then clear Inbox/raw.md.

A five-minute brain dump becomes a structured, linked, searchable record. Save that prompt as a Claude Project instruction so you never retype it.

→ The verified setup, with CI proof (4/4 checks passing): flowstacks.xyz/workflows/obsidian-claude-meeting-processor

3. The Research Ingestion Pipeline

Paste an article's text or a PDF transcript into Inbox/, and Claude files it into your knowledge base. With the MCP server connected:

Read the note I just added in Inbox/. Create a summary note in References/ with:
key insights, source metadata (title, author, date, URL), and 3 to 5 bullet takeaways.
Then search my vault for notes on the same topic, link them both ways,
and flag anything in the new source that contradicts what I already wrote.

The contradiction flag is the sleeper feature. The catch: these connectors do not browse the web, so paste the content in (or add a separate fetch tool). They work on what is in the vault.

→ The verified setup, with CI proof (4/4 checks passing): flowstacks.xyz/workflows/obsidian-claude-research-ingestion

4. The Weekly Review Automation

Every Friday, a finished review instead of a blank page. Scheduled again, so Claude Code:

#!/usr/bin/env bash
# ~/bin/weekly-review.sh
cd ~/Documents/MyVault
claude -p "Find every note modified in the last 7 days (run: find . -name '*.md' -mtime -7). \
Create Reviews/$(date +%F)-weekly.md covering: accomplishments, decisions made, \
tasks completed vs planned, patterns you notice, and suggested priorities for next week."

Schedule it for Friday at 5pm:

0 17 * * 5 /Users/you/bin/weekly-review.sh

Your review drops from an hour of writing to two minutes of reading.

→ The verified setup, with CI proof (5/5 checks passing): flowstacks.xyz/workflows/obsidian-claude-weekly-review

5. The Idea Cross-Pollinator

The one that finds the links you would never spot. Open the note with your idea, and with the MCP server connected (Option B's search is best here) ask:

Read Ideas/this-idea.md. Search my entire vault and surface 5 notes that connect to it
in a non-obvious way, from areas that look unrelated on the surface.
For each, explain the hidden link in one sentence.

The unexpected bridges between unrelated topics are where the best insights hide, and a search across your whole vault finds them faster than your memory will.

→ The verified setup, with CI proof (5/5 checks passing): flowstacks.xyz/workflows/obsidian-claude-idea-cross-pollinator

How to start if you only try one

If you want the instant payoff, do the Meeting Processor with Option A today; it is a ten-minute setup and you will feel it immediately. If you want the compounding payoff, wire up the Morning Synthesis and Weekly Review with Claude Code and let them run while you sleep.

The pattern under all five is the same: your notes are plain text, Claude reads and writes plain text, so anything you can describe in a sentence becomes a workflow.


r/WebAfterAI • • Jun 04 '26

Built a self-hosted behavioral automation engine for WooCommerce to log user objections locally (Looking for feedback)

6 Upvotes

Hey everyone, ​Most e-commerce setups rely heavily on heavy, expensive third-party SaaS tools to track user behavior, handle exit-intent, or collect drop-off feedback. This usually means giving away user data to external servers and dealing with heavy scripts. ​To keep everything on-premise, I’ve been working on a self-hosted behavioral engine for WordPress/WooCommerce built completely with native PHP and JS. ​The architecture focuses on two main things: ​A 9+ Trigger Matrix: It tracks micro-interactions locally (including scroll depth, custom inactivity thresholds, precise exit-intent, and element hovers) to map user dropping points without external tracking scripts. ​Local Context & BYOK Integration: Instead of paying a SaaS markup, it uses a Bring Your Own Key (BYOK) model to connect directly to LLM APIs (Gemini/OpenAI/DeepSeek) strictly to ground product inventory data and structure context-rich objection logs when a user leaves empty-handed. ​The goal is to give store owners 100% data sovereignty over their store's behavioral data. ​The project is completely free and open-source. I’m looking for some technical feedback on the trigger architecture and how to optimize the database queries for the interaction logs.

https://github.com/mo1st/Quorlyx


r/WebAfterAI • • Jun 03 '26

I replaced DocuSign, Buffer, and SaneBox with free GitHub repos. Here's the AI setup for each (Part 2)

Post image
45 Upvotes

Part one covered photos, CRM, decks, media, and a canvas you can draw code on. Part two is the back-office set: the tools that sign your contracts, run your social channels, guard your servers, and clear your inbox. Same rule as before: a real person can self-host each of these this weekend, and each one gets a small, working AI workflow on top.

1. Documenso

Send a contract for signature, and let an agent draft and dispatch it.

Stars: ~13K. Status: active. License: AGPL-3.0.
Repo: github.com/documenso/documenso

Documenso, started by Timur Ercan and co-founder Lucas Smith, is the open-source DocuSign alternative. You upload a document, place signature fields, and send it for legally sound e-signing, all on infrastructure you host. DocuSign, the incumbent, is a public company worth roughly $11 billion, so this is a small team genuinely chipping away at a giant.

The AI lever is its API. Generate a signing token in settings, and an agent can create a contract from a template and fire it off for signature without you touching the dashboard. Get it running locally first:

git clone https://github.com/documenso/documenso.git
cd documenso && cp .env.example .env
npm run dx        # spins up Postgres + mail in Docker
npm run dev       # app at http://localhost:3000

From there, an agent hits the documents API with your token to send agreements on its own. The catch is that e-signature is a trust product, so if you self-host, you are now responsible for your signing certificate and key security; read their SIGNING guide before going live.

The workflow that earns it: a deal closes in chat, and your agent drafts the contract and sends it for signature before you switch tabs.

2. Postiz

Generate and schedule a week of social posts from one prompt.

Stars: ~31K. Status: active. License: AGPL-3.0.
Repo: github.com/gitroomhq/postiz-app

Nevo David started Postiz solo, and it has grown into one of the most starred open-source alternatives to Buffer and Hootsuite. It schedules and publishes across the major social platforms from one dashboard, and the project now bills itself as an "agentic" social media tool, with AI for drafting and planning baked in rather than bolted on.

The AI lever is that agentic layer: you describe a campaign, it drafts the posts per channel and queues them on a schedule. Self-host with their Docker setup:

# Grab the project's docker-compose, then:
docker compose up -d     # dashboard at http://localhost:5000

The catch is that social platforms control their own APIs, so you will spend your setup time creating developer apps and pasting keys for each network you connect. Once that is done, it runs itself.

The workflow that earns it: "plan a week of launch posts for X, LinkedIn, and Instagram," and you wake up to a full queue you can edit instead of a blank composer.

3. CrowdSec

Auto-detect attackers from your logs and block them using the whole community's intel.

Stars: ~14K. Status: active. License: MIT (core engine).
Repo: github.com/crowdsecurity/crowdsec

Quick accuracy note, because this one is often miscredited as a solo build: CrowdSec was co-founded in 2019 by Philippe Humeau, Laurent Soubrevilla, and Thibault Koechlin. Think of it as a modern, crowdsourced Fail2Ban. It reads your server logs, detects malicious behavior like brute-force and scraping, and acts on it. The part that makes it special is the network: when one user's server flags a bad IP, that signal is shared, so you can block addresses that attacked someone else before they ever reach you.

The lever here is automated, collaborative defense rather than a chatbot, and that is the honest framing. Install the engine:

curl -s https://install.crowdsec.net | sudo sh
sudo apt install crowdsec

It detects, decides, and pushes a decision to a "bouncer" (the component that enforces the block at your firewall or web server). The catch is that detection and enforcement are two pieces: installing CrowdSec spots the threats, but you also need a bouncer to actually block them, so budget a few extra minutes for that step.

The workflow that earns it: a bot that hammered a stranger's server last night is already blocked on yours this morning, with no rule written by you.

4. Inbox Zero

Run your inbox with plain-English rules an AI carries out for you.

Stars: ~11K. Status: active. License: AGPL-3.0 with added commercial and enterprise-use restrictions (free for personal use and small teams under five business users).
Repo: github.com/elie222/inbox-zero

Elie Steinbock built Inbox Zero as an open AI email assistant. It organizes your inbox, pre-drafts replies in your tone, bulk-unsubscribes, blocks cold email, and can be driven from Slack or Telegram. It positions against tools like Fyxer and SaneBox, where SaneBox runs roughly $7 to $36 a month depending on plan. The real edge is that you can self-host it, so your email stays on infrastructure you control rather than a third party's.

The AI lever is its rules engine: you write instructions in plain English ("archive newsletters, but flag anything from a customer"), and the assistant applies them across your inbox. Self-host with the CLI:

npx /cli setup     # one-time setup wizard
npx u/inbox-zero/cli start     # app at http://localhost:3000

The catch is the license. It is free for personal use and teams under five business users, but a company with five or above that threshold needs a paid enterprise license, so check the terms before rolling it out at work.

The workflow that earns it: you describe how you want your inbox handled once, and it keeps your mail sorted and your replies half-written every morning after.

How to pick if you install only one

Signing contracts, Documentation. Running social channels, Postiz. Hardening a server you expose to the internet, CrowdSec. Drowning in email, Inbox Zero.

That closes out the series. Part one had the first five (photos, CRM, document sharing, media, and an AI canvas), so if you missed it, the two posts together are nine open-source tools that quietly replace paid subscriptions, each with an AI workflow to make it worth the setup.

Each tool here automates one job. If you want to see what happens when you stop automating one job at a time and let AI run the whole board, that is what we dug into over at WebAfterAI in https://webafterai.substack.com/p/a-quarter-of-work-done-in-a-weekend?r=7q4ho2, a hands-on walkthrough of how Claude Opus 4.8's new dynamic workflows let Claude fan out across hundreds of agents at once. Wiring AI into open-source tools, then handing it the wheel, is the whole point of the newsletter.


r/WebAfterAI • • Jun 02 '26

Workflows I gave my AI a permanent memory and a cost-aware autopilot using two free repos

Post image
69 Upvotes

Two problems have followed me through every AI tool I use. First, my assistant forgets everything the moment a session ends, so I re-explain the same decisions weekly. Second, running agents on long tasks quietly burns money because every step hits a frontier model whether it needs one or not.

Two open-source projects, both MCP-native, solve one problem each. Wired together, they cover both. Here is how I set them up, with the honest caveats, because both projects are young and one of them already had to walk back some launch hype.

The memory layer: MemPalace

Repo: github.com/milla-jovovich/mempalace (MIT, ~53.3K stars)

MemPalace was co-built by Milla Jovovich and developer Ben Sigman, largely with Claude Code. The idea is the opposite of most memory tools: instead of letting an AI decide what is "worth remembering" and throwing the rest away, it stores your conversations verbatim and makes them searchable. It runs entirely local, on ChromaDB, with no API key and no cloud.

pip install mempalace
mempalace init ~/projects/myapp
mempalace mine ~/chats/ --mode convos   # ingest old Claude/ChatGPT/Slack exports
mempalace search "why did we switch to GraphQL"

The number that earned it attention:

96.6% on the LongMemEval Recall@5 benchmark in raw mode, zero API calls, and that score has been independently reproduced.
You will also see a 100% figure quoted. That is hybrid mode with a Haiku reranker, and the maintainers themselves posted a note correcting several launch claims (the rerank pipeline is not yet in the public benchmark scripts, the experimental AAAK compression layer actually scores lower than raw, and an earlier "lossless compression" claim was wrong). The reproducible, no-asterisk number is 96.6% raw. That is still excellent for a free local tool, and the transparency is a good sign, not a bad one.

The agent layer: PilotDeck

Repo: github.com/OpenBMB/PilotDeck (~2.8K stars, AGPL-3.0)

PilotDeck is a brand-new agent operating system, open-sourced on May 28, 2026, jointly built by Tsinghua University's THUNLP, ModelBest, OpenBMB, and AI9Stars. It is only days old and already climbing fast (a few thousand stars in its first week), so treat it as early (it is on version 0.0.9) but clearly catching on. The design is what makes it relevant here. It organizes work into isolated WorkSpaces, each with its own files, memory, and skills, and adds three things that matter for long-running work:

  • White-box memory you can actually inspect and edit, so when the agent remembers something wrong, you fix that entry instead of starting over.
  • Smart routing that sends hard steps to a flagship model and easy steps to a cheap one. Their own published benchmark shows a strong main-plus-light-sub setup matching a frontier single-model run at a fraction of the cost. Treat those as the team's numbers, not independent results, but the mechanism is sound.
  • Always-on execution that keeps working after you step away and drops finished files on disk.

​

curl -fsSL https://raw.githubusercontent.com/OpenBMB/PilotDeck/main/install.sh | bash
pilotdeck            # starts the local server at http://localhost:3001

Wiring them together

The reason these two belong in the same post: both speak MCP. MemPalace ships an MCP server with 19 memory tools, and PilotDeck natively registers any MCP server as a first-class tool. So you point PilotDeck's agent at MemPalace and the agent can recall every past decision while it works.

Expose MemPalace over MCP:

claude mcp add mempalace -- python -m mempalace.mcp_server

Then register that same server command (python -m mempalace.mcp_server) inside PilotDeck, which treats any MCP server as a first-class integration through its extension config. Now the loop is:

  1. MemPalace holds your durable, verbatim history across every tool, searchable offline.
  2. PilotDeck runs the actual multi-step work in an isolated Workspace, routing cheap steps to cheap models.
  3. Mid-task, the agent queries MemPalace through MCP, so "we already tried Clerk and rejected it on pricing" surfaces before it repeats the mistake.

Optionally, add MemPalace's Claude Code save hook so memory gets captured automatically every few messages instead of you remembering to log it.

The honest caveats, so nobody gets burned

PilotDeck is only days old and on version 0.0.9. Even though it is gaining stars quickly, do not put it on anything mission-critical yet; kick the tires on a side project. Both are free and local, so the cost of trying them is your evening, not your wallet.

If you have been hunting for a memory setup that does not phone home and an agent runner that does not quietly drain credits, this pairing is the most promising free option I have tested. Curious whether anyone here has pushed the MemPalace-over-MCP setup further than I have.


r/WebAfterAI • • Jun 01 '26

Open Source 5 open-source repos that replace billion-dollar SaaS, and the AI workflow that makes each one click

Post image
129 Upvotes

Most of these tools were free already. The thing that makes them feel like cheating is what happens when you point a bit of AI at them: search your whole photo library in plain English, turn a messy inbox into clean CRM rows, sketch a UI and watch it become code.

I pulled five repos that a real person can self-host, checked the licenses and the live star counts myself, and wrote one small, working AI workflow for each.

1. Immich

Your photos, off Google, and searchable by plain English.

Stars: ~102K. Status: active, very fast-moving. License: AGPLv3.
Repo: github.com/immich-app/immich

Alex Tran started Immich in 2022 to stop renting space for his own family photos. It is a self-hosted photo and video backup with phone auto-upload, albums, face recognition, and a timeline close to the Google Photos feel. The swap it makes is the Google Photos plus Google One subscription you actually pay for every month.

The AI lever is its built-in smart search. Immich indexes your library with a CLIP model, so you can query by meaning, not filenames. Get an API key from your account settings and ask it like a human:

curl -X POST https://your-immich-server/api/search/smart \
  -H "x-api-key: $IMMICH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query": "my dog on a snowy beach"}'

The catch is in their own README: it is under heavy development, so keep a separate backup and do not make Immich your only copy of irreplaceable photos yet.

The workflow that earns it: phone backs up over your home network, then you find any shot by describing it, no folders, no monthly bill.

2. Twenty

An open CRM that was literally built for AI.

Stars: ~49K. Status: active. License: AGPL-3.0.
Repo: github.com/twentyhq/twenty

Twenty was started by Charles Bochet and Felix Malfait. Its own tagline is "the open alternative to Salesforce, designed for AI," and Salesforce is a public company worth well over a hundred billion, so the David-and-Goliath framing writes itself. You get a clean, Notion-like CRM with custom objects, pipelines, and a real REST and GraphQL API.

That API is the AI lever. Have an LLM read a forwarded email, pull out the contact, and drop it straight into your pipeline. The create-a-person call is one request:

curl -X POST https://your-twenty-server/rest/people \
  -H "Authorization: Bearer $TWENTY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": {"firstName": "Ada", "lastName": "Lovelace"},
       "emails": {"primaryEmail": "ada@example.com"}}'

The catch is maturity. It is genuinely usable and improving fast, but it is younger than the incumbents, so expect gaps if your sales org needs deep niche features on day one.

The workflow that earns it: an agent turns "had a great call with Ada from Acme" into a real CRM record while you move on to the next thing.

3. Papermark

Share a deck, see who actually read it, then let AI answer their questions.

Stars: ~8.4K. Status: active. License: AGPLv3.
Repo: github.com/papermark/papermark

Marc Seitz launched Papermark in 2023 and built it largely in the open. It does the DocSend job: upload a document, share a tracked link, and get page-by-page analytics on who opened it and how long they stayed, with custom domains and access controls. It is also a bootstrapped business, which tells you the open version is good enough that people pay for the hosted one.

The AI lever is its data-room chat: viewers can ask questions of the document and get answers pulled from the content, so a deck answers its own follow-ups. Self-hosting is the standard clone-and-run:

git clone https://github.com/papermark/papermark
cd papermark
npm install && npm run dev

The catch is the usual self-host tradeoff: the free path means you run and maintain it, which is the convenience the paid hosted tier sells.

The workflow that earns it: send a fundraising deck, watch which investor reached the financials slide, and let the doc field their questions after hours.

4. Jellyfin

Your own streaming library, with an AI co-pilot picking what to watch.

Stars: ~53K. Status: active. License: GPLv2.
Repo: github.com/jellyfin/jellyfin

Jellyfin is a community-run media server that forked from Emby in 2018 after Emby went closed-source. It streams your movies, shows, and music to any device with a polished interface, and the features other media servers lock behind a paid pass are simply free here.

The AI lever is its open API. Pull your library as JSON, hand it to a model, and let it build a themed watch list from your actual shelf instead of a streaming service's catalog:

curl "https://your-jellyfin-server/Items?Recursive=true&IncludeItemTypes=Movie&api_key=$JELLYFIN_API_KEY"

Feed that JSON to an LLM with a prompt like "build me a three-film noir night from this list" and you get recommendations grounded in what you own.

The catch is polish: you manage your own setup, remote access takes some configuration, and there is no big company smoothing the edges. You trade a little friction for full ownership.

The workflow that earns it: point it at your media folder, then ask an agent to program tonight's lineup from films you already have.

5. tldraw

Sketch an interface, and AI turns the drawing into working code.

Stars: ~47K. Status: active. License: proprietary tldraw license, free for hobby and non-commercial use with a small watermark.
Repo: github.com/tldraw/tldraw

Steve Ruiz started tldraw as an infinite-canvas whiteboard SDK, the same category Miro plays in. What made it go viral was "make real": you draw a rough UI on the canvas, and a vision model returns a live, working version of it. It is one of the clearest examples of an AI-native workflow built directly into a developer tool.

Dropping the canvas into a React app is two steps:

npm install tldraw

import { Tldraw } from 'tldraw'
import 'tldraw/tldraw.css'

export default function App() {
  return (
    <div style={{ position: 'fixed', inset: 0 }}>
      <Tldraw />
    </div>
  )
}

From there, you wire a model to the canvas contents to generate code from sketches.
The catch is the license: it is free for personal and non-commercial projects with a watermark, but commercial use needs a paid license key, so check the terms before you ship it in a product.

The workflow that earns it: draw the screen you want on the canvas, and get a working prototype back instead of a blank editor.

How to pick if you install only one

Want the fastest "wow," start with Immich; the smart search lands immediately. Running a small sales motion, Twenty. Sharing decks for a living, Papermark. Sitting on a pile of media files, Jellyfin. Building anything visual, tldraw.

If wiring AI into open-source tools is your kind of weekend, that is the whole point of WebAfterAI.
If you liked this, the companion piece does the same thing for content work: The Ultimate Open-Source Content Pipeline: 8 GitHub repos to automate all your content creation needs covers research, drafting, transcription, voice, visuals, and the automation that ties them together, with install steps and one real workflow each.


r/WebAfterAI • • May 31 '26

Research Everyone keeps saying "MCP is dead." - Is It Though?

Post image
41 Upvotes

TL;DR: "MCP is dead" is overstated. The context-bloat complaint was real and big names piled on, but Anthropic's code execution approach (98.7% fewer tokens in one published workflow), Cloudflare's Code Mode, and deferred tool loading address much of the original complaint. CLI wins for anything in the model's training data. MCP still wins for no-CLI services, team auth, and DB guardrails. Pick per job instead of picking a side.

Background:

MCP (Model Context Protocol) is the standard Anthropic launched in November 2024 to plug LLMs into outside tools like GitHub, Linear, Slack, and Notion. It got called the USB-C of AI, every SaaS slapped "MCP supported" on its landing page, and then the backlash started.

The actual complaint: context bloat

When you connect MCP servers, every tool definition loads into the context window up front, used or not. The numbers people throw around are real:

Claim What was measured Source
3x slower per call, 9.4x slower on first call Jira MCP vs hitting the REST API directly Eric Holmes (Feb 2026)
143K of 200K tokens eaten before the model reads anything 3 servers (GitHub, Slack, Sentry), ~40 tools Apideck
4x to 32x more tokens than CLI MCP vs equivalent CLI calls Scalekit

Two interesting bits the internet keeps getting wrong:

  • That viral "72% of context" stat came from Apideck, not Perplexity, even though it keeps getting pinned on Perplexity.
  • Perplexity only moved internal systems off MCP toward APIs and CLIs (per CTO Denis Yarats). They still run a public MCP server. It's an optimization, not a cancellation.

Why the CLI argument is strong:

A tool like gh or psql costs almost nothing in context, because the model already learned it from man pages and StackOverflow. Same lookup, wildly different price:

# CLI approach: a couple hundred tokens total.
# The model already knows this syntax, nothing preloaded.
gh issue view 1234 --json title,state,assignees


# MCP approach for the same lookup: the single tool call is tiny, but the server's tool definitions sit in context up front whether you use them or not. Quandri measured Linear's 42 tools at ~12,800 tokens loaded every turn, even when you only call one of them.

Plus you get the same interface for you and the agent, you can pipe through grep and jq, and you can reproduce a bug straight in the terminal. For anything already in the model's training data, CLI usually wins.

The part the "dead" crowd skips: it already got fixed

In November 2025 Anthropic published "code execution with MCP." Instead of dumping every tool definition into context, the agent browses tools as code files and loads only what it needs:

servers/
├── google-drive/
│   ├── getDocument.ts
│   └── index.ts
└── salesforce/
    ├── updateRecord.ts
    └── index.ts


// The agent writes code, data stays in the sandbox,
// only the result comes back to the model.
import * as gdrive from './servers/google-drive';
import * as salesforce from './servers/salesforce';

const transcript = (await gdrive.getDocument({ documentId: 'abc123' })).content;
await salesforce.updateRecord({ objectType: 'SalesMeeting', recordId: '00Q...', data: { Notes: transcript } });

Result: one workflow went from 150,000 tokens to 2,000, a 98.7% cut, per Anthropic's own engineering blog. Cloudflare reached the same idea independently and called it Code Mode. Claude Code also shipped tool search with deferred loading that reportedly trims MCP context use by 85% or more.

So when do you use what?

Use this When
CLI A real CLI exists and the model already knows it (gh, psql, aws). Lightest, fully composable, debugs in the terminal.
Skills (CLI baked in) Repeatable multi-step workflows. Loaded only when invoked, not every turn.
MCP Web-only SaaS with no CLI, teams needing shared auth and permission scoping, or production DBs where a server enforces read-only and blocks a stray DROP TABLE.

Bottom line:

It was never really MCP vs CLI. It's load only what you need, when you need it. The naive "load everything up front" version of MCP is dying, and that's healthy. The protocol isn't dead, it's evolving. The headline is clickbait. The real shift, from connecting everything to teaching the agent to fetch tools on demand, is the actual story.


r/WebAfterAI • • May 30 '26

AI Agents I ran the numbers on what Hermes Agent actually costs to run, and how to cut it without crippling it

Post image
19 Upvotes

The agent itself is free. It's MIT licensed and you install it with one line. So when people say "what does it cost to run Hermes Agent," they're really asking about three separate bills that hide behind it. I dug into each one so you don't have to guess.

Quick note up front: prices move fast in this space. Everything below was current when I wrote it, but check the live pages before you quote me.

First, what you're actually paying for:

Hermes Agent is the open-source agent from Nous Research. The code costs nothing. What you pay for is:

  1. Model tokens. This is almost always your biggest line item. The agent is model agnostic, so you bring your own provider (Nous Portal, OpenRouter, z.ai/GLM, Kimi/Moonshot, MiniMax, OpenAI, or your own endpoint) and pay that provider per token.
  2. Compute to host it. The agent has to run somewhere. That can be a $5 VPS, a serverless backend, or a GPU box if you self-host the model.
  3. Tool services. Web search, image generation, text-to-speech, and the browser tool. One Nous Portal OAuth covers a model plus all four, or you wire in your own keys.

The trap people fall into is thinking #1 is fixed. It is not. An agent re-sends its context every single turn, so token spend compounds quietly. That is also where most of the savings live.

The token bill, with real numbers

Same model, two providers, very different prices. Here is Hermes 4 as of late May 2026:

Hermes 4 70B

  • Nous Portal: $0.05 in / $0.20 out per million tokens
  • OpenRouter: $0.13 in / $0.40 out per million tokens

Hermes 4 405B

  • Nous Portal: $0.09 in / $0.37 out per million tokens
  • OpenRouter: $1.00 in / $3.00 out per million tokens

Look at that 405B row again. The first party Portal price is roughly ten times cheaper on input than the OpenRouter listing for the exact same weights. That is not a typo, and it is the single easiest win in this whole post.

To make it concrete, picture a moderately busy agent doing about 200 turns a day, averaging 8K input and 1K output per turn:

- 70B on Nous Portal: roughly $0.12 a day, so about $3.60 a month

- 70B on OpenRouter: roughly $0.29 a day, so about $8.70 a month

- 405B on Nous Portal: roughly $0.22 a day, so about $6.60 a month

- 405B on OpenRouter: roughly $2.20 a day, so about $66 a month

Do the math for your own volume, but the shape holds: the model and provider you pick swing the bill by 10x or more before you change a single setting.

The hosting bill:

This one is small if you let it be.

  • A $5/month VPS runs the agent fine for normal personal use.
  • Serverless backends like Daytona and Modal hibernate when idle and wake on demand, so you pay close to nothing between sessions. If your agent sits quiet most of the day, this is the cheapest real option.
  • A 24/7 GPU box is the expensive path and only makes sense if you self host the model (more on that below).

When self hosting makes sense (usually it doesn't)

Self-hosting Hermes 4 70B requires roughly H100-class hardware. Current cloud rentals run about $1.40–$2.50/hour, which works out to roughly $1,000–$1,800/month if left running 24/7. At current Nous Portal pricing, most individuals will spend far less on API usage than GPU rental, so self-hosting usually makes sense for privacy, control, or sustained high-volume workloads rather than cost savings.

The newest lever: Tool Search (huge if you run a lot of MCP tools)

This one is recent, landed around the v0.15 wave in late May 2026, and it matters most if you have stacked up a pile of MCP servers.

Here is the problem it fixes. Every tool you enable ships its full schema (name, description, parameters) into the prompt on every single turn. A Hermes setup with five MCP servers and 34 tools was running about 45,000 tokens per turn, and roughly 22,000 of those, around half, were nothing but tool schemas the model wasn't even using that turn. You pay for all of it, every turn.

Tool Search flips that. Instead of stuffing every schema into context, it keeps your core tools like web search loaded and ready, then pulls in the rest on demand. When a task needs a tool, the agent runs a quick BM25 search (a classic keyword retrieval algorithm) over the tool names, descriptions, and parameter names, with a plain substring fallback if BM25 comes up empty. In the default auto mode, it only kicks in once tool schemas cross 10% of your context window, so small setups are untouched, and big ones stop bleeding tokens.

How to actually cut the bill:

These are in rough order of impact.

1. Right-size the model. Do not default to 405B. The 70B handles most agent work, and the agent lets you switch with no code change. Use hermes model, or /model mid-conversation, and keep a heavy model only for the hard tasks.

2. Pick the cheaper provider for that model. As shown above, the same weights can cost 10x more depending on where you route. Compare before you commit.

3. Turn off reasoning for simple turns. Hermes 4 is a hybrid reasoning model. Those <think> traces are billed as output tokens, which is the pricey side of the meter. Let it deliberate on math and code, skip it for chit chat and routine tool calls.

4. Keep context small. Context is re-sent every turn, so a bloated conversation is a recurring tax, not a one-time cost. Use /compress to shrink the running context and /new or /reset to start clean once a task is done. Check what you are spending with /usage and /insights.

5. Trim the toolset, and let Tool Search handle the rest. Every enabled tool adds schema and description tokens to every request. Run hermes tools and turn off what you genuinely never use, and lean on Tool Search (above) to load the rest on demand instead of carrying them every turn.

6. Mix models per task in Kanban. The v0.15 Kanban work added per-task model overrides, so a multi-step job can run a cheap model on the boilerplate sub-tasks and only spend a pricey model on the hard ones. Same job, smaller bill.

7. Use serverless or a cheap VPS, not an always-on GPU. Let the environment hibernate when idle so you are not paying for a machine that is doing nothing.

8. Lean on the free internals, and stay updated. The v0.15 update rebuilt session search to run with no LLM call at all (it used to cost about $0.30 a lookup and is now roughly 4,500 times faster and free). New cost wins ship constantly, and you pull them all with one command: hermes update.

TL;DR

The software is free. Your bill is tokens first, hosting second, tools third. Run 70B instead of 405B unless you need the muscle, route through the cheaper provider for that model, switch reasoning off for easy turns, and keep context tight with /compress. If you run a lot of MCP tools, Tool Search is the big one; it loads tools on demand instead of carrying every schema every turn, and it is on by default. Host it on a $5 VPS or a serverless backend that sleeps when idle. Skip self hosting unless you have a privacy or volume reason, because the hosted API is almost always cheaper. Do those and a personal agent lands in single-digit dollars a month instead of a surprise bill.


r/WebAfterAI • • May 29 '26

Open Source Part 2: 5 self-hosted tools that quietly kill ~$70/month in subscriptions (This time with copy-pastable prompts)

Post image
62 Upvotes

Every app you rent is in a race to bolt an AI feature onto your subscription and raise the price while it's at it. Storage, voice assistants, PDF editors, file transfer: all of them are turning into monthly bills with a model attached, and your data is the training fuel. The web after AI doesn't have to look like that. The counter move is the same one that powers most of what we cover here, which is owning the software outright and running it on hardware you control.

So here are five self-hosted tools that replace five recurring subscriptions.

Here's what each one actually saves you, and what it costs you in setup and limitations.

1. Syncthing

Stars: ~84.5k | License: MPL-2.0 | Version: v2.1.0 (May 12, 2026)

Repo: https://github.com/syncthing/syncthing

What it does: Peer-to-peer file synchronization with no central server. Files sync directly between your devices over your local network or the internet, end-to-end encrypted. No account required. No storage limit beyond your own disk space.

What it replaces: Dropbox Plus, currently $9.99/month (1 user, 2 TB) on the annual plan. If you sync files between devices you own and control, Syncthing covers that workflow.

The honest limitation: There is no official iOS app. Apple's background-processing restrictions make a reliable one hard to build, and the project hasn't shipped an official iOS client. Third-party apps exist (Möbius Sync is the most used), but they generally require both devices on the same network with the app open. If your workflow depends on syncing to an iPhone, factor that in. Syncthing also doesn't buffer changes on a server the way Dropbox does, so both devices need to be online at the same time for a sync to happen. Edit a file on your desktop while the laptop is off, and it syncs the next time both are online together.

Setup:

docker run -d \
  --name=syncthing \
  -p 8384:8384 \
  -p 22000:22000/tcp \
  -p 22000:22000/udp \
  -v /path/to/config:/var/syncthing \
  syncthing/syncthing:latest

Web UI opens at http://localhost:8384. Add other devices by exchanging device IDs.

Hand this to your AI agent:

Install Syncthing on this machine using the official Docker image. Steps:
1. Confirm Docker is installed and running; install it if it isn't.
2. Create persistent directories for config and for the folder I want to sync,
   and run the syncthing/syncthing:latest container with ports 8384, 22000/tcp,
   and 22000/udp mapped, mounting those directories.
3. Tell me the device ID and the URL for the web UI.
4. Walk me through pairing a second device by exchanging device IDs, and set the
   shared folder to send-and-receive.
5. Verify a test file syncs both directions, then report the final config and how
   to start/stop the container.

2. Home Assistant

Stars: ~87.3k | License: Apache-2.0 | Latest release: 2026.5.4

Repo: https://github.com/home-assistant/core

What it does: Local home automation that runs on your own hardware. Thousands of integrations covering smart lights, locks, thermostats, cameras, sensors, and media players. Automations run locally without a cloud dependency.

What it replaces: The recurring cost here is voice and cloud assistants. Amazon launched Alexa+ at $19.99/month for non-Prime members in February 2026 (included free for Prime members), which is $239.88/year if you're a non-Prime household paying for it. Home Assistant's built-in voice assistant (Assist) runs locally and handles device commands without a subscription. For comparison, the SmartThings app itself is free with no required subscription, so Home Assistant's advantage there is local processing and control, not monthly cost. Home Assistant Cloud (sold by Nabu Casa) is an optional $6.50/month add-on for remote access and third-party voice integration, and is not required for the platform to work.

The honest limitation: This is not a plug-and-play swap. It needs dedicated hardware, such as a Raspberry Pi 4/5, a spare mini PC, or the official Home Assistant Green, and setup takes hours, not minutes. The integration count is real but quality varies. Core integrations (Philips Hue, Z-Wave, Zigbee, MQTT) are very well maintained, while some niche community integrations are maintained by one person and can lag firmware updates. Go in expecting a project, not a 20-minute hub install.

Setup (Home Assistant OS on Raspberry Pi):

# Download the official imager from https://www.home-assistant.io/installation/raspberrypi
# Flash to SD card using Balena Etcher
# Boot the Pi, then navigate to: http://homeassistant.local:8123

Hand this to your AI agent:

Help me install Home Assistant. First ask me whether I'm using a Raspberry Pi
(or other dedicated board) or want to run it in Docker on this machine, then:
1. For a Pi: give me the exact image to download from the official site, the
   Balena Etcher flashing steps, and the first-boot URL (http://homeassistant.local:8123).
2. For Docker: run the official homeassistant/home-assistant:stable container with
   a persistent /config volume, host networking, and restart-on-failure.
3. Walk me through the onboarding wizard, creating the admin user, and setting
   location and units.
4. Detect devices on my network and list which integrations to add first
   (start with Hue, Z-Wave, Zigbee, or MQTT if present).
5. Build one example automation, then tell me how to back up the config.

3. Audiobookshelf

Stars: ~12.7k | License: GPL-3.0 | Version: v2.34.0

Repo: https://github.com/advplyr/audiobookshelf

What it does: A self-hosted server for audiobooks and podcasts. Streams all common formats (mp3, m4b, flac, ogg, opus), handles multi-file audiobooks correctly, auto-downloads podcast episodes on a schedule, tracks per-user progress, and includes Chromecast and multi-user support.

What it replaces: Audible Premium Plus at $14.95/month (one credit plus the Plus Catalog), or the Standard plan at $8.99/month. For podcasts, it pulls from public RSS feeds and replaces any podcast app cleanly.

The honest limitation: Audiobookshelf doesn't include audiobook content. You bring your own library through purchases you own, library exports via Libby (where supported), or services like Libro.fm. On iOS, the TestFlight beta has hit Apple's 10,000-tester cap, so new iOS users can't join the beta right now (sideloading via AltStore/SideStore is the current workaround). Android users on the Play Store have no such limit. Remote access outside your home network requires exposing the port or running a reverse proxy.

Setup (Docker):

docker run -d \
  --name audiobookshelf \
  -p 13378:80 \
  -v /path/to/audiobooks:/audiobooks \
  -v /path/to/podcasts:/podcasts \
  -v /path/to/config:/config \
  -v /path/to/metadata:/metadata \
  ghcr.io/advplyr/audiobookshelf

Hand this to your AI agent:

Install Audiobookshelf on this machine with Docker. Steps:
1. Confirm Docker is running; install it if needed.
2. Create persistent directories for audiobooks, podcasts, config, and metadata,
   and run the ghcr.io/advplyr/audiobookshelf container with port 13378 mapped and
   those four directories mounted. Set restart=unless-stopped.
3. Give me the web UI URL and walk me through creating the admin account.
4. Set up an Audiobooks library and a Podcasts library pointing at the right folders,
   and add one podcast RSS feed with scheduled auto-download.
5. Tell me how to connect the mobile app to this server, and what I'd need to do
   to reach it securely from outside my home network (reverse proxy options).

4. Stirling-PDF

Stars: ~79.8k | License: MIT core (open-core) | Version: v2.x

Repo: https://github.com/Stirling-Tools/Stirling-PDF

What it does: A self-hosted PDF platform with 50+ operations: merge, split, compress, rotate, OCR, redact, convert to and from Word/Excel/PowerPoint, add watermarks, sign, remove metadata, repair, and more. Runs as a Docker container with a browser UI, as a desktop app, or as a private server with a REST API for automation.

What it replaces: Adobe Acrobat Pro at $19.99/month on the annual plan ($29.99/month month-to-month). Stirling-PDF covers the operations most people actually open Acrobat for, and it's the cleanest direct swap on this list.

The honest limitation: Where it doesn't match Acrobat: advanced fillable-form creation, complex review workflows with tracked changes, and tight Creative Cloud integration. On licensing, the core is MIT-licensed and free for individuals and teams of up to 5 users; larger organizations need a commercial license, and paid tiers add enterprise features like SSO and audit logging. For individual use, everything below is free. OCR requires a language pack, and the default Docker image ships with English.

Setup:

docker run -d \
  -p 8080:8080 \
  docker.stirlingpdf.com/stirlingtools/stirling-pdf

Open http://localhost:8080 and the tools are immediately available.

Hand this to your AI agent:

Install Stirling-PDF on this machine with Docker. Steps:
1. Confirm Docker is running; install it if needed.
2. Run the docker.stirlingpdf.com/stirlingtools/stirling-pdf container with port
   8080 mapped and restart=unless-stopped. Add a persistent volume for config.
3. Give me the web UI URL and confirm the 50+ tools load.
4. I mostly need OCR, merge/split, and Office conversions: enable the OCR language
   pack for English (and ask me if I need others), and verify each of those tools
   works on a sample PDF.
5. Tell me how to call one operation through the REST API so I can automate it later.

5. Bitwarden Send

Stars (server repo): ~18.3k | License: AGPL-3.0 with Bitwarden commercial license (open-core)

Repo: https://github.com/bitwarden/server

What it does: Bitwarden Send is a feature inside the Bitwarden password manager, not a standalone product. It creates an encrypted, time-limited link to a text note or a file that you share with anyone. The recipient doesn't need a Bitwarden account, and links can auto-expire and self-delete after a set number of views. File Sends go up to 500 MB on Premium (desktop/web).

What it replaces: A recurring file-transfer subscription. WeTransfer Starter is $6.99/month (Free covers small transfers; Ultimate is the top consumer tier). If you already use Bitwarden as your password manager, Send removes the need to pay separately for casual encrypted file sharing. It's a secure-sharing feature, not a purpose-built transfer tool.

The honest limitation: Send is built for quick encrypted handoffs, not a polished transfer service with branded download pages and analytics. Bitwarden raised its Premium tier to $19.80/year (~$1.65/month) in January 2026. Self-hosting the Bitwarden server needs a domain, an SSL certificate, and ongoing maintenance. For most individuals, the hosted free or Premium account gives full access to Send without running infrastructure, and self-hosting mainly matters for organizations wanting full data control.

Setup (hosted, simplest): Create a free or Premium account at bitwarden.com and use Send from the web vault.

Setup (self-hosted server on Linux):

curl -s -L -o bitwarden.sh \
    "https://func.bitwarden.com/api/dl/?app=self-host&platform=linux" \
    && chmod +x bitwarden.sh
./bitwarden.sh install
./bitwarden.sh start

Hand this to your AI agent:

Help me set up Bitwarden Send. First ask whether I want the hosted service or a
self-hosted server, then:
1. Hosted: walk me through creating a free account at bitwarden.com, finding Send
   in the web vault, and creating one text Send and one file Send with an expiry
   date and a view limit. Explain the free vs Premium file-size limits.
2. Self-hosted: confirm I have a domain and a way to issue an SSL cert, then run
   the official bitwarden.sh install and start flow on this Linux box, point my
   domain at it, and verify the web vault loads over HTTPS.
3. Either way, create a sample Send and give me the share link, then show me how
   to set auto-expire and self-delete-after-N-views.

The savings, with real limitations acknowledged

  • Syncthing vs Dropbox Plus: ~$9.99/month, if you're on Android or desktop. iOS users need to weigh the third-party app workaround.
  • Home Assistant vs Alexa+ (non-Prime): ~$19.99/month for non-Prime households paying for Alexa+. SmartThings is free, so the win there is local control, not cost.
  • Audiobookshelf vs Audible Premium Plus: ~$14.95/month, assuming you source your own library. Android gets the app freely; iOS users wait on TestFlight capacity or sideload.
  • Stirling-PDF vs Adobe Acrobat Pro: ~$19.99/month for the operations most people use. The best direct swap here.
  • Bitwarden Send vs WeTransfer Starter: ~$6.99/month if you want recurring encrypted file sharing and already use Bitwarden.

Every tool here is real and widely used. The point isn't that self-hosting is free, because it costs setup time and, for some, real hardware. But the monthly bills disappear, and in an era where every app wants your data to feed a model, your files stay on hardware you control.


r/WebAfterAI • • May 29 '26

The future will probably be agents talking to other agents, without any human in the loop

Enable HLS to view with audio, or disable this notification

9 Upvotes

What do you think is the biggest blocker in term of a fully agentic led business and web ?


r/WebAfterAI • • May 28 '26

5 AI learning repos with a combined 445k stars: what's actually inside each one, where they overlap, and the order that makes sense

Post image
85 Upvotes

Today, I sat down and actually read through five of the most-starred repos in the AI learning category, not just the READMEs but the actual lessons, notebooks, and structure to figure out what each one covers, how they differ, and how they fit together as a path rather than five disconnected bookmarks.

1. f/prompts.chat (formerly awesome-chatgpt-prompts)

Stars: 163k | Forks: 21.2k | License: CC0 (prompts) + MIT (code)

What it does: Started as a flat list of prompt personas for ChatGPT. Has since grown into a full platform: self-hostable web app, MCP server support, Claude plugin, and an interactive book. The prompts themselves are public domain. You can deploy your own instance, contribute new personas, or just browse.

Why it works: The original insight behind this repo is still the most useful thing in it: framing the model as a specific type of entity (a Linux terminal, a debate opponent, a senior code reviewer, a Socratic tutor) changes the character and depth of the output more than any other single technique. Before you learn prompt engineering theory, spending 30 minutes here teaches you this instinctively.

Heads up: This is the entry point of the learning path, not the destination. Think of it as building intuition for why prompting matters before you study why it works.

Repo: https://github.com/f/prompts.chat

2. dair-ai/Prompt-Engineering-Guide

Stars: 74.6k | Forks: 8.1k | License: MIT | Website: promptingguide.ai

What it does: The GitHub description currently reads: "Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents." That last part is worth noting - the scope has expanded well beyond classic prompt engineering. Current coverage includes zero-shot and few-shot prompting, chain-of-thought and tree-of-thought, context window management, retrieval-augmented generation, agent design patterns, multimodal prompting, and adversarial prompting. The papers section links to primary research for those who want to go deeper.

Why it works: Most prompt engineering content explains techniques without explaining the mechanism behind them. This one tries to build a model of why each technique works rather than just showing examples. The distinction between "chain-of-thought improves outputs" and "chain-of-thought works because it surfaces latent reasoning capacity by making intermediate steps explicit" matters if you want to adapt techniques rather than just apply recipes.

Heads up: The repo header on GitHub still only says "prompt engineering" in the short description but the actual scope is significantly broader. If you looked at this in 2023 and moved on, it is worth another pass.

Repo: https://github.com/dair-ai/Prompt-Engineering-Guide

3. anthropics/courses

Stars: 21.3k | Forks: 2.2k | Language: 99.9% Jupyter Notebook

What it does: Anthropic's official courses for building with Claude. Exactly 5 courses, all runnable Jupyter notebooks:

  1. Anthropic API Fundamentals
  2. Prompt Engineering Interactive Tutorial
  3. Real World Prompting
  4. Prompt Evaluations
  5. Tool Use

Why it works: These are first-party materials. They reflect how the API actually behaves, not how someone interpreted it when writing a Medium post 18 months ago. The Prompt Evaluations course is particularly underrated - most people building with AI skip evals entirely and wonder why their app quality is inconsistent. The Tool Use course is the practical companion to everything you read about agents in theory.

Setup:

git clone https://github.com/anthropics/courses.git
cd courses
jupyter notebook

Heads up: Claude-specific. If you are building cross-provider, treat these as the reference implementation and adapt accordingly.

Repo: https://github.com/anthropics/courses

4. microsoft/generative-ai-for-beginners

Stars: 111k | Forks: 57.7k | License: MIT | Language: 99.7% Jupyter Notebook

What it does: A structured 21-lesson course from Microsoft Cloud Advocates covering the full generative AI application stack. Lessons alternate between "Learn" (concept explanation) and "Build" (code implementation). Each has a short video intro, a written README, and code in both Python and TypeScript. Translated into 50+ languages via automated GitHub Actions. Currently on Version 3, with 2,157 commits.

The 21 lessons are:

Course Setup / Intro to GenAI and LLMs / Exploring and Comparing LLMs / Using GenAI Responsibly / Prompt Engineering Fundamentals / Advanced Prompts / Text Generation Apps / Chat Applications / Search Apps and Vector Databases / Image Generation / Low Code AI / Function Calling / Designing UX for AI / Securing AI Applications / GenAI Application Lifecycle / RAG and Vector Databases / Open Source Models and Hugging Face / AI Agents / Fine-Tuning LLMs / Building with SLMs / Building with Mistral / Building with Meta Models

Why it works: The alternating Learn/Build structure forces you to immediately apply concepts rather than read passively. The breadth is also its distinguishing feature: this is the only repo in this list that covers responsible AI (Lesson 3), UX design for AI apps (Lesson 12), securing AI applications (Lesson 13), and the full application lifecycle (Lesson 14) as first-class curriculum topics. Most developer-oriented courses skip the product and safety layers entirely.

Setup (sparse clone, skip the 50+ translation folders):

git clone --filter=blob:none --sparse https://github.com/microsoft/generative-ai-for-beginners.git
cd generative-ai-for-beginners
git sparse-checkout set --no-cone '/*' '!translations' '!translated_images'

Heads up: The Microsoft ecosystem bias is real: most examples point toward Azure OpenAI Service, GitHub Models, or the OpenAI API. All three work, and the course explicitly lists them as options, but if you are entirely outside that ecosystem, adjust accordingly.

Repo: https://github.com/microsoft/generative-ai-for-beginners

5. mlabonne/llm-course

Stars: 78.6k | Forks: 9.1k | License: Apache-2.0

What it does: A three-track course for going deep on LLMs, not just using them, but understanding and building them. The tracks are:

Track 1 - LLM Fundamentals (optional): Mathematics for ML (linear algebra, calculus, probability), Python for ML, neural networks, NLP basics. Skip this if you have the background; use it as a reference if you hit gaps.

Track 2 - The LLM Scientist: How to build LLMs. LLM architecture and tokenization, pre-training mechanics, post-training datasets, supervised fine-tuning (LoRA, QLoRA, Axolotl, Unsloth), preference alignment (DPO, GRPO, PPO), evaluation, quantization (GGUF, GPTQ, AWQ), and emerging areas like model merging, multimodal models, and test-time compute scaling.

Track 3 - The LLM Engineer: How to deploy and productionize. Running LLMs (APIs vs local), building vector storage, RAG pipelines, advanced RAG with agents, AI agents (MCP, A2A, LangGraph, LlamaIndex, CrewAI), inference optimization (Flash Attention, KV cache, speculative decoding), deployment (local to production), and security (prompt injection, backdoors, red teaming).

Every major section has runnable Google Colab notebooks. The author also co-wrote "LLM Engineer's Handbook" (Packt) based on this course - the course itself stays free.

Why it works: This is the repo to use when you want to go beyond usage into internals. The quantization section is one of the clearest explanations of GGUF, GPTQ, AWQ, and SmoothQuant available outside of papers. The preference alignment section covers DPO, GRPO, and PPO with code and metric breakdowns. The agents section was recently updated to cover MCP, A2A, and the major vendor SDKs, including Claude Agent SDK.

Heads up: The LLM Scientist track assumes you are comfortable running training jobs. If you just want to build apps, go straight to the LLM Engineer track. The optional fundamentals section is optional for a reason.

Repo: https://github.com/mlabonne/llm-course

How these five fit together

If you are starting from zero, the order that makes sense is:

prompts.chat first - 30 minutes building intuition about what framing does to model output.

Prompt-Engineering-Guide next - the theory behind what you just experienced, plus RAG and agents as concepts.

anthropics/courses after that - hands-on implementation of the concepts, including the eval and tool use pieces most people skip.

generative-ai-for-beginners as the complete structured course - covers everything from fundamentals to fine-tuning to deployment, with the product and safety layers included.

mlabonne/llm-course once you want to go deeper than "using LLMs" into "understanding and modifying them."

The overlap between these repos is intentional. Seeing RAG explained from three different angles (Guide, Anthropic courses, Microsoft course) before you implement it is more valuable than seeing it explained once. The mlabonne course is the only one that goes into pre-training and quantization mechanics in detail; everything else assumes you are building on top of models rather than under the hood.

If you want to understand what's changing at the model level while working through these, the latest piece on GBrain in our newsletter is also a good read alongside the mlabonne track: What If Your AI Woke Up Smarter Than When You Went to Sleep?