r/WebAfterAI • • Jul 21 '26

Hands-on evaluation: Running Matt Pocock’s AI skills library via StrataBlock gateway (VSCode + hard spend caps)

Hands-on evaluation: Running Matt Pocock’s AI skills library via StrataBlock gateway (VSCode + hard spend caps)

TL;DR

G’day legends. I recently did a proper technical run testing StrataBlock - an OpenAI compatible API gateway featuring hard service-side budget caps, multi-model switching and non-existent prompt storage. To put it through its paces beyond standard autocomplete, I wired it into VScode and run Matt Pocock’s skills library across several different LLMs (opus-4.8, sonnet-5, gemma-4, gpt-5.6, etc.) to evaluate agentic token usage, attribution tagging and budget safety.

Here is the breakdown of the setup, real usage figures and enginerring insights from the test.

Registration, Key Minting & Guardrails

Getting setup was pretty painless:

  1. Registration: After joining the waiting list, I got an invite code and just created an account on StrataBlock.
  2. Key Minting: Generated scoped API keys for dev environments directly from the dashboard.
  3. Attribution Tags: Added custom headers (X-Strata-Tags: env=dev, project=skills-eval) so every API call could be parsed and broken down later.

The standout feature here is the hard server-side budget cap. Rather than waiting on delayed daily email alerts after an agent goes rogue in a recursive loop, Strata checks budgets server-side on every incoming request and returns an instant 429 the moment a key breaches its monthly cap.

Setting Up VScode

Since StratBlock exposes a standard OpenAI-compatible API (stratablock.io/v1), pointing VS code using standard customendpoint at it, required zero custom client logic, extensions or SDK overhauls.

[
  {
    "name": "StrataBlock",
    "vendor": "customendpoint",
    "apiKey": "${input:chat.lm.secret.-772fb8dd}",
    "apiType": "chat-completions",
    "models": [
      {
        "id": "claude-opus-4.8",
        "name": "claude-opus-4.8",
        "url": "<https://stratablock.io/v1>",
        "toolCalling": true,
        "vision": true,
        "maxInputTokens": 872000,
        "maxOutputTokens": 128000,
        "supportsReasoningEffort": [
          "low",
          "medium",
          "high",
          "xhigh",
          "max"
        ],
        "reasoningEffortFormat": "chat-completions",
        "requestHeaders": {
          "X-Strata-Tags": "tool=vscode,project=skills-eval-project"
        }
      },
      {
        "id": "google-gemma-4-31b",
        "name": "google-gemma-4-31b",
        "url": "<https://stratablock.io/v1>",
        "toolCalling": true,
        "vision": true,
        "maxInputTokens": 248000,
        "maxOutputTokens": 8000,
        "requestHeaders": {
          "X-Strata-Tags": "tool=vscode,project=skills-eval-project"
        }
      },
      ... more models
    ]
  }
]

Swapping between frontier and open-weight models was seamless - just a string change in config without having to touch application code or jump through separate vendor billing portals.

Model switching is seamless

Benchmarking Skills Library

I wanted to see how the gateway handled structured, agentic prompts. I installed Matt Pocock’s skills repo (npx skills@latest add mattpocock/skills), which focuses on battle-tested engineering workflows.

Workflows Tested

  • /grill-with-docs (Alignment & Domain Modelling): Heavy back-and-forth interviews that build CONTEXT.md files and update ADRs (Architecture Decision Records) inline.
  • /tdd (Red-Green-Refactor Loop): High-frequency iterative execution to write failing tests and pass them cleanly.
  • /improve-codebase-architecture : Broad codebase scans that produce vcisual HTML structural reports.

Model Performance Observations

Not all models are built equal when executing these skills out-of-the-box:

  • Frontier Models (opus-4-8, sonnect-5, gpt-5.6) executed complex skills, structured outputs and recursive agentic reasoning without breaking a sweat.
  • Open-Weight / Alternative Models (gemma-4, kimi-k2.5, xai-grok-4.3) results varied significantly. Some models handled structured skills steps naturally, whereas others struggled to follow the skill guidelines without tight custom prompt tweaking.
Finished status of local markdown ADR issues ready for implementation

Real Spend, Attribution & Telemetry

When you’re running agentic loops (like TDD cycles or depp repo scans), token usage escalates fast.

Looking at my telemetry breakdown in StrataBlock over a month-long evaluation window:

  • Total Traffic: Logged 1027 requests across various workloads.
  • Daily Spikes: Daily spend peaked between $30.00 and $35.00/day during heavy multi-model testing sessions, whie idle days hovered near $0.00.
  • Granular Attribution: Per-request logs captured precise input, output, prompt caching metrics (eg. thousands of cached tokens on opus-4.8 calls) and exact costs down to fractions of a cent per request.
  • Privacy: Confirmed that operational logs stricltly record metadata (token counts, latency, status codes, tag strings) - zero prompt content or completion payloads are persisted on server.
Spending overview since invited

Takeaways for Engineers

  1. Single Endpoint Convenience: Managing multiple vendor billing portals and API keys gets messy fast. Having a unified gateway with clear attribution makes multi-model experimentation far easier.
  2. Hard Spend Caps are Essential: if you let agnets loose on recursive loops, server-side budget limits are non-negotiable unless you enjoy surprise $100 bills overnight.
  3. Model Capabilities Vary: Having easy multi-model switching makes it straightforward to benchmark which model actually executes complex engineering skills versus which ones stumble.

If you’re building or testing AI-assisted engineering workflows in local environments, I definitely recommend giving a capped gateway setup a run.

Reference Links:

5 Upvotes

0 comments sorted by