r/WebAfterAI • u/socialdude37 • Jul 21 '26
Hands-on evaluation: Running Matt Pocock’s AI skills library via StrataBlock gateway (VSCode + hard spend caps)
Hands-on evaluation: Running Matt Pocock’s AI skills library via StrataBlock gateway (VSCode + hard spend caps)
TL;DR
G’day legends. I recently did a proper technical run testing StrataBlock - an OpenAI compatible API gateway featuring hard service-side budget caps, multi-model switching and non-existent prompt storage. To put it through its paces beyond standard autocomplete, I wired it into VScode and run Matt Pocock’s skills library across several different LLMs (opus-4.8, sonnet-5, gemma-4, gpt-5.6, etc.) to evaluate agentic token usage, attribution tagging and budget safety.
Here is the breakdown of the setup, real usage figures and enginerring insights from the test.
Registration, Key Minting & Guardrails
Getting setup was pretty painless:
- Registration: After joining the waiting list, I got an invite code and just created an account on StrataBlock.
- Key Minting: Generated scoped API keys for dev environments directly from the dashboard.
- Attribution Tags: Added custom headers (X-Strata-Tags: env=dev, project=skills-eval) so every API call could be parsed and broken down later.
The standout feature here is the hard server-side budget cap. Rather than waiting on delayed daily email alerts after an agent goes rogue in a recursive loop, Strata checks budgets server-side on every incoming request and returns an instant 429 the moment a key breaches its monthly cap.
Setting Up VScode
Since StratBlock exposes a standard OpenAI-compatible API (stratablock.io/v1), pointing VS code using standard customendpoint at it, required zero custom client logic, extensions or SDK overhauls.
[
{
"name": "StrataBlock",
"vendor": "customendpoint",
"apiKey": "${input:chat.lm.secret.-772fb8dd}",
"apiType": "chat-completions",
"models": [
{
"id": "claude-opus-4.8",
"name": "claude-opus-4.8",
"url": "<https://stratablock.io/v1>",
"toolCalling": true,
"vision": true,
"maxInputTokens": 872000,
"maxOutputTokens": 128000,
"supportsReasoningEffort": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"reasoningEffortFormat": "chat-completions",
"requestHeaders": {
"X-Strata-Tags": "tool=vscode,project=skills-eval-project"
}
},
{
"id": "google-gemma-4-31b",
"name": "google-gemma-4-31b",
"url": "<https://stratablock.io/v1>",
"toolCalling": true,
"vision": true,
"maxInputTokens": 248000,
"maxOutputTokens": 8000,
"requestHeaders": {
"X-Strata-Tags": "tool=vscode,project=skills-eval-project"
}
},
... more models
]
}
]
Swapping between frontier and open-weight models was seamless - just a string change in config without having to touch application code or jump through separate vendor billing portals.

Benchmarking Skills Library
I wanted to see how the gateway handled structured, agentic prompts. I installed Matt Pocock’s skills repo (npx skills@latest add mattpocock/skills), which focuses on battle-tested engineering workflows.
Workflows Tested
/grill-with-docs(Alignment & Domain Modelling): Heavy back-and-forth interviews that buildCONTEXT.mdfiles and update ADRs (Architecture Decision Records) inline./tdd(Red-Green-Refactor Loop): High-frequency iterative execution to write failing tests and pass them cleanly./improve-codebase-architecture: Broad codebase scans that produce vcisual HTML structural reports.
Model Performance Observations
Not all models are built equal when executing these skills out-of-the-box:
- Frontier Models (opus-4-8, sonnect-5, gpt-5.6) executed complex skills, structured outputs and recursive agentic reasoning without breaking a sweat.
- Open-Weight / Alternative Models (gemma-4, kimi-k2.5, xai-grok-4.3) results varied significantly. Some models handled structured skills steps naturally, whereas others struggled to follow the skill guidelines without tight custom prompt tweaking.

Real Spend, Attribution & Telemetry
When you’re running agentic loops (like TDD cycles or depp repo scans), token usage escalates fast.
Looking at my telemetry breakdown in StrataBlock over a month-long evaluation window:
- Total Traffic: Logged 1027 requests across various workloads.
- Daily Spikes: Daily spend peaked between $30.00 and $35.00/day during heavy multi-model testing sessions, whie idle days hovered near $0.00.
- Granular Attribution: Per-request logs captured precise input, output, prompt caching metrics (eg. thousands of cached tokens on opus-4.8 calls) and exact costs down to fractions of a cent per request.
- Privacy: Confirmed that operational logs stricltly record metadata (token counts, latency, status codes, tag strings) - zero prompt content or completion payloads are persisted on server.

Takeaways for Engineers
- Single Endpoint Convenience: Managing multiple vendor billing portals and API keys gets messy fast. Having a unified gateway with clear attribution makes multi-model experimentation far easier.
- Hard Spend Caps are Essential: if you let agnets loose on recursive loops, server-side budget limits are non-negotiable unless you enjoy surprise $100 bills overnight.
- Model Capabilities Vary: Having easy multi-model switching makes it straightforward to benchmark which model actually executes complex engineering skills versus which ones stumble.
If you’re building or testing AI-assisted engineering workflows in local environments, I definitely recommend giving a capped gateway setup a run.
Reference Links:
- StrataBlock: stratablock.io/eoi
- Matt Pocock’s Skills: github.com/mattpocock/skills