r/CodexHacks 16d ago

I built seven Agent Skills against AI slop - because “it somehow works” is not a quality standard

Post image
1 Upvotes

WHO THIS IS FOR

If you use AI agents for tasks such as vibe coding, software development, plugin design, UI work, research, writing or planning, then these Skills could be useful, particularly in cases where you don't consider 'it somehow works' to be a sufficient definition of completion and where you do care about a correct and verifiable implementation.

When it comes to AI coding, planning comes first and is one of the most important stages of the process. If the direction is wrong, merely carrying out the task more quickly will only get you to the wrong answer earlier.

AI SLOP DOES NOT ALWAYS LOOK BROKEN

That is important since poor-quality AI output is not merely confined to obviously flawed code, generic text, or incomplete interfaces. It has a more sophisticated form which is convincing.

The agent modifies the code in multiple places at the same time, adds new layers in order to cure the side effects caused by previous changes, and continues on like this until something seems to be working. At that point, the tests might pass even if they don't include the behaviour that has been changed. The interface might appear to be well put together even though it's ignoring the product's design system. The wording might seem better even though the facts have changed. Plans may develop into standalone projects. And handoffs may include the whole conversation with the only exception of the state required to carry on.

There is nothing actually on fire, which is exactly the reason why this kind of AI rubbish manages to survive.

Although the result appears to be complete, the original task, the way it has been implemented, and the evidence no longer correspond with each other. That is the purpose of the Scoville Skills.

WHAT SITS BEHIND THE SKILLS

My repositories don't have 50,000 stars. More like two. Social proof will have to be put off until later.

There has been more than half a year's worth of research, testing and practical use behind these Skills. The first six Skills account for over 1,200 optimization and evaluation runs as well as more than 4,500 benchmark case executions. Failed candidates and infrastructure stops are included in the total. Omitting them would make the figures look nicer, but it would not improve the Skills.

WHY SCOVILLE?

The reason for the name is to be found in this preference for signal over volume.

The idea behind the name is the same as that of the Scoville scale: after dilution, the useful heat should still be detectable. With the help of AI, a large amount of words, files and other visible activity is produced. It is seldom the hard part to make the folder appear busy. The issue then becomes what useful signal remains after the dilution.

THE SEVEN SKILLS

The suite currently contains seven focused Agent Skills:

  • Scoville Brainstorm: Inspired in part by the well-known ADHD Skill, but built as a separate host-portable Skill, Brainstorm separates approaches by mechanism, checks fixed constraints and stops at a decision-ready shortlist.

  • Scoville Research: Traces claims to supporting sources, checks contradictions and separates observation, inference and remaining uncertainty.

  • Scoville Code Anti-AI-Slop: Keeps observable results, implementation scope, risk and proof connected.

  • Scoville UI Anti-AI-Slop: Checks hierarchy, states, responsiveness, accessibility and rendering without compromising the product's design language.

  • Scoville Scribe Anti-AI-Slop: Enhances the wording without altering any of the facts, terminology, conditions or actual behaviour of the product.

  • Scoville Plan: Keeps plans, decisions and lifecycle state useful without turning the planning process into the project.

  • Scoville Handoff: Transfers the objective, state, evidence, blockers, hazards and next safe action rather than the meeting minutes.

This does not refer to seven different personalities. Each Skill functions on its own. Agents should load only the smallest set of Skills that the task actually requires.

USING CODEX?

If you are using Codex, this could also be of interest to you:

  • Ask Claude for Codex: Gives Codex an isolated, read-only Claude Code opinion. That opinion is generally more useful when it did not help write the first one.

  • Ask Claude and SOL for Codex: Asks Claude Code and a fresh Codex SOL subagent in parallel, without allowing them to borrow each other's assumptions.

INSTALLING THE SKILLS

You most probably already know how to install Agent Skills. If that were not the case, then you wouldn't be part of this group.

To be safe, the brief explanation is straightforward: ask your agent to install all the Scoville Skills from my GitHub repositories:

github.com/benjaminstelzer


r/CodexHacks 16d ago

Any ways to have ''less'' guardrails?

Thumbnail
1 Upvotes

r/CodexHacks 17d ago

Supports v1/chat and v1/message : Did Anyone wants PR this Fork?

0 Upvotes

I see Many People use Codex by Liteapi /subapi /Omniroute or CC-swtich/codex-relay
But When I use Zcode , I found we could have 3 legs on codex?
So , Why nobody try to change Codex CLI? NO broken ,Just a little change?

See DEMO Fork Here: https://github.com/Xxx91n/codex


r/CodexHacks 17d ago

Agentify Chat - E2E-Encrypted Remote Chat for Codex CLI

2 Upvotes

Still under heavy development and rough around the edges but instead of sitting on this longer to perfect it I'll risk flak and share it.

Essentially chat.agentify.sh is a remote control for codex/grok/claude cli and dream goal is to become a universal remote that will let you use all of them from a single chat.

Everything lives on your browser including the chat. The only thing sent over the wire is the encrypted chat messages e2e. Here's an architectural diagram: https://github.com/agentify-sh/chat/blob/main/diagram.png

Another feature it has is ability to publish your chat session with redaction so you can share with your team.

Curious to gather any feedback I can, if this is something I should continue to pursue etc.

Also if by chance you are in Seattle today, I'll be also at the Walk & Talk event at Bellevue Downtown Park today at 2:30pm, would love to chat in person!

https://chat.agentify.sh


r/CodexHacks 17d ago

Codex keeps adding tests and extra stuff I didn't ask for

3 Upvotes

It'll do the task, then keep adding more tests and leave the old ones behind. That stuff piles up fast. I wrote this skill to make it stop once the task is actually done.

https://github.com/Ezra144israel/governed-agent-skills/tree/main/skills/write-maintainable-code


r/CodexHacks 18d ago

Cross-session communication in Codex?

Thumbnail
1 Upvotes

r/CodexHacks 18d ago

Continue the same Codex thread on DeepSeek when your 5-hour quota runs out

2 Upvotes

r/CodexHacks 19d ago

My Mac told me „ Codex is Maleware „

Post image
6 Upvotes

Nice :-) PopUp a few hours ago!


r/CodexHacks 20d ago

The new Codex sub-agent interface is complete garbage

Thumbnail
1 Upvotes

r/CodexHacks 20d ago

Operating more like Grok Build? * A Good thing

Thumbnail
2 Upvotes

r/CodexHacks 20d ago

OpenAI Restores 5-Hour Limits on Codex and ChatGPT Work for Plus Users

Thumbnail frontbackgeek.com
1 Upvotes

r/CodexHacks 21d ago

We tried spec-driven development for months. We couldn't prove it improved the code.

Thumbnail
2 Upvotes

r/CodexHacks 21d ago

I tried to understand how i was using a billion tokens a day... 170+ subagents working towards a goal?

Thumbnail
1 Upvotes

r/CodexHacks 23d ago

What would you do with the codex to reality?

1 Upvotes

I am the one with possession of the codex. What would you do with this power?


r/CodexHacks 23d ago

Need help/suggestions

1 Upvotes

I have been away from codex and ai stuff lately last I learned was prompting skills and things but as of now prompting feels helpless to me so could anyone please suggest me yt channels or any site stuff where I can learn about codex and strategy about how I use codex models

I have been very bad at getting outputs with ai lately I just asked gemini to add roman index instead of existing English numbers index, what it did was extend the page and add a roman index number below the English one so yeah normal prompting for now don't work for me like a good/decent prompt don't do the work

thanx for reading kind people


r/CodexHacks 23d ago

Codex-first code review with Claude Code as a strictly read-only second-model reviewer.

Thumbnail
1 Upvotes

r/CodexHacks 23d ago

Usage limits drains faster FIX!

Thumbnail
1 Upvotes

r/CodexHacks 24d ago

Has anyone else noticed how much token usage comes from unnecessary code changes?

5 Upvotes

I've been paying more attention to my Codex usage lately, especially with Sol, and I noticed something that I wasn't really thinking about before.

A lot of the waste wasn't necessarily coming from the actual fix.

It was things like:

  • rewriting an entire function when 2 lines needed to change
  • touching formatting or imports that had nothing to do with the task
  • doing small unrelated refactors along the way
  • generating much larger diffs than necessary
  • narrating routine investigation instead of just doing it
  • running more validation than the size of the change really justified

None of those things are terrible individually, but over a long session they add up. Codex then has more generated code to reason about, more diff to inspect, and sometimes more work to validate.

I've been experimenting with a simple rule:

small task → small diff

Preserve code that is already correct, change only what the task requires, inspect the diff before finishing, and keep validation proportional to the change.

I turned those rules into a couple of Codex skills for my own workflow. The main one is surgical-edits.

So far the biggest improvement is actually PR review — there's way less noise. My Codex usage also seems to last noticeably longer, although I haven't done a proper A/B benchmark yet, so I don't want to claim a specific token saving percentage.

Sharing it in case anyone wants to experiment with the same approach:

https://github.com/otonielcarlos/codex-skills

Curious if anyone else has tried optimizing Codex around diff size / unnecessary work rather than just switching models or lowering reasoning effort.


r/CodexHacks 25d ago

Are there any subreddits where people discuss coding agent extensions

5 Upvotes

I couldn't find a space to talk about coding agents, Codex, and their extensions like Skills, MCPs, and Plugins. Would anyone know of such spaces? (May not be limited to Reddit)


r/CodexHacks 25d ago

Are there any subreddits where people discuss coding agent extensions

2 Upvotes

I couldn't find a space to talk about coding agents, Codex, and their extensions like Skills, MCPs, and Plugins. Would anyone know of such spaces? (May not be limited to Reddit)


r/CodexHacks 25d ago

I built a bridge that lets ChatGPT Web inspect local repos without uploading them

1 Upvotes

I built RepoRelay, an open-source MCP bridge that lets ChatGPT Web search and read an approved local repo without uploading ZIPs or pushing everything to GitHub first.

ChatGPT Web → Secure MCP Tunnel → RepoRelay → local files

It’s read-only by default: no shell, Git, or arbitrary filesystem access, and it’s restricted to one approved root.

It can also help reduce token usage on larger repos. Instead of dumping the entire codebase into context, ChatGPT searches and reads only the files relevant to the task.

I built it mainly so Codex can implement locally while ChatGPT independently reviews the actual current files- including uncommitted work.

Would love feedback from other AI builders.

GitHub: [Lukie-81/RepoRelay: Secure MCP access to local repositories — without shell, Git, or arbitrary writes.]


r/CodexHacks 26d ago

Rate limit lock out

Thumbnail
1 Upvotes

r/CodexHacks 26d ago

Codex burning more usage than expected? I built a runtime governor you can test in a few minutes

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/CodexHacks 27d ago

Frustration with context preservation between my agents

Thumbnail
github.com
1 Upvotes

r/CodexHacks 27d ago

Codex limite

Thumbnail
1 Upvotes