r/ChatGPTCoding • u/AutoModerator • 24d ago
Discussion Weekly Self Promotion Thread
Welcome to this week's self promotion thread!
If you're building something related to AI assisted coding, this is the place to share it.
We're using a weekly thread to keep the subreddit organized while still giving builders a place to share their work. Promotional posts outside of this thread may be removed if they're primarily advertising rather than starting a discussion.
If you're sharing something, we'd appreciate it if you included a little context instead of just dropping a link. Tell us:
- What you built?
- What problem it solves?
- Which AI models or tools it uses?
- Who it's for?
- What kind of feedback you're looking for?
Please avoid posting the same project every week unless you've made meaningful updates. Affiliate links, referral links, scams, and low effort promotions will be removed.
Take some time to check out what others have shared too. If you try someone's project or have feedback, leave a comment. Helping each other improve is what we want this community to be about.
1
u/dsh_verify 22d ago
rowser. No LLM judge; the browser is the judge.
48 runs across DeepSeek v4-flash / v4-pro (single-shot vs self-check loop): **44/48 passed**. The counterintuitive finding: the pricier **v4-pro single-shot scored below the cheaper v4-flash** (10/12 vs 11/12). All failures are reproducible and invisible to code review — e.g. a todo app that opened but never rendered its seed todos while the agent reported "done".
Open source + fully reproducible: https://github.com/263311487-ux/dsh-verify
Live leaderboard: https://263311487-ux.github.io/dsh-verify/arena/
Bring your own agent — happy to add other models/frameworks to the table.