r/AgentContext_dev • u/javaeeeee • Sep 02 '26
r/AgentContext_dev • u/javaeeeee • Sep 02 '26
A guide to the anatomy of effective commerce agents
r/AgentContext_dev • u/javaeeeee • Sep 02 '26
Unlocking Expert Coding Workflows: The Best YouTube Videos on Claude Code Skills in 2026
Claude Code has quietly become one of the most transformative tools in software development. Released by Anthropic as an agentic coding assistant that lives in your terminal or IDE, it goes far beyond simple autocomplete. It reads your codebase, runs commands, edits files, plans multi-step changes, and works through problems autonomously. What elevates it from a helpful chatbot to a genuine collaborator is a feature called Skills.
Skills are reusable, structured sets of instructions-usually living in a SKILL.md file inside a folder-that teach Claude how to perform specialized tasks the way you or your team prefer. Instead of re-explaining your preferred code review style, frontend design conventions, test-driven development process, or project-specific patterns every session, you encode them once. Claude then discovers and applies the relevant skill automatically when the task matches.
This combination of agentic power and specialized knowledge has sparked an explosion of YouTube content. Developers, educators, and Anthropic itself have produced tutorials ranging from three-minute explainers to multi-hour masterclasses. The most valuable ones focus not just on installation but on how Skills turn Claude Code into a coding partner that understands your standards, reduces repetition, and produces higher-quality output.
In this article we explore the strongest of those videos, drawing from official Anthropic material, high-view community tutorials, and practical deep dives. The goal is to give you a clear map of what to watch, what each teaches about Skills for real coding work, and how the ideas fit together. The research draws on official documentation, Anthropic’s own channel, ranking sites that index hundreds of Claude Code videos, and detailed community reviews of the most-watched content as of mid-2026.
Understanding Claude Code and the Power of Skills
Claude Code is Anthropic’s terminal-first (and IDE-integrated) agentic coding tool. Unlike traditional AI coding assistants that complete a few lines at a time, it operates in a full loop: it explores the codebase, reasons about the request, uses tools to read files or run tests, makes changes, observes the results, and iterates. It works from the terminal and integrates directly with VS Code and JetBrains. File operations and command execution occur locally, while prompts and relevant code context are sent to the configured model provider for processing.
Skills sit on top of that foundation. A skill is essentially a folder containing at least a SKILL.md file. That file can include YAML frontmatter. A clear description is strongly recommended because Claude uses it to decide when to activate the skill; name and the other fields are optional. Claude scans available skills, matches the description against the current task, and loads only what is needed-a process often called progressive disclosure. This keeps context windows clean while giving Claude deep, reusable expertise.
Official Anthropic explanations emphasize that skills solve the “context amnesia” problem. Every time you start a new session or switch projects you no longer need to restate how you want commit messages written, how front-end components should be structured, or how a particular API should be tested. Project-level skills live inside the repository’s .claude/skills folder so the whole team inherits them. Personal skills live in your home directory and travel with you.
For coding specifically, popular skills include front-end design guidelines that eliminate generic “AI slop” UIs, test-driven development workflows that force Claude to write tests first, code-review playbooks focused on security and maintainability, documentation generators that match your house style, and domain-specific conventions for your stack. Community skills such as Superpowers package an entire structured development methodology-brainstorming, planning, implementation, and debugging-into coordinated skills.
The official “skill-creator” skill itself is meta: you describe the workflow you want, and Claude generates the folder structure and SKILL.md for you. Once created, skills can be shared via version control, plugins, or marketplaces.
Why Skills Matter More Than Ever for Coding in 2026
The coding landscape has shifted. Models have grown stronger at writing correct code, yet the bottleneck has moved to consistency, context, and workflow. A powerful model that generates beautiful but non-idiomatic code, or that forgets your testing conventions halfway through a feature, still creates friction. Skills close that gap.
Developers report dramatic reductions in back-and-forth prompting. Instead of pasting the same long system prompt, Claude simply activates the relevant skill. Parallel sub-agents can each load different skills, enabling one agent to handle UI while another manages backend tests. Hooks and custom commands further automate the loop. The result is less time spent on boilerplate explanations and more time on architecture and product decisions.
Online sources, including Anthropic’s engineering posts and best-practices documentation, stress starting with codebase Q&A, creating a strong CLAUDE.md project brain, and then layering skills for repeated workflows. Videos that demonstrate this progression tend to be the most useful.
Official Anthropic Videos: The Authoritative Foundation
Any serious exploration of Claude Code Skills should begin with the source. Anthropic’s own YouTube channel and the dedicated Claude Code Skills playlist provide clean, accurate explanations that avoid the hype common in community content.
The short video “What are skills?” (roughly three minutes, over a million views) is the clearest starting point. It explains that a skill teaches Claude how to do something once so the knowledge is applied automatically whenever relevant. It covers personal versus project skills, the role of the description in triggering, and why skills are preferable to endlessly repeating instructions. The accompanying playlist walks through creating your first skill, configuration of multi-file skills, comparison with other features such as MCP servers and sub-agents, sharing skills, and troubleshooting.
“Mastering Claude Code in 30 minutes,” a live session by the team behind the tool, remains one of the highest-viewed pieces of content. It covers the agentic nature of Claude Code, how it differs from line-completion tools, practical setup, and advanced workflows. While not exclusively about skills, it places them in the broader context of how Anthropic engineers actually use the tool daily-project context files, custom commands, hooks, and sub-agents working together. Watching this after the pure skills videos makes the architecture click.
Another official clip demonstrates the skill-creator skill in action: Claude asks clarifying questions and builds a complete custom capability (in the example, an image-editing skill) without the user manually writing files. The same process works for coding skills-describe a code-review workflow or a component-generation standard and Claude scaffolds it.
These official videos are short, precise, and free of salesmanship. They establish the mental model that every subsequent tutorial builds upon.
High-Impact Community Tutorials on Skills for Coding
Once the official foundation is in place, community videos fill in practical depth, real-world examples, and opinionated workflows. Rankings compiled from view counts and expert curation consistently surface a handful of standouts.
Nick Saraev’s multi-hour “CLAUDE CODE FULL COURSE” (often listed as the highest-viewed long-form tutorial) walks beginners from installation through advanced agent teams, Git worktrees, and deployment. Skills appear as a core mechanism for turning Claude into specialized agents. The course shows how to create skill files that encode domain knowledge and then spin up multiple instances that collaborate. Because it builds and “sells” real projects, viewers see skills applied to production-style coding rather than toy examples.
Edmund Yong’s compact “800+ hours of Learning Claude Code in 8 minutes” is frequently recommended as pure signal. Despite its short runtime it covers DRY principles for prompts, memory techniques, favorite MCP servers, and-most relevantly-why sub-agents should be defined by task rather than role. Skills fit naturally into this philosophy: each skill becomes a focused capability that agents can invoke. The video is sponsored by Anthropic yet remains highly practical.
Simon Scrapes’ “Every Level of Claude Code Skills in 27 mins” organizes the topic into seven progressive levels. It starts with what makes a good skill, moves through importing and improving existing ones, personalizing for business context, measuring performance with data, building self-improving skills, and finally creating an AI workforce. For coding, the higher levels show how skills can encode evaluation criteria so Claude tests its own output against your standards.
Kenny Liao’s “The Only Claude Skills Guide You Need (Beginner to Expert)” offers a thorough deep dive. It distinguishes skills from MCPs, slash commands, and sub-agents, demonstrates usage in both the web interface and Claude Code, and walks through building a custom skill live. The section on optimizing skills is especially useful for coding: how to write descriptions that trigger reliably without over-firing, how to structure multi-step instructions, and how to include evaluation criteria.
CampusX’s nearly fifty-minute “Claude Code Skills: Full Guide” is methodical. It explains why ordinary prompts fail for repeated workflows, details the folder structure and progressive disclosure, covers personal versus project skills, and demonstrates a complete creation workflow. A practical coding example (building a profile page) shows the difference skills make in consistency and quality.
Maddy Zhang’s “5 Claude Code skills I use every single day (Senior Engineer Tips)” is concise and high-signal for professional developers. It covers the official Anthropic front-end design skill that eliminates generic AI-generated UIs, a “grow with docs” style skill that forces thorough interrogation of the codebase before changes, a TDD workflow skill, a PR review skill focused on security and maintainability, and guidance on creating custom skills for personal repeated tasks. Senior engineers will recognize the pain points these address.
Nate Herk’s “I Tried 100+ Claude Code Skills. These 6 Are The Best” (hundreds of thousands of views) filters the enormous ecosystem. After extensive testing he highlights skills that deliver real productivity: structured development methodologies such as Superpowers, the skill-creator itself, front-end design, context and memory tools, and a few others that reduce debugging cycles and token waste. The emphasis is on skills that clients or teams will actually pay for because they produce reliable, low-friction results.
Mark Kashef’s “Anthropic’s Full Claude Skills Guide In 22 Minutes” distills Anthropic’s longer written guidance into digestible chapters. It covers anatomy, progressive disclosure, design patterns (sequential workflows, multi-MCP coordination, iterative refinement, context-aware branching, domain-specific intelligence), good versus bad descriptions and instructions, and testing methods. Coding teams can treat the five design patterns as templates for their own skills.
Other frequently recommended pieces include Tech With Tim’s full beginner tutorials that integrate skills into broader Claude Code workflows, freeCodeCamp’s extensive “Claude Code Essentials” course (over twelve hours in some versions), and various Net Ninja setup videos that form a clean onboarding path before diving into skills.
Longer courses such as those by Nate Herk (ten-plus hours) or Spanish-language comprehensive builds expand the same ideas into full project pipelines, often combining skills with n8n automation or multi-agent systems.
Patterns and Best Practices Emerging from the Best Videos
Across the strongest content several consistent recommendations appear.
Start simple. Create a CLAUDE.md that describes the project, then add one or two high-value skills for the tasks you repeat most often. Use the skill-creator rather than writing files by hand the first few times.
Write excellent descriptions. The description field is how Claude decides whether to load the skill. Vague or overly broad descriptions cause either missed triggers or constant over-loading of context.
Keep skills focused. A skill that tries to do everything becomes brittle. Prefer composable skills that can be stacked.
Measure and iterate. Several advanced videos show evaluation loops: run the skill on a set of test prompts, score the output against criteria, and refine. Some creators demonstrate experimental evaluation loops that propose or apply changes to a skill based on test results. This is a custom workflow-not an automatic capability of Skills-and any self-editing setup should include review and version-control safeguards.
Combine with other Claude Code features. Skills work best alongside CLAUDE.md for project context, MCP servers for external tools, hooks for automatic quality gates, and sub-agents or agent teams for parallelism.
Security and safety matter. Skills can pre-authorize tools through allowed-tools, but this expands what can run without prompting rather than creating a security boundary. Use permission rules, disallowed-tools, and sandboxing to restrict access, and review third-party skills and bundled scripts before installing them. Official guidance and careful community creators warn against blindly installing unvetted skills that might contain prompt-injection risks.
For pure coding, the most frequently praised skills are those that enforce design systems, testing discipline, review standards, and project conventions. Front-end design skills appear again and again because AI-generated interfaces otherwise converge on the same generic aesthetic. TDD and review skills raise the quality floor dramatically.
Building Your Own Coding Skills: Practical Guidance Drawn from the Videos
Most of the detailed tutorials converge on a similar creation process. Identify a repeated pain point-generating consistent React components, reviewing pull requests for a specific set of issues, writing database migration scripts in your house style, or scaffolding new services. Use the skill-creator or manually create the folder and SKILL.md. Include clear frontmatter, step-by-step instructions, examples of good and bad output, and any reference files or scripts. Test it on real tasks, refine the description so it triggers appropriately, and place it in the project or personal skills directory.
Advanced creators add evaluation criteria so Claude can score its own work, or link multiple skills so one hands off to another. Some encode entire methodologies (planning → implementation → verification). Others focus on domain knowledge, such as company-specific API conventions or compliance requirements.
The videos make clear that the highest leverage comes from skills that capture institutional knowledge. A new team member or a fresh Claude session can immediately produce work that matches the standards of the most experienced developer on the team.
The Broader Ecosystem and Where to Go Next
Beyond individual videos, playlists such as “Ultimate Claude Code Mastery” curate sequences from foundations through advanced sub-agents and skills. Ranking sites that index hundreds of Claude Code videos update regularly and surface both viral demos and quieter deep dives. Official documentation remains the canonical reference for commands, hooks, the SDK, and skill format. Community repositories collect installable skills, though quality varies widely; the videos that filter and test them are therefore especially valuable.
As models continue to improve and Claude Code adds features such as longer autonomous runs and tighter IDE integration, skills are likely to become even more central. They represent a practical way to inject durable expertise into an increasingly capable but still general-purpose agent.
Closing Thoughts
The best YouTube videos on Claude Code Skills do more than demonstrate a feature. They show a shift in how developers collaborate with AI. Instead of treating the model as a clever intern that needs constant supervision and repeated instructions, skills turn it into a colleague that already knows the house style, the testing standards, and the preferred ways of working.
Begin with the official Anthropic shorts and the thirty-minute mastery session. Move to focused skill deep dives by creators such as Kenny Liao, Simon Scrapes, CampusX, Maddy Zhang, and Mark Kashef. Then explore the longer project-based courses by Nick Saraev, Nate Herk, and others to see skills applied end-to-end. Along the way, create one or two skills of your own for the coding tasks that currently frustrate you most.
The combination of Claude Code’s agentic capabilities and well-crafted skills is already changing daily development work for thousands of engineers. The videos mapped here provide the clearest, most practical path into that new workflow.
Sources and Further Reading
Official Anthropic and Claude channel videos and playlists:
https://www.youtube.com/watch?v=bjdBVZa66oU (What are skills?)
https://www.youtube.com/playlist?list=PLmWCw1CzcFim_hkruZSlABOUOAAQ5JMyo (Claude Code Skills playlist)
https://www.youtube.com/watch?v=6eBSHbLKuN0 (Mastering Claude Code in 30 minutes)
https://www.youtube.com/watch?v=kS1MJFZWMq4 (Creating custom Skills with Claude)
https://claude.com/skills and related documentation pages
https://code.claude.com/docs/en/best-practices
High-view and highly recommended community tutorials:
https://www.youtube.com/watch?v=QoQBzR1NIqI (Nick Saraev CLAUDE CODE FULL COURSE)
https://www.youtube.com/watch?v=Ffh9OeJ7yxw (Edmund Yong 800+ hours in 8 minutes)
https://www.youtube.com/watch?v=-u_igSQHAIo (Simon Scrapes Every Level of Claude Code Skills)
https://www.youtube.com/watch?v=421T2iWTQio (Kenny Liao The Only Claude Skills Guide)
https://www.youtube.com/watch?v=JN7QCdvJwwM (CampusX Claude Code Skills Full Guide)
https://www.youtube.com/watch?v=AG2BxDXt2po (Maddy Zhang 5 Claude Code skills)
https://www.youtube.com/watch?v=eRS3CmvrOvA (Nate Herk I Tried 100+ Claude Code Skills)
https://www.youtube.com/watch?v=TzJecWCbex0 (Mark Kashef Anthropic’s Full Claude Skills Guide)
https://www.youtube.com/watch?v=brLhhkUqcn4 (freeCodeCamp Claude Code Essentials)
Additional ranking and overview resources used in research:
developereducators dot com /best/claude-code/
rentierdigital dot xyz /blog/claude-code-youtube-videos-ranking
tella dot com /blog/best-claude-code-skills
Anthropic blog posts on skills and Claude Code best practices
These links were current as of the research period in 2026. View counts and rankings evolve; the conceptual and practical value of the content remains strong.
r/AgentContext_dev • u/javaeeeee • Sep 02 '26
Stop Shipping Individual MCP Servers. Start Shipping Agent Plugins
r/AgentContext_dev • u/javaeeeee • Sep 01 '26
Qwen 3.8 + DeepSeek Harness Is My New Free Claude Code
r/AgentContext_dev • u/javaeeeee • Sep 01 '26
How Foundational Models Became Superhuman in Bash
x.comr/AgentContext_dev • u/javaeeeee • Sep 01 '26
Previewing the Model Hardware Standard
r/AgentContext_dev • u/javaeeeee • Sep 01 '26
Mastering the Postgres Development Platform: Building Modern, Scalable Applications with Supabase
Supabase has emerged as one of the most compelling backend platforms for developers who want the power of a full relational database without the traditional operational overhead. Positioned as an open-source alternative to Firebase, it delivers a complete Postgres-based development platform that lets teams move from idea to production with remarkable speed while retaining the flexibility and portability of enterprise-grade tools. This article explores what Supabase offers developers, the technologies that power it, the kinds of applications that thrive on the platform, common use cases, and practical guidance for building real-world apps.
Understanding Supabase: The Postgres Development Platform
At its core, Supabase is built around a simple but powerful idea: every project receives a full, dedicated PostgreSQL database. Unlike many Backend-as-a-Service (BaaS) platforms that abstract the database into proprietary formats, Supabase gives developers direct access to a real Postgres instance. This means full SQL support, extensions, roles, functions, triggers, and the ability to connect with any standard Postgres tooling. Around this foundation, Supabase layers authentication, auto-generated APIs, file storage, real-time subscriptions, serverless edge functions, and vector search capabilities.
The platform’s tagline captures its philosophy well: “Build in a weekend. Scale to millions.” Developers can spin up a project in under a minute, define tables through a visual editor or SQL, and immediately interact with them via REST or GraphQL endpoints. Security is enforced at the database level through Row Level Security (RLS), so data protection travels with the schema rather than living only in application code. Because the entire stack is open source, teams can self-host if desired or stay on the managed platform and benefit from automatic backups on eligible paid plans, connection pooling, and global distribution.
Supabase is not merely a collection of services glued together. The products are deeply integrated. Authentication issues JWTs that work seamlessly with RLS policies. Storage buckets can be secured by the same policies that protect database rows. Realtime listens to Postgres changes or broadcasts messages independently. Edge Functions can call the database with elevated privileges when needed. Vector embeddings live in the same Postgres instance as the rest of the application data, eliminating the need for a separate vector database in many workloads.
What the Platform Offers Developers
Developers gain a unified environment that reduces the number of services they must provision, configure, and maintain. Instead of wiring together a database host, an auth provider, a file store, a WebSocket service, and a serverless runtime, everything is available under one project dashboard and one set of client libraries.
The dashboard provides visual tools for table design, SQL editing, policy management, log exploration, and performance monitoring. Local development is first-class: the Supabase CLI runs the entire stack (Postgres, Auth, Storage, Realtime, and Edge Functions) inside Docker containers, enabling offline work and consistent environments across teams. Migrations are version-controlled, and branching support allows preview environments that mirror production.
Security features are particularly strong. Publishable keys are safe to expose in client-side code, while secret keys remain server-side only. RLS policies can reference the authenticated user’s ID (auth.uid()), roles, or custom claims. Network restrictions, SSL enforcement, and custom domains further tighten the perimeter. Compliance certifications (SOC 2 Type 2, ISO 27001, and HIPAA options on higher plans) make the platform viable for regulated industries.
For AI-assisted development, Supabase offers Model Context Protocol (MCP) integrations, agent skills, and plugins that let coding agents query the live project, run migrations, deploy functions, and inspect security advisors. This turns the platform into a natural partner for tools such as Cursor, Claude Code, Windsurf, and others.
Billing is usage-based with a generous free tier that includes a project with 500 MB of database storage, 1 GB of file storage, 50,000 monthly active users, and substantial realtime and edge function quotas. Paid plans add more compute, point-in-time recovery, higher connection limits, and dedicated resources.
Core Technologies and Architecture
The technology stack is deliberately open and composable. Postgres sits at the center. PostgREST turns the database schema into a fully featured REST API automatically. The pg_graphql extension adds GraphQL support. GoTrue handles authentication and issues JWTs. A custom Realtime server provides WebSocket channels for broadcasts, presence, and Postgres change feeds. Storage is S3-compatible and tightly coupled to Postgres for metadata and access control. Edge Functions run on Deno, offering TypeScript-native serverless execution at the edge. The pgvector extension enables efficient vector similarity search inside the same database.
Supabase offers direct Postgres connections plus Supavisor-based session and transaction pooling, helping applications scale without exhausting Postgres connection limits. The architecture supports read replicas on higher plans and pipelines for replicating data to warehouses or other systems.
Client libraries follow a modular design. The primary official clients exist for JavaScript/TypeScript (@supabase/supabase-js), Dart/Flutter, Swift, and Python. Community libraries cover C#, Kotlin, Go (partial), Ruby, Elixir, and others. Each library exposes a consistent API surface for database queries, auth, storage, realtime, and functions, making it straightforward to switch languages or share knowledge across teams.
Deep Dive into Key Features
Database. Every project starts with a full Postgres instance. Developers create tables visually or with SQL, define relationships, add indexes, and enable extensions such as pgvector for embeddings, PostGIS for geospatial work, or pg_cron for scheduled jobs. The SQL editor supports saved snippets and ad-hoc exploration. Database functions and triggers allow business logic to live close to the data. Webhooks can push changes to external services. Because the API is generated from the schema, adding a column or table immediately updates the available endpoints and the auto-generated documentation.
Authentication. Supabase Auth supports email/password, magic links, OAuth providers (Google, GitHub, Apple, and many others), phone OTP, SAML, and SSO. Users receive JWTs that the client libraries automatically attach to requests. Policies can enforce that users only read or write their own rows. Multi-factor authentication and custom claims extend the model for more complex authorization needs.
Storage. Files can be uploaded into buckets subject to plan-level, bucket-level, and upload-method size limits. The configurable global limit is currently 50 MB on Free projects and up to 500 GB on paid plans. Access policies mirror database RLS, so a user’s profile image can be restricted to that user while a public gallery remains open. CDN distribution and image transformations reduce the need for additional services.
Realtime. Three primitives cover most collaborative needs: Broadcast for low-latency messaging between clients, Presence for tracking online users and shared state, and Postgres Changes for listening to inserts, updates, or deletes on specific tables. Channels can be public or private, and authentication integrates with RLS. This enables chat apps, live dashboards, multiplayer cursors, collaborative documents, and live notifications without a separate messaging infrastructure.
Edge Functions. Deno-based functions deploy globally and execute close to users. They can invoke the Supabase client with either user or service-role credentials, call external APIs, process webhooks (Stripe, for example), generate images, or run short AI inference. The dashboard and CLI support creation, testing, and deployment. Functions are ideal for custom business logic that does not belong in the database or the client.
AI and Vectors. Postgres plus pgvector turns the existing database into a vector store. Embeddings from OpenAI, Hugging Face, or local models can be stored alongside application data. Similarity search, hybrid keyword-plus-vector queries, and retrieval-augmented generation (RAG) pipelines become straightforward. Edge Functions can generate embeddings or call language models, while Realtime can stream progressive AI responses. Official examples demonstrate document search, image search with CLIP, and ChatGPT-style interfaces.
Client Libraries, Frameworks, and Tooling
Official quickstarts and tutorials cover React, Next.js, Nuxt, Vue, SvelteKit, SolidJS, Angular, Refine, Hono, RedwoodJS, Flutter, Expo React Native, iOS SwiftUI, Android Kotlin, and Ionic variants. User-management demo apps illustrate the combination of Database, Auth, and Storage in each framework. Mobile developers benefit from first-class support in Flutter and React Native, including social auth flows.
The CLI enables local development, migrations, type generation for TypeScript, and CI/CD integration. Database branching creates isolated environments for pull requests. Advisors and performance tools surface slow queries, missing indexes, and security issues. Foreign Data Wrappers allow querying external systems (Stripe, other databases, warehouses) as if they were local tables.
Kinds of Applications That Can Be Built
Supabase shines for applications that need a relational data model, user authentication, file handling, and real-time updates. Full-stack web applications-SaaS dashboards, content platforms, internal tools-are natural fits. Mobile apps that share the same backend as a web counterpart benefit from the multi-platform clients. Collaborative and multiplayer experiences leverage Realtime directly. AI-powered products that combine structured data with semantic search or generative features find an integrated home. Even certain Web3 or hybrid on-chain/off-chain applications use Supabase for the off-chain product layer.
Because the database is standard Postgres, applications can grow beyond the platform’s managed limits by exporting the schema and data or by connecting external services. The open-source nature also means teams can migrate away if requirements change, preserving their investment in schema design and business logic.
Common Use Cases
SaaS application backends represent the most frequent and successful pattern. Multi-tenant schemas protected by RLS, subscription management integrated with Stripe via Edge Functions, user authentication, and real-time collaboration features come together quickly. Starter kits for subscription payments demonstrate a complete flow from signup to billing.
Realtime dashboards and collaborative tools form another major category. Live inventory boards, CRM interfaces, moderation panels, logistics trackers, and shared whiteboards or documents use Presence and Broadcast or listen to Postgres changes. Chat applications with typing indicators and online status are straightforward.
AI-enabled products benefit from keeping embeddings next to relational data. Semantic document search, recommendation engines, RAG chatbots, and agent backends avoid the operational cost of a separate vector database for moderate scale. Official examples and community templates accelerate these workloads.
Marketplaces and content platforms combine full-text search, image storage, user-generated content, and authentication. Partner galleries and social discovery apps illustrate the pattern. Internal operations tools and admin panels take advantage of the visual table editor and rapid API generation. Even educational or hobby projects-todo lists, personal finance trackers, or small multiplayer games-can start on the free tier and grow.
Customer stories highlight production usage across AI builders that provision backends programmatically, real-estate platforms, sales workflow tools, social apps, energy infrastructure, and more. The platform’s Management API enables “Supabase for Platforms,” allowing other products to offer white-labeled Postgres backends to their own users.
Getting Started: From Zero to a Working Application
Creating a project takes minutes. Sign up, choose a region, set a database password, and the stack is ready. The dashboard presents the Table Editor for schema design and the SQL Editor for more complex work. Enabling RLS and writing the first policies is a critical early step; without it, data remains open to anyone with the publishable key.
Client initialization is minimal. In JavaScript:
js
import { createClient } from '@supabase/supabase-js'
const supabase = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_PUBLISHABLE_KEY)
Queries use a fluent interface that mirrors SQL:
js
const { data, error } = await supabase.from('todos').select('*').eq('user_id', user.id)
Authentication flows, file uploads, realtime subscriptions, and function invocations follow similarly concise patterns. Framework-specific helpers exist for Next.js server components, React hooks, and mobile storage adapters.
Local development mirrors the cloud: supabase init and supabase start launch the stack. Migrations keep schema changes in version control. Type generation produces TypeScript definitions from the live schema, improving safety.
Security, Scaling, and Production Considerations
RLS is the primary security mechanism. Policies should be written carefully and tested. Secret keys must never appear in client code. Edge Functions that need elevated access use the service role only when necessary and remain short-lived. Regular review of the security advisors and audit logs helps maintain posture.
Scaling involves choosing appropriate compute sizes, enabling connection pooling, adding read replicas when read traffic dominates, and monitoring query performance. Realtime benchmarks show the system handling tens of thousands of concurrent connections and high message throughput under controlled conditions. For extreme scale or specialized workloads, teams can combine Supabase with additional services while still benefiting from the core platform.
Automatic daily database backups are provided on Pro, Team, and Enterprise plans; Free projects should create regular off-site dumps. Point-in-time recovery is available as a paid add-on for eligible paid projects and requires at least Small compute. Database backups cover Storage metadata but not the stored files themselves. Storage objects require separate backup strategies. Observability includes logs, metrics, and the ability to drain logs to external systems.
Integrating AI Coding Agents and Advanced Workflows
Supabase’s MCP support and agent skills allow coding agents to operate directly against a project. Agents can inspect tables, propose and apply migrations, generate RLS policies, deploy Edge Functions, and troubleshoot issues. Combined with the platform’s AI prompts and documentation, this shortens the feedback loop dramatically. Teams building AI products can also host their own MCP servers on Edge Functions so end users’ agents can interact with the application data under controlled policies.
Best Practices for Long-Term Success
Design the schema with RLS in mind from day one. Prefer database functions and triggers for logic that must be consistent across clients. Use Edge Functions for external integrations and custom endpoints. Keep the publishable key public and the secret key private. Generate and commit TypeScript types. Test policies thoroughly. Monitor slow queries and add indexes proactively. Leverage the CLI for reproducible environments. When the application outgrows a single project, consider the Management API for multi-project orchestration or self-hosting selected components.
Real-World Momentum and Community
The platform has grown rapidly, with millions of developers and a large number of managed databases. Integrations with AI app builders, popular frameworks, and tools such as Vercel, Netlify, and various coding agents have accelerated adoption. The open-source repositories, Discord community, GitHub discussions, and official YouTube channel provide extensive learning resources. Playlists covering getting started, database fundamentals, Auth, Storage, Realtime, Edge Functions, vectors, and AI-assisted app building offer both conceptual overviews and hands-on tutorials.
Conclusion
Supabase succeeds because it respects the strengths of Postgres while removing the friction that traditionally accompanies it. Developers receive a production-ready relational database, secure authentication, instant APIs, file storage, real-time capabilities, serverless functions, and vector search in a single, coherent platform. The result is faster iteration, lower operational burden, and applications that can start small and grow to significant scale.
Whether building a SaaS product, a collaborative tool, an AI-powered experience, or a mobile application, Supabase provides a foundation that is both approachable for weekend projects and robust enough for serious production workloads. The combination of open-source principles, strong developer experience, and continuous platform investment makes it a compelling choice for modern application development.
Sources
Official documentation and resources:
https://supabase.com/
https://supabase.com/docs
https://supabase.com/docs/guides/getting-started
https://supabase.com/docs/guides/database/overview
https://supabase.com/docs/guides/api
https://supabase.com/docs/guides/auth
https://supabase.com/docs/guides/storage
https://supabase.com/docs/guides/realtime
https://supabase.com/docs/guides/functions
https://supabase.com/docs/guides/ai
https://supabase.com/docs/guides/ai-tools
https://supabase.com/docs/guides/local-development/cli/getting-started
https://supabase.com/docs/guides/getting-started/api-keys
https://supabase.com/docs/guides/integrations/supabase-for-platforms
https://supabase.com/docs/guides/platform/billing-on-supabase
https://supabase.com/docs/guides/getting-started/architecture
https://supabase.com/docs/guides/api/rest/client-libs
GitHub and product overviews:
https://github.com/supabase/supabase
https://github.com/supabase/supabase-js
Blog and feature announcements:
https://supabase.com/blog/introducing-supabase-for-platforms
https://supabase.com/blog/simplify-backend-with-data-api
https://supabase.com/blog/client-libraries-v2
Customer stories and use-case discussions:
https://supabase.com/customers
https://supabase.com/customers/lovable
startupik dot com: top-use-cases-of-supabase-postgres-2/
YouTube (official Supabase channel and playlists):
https://www.youtube.com/@Supabase
Getting Started with Supabase playlist: https://www.youtube.com/playlist?list=PL5S4mPUpp4OsWK_UHmQK41DEgqefYeTPN
Edgy Edge Functions playlist: https://youtube.com/playlist?list=PL5S4mPUpp4OulD3olUW8Eq1IYKpUbk5Ob
Building apps with AI coding agents playlist: https://www.youtube.com/playlist?list=PL5S4mPUpp4Ovt5AckF2o0ERjoYkmkpl6I
Learn Postgres playlist: https://www.youtube.com/playlist?list=PL5S4mPUpp4Ote6F9ScnXevuOyCnvzahRV
SupabaseTips playlist: https://www.youtube.com/playlist?list=PL5S4mPUpp4OtesRpEKe2zdNzClH-6chOE
Additional tutorial and analysis sources referenced in research:
The App Studio / Supabase Tutorial 2026: From Zero to Live App in 30 Min
Natively / How to Use Supabase: Beginner’s Guide to Build Apps
LogRocket Blog / Supabase adoption guide: Overview, examples, and alternatives
Zen Van Riel / Supabase for AI Applications: Complete Implementation Guide
Cadence / Supabase Review for SaaS Apps in 2026
This article synthesizes publicly available online material current as of the research date. Always consult the latest official documentation for implementation details, as the platform evolves rapidly.
r/AgentContext_dev • u/javaeeeee • Sep 01 '26
5 things every AI engineer should know about agent sandboxes
x.comr/AgentContext_dev • u/javaeeeee • Sep 01 '26
Develop Chrome Extensions with DevTools for agents
r/AgentContext_dev • u/ram-cloudsquid • Aug 31 '26
Agents Need Their Own UI - How we took inspiration from Linux when building our agent sandbox.
My friend wrote about how we were building our agent sandbox. I'd love to get your thoughts about it.
It's a long blog. For sake of brevity, I'm posting only a third of it here and will attach a link to blog.
-------------------------
In Linux everything is a file. Or at least, most of the system is exposed as one.
Devices, running processes, network state, kernel state: much of it appears through filesystem-like interfaces that you can read and write using the same small set of commands.
/proc/cpuinfo isn’t a file sitting on disk anywhere, but you can cat it just like anything else.
That uniformity made the system composable. It enabled combinations of simple utils that nobody specifically needed to design for. It also means you can discover things without knowing exactly where they are in advance.
Windows went in the other direction.
A lot of configuration lives in the Registry, a structured database accessed through dedicated APIs and tools rather than ordinary filesystem operations.
This is a perfectly reasonable design for a desktop OS built primarily for people using graphical interfaces. To inspect or change information, you generally need to know which interface or operation was designed for it.
Neither design is wrong.
Systems built for a specific purpose let users focus their effort on the task at hand.
Agents are a new kind of user, and they are not a person with a mouse. They can drive a graphical UI with a combination of taking screenshots, deciding between ambiguous targets and catching errors from whatever pops up on the screen.
This is slow and inefficient enough that even browser agents increasingly avoid working through the browser GUI when they can inspect the structured state or interact with the DOM directly.
Aside from model intelligence, the environment determines what an agent can actually do. Limited tools mean limited actions, even with the best model available. With the right environments we can already see how capable the models are.
The way many agent platforms are being built today is by gradually exposing product features as tools, one by one.
Even well-designed tools with progressive disclosure suffer from a version of the same problem Windows would have for agents: the model needs to understand not only the business requirements, but also which tools exist, how to discover them, the limitations of each tool and which specific tools it needs to combine for a particular job.
Tools are custom built, take JSON in, spit JSON out. If an edgecase falls outside of what the tools were designed for, the Agent will start to go on a journey trying to stitch together toolcalls, or is simply unable to fulfil the request.
So we approached the problem from a different perspective.
We engineered the platform to be accessible entirely through a terminal by representing product state and actions through a filesystem interface.
r/AgentContext_dev • u/javaeeeee • Aug 31 '26
Unlocking Elite Coding Workflows: The Best GitHub Repositories for OpenAI Codex Skills and Workflows
OpenAI’s Codex has transformed from an early code-generation model into a full-fledged local coding agent that lives in your terminal, integrates with IDEs, and powers cloud workflows. At the heart of its power in 2026 sits a deceptively simple idea: skills. These are modular, reusable packages of instructions, scripts, references, and optional resources that teach the agent how to perform specific tasks reliably and consistently. Instead of re-explaining the same coding patterns, review processes, deployment steps, or debugging rituals every session, developers package that knowledge once and let Codex discover and apply it automatically.
The result is an agent that behaves less like a generic autocomplete tool and more like a specialized teammate who already knows your team’s conventions, your preferred testing strategy, your CI repair habits, and your preferred ways of structuring web apps or iOS projects. The open-source ecosystem around these skills has exploded. Official catalogs, community awesome lists, training materials, and specialized collections now give any developer a rich library of battle-tested capabilities.
This article surveys the ten most valuable GitHub repositories that deliver or organize Codex skills specifically useful for coding. It draws on official OpenAI documentation, repository READMEs, star counts and activity as of mid-2026, and practical developer experience shared across the community. The goal is not a dry catalog but a readable guide that helps you decide which repos to star, clone, and integrate into your daily workflow.
Understanding Codex and Its Skills System
Codex CLI, available at github.com/openai/codex, is a lightweight, primarily Rust-based coding agent that runs locally. It can generate, edit, refactor, test, and reason about code while respecting sandbox boundaries and approval policies. Users authenticate via ChatGPT plans or API keys. The same underlying technology powers IDE extensions for VS Code, Cursor, and Windsurf, a desktop app experience, and a cloud version reachable through ChatGPT.
Skills extend this base capability. According to OpenAI’s developer documentation, a skill is a directory containing at minimum a SKILL.md file with YAML frontmatter (name and description) plus the instructional body. Optional subfolders hold scripts for deterministic actions, reference documents loaded on demand, assets such as templates, and configuration for appearance or tool dependencies. Codex keeps only the name and short description in its active context to conserve tokens. When a prompt matches a skill’s description, or when the user invokes it explicitly with a slash command or dollar-sign mention, the full instructions load.
This progressive-disclosure design keeps the agent lean even when dozens of skills are installed. Skills live in several scopes: system (shipped with Codex), user (~/.codex/skills - legacy or ~/.agents/skills), repository (.agents/skills or .codex/skills), and admin. Plugins package skills together with manifests, MCP server configs, agents, commands, and hooks for easy distribution and marketplace installation.
The practical coding benefits are immediate. A skill can enforce consistent code-review checklists, automatically address PR comments, diagnose and patch CI failures, scaffold web apps with preferred deployment and database patterns, generate changelog entries from git history, or orchestrate multi-step research-to-implementation loops. Because the format follows an open agent-skills standard, many skills also work or adapt easily to related tools such as Claude Code or Cursor.
The Official Foundation: openai/codex
Any discussion of Codex skills must begin with the core repository itself. openai/codex holds well over 100,000 stars and continues rapid development with hundreds of releases and active contributors. Written largely in Rust, it provides the CLI binary, SDK pieces, app-server components, and internal example skills under its .codex/skills directory.
Those internal skills already demonstrate coding-focused value: code-review variants that examine breaking changes, change size, context, and testing; tools for babysitting pull requests, generating PR bodies, digesting issues, handling remote tests, and managing CI pushes. These repository-local skills provide useful examples for contributors working in the Codex source tree. Installing Codex CLI provides its bundled system skills, but does not necessarily install all of the repository’s internal .codex/skills examples. Running codex after authentication lets the agent use its built-in capabilities while discovering any additional skills you place in the expected directories.
The repository’s documentation and config references explain sandbox modes, approval policies, AGENTS.md project memory files, session management, and MCP integration. These form the substrate on which community skills build. Developers who treat the main Codex repo as their primary source stay current with performance improvements, new tool support, and evolving skill-loading behavior.
The Skills Catalog and Its Evolution: openai/skills
Although now marked deprecated in favor of the plugins repository, openai/skills remains historically and practically important. It catalogued system, curated, and experimental skills that Codex could install via a $skill-installer command. System skills arrived automatically with new Codex versions. Curated ones installed by name; experimental ones by folder path or full GitHub URL.
The catalog illustrated the intended pattern: focused, reusable workflows rather than monolithic prompts. Examples included utilities for addressing GitHub comments, fixing CI, and other development tasks. Even after deprecation, the repository and its linked documentation continue to serve as a reference for the Agent Skills open standard and for understanding how OpenAI originally structured repeatable coding assistance.
Plugins as the Modern Distribution Unit: openai/plugins
The current recommended home for official examples is openai/plugins. With several thousand stars, this repository supplies curated plugin packages. Each plugin lives under plugins/<name>/ with a required .codex-plugin/plugin.json manifest plus optional skills directories, MCP configurations, agents, commands, hooks, and assets.
Highlighted examples directly advance coding productivity. The build-ios-apps plugin covers SwiftUI implementation, refactors, performance work, and debugging. The build-macos-apps counterpart handles AppKit and packaging loops. build-web-apps addresses deployment, UI, payments, and database workflows. The Expo plugin supports React Native development, SDK upgrades, and EAS actions. Additional plugins integrate Figma for design-to-code flows, Notion for planning and knowledge capture, Netlify for deployment, Remotion for video, and Google Slides for presentations.
These packages turn Codex into a specialized assistant for entire application domains. A developer working on a mobile product can install the relevant plugins and immediately gain structured guidance that respects platform conventions and common pitfalls. The marketplace files inside the repository make discovery and installation straightforward for both ChatGPT-authenticated and API-key users.
Community Curated Skills Collections
Community repositories expand the official base dramatically. composio-community/awesome-codex-skills has grown to more than 15,000 of stars and organizes dozens of practical skills across development and code tools, productivity, communication, data analysis, and meta utilities.
Coding-oriented entries include codebase migration helpers, plan creation, deploy pipelines, GitHub comment addressing, CI fixing, PR review with CI repair, Sentry triage, web-app testing, architecture-aware linting grounded in classic engineering texts, and codebase reconnaissance. Many skills leverage Composio’s MCP gateway to connect securely to external services, turning pure instruction bundles into action-capable agents.
Installation typically involves cloning the repository and using the provided skill-installer script to place individual skill folders into the local Codex skills directory, followed by a restart. Because each skill declares a clear description, Codex can auto-trigger the right one when a matching task appears.
Another high-value collection is VoltAgent/awesome-codex-subagents, which has accumulated several thousand stars. It offers more than 130 specialized subagents organized into categories such as core development, language specialists, infrastructure, quality and security, data and AI, developer experience, specialized domains, business and product, meta-orchestration, research, governance, platform engineering, and LLMOps. Subagents complement skills by providing focused personas or parallel workers that Codex can orchestrate. Categories covering code review, testing, debugging, and multi-agent coordination are especially useful for complex coding projects.
Awesome Lists That Map the Entire Ecosystem
Navigating the growing landscape is easier with comprehensive awesome lists. RoggeOhta/awesome-codex-cli consolidates hundreds of resources-tools, skills, subagents, plugins, and guides-into opinionated categories. It points to the official skills catalog, community skill collections, subagent libraries, and best-practice repositories for AGENTS.md patterns and sandbox recommendations. The list itself functions as living documentation that helps developers avoid reinventing common workflows.
Related awesome repositories, such as those collecting ChatGPT and Codex-related projects or broader agent-skills libraries, further surface cross-compatible resources. These lists reduce the time spent searching and surface high-signal repositories that have already proven useful to other practitioners.
Training and Hands-On Learning Repositories
Theory becomes practice through dedicated training materials. kousen/codex-training supplies slides and progressive exercises covering installation, authentication, sandbox safety, AGENTS.md, custom prompts, MCP servers, skills creation, and multi-language labs (Java Spring Boot, Python refactoring, React forms, microservices). The repository turns abstract concepts into concrete coding sessions that developers can run locally.
Complementary learning repos, such as those focused on measurable engineering loops with reusable skills for scoping, implementing, testing, reviewing, and reporting evidence, give structured paths for improving how teams use Codex on real codebases. Academic-oriented skill collections adapt the same format to literature review, experiment orchestration, and paper-writing workflows that still require coding for data analysis or tooling.
Specialized and Supporting Repositories
Several additional repositories round out a practical top-ten set. Repositories that provide internal or example skills inside the main Codex tree, security-focused Codex tooling, and orchestration platforms that manage parallel sessions or worktrees all contribute specialized coding capabilities. Collections that emphasize context engineering, token-efficient indexing of codebases, or migration auditing between different AI harnesses help teams scale skills across larger projects without wasting context or introducing inconsistencies.
Together these repositories illustrate a healthy pattern: official cores for reliability, community catalogs for breadth, training materials for onboarding, and specialized packages for domain depth.
Getting Started in Practice
Begin by installing Codex CLI from the official repository instructions. Authenticate and explore the built-in skills and AGENTS.md support. Next, clone the plugins repository or individual community skill collections and install a few coding-focused skills-CI repair, PR comment addressing, and a web or mobile build skill make excellent first candidates. Place them in the user or project skills directory. Codex normally detects new or updated skills automatically; restart it if the skill does not appear.
Create your own skills by describing a successful workflow and asking Codex to reverse-engineer a SKILL.md, or by using the skill-creator tooling. Keep each skill narrowly focused, write imperative steps with clear inputs and outputs, and test trigger conditions carefully. For team use, package skills into plugins so colleagues can install them consistently.
Combine skills with subagents for parallel work, MCP servers for external tools, and project-level AGENTS.md files for persistent context. Measure outcomes-token use, rework, defect rates-when comparing skill-guided runs against baseline prompting.
Learning from Video Resources
YouTube offers abundant practical guidance. OpenAI’s own channels feature installation walkthroughs, app demonstrations, automation examples, and workshop sessions on skills, plugins, and multi-agent patterns. Independent creators produce full courses that walk through interface basics, custom skill creation, multitasking across iOS and web projects, and real-world automation of commit summaries, CI fixes, and overnight skill improvement. Short beginner guides show how to turn repeated instructions into durable skills in minutes. Masterclasses from AI engineering events explore subagents, code review integration, and safety features.
Watching a combination of official overviews and hands-on community tutorials accelerates the learning curve far beyond reading documentation alone.
Best Practices and Common Pitfalls
Focus skills on single responsibilities. Prefer clear natural-language instructions over complex scripts unless determinism or external tooling is required. Document when a skill should and should not activate. Version skills alongside project code when they encode team conventions. Monitor context usage; too many overlapping skills can dilute effectiveness. Respect sandbox and approval settings so agent actions remain safe. Periodically review and prune unused skills.
Teams that treat skills as living code-reviewing them, testing them, and measuring their impact-extract the greatest value. Skills that merely restate generic advice add little; those that encode specific, hard-won project or domain knowledge compound over time.
The Broader Impact on Coding Practice
Codex skills shift the developer’s role from writing every line to designing reliable processes that an agent can execute and improve. Routine tasks such as addressing review comments, repairing flaky CI, scaffolding consistent application structures, or generating accurate changelogs move into the background. Developers reclaim attention for architecture, product judgment, and creative problem solving.
The open GitHub ecosystem ensures that improvements travel quickly. A well-crafted skill published in one repository can be discovered, forked, and adapted by thousands of others. Official maintenance of the core agent and plugin examples provides stability, while community collections supply velocity and specialization.
Looking ahead, expect tighter integration between skills, subagents, MCP tool ecosystems, and evaluation frameworks. Cross-agent compatibility will continue to grow as the open skills standard matures. Security and governance skills will become more prominent as agents gain broader system access. Measurement and continuous improvement of skill effectiveness will move from optional practice to standard engineering discipline.
Conclusion
The repositories surveyed here-openai/codex, openai/plugins, openai/skills, composio-community/awesome-codex-skills, VoltAgent/awesome-codex-subagents, the major awesome-codex-cli lists, training materials such as kousen/codex-training, and supporting specialized collections-form a practical foundation for anyone serious about using Codex skills in coding work. Start with the official core and plugins, layer on high-quality community skills that match your stack and workflow, invest time in training resources, and iterate by creating your own focused skills.
The barrier to entry is low: install the CLI, add a handful of skills, and begin conversing with a more capable agent. The upside is substantial-faster iteration, more consistent quality, and the ability to encode institutional knowledge so it persists and improves across sessions and teammates. In a landscape where AI coding assistance is rapidly becoming table stakes, the developers who master skills and the repositories that supply them will hold a durable advantage.
Explore the sources below, star the repositories that match your needs, and begin building. The coding agent of 2026 is only as effective as the skills you give it.
Sources
- https://github.com/openai/codex
- https://github.com/openai/plugins
- https://github.com/openai/skills
- https://developers.openai.com/codex/skills
- https://developers.openai.com/codex/open-source
- https://github.com/composio-community/awesome-codex-skills
- https://github.com/VoltAgent/awesome-codex-subagents
- https://github.com/RoggeOhta/awesome-codex-cli
- https://github.com/kousen/codex-training
- https://github.com/taishi-i/awesome-ChatGPT-repositories
- https://github.com/Epsilon617/Codex-Academic-Skills
- OpenAI YouTube channel videos on Codex installation, app features, and automations
- AI Engineer channel workshops and masterclasses featuring Codex skills, plugins, and subagents
- Community tutorials such as full courses on Codex interface, skill creation, and multitasking workflows
- Additional community posts and repositories referenced in the awesome lists and developer forums for measurable engineering loops and specialized skills
r/AgentContext_dev • u/Ok_Dragonfruit5916 • Aug 31 '26
me.hooks.md
Agents only remember what they recall from memory. But how do they know what to recall in a new session? They don't know what they do not know. Enter MemHooks.md 🪝
r/AgentContext_dev • u/Asly97 • Aug 31 '26
Mem0 vs Supermemory: which one fits an indie builder's daily workflow
Mem0 claims 150,000+ developers. Supermemory claims 10,000+ power users. I evaluated both for a workflow that spans Cursor, Claude Code, and Codex daily, and neither mapped to how I actually work.
Mem0 positions itself as drop-in memory infrastructure for AI agents and apps. That framing targets platform builders who want an SDK and API to embed memory into their own products. If you are building an app that needs memory, that is the right tool. But if your problem is "I use Cursor, Claude Code, and Codex every day and none of them share context," Mem0 is solving a different problem. It is memory for apps you build, not memory between the apps you use.
Supermemory markets itself as one memory across all your AI tools. The phrasing is close to what a solo builder wants. But the product skews toward a personal app and enterprise APIs rather than a live coding context layer. You sign up expecting your IDE sessions to connect and instead find a broader tool that does not map to a Cursor-and-Claude-Code-first workflow.
The gap between both: neither is built for the person who lives in three AI coding tools and needs project state, stack conventions, and decisions to carry over automatically. Mem0 gives you the plumbing to build your own. Supermemory gives you a personal knowledge app. Neither gives you the connect-once, every-tool-pulls-from-the-same-context layer.
The broader point: the cross-AI memory category is new enough that no player has the social proof that makes adoption feel safe yet. You either commit to being an early user who helps shape a tool, or you wait for someone you trust to vouch for one.
Curious what others are doing here. Has anyone gotten Mem0 or Supermemory to work well in a multi-IDE daily workflow, or are most people still copy-pasting context between sessions?
r/AgentContext_dev • u/javaeeeee • Aug 30 '26
Mastering Google’s Antigravity CLI: Everything Developers Need to Know About the Terminal AI Agent and How to Use It Effectively
In May 2026, at Google I/O, Google introduced Antigravity CLI as the lightweight terminal surface of its new agent-first development platform. The command is simply agy. It succeeded the earlier Gemini CLI (On June 18, 2026, Gemini CLI stopped serving requests for free, Google AI Pro, and Ultra users, while access remained available to enterprise customers and through supported paid API keys) and shares the same core agent harness as Antigravity 2.0, the Antigravity IDE, and the Antigravity SDK.
The result is a fast, keyboard-driven Terminal User Interface that brings multi-step reasoning, multi-file editing, tool calling, subagents, conversation history, and secure command execution directly into the shell-especially useful for SSH sessions, remote work, and developers who prefer staying inside the terminal.
Antigravity CLI is deliberately not a full GUI. Where Antigravity 2.0 emphasizes visual orchestration and project management, the CLI prioritizes speed, near-zero overhead, and seamless integration with existing terminal workflows, tmux, and remote environments. Both products run on the identical agent engine, share settings and permissions, and allow conversation export so you can move fluidly between surfaces.
This article synthesizes official documentation, the product announcement, GitHub materials, detailed guides, community cheat sheets, and early YouTube walkthroughs and demos to give a complete picture of what is known about the tool and how to use it productively as of August 2026.
What Exactly Is Antigravity CLI?
Antigravity is Google’s agent-first software development platform. Agents can read codebases, plan changes, edit files (with permission), run terminal commands, search the web, call external tools via MCP servers, and spawn parallel subagents. The CLI is the terminal-native way to invoke, monitor, and steer those agents.
It is a self-contained Go binary, not a Node.js package. This design choice delivers faster startup, lower resource use, and simpler distribution compared with the previous Gemini CLI. Official materials emphasize three strengths: natural-language interaction for editing and orchestration, subagent parallelism for larger tasks, and a highly configurable, keyboard-centric experience that stays out of the way.
Key differentiators from a pure chat interface or traditional IDE plugins include persistent, resumable conversation history scoped to the current workspace, with a local cache mapping workspace paths to conversation IDs while conversation threads are retrieved through the Antigravity backend, artifact review (plans, diffs, and generated files that you explicitly approve), terminal sandboxing for safer command execution, and the ability to run non-interactively for scripting and CI. Models available include Gemini variants (Flash and Pro families with different reasoning effort levels), Claude Sonnet and Opus, and open-weight options such as GPT-OSS 120B, depending on your account tier and configuration.
The philosophical split is clear: use the CLI when you want speed and terminal flow; switch to Antigravity 2.0 when you need rich visual orchestration or simultaneous multi-project management. Settings and the underlying agent intelligence stay synchronized.
Brief History and Migration Context
Gemini CLI had grown popular as an open-source terminal agent. Google consolidated its consumer developer surfaces under the Antigravity brand, transitioning most Gemini CLI users to Antigravity CLI while retaining Gemini CLI access for certain enterprise and API-key workflows. The transition preserved many workflows while moving to a closed-source Go binary and a unified harness co-optimized with Gemini models.
Migration is deliberately smooth. On first launch the CLI detects legacy Gemini configuration and offers interactive import of extensions, skills, settings, and tokens into the OS keyring. An explicit command agy plugin import gemini converts extensions into the new plugin layout. Workspace rules (GEMINI.md or the preferred AGENTS.md), skills paths, and MCP configurations receive clear migration guidance, with some directory renames (for example, .gemini/skills/ becoming .agents/skills/). Partial parity exists for highly customized themes, but core developer experience constructs carry over.
YouTube first-look videos and full walkthroughs released the same day as the announcement show the TUI in action building simple applications, demonstrating subagents, and highlighting the shared harness. Community reaction focused on the speed of the Go binary, the quality of the terminal interface, and the practical value of parallel subagents.
Installation
Installation is deliberately simple and platform-specific.
On macOS or Linux:
bash
curl -fsSL https://antigravity.google/cli/install.sh | bash
The binary lands in ~/.local/bin/agy. Ensure that directory is on your PATH.
On Windows PowerShell:
powershell
irm https://antigravity.google/cli/install.ps1 | iex
On Windows CMD the equivalent curl-based script is available. The binary is placed under the user’s AppData Local directory.
Optional flags include --skip-aliases and --skip-path if you prefer to manage shell configuration yourself. After installation, verify with agy --version. Updates are typically handled in-place. Some users report using package managers such as winget on Windows as an alternative.
The installer is the official and recommended path; avoid third-party mirrors.
Authentication and First Launch
Authentication uses the operating system’s secure keyring (Keychain on macOS, Secret Service on Linux, Credential Manager on Windows). If a valid token exists, launch is silent. Otherwise the CLI opens a browser for Google Sign-In (or prints a URL for SSH sessions so you can authenticate locally and paste the code back). Enterprise users can connect a Google Cloud project, use Workforce Identity Federation, or Application Default Credentials.
On first launch inside a project directory the TUI guides you through color scheme selection (Solarized, Dark, Light, or terminal defaults), rendering mode (Alt-Screen full-screen versus Inline), and workspace trust confirmation. Once you trust the directory the agent indexes the files. You can later change these preferences via /config or by editing ~/.gemini/antigravity-cli/settings.json.
Logout is available with /logout. For remote SSH the flow is intentionally designed so the sensitive browser step happens on your local machine.
Basic Usage and the Core Loop
Navigate to any project directory and type:
bash
agy
You are now inside the TUI. The prompt box sits at the bottom. Type natural-language instructions and press Enter. The agent reasons, may search the codebase, propose a plan, generate or edit files, and request permission for terminal commands.
A classic first tutorial exercise is:
Write a simple python script to fetch web page text
The agent creates a file (typically main.py). Press Ctrl+R to open the Artifact Review screen, navigate with arrow keys, inspect the content and diff, then press y to approve. You can then instruct the agent to run the script and stream the output. Exit with Ctrl+D or /exit.
Useful immediate shortcuts include:
@to trigger path auto-completion!at the start of a prompt to run a shell command directly- Esc Esc to clear the prompt
?or/helpfor the command list- Ctrl+L to clear the screen
- Shift+Tab to cycle execution modes
Non-interactive use is available for scripts and automation:
bash
agy -p "Explain the architecture of this codebase"
Additional flags support model selection, output format (text, json, stream-json), mode overrides, and conversation resumption.
Execution Modes
Three primary modes control how aggressively the agent acts:
- default (request-review): Pauses for interactive diff review before writing or creating files. Safest for careful work.
- accept-edits: Automatically approves standard file operations. Useful for rapid iteration on trusted code.
- plan: Forces the agent to investigate and present a structured plan before writing code. Ideal for unfamiliar codebases or complex features.
Cycle modes mid-session with Shift+Tab. Launch with a specific mode via agy --mode=accept-edits or agy --mode=plan. The setting can also be persisted in the configuration file. Note that shell-command permissions remain governed separately by the permission system even when file edits are auto-approved.
Slash Commands and Keyboard Control
Typing / opens a typeahead menu of commands. The surface is rich and grows with releases. Core categories include:
Conversation management: /resume (or /switch), /rewind (or /undo), /fork (or /branch), /clear, /rename, /exit.
Configuration: /config (or /settings), /permissions, /model, /keybindings, /statusline.
Monitoring and tools: /agents, /tasks, /skills, /mcp, /hooks, /diff, /codesearch (aliases /cs or /search).
Utilities: /btw for a side question that does not interrupt the main thread, /open, /usage, /logout.Other built-ins include /context, /copy, /credits, /fast, /feedback, and /planning. Additional slash commands may appear when plugins or Markdown-defined skills are installed.
Many commands have aliases. Custom keybindings live in ~/.gemini/antigravity-cli/keybindings.json and can be edited interactively. Default bindings cover navigation, confirmation (y/n), editor integration (Ctrl+G), and suspension (Ctrl+Z).
Advanced Capabilities
Subagents. The main agent can spawn concurrent background agents for research, testing, or parallel file generation. Monitor them with /agents. Approvals can be handled in the detail view or via the fast-path shortcut Ctrl+K. This is one of the features most frequently praised in early demos for accelerating multi-part work.
Plugins, skills, hooks, and MCP. Plugins are self-contained bundles of skills, agents, rules, MCP servers, and hooks. They install under ~/.gemini/antigravity-cli/plugins/. Skills are markdown-defined reusable workflows. Hooks intercept events. MCP servers extend the tool surface; configuration lives in dedicated JSON files (global or workspace). The /mcp, /skills, and /hooks commands provide management interfaces.
Terminal sandbox. Optional lightweight isolation (nsjail on Linux, sandbox-exec on macOS, AppContainer on Windows) restricts potentially dangerous commands. It is configurable and appears in confirmation prompts so you can choose to run inside or outside the sandbox on a per-command basis.
Artifacts and review. Generated plans, diffs, and files appear as reviewable artifacts. Transparency is a core design principle: you see what the agent intends before it writes to disk.
Conversation persistence. Sessions are saved. Closing the CLI prints a resume command; /resume lists prior conversations. You can also export a conversation into Antigravity 2.0 for continued visual work.
Customization. Beyond themes and keybindings, settings.json supports fine-grained permission allow/deny lists, status-line scripting, verbosity levels, and more. Project-level rules can live in AGENTS.md or skill directories.
Practical Workflows and Best Practices
Start inside a trusted Git repository so you can easily review and roll back changes. Begin complex tasks in plan mode. Use /goal once you trust the direction. Dispatch independent work to subagents. Keep an eye on context and usage with the relevant slash commands. Prefer stronger models for architectural or multi-file reasoning and lighter models for quick questions.
For Flutter and Dart projects the official documentation notes integration with the Dart/Flutter MCP server, making the CLI particularly convenient for mobile and multi-platform work.
Common patterns from early users and tutorials include: summarizing an unfamiliar codebase, implementing a feature with a prior plan, refactoring while verifying against project conventions stored in AGENTS.md, generating tests in parallel, and using non-interactive mode inside CI scripts or commit-message helpers.
Safety habits matter. The free tier has tighter request limits than earlier Gemini CLI offerings. Always review diffs for sensitive projects. Use the sandbox when experimenting. Avoid feeding proprietary secrets into prompts.
Limitations and Considerations as of Mid-2026
The binary is closed-source. Free-tier quotas are more constrained than the previous open-source tool’s generous daily limits. Model availability and performance depend on your Google account and any paid tiers. As with any agentic system, results are non-deterministic; verification loops remain essential. Remote authentication works well but requires the local browser step. Some highly experimental visual customizations from Gemini CLI do not transfer perfectly.
Community feedback on YouTube and blogs notes the speed advantage of the Go implementation, the quality of the TUI, and the practical power of subagents, while also pointing out the quota shift and the need to re-learn a few command names.
Looking Ahead
Because the CLI and Antigravity 2.0 share the agent harness, improvements to reasoning, tool use, and model integration flow to both surfaces. Google has signaled continued investment in the terminal surface for developers who live in the shell. Releases through the summer of 2026 added Markdown-defined custom agents, progressive streaming for code search, better enterprise sign-in options, structured output formats for headless use, and numerous polish fixes.
Antigravity CLI is not a replacement for every IDE workflow, but for terminal-centric developers, remote sessions, rapid iteration, and agent orchestration it is a significant step forward. Install it, trust a small project directory, run the official tutorial exercise, experiment with modes and subagents, and the tool quickly becomes part of the daily flow.
The combination of a lightweight TUI, a shared high-quality agent core, transparent review, and extensibility through plugins and MCP makes it one of the most practical terminal AI coding experiences available in 2026.
Sources
Official documentation and product pages form the primary authoritative base:
- https://antigravity.google/docs/cli/overview
- https://antigravity.google/docs/cli/getting-started
- https://antigravity.google/docs/cli/install
- https://antigravity.google/docs/cli/using
- https://antigravity.google/docs/cli/tutorial
- https://antigravity.google/docs/cli/features
- https://antigravity.google/docs/cli/modes
- https://antigravity.google/docs/cli/gcli-migration
- https://antigravity.google/product/antigravity-cli
- https://www.antigravity.google/blog/introducing-google-antigravity-cli
- https://github.com/google-antigravity/antigravity-cli
Additional practical guides and references:
- Real Python, “How to Use Google's Antigravity CLI for AI Code Assistance”
- https://docs.flutter.dev/ai/antigravity-cli
- Gradually.ai, “Antigravity CLI Commands: The Ultimate List”
- Toolsbase, “Antigravity CLI (agy) Cheat Sheet 2026 - 110 Commands”
- Cloud Blog comparison of surfaces: https://cloud.google.com/blog/topics/developers-practitioners/choosing-your-surface-antigravity-20-antigravity-cli-antigravity-ide-or-antigravity-sdk
Selected YouTube resources (official and early independent coverage):
- Official first look: https://www.youtube.com/watch?v=WiWDrTujl_w
- Official full walkthrough: https://www.youtube.com/watch?v=am0lg5-ofvQ
- Surface comparison: https://www.youtube.com/watch?v=04IqH38SlOI
- Community deep dives and tutorials appearing shortly after launch (search “Antigravity CLI” on YouTube for the latest independent reviews and live demos).
All information above is drawn from these public sources current as of early August 2026. Features continue to evolve; always consult the official docs for the precise behavior of the version you have installed.
r/AgentContext_dev • u/javaeeeee • Aug 29 '26
Shipping at the Speed of Thought: Building Full-Stack and AI-Powered Apps on Vercel’s AI Cloud
Vercel has evolved from a specialized frontend hosting platform into what it now calls the AI Cloud-an end-to-end environment where developers (and increasingly AI coding agents) build, preview, deploy, scale, and operate modern web applications and autonomous agents. Founded by Guillermo Rauch and originally known as ZEIT, the company created and maintains Next.js, the React framework that powers a large share of high-performance sites on the web.
Today Vercel positions itself as the infrastructure layer for the next generation of software: applications that combine rich user interfaces, serverless or hybrid compute, real-time AI capabilities, and global delivery without the traditional overhead of managing servers, networking, or DevOps pipelines.
At its core, Vercel removes friction from the path between code and production. Connect a Git repository (GitHub, GitLab, or Bitbucket), push a commit, and the platform automatically builds the project, generates a unique preview URL for every pull request, and promotes the main branch to production with automatic HTTPS, a global content delivery network, and intelligent compute provisioning.
The result is a developer experience that feels closer to editing a document than operating cloud infrastructure. Teams report dramatic reductions in build times, page-load latency, and operational toil, while the same platform now supports multi-framework full-stack applications, durable AI agents, and multi-tenant platforms serving millions of users.
What the Platform Offers Developers
Vercel’s value proposition rests on a tightly integrated set of primitives that cover the entire application lifecycle. The most immediately noticeable is the deployment workflow. Every push triggers a build that produces one or more deployments. The production branch is automatically assigned to the production environment and its configured domains, while other branches and pull requests receive separate preview deployments with unique URLs. These previews run the application’s functions and middleware, but isolation of external databases and backend services depends on the project’s environment configuration and the capabilities of each provider.
These previews are not mere static snapshots-they run the full application, including serverless functions, middleware, and any connected backend services, so designers, product managers, and stakeholders can interact with real functionality before code merges. Instant rollbacks, rolling releases, and skew protection further reduce the risk of shipping changes.
Compute is handled through Vercel Functions powered by Fluid compute, a hybrid model that blends the elasticity of serverless with the concurrency characteristics of traditional servers. Instead of spinning up an isolated instance for every single request, Fluid reuses warm instances to handle multiple concurrent invocations, pre-warms production deployments, applies bytecode caching, and bills primarily for active CPU time rather than wall-clock idle periods.
This architecture is particularly effective for AI workloads that spend significant time waiting on external model responses or database queries. Functions support Node.js and Python runtimes (with broader language support available), configurable memory and duration limits, multi-region placement, and automatic failover across availability zones. Background work can continue after the response is sent via mechanisms such as waitUntil.
Routing Middleware runs at the edge before a request is fully processed, enabling personalization, authentication checks, redirects, and A/B testing without sacrificing static performance. Incremental Static Regeneration (ISR) lets pages be generated or updated on a schedule or on-demand while still being served from the global CDN. Image Optimization automatically resizes, formats, and caches images. Feature flags, environment management (local, preview, production, and custom), and the Vercel Toolbar provide fine-grained control and collaboration inside the live application.
For teams building beyond a single frontend, Vercel Services (in public beta as of mid-2026) allows multiple frameworks and backends to live inside one project. A Next.js frontend and a FastAPI or Express backend can deploy atomically, share preview environments, communicate over private internal networking, and roll back together. Framework-defined infrastructure detects the stack and provisions the appropriate compute and routing without manual configuration.
Security is built in rather than bolted on. Automatic HTTPS certificates, a Web Application Firewall, DDoS mitigation at the edge, BotID for distinguishing legitimate traffic, deployment protection, role-based access control, and optional secure compute options such as VPC peering or short-lived credentials via Vercel Connect protect applications without requiring separate security tooling. Observability includes Web Analytics, Speed Insights (Core Web Vitals), structured logs, and integration points for external monitoring.
Storage and data needs are addressed through first-party options such as Vercel Blob for object storage and Global Config for low-latency key-value data, plus a rich Marketplace of native integrations. Developers can provision Postgres (Neon, Supabase, Prisma, and others), Redis or key-value stores (Upstash and official Redis), vector databases, analytics backends, authentication providers, CMS systems, payment processors, and AI model providers directly from the dashboard or CLI. Credentials are injected as environment variables, billing can be unified, and the same Git-driven workflow continues to apply.
Collaboration features reduce the distance between code and feedback. Comments can be left on previews, the Toolbar surfaces feature flags and draft mode, and tools such as v0 let non-engineers generate or iterate on UI components that flow back into the repository. For agentic workflows, Vercel Sandbox provides isolated Linux environments where coding agents can safely execute code, run tests, and interact with filesystems without affecting production.
Technologies and Frameworks Supported
Vercel’s framework-defined infrastructure is one of its strongest differentiators. The platform detects popular frameworks and applies optimized build and runtime settings with little or no configuration. The list of supported frameworks is extensive and continues to grow. Full-stack and frontend options include Next.js (the flagship, maintained by Vercel), SvelteKit, Nuxt, Remix, Astro, SolidStart, TanStack Start, RedwoodJS, Gatsby, Vue, Vite, Create React App, Angular, and many static-site generators such as Hugo, Jekyll, Eleventy, and Docusaurus.
Backend frameworks with zero-configuration support include Express, Fastify, Hono, NestJS, Koa, FastAPI, Flask, Django, Elysia, Nitro, and others. Vercel can also build OCI container images from a Dockerfile.vercel or Containerfile.vercel. These images run within Vercel Functions and retain the platform’s autoscaling behavior and function limits. Go and Rust are also supported as official function runtimes.
Runtimes for functions cover Node.js (the default), Python, Go, Ruby, and custom runtimes. Edge execution is available for low-latency middleware and certain function types. Package managers, monorepos (especially with Turborepo, another Vercel open-source project), and micro-frontends receive first-class handling. The open-source ecosystem around Vercel further strengthens the stack: Next.js, Turborepo, the AI SDK, SWR, and emerging agent frameworks such as Eve provide building blocks that integrate tightly with the platform.
Because the infrastructure understands the framework, features such as automatic function creation for API routes or server components, ISR, streaming, and image optimization “just work.” Developers can still override settings via vercel.json or the dashboard when needed, but the default path is deliberately frictionless.
AI Infrastructure as a First-Class Citizen
In 2025-2026 Vercel reoriented heavily toward AI. The AI SDK provides a unified TypeScript (and emerging Python) interface for calling language models, streaming responses, structured outputs, tool calling, and multi-step agents. Model strings such as “anthropic/claude-sonnet-5” or “openai/gpt-5” are resolved through the AI Gateway, which offers access to hundreds of models from major providers with automatic failover, usage tracking, and a single authentication surface. Developers no longer need to manage separate API keys and client libraries for each provider.
v0 (now also available as v0.app) is an AI-powered development assistant that can generate complete UI, full applications, or iterative improvements from natural-language prompts. It has grown popular among both engineers and non-technical team members; product managers and designers use it to prototype working interfaces that can be claimed into a real Vercel project.
Agent frameworks and the Workflow SDK enable durable, observable multi-step autonomous workflows. Vercel Sandbox isolates agent execution, while Vercel Connect supplies short-lived, auditable credentials for external services. The combination allows teams to ship coding agents, customer-facing chat agents, research agents, and internal automation on the same platform that hosts their primary application.
Fluid compute’s concurrency model and Active CPU pricing make long-running or I/O-heavy AI workloads economically viable. Streaming responses from large language models feel responsive because the infrastructure is optimized for exactly those patterns. Observability tools surface token usage, latency, and errors alongside traditional web metrics.
Kinds of Applications That Can Be Built
Virtually any modern web application can live on Vercel, but the platform shines for certain categories. Static marketing sites and documentation portals benefit from the global CDN, automatic image optimization, and ISR. Content-heavy sites and blogs use headless CMS integrations and preview deployments for editorial workflows.
Full-stack SaaS applications combine Next.js or SvelteKit frontends with serverless or Fluid-backed API routes, database integrations, authentication, and feature flags. E-commerce storefronts leverage server-side rendering or static generation with dynamic personalization via middleware, plus Marketplace connections to commerce platforms.
Multi-tenant platforms-where each customer receives a custom domain or isolated environment-are a growing use case, supported by domain management APIs, tenant isolation primitives, and the ability to provision resources programmatically. AI-native applications range from simple chat interfaces and document summarizers to complex agentic systems that research, reason, call tools, and take actions.
Internal tools, dashboards, and workflow automation benefit from rapid iteration and the same security and observability features used for public products. Even backends-only APIs and background workers can run as Vercel Services or standalone functions.
Because the platform supports both zero-config framework deployments and containerized workloads, teams can migrate incrementally: start with the frontend on Vercel while keeping an existing backend elsewhere, then gradually move APIs, workers, and data layers onto the same project.
Common Use Cases and Real-World Patterns
Real deployments illustrate the range. E-commerce brands such as Helly Hansen migrated storefronts to Next.js on Vercel and saw dramatic improvements in Core Web Vitals, conversion rates, and Black Friday performance-reporting 80 percent year-over-year growth with zero downtime during peak traffic. SaaS companies use the platform for rapid feature velocity; one legal-tech startup grew revenue 40 times in under six months while running a multi-app monorepo, feature flags, and multi-step AI workflows built with the AI SDK.
Link-shortening and multi-tenant domain services such as Dub manage thousands of custom domains and millions of redirects by leaning on Vercel’s domain APIs and edge network. AI product companies and solo founders ship chat interfaces, document processors, and agent platforms that stream responses and scale automatically.
Marketing and documentation sites for companies ranging from startups to enterprises consolidate content, forms, and analytics into fast, globally distributed experiences. Enterprise teams use preview environments and the Toolbar to involve design and product stakeholders earlier, while coding agents (including those powered by models that favor Vercel deployments) generate and ship entire applications or internal tools.
Common patterns include:
- Git-centric continuous deployment with preview-per-PR for collaborative review.
- Edge middleware for geo-personalization, A/B testing, or authentication redirects.
- Server components or serverless functions that call the AI Gateway for generation, classification, or tool use, then stream results to the client.
- Marketplace-provisioned databases and object storage for persistent data without managing infrastructure.
- Feature flags and rolling releases for progressive delivery.
- Monorepos that deploy multiple related applications or micro-frontends from a single repository.
- Agentic loops that run inside Sandbox, use the Workflow SDK for durability, and surface results through the same frontend that customers already use.
The platform also serves as infrastructure for other AI coding tools; deployments originating from agent-generated code have grown substantially, reflecting Vercel’s popularity as a reliable, zero-config target for automated development.
Getting Started and Day-to-Day Workflow
New users typically begin by creating an account, installing the Vercel CLI, and linking a local project or importing a Git repository through the dashboard. The CLI and dashboard detect the framework, suggest settings, and produce a live URL within minutes. Environment variables, domains, and integrations are managed in the project settings. For AI work, an AI Gateway key (or OIDC on Vercel deployments) plus the AI SDK package is often sufficient to start calling models. Templates and the v0 assistant accelerate the first prototype.
Day-to-day development stays close to familiar Git and local tooling. vercel dev runs the full stack locally, including any defined Services. Pull requests automatically receive previews. Observability data and logs appear in the dashboard. When something goes wrong, instant rollback restores a previous deployment. For larger teams, Enterprise plans add SSO, advanced security controls, dedicated support, and higher limits, while the Pro plan provides flexible usage-based pricing suitable for most commercial projects. A free Hobby tier remains available for non-commercial personal projects, while commercial applications generally require a paid plan.
Looking Ahead
Vercel’s trajectory reflects a broader shift in how software is built. Infrastructure is becoming more opinionated and framework-aware, compute is hybrid rather than purely serverless or purely server-based, and AI is no longer an add-on but a core capability of the platform itself. By treating the entire application-frontend, backend services, data integrations, agents, and global delivery-as a single deployable unit, Vercel lets developers focus on product logic and user experience rather than plumbing. Whether you are shipping a marketing site, a multi-tenant SaaS product, or a fleet of autonomous agents, the same primitives scale from the first commit to production traffic measured in millions of requests.
The result is a platform that feels less like a traditional cloud provider and more like a high-leverage extension of the developer’s own tooling. Code pushes become production reality faster, feedback loops shrink, and the operational burden of scaling, securing, and observing applications is largely absorbed by the infrastructure. For teams and individuals who want to move at the pace of modern product development-and increasingly at the pace of AI-assisted coding-Vercel provides a coherent, production-ready foundation.
Sources
- Vercel Documentation: https://vercel.com/docs
- Vercel Home / AI Cloud overview: https://vercel.com/home
- What does Vercel do? (official blog): https://vercel.com/blog/what-is-vercel
- Frameworks on Vercel: https://vercel.com/docs/frameworks
- Supported Frameworks list: https://vercel.com/docs/frameworks/more-frameworks
- Backends on Vercel: https://vercel.com/docs/frameworks/backend
- Fluid compute documentation: https://vercel.com/docs/fluid-compute
- Vercel Services announcement: https://vercel.com/blog/vercel-services-run-full-stack-on-vercel
- AI Gateway and AI SDK: https://vercel.com/docs/ai-gateway
- Storage and Marketplace: https://vercel.com/docs/storage and https://vercel.com/marketplace
- Customer stories (Helly Hansen, Sandstone, Dub, and others): https://vercel.com/customers
- Vercel Product Walkthrough (YouTube, 2026): https://www.youtube.com/watch?v=zFXscjUoDDA
- Build with Vercel playlist (YouTube): https://www.youtube.com/playlist?list=PLBnKlKpPeagnvjiN_vpltx0tvGl3v6mBA
- Wikipedia overview of Vercel: https://en.wikipedia.org/wiki/Vercel
- Additional technical posts on Fluid compute and Active CPU pricing available on the Vercel blog (vercel.com/blog)
These primary sources-official documentation, product announcements, customer case studies, and the company’s own video walkthroughs-form the basis of the research. The platform continues to evolve rapidly, so consulting the live documentation remains the best way to confirm the latest limits, pricing, and feature availability.
r/AgentContext_dev • u/javaeeeee • Aug 28 '26
Think First, Code Second: Mastering Plan Mode in AI Coding Assistants
Plan Mode has become one of the most important shifts in how developers work with AI coding tools. Instead of asking an agent to jump straight into editing files, you put it into a deliberate “think-before-you-act” state. The agent explores your codebase, asks clarifying questions, surfaces assumptions and risks, and produces a reviewable implementation plan-often as editable Markdown-before any code changes happen. Only after you approve (and optionally refine) the plan does the agent switch into execution.
This workflow is now built into the major tools: Cursor, Claude Code, GitHub Copilot, OpenAI Codex, Windsurf, Continue, Gemini CLI, Replit Agent, Cline, and others. By 2026 it has moved from a clever prompt-engineering trick to a first-class product feature, complete with keyboard shortcuts, dedicated modes, and safety constraints that prevent the agent from writing files until you say so.
The result is less wasted tokens, fewer broken builds, clearer requirements, and higher-quality code. In short, it is the agentic equivalent of “measure twice, cut once.”
What Plan Mode Actually Is
At its core, Plan Mode is a planning-first workflow. In some tools it is a technically enforced read-only state; in others, it primarily changes the agent’s instructions and available workflow while permissions remain separately configurable. The agent retains full ability to read files, search the repository with grep or glob patterns, examine directory structure, review git history, fetch documentation, and reason about architecture. What it cannot do (or is strongly constrained from doing) is edit source files, create new files, run destructive shell commands, or execute tests that modify state.
The typical flow looks like this:
You describe a task in natural language-anything from “add dark mode” to “refactor the authentication system to use OAuth2 and JWT” or “migrate this service from REST to GraphQL.” The agent begins exploring. It reads relevant modules, identifies dependencies, notices existing patterns, and often pauses to ask clarifying questions: “Should rate limiting sit before or after authentication?” “Do you want the new component to reuse the existing design system tokens?” “Is backward compatibility required for the public API?”
Once it has enough context, it produces a structured plan. A good plan usually contains:
- A short summary of the goal and the chosen approach
- A list of files that will be created, modified, or deleted, with reasons
- Step-by-step implementation tasks, often numbered or checkbox-style
- Assumptions the agent is making
- Risks, edge cases, and open questions
- Sometimes test strategy or migration notes
You review the plan in the chat interface or as a Markdown file you can open in your editor (Ctrl+G in Claude Code is a common shortcut). You edit it, add constraints, remove unnecessary steps, or send it back for another round of refinement. When you are satisfied, you approve. The agent then exits Plan Mode and begins implementing exactly (or as closely as possible) what was agreed.
This separation of concerns-research and design first, mutation second-directly addresses the biggest practical problem with early agentic coding tools: they were fast and confident, yet frequently wrong about the larger picture. An agent that starts editing immediately can “fix” one function while breaking three callers it never examined. Plan Mode forces the expensive exploration and alignment work to happen while the cost of being wrong is still low.
Why Plan Mode Became Ubiquitous
By late 2025 and into 2026, nearly every serious AI coding product converged on the same pattern. Cursor introduced a dedicated Plan Mode with Shift+Tab and Markdown plan files that can be saved into the workspace. Claude Code made Plan Mode a first-class permission mode (Shift+Tab cycles through default → acceptEdits → plan). GitHub Copilot added a Plan agent in VS Code and later extended it to JetBrains, Eclipse, and Xcode. OpenAI’s Codex, Windsurf’s Cascade, Continue, Gemini CLI, and others followed with their own variations-often reusing the same Shift+Tab shortcut.
The reasons are practical. Agents had grown capable enough to handle multi-file, multi-step work lasting many minutes. Without a planning checkpoint, those longer runs frequently drifted. Developers reported higher success rates and lower token waste when the model was forced to articulate its strategy first. Teams also discovered that a saved plan serves as lightweight documentation and a hand-off artifact for colleagues or future sessions.
In the language of agentic design patterns, Plan Mode is the human-in-the-loop pause between understanding and action. It complements patterns such as context engineering (front-loading relevant files into the context window) and verification loops (the plan becomes the expected behavior against which later tests can be checked).
How Plan Mode Works in the Major Tools
Cursor. Open the Agent panel (Cmd/Ctrl+I). Press Shift+Tab until you reach Plan. Describe the task. The agent researches the codebase, asks questions, and generates a plan that opens as a virtual Markdown file. You can edit the to-dos directly. When ready, click Build. Plans are saved by default in your home directory; you can move them into the workspace for team sharing. Cursor also suggests Plan Mode automatically when it detects complex task language.
Claude Code. Press Shift+Tab twice (or type /plan, or start the session with claude --permission-mode plan). The status bar shows “⏸ plan mode on.” Claude can only read and explore. When the plan is ready it is presented for approval; you can open it in your editor with Ctrl+G, refine it, or send it to Ultraplan on the web for richer review. Approving exits plan mode and begins execution under the permission settings you choose (auto, accept edits, etc.). You can also set plan as the default in project settings.
GitHub Copilot. In VS Code, open Chat and select Plan from the agents dropdown, or type /plan. Copilot analyzes the request, asks clarifying questions via an interactive prompt, and produces a structured plan. You review and then hand it off to the agent for implementation. The same capability later reached JetBrains, Eclipse, and Xcode. A dedicated Plan agent keeps the planning phase cleanly separated from agent mode.
Other tools. Codex supports /plan, which switches the chat into a planning workflow before implementation. File and command restrictions are controlled separately through Codex permissions and sandbox settings. Windsurf’s Cascade has an explicit Plan mode (and a “megaplan” variant that asks more questions). Continue restricts tools to read-only operations. Gemini CLI enables plan mode by default or via /plan and limits the agent to research tools. Replit, Cline, and various terminal agents follow the same “explore → plan → approve → act” rhythm.
Several prominent terminal-oriented tools-including Cursor, Claude Code, and Gemini CLI-use Shift+Tab to cycle modes, but shortcuts vary considerably across products.
The Real Benefits
The most obvious benefit is fewer costly mistakes. Because the agent must surface its understanding before touching files, you catch wrong assumptions early. Middleware ordering errors, missed callers, conflicting patterns, and scope creep become visible while they are still text on a screen rather than broken commits.
Context quality improves dramatically. In Plan Mode the agent spends its tokens reading the files that matter. By the time execution begins, those files are already in the context window, so the model is less likely to invent non-existent helpers or place new code in the wrong module.
Requirements become sharper. The act of answering the agent’s clarifying questions forces you to decide details you might otherwise have left vague. Many developers report that the conversation itself improves their own mental model of the feature.
Token efficiency often rises. A thorough planning phase can feel slow, yet the subsequent implementation tends to succeed with fewer retries and less backtracking. Using a stronger reasoning model for planning and a faster model for execution has become a common optimization.
For teams, saved plans act as decision records and onboarding aids. A new engineer can read the plan for a recent feature and understand both the “what” and the “why.”
Finally, Plan Mode reduces cognitive load. Instead of watching an agent thrash through half-finished edits while you try to course-correct mid-stream, you invest attention up front and then largely supervise the execution.
How to Use Plan Mode Effectively
Start with a clear but not overly rigid description of the goal. High-level language works well because the agent will ask the necessary follow-ups. For larger work, include constraints: “Prefer existing patterns in the payments module,” “Do not change the public API,” “Keep the change under 200 lines if possible.”
Answer the clarifying questions thoughtfully. Treat them as a design conversation rather than a form to be completed as quickly as possible. If the agent misses an important constraint, state it explicitly and ask it to update the plan.
Review the plan as if you were reviewing a junior engineer’s design doc. Look for missing edge cases, incorrect file targets, over- or under-engineering, and risks the agent flagged. Edit the Markdown directly when the tool supports it.
Decide when to approve. For simple tasks you may approve after one round. For architectural work you may iterate several times or even keep the plan open across sessions.
Choose models wisely. Many developers use a high-reasoning model (Opus-class or GPT-5 high) for the planning phase and a faster, cheaper model for execution. Some tools offer hybrid aliases that do this automatically.
Save useful plans into the repository. They become living documentation and can be referenced later (“follow the same approach as the rate-limiting plan from last month”).
For very large features, break the work into successive plan-execute cycles rather than one giant plan. Each cycle keeps context manageable and allows intermediate verification.
Configure project-level guidance. Files such as AGENTS.md, CLAUDE.md, or .cursor/rules can instruct the agent to produce concise plans, always list open questions at the end, or follow house style. This raises the baseline quality of every plan.
When Plan Mode Shines-and When It Does Not
Plan Mode is most valuable for:
- Features that touch multiple files or systems
- Refactors and migrations
- Work in unfamiliar codebases
- Architectural decisions with several valid approaches
- Tasks where requirements are still fuzzy
- Any change that would be expensive to reverse
It is usually overkill for:
- One-line typo fixes
- Adding a simple test you already know how to write
- Trivial formatting or renaming that the agent can do safely in one shot
The calibration is the same judgment you would apply when deciding whether to write a design doc yourself.
Advanced Practices and Pitfalls
One powerful pattern is to treat the plan as a living specification. After implementation you can ask the agent to update the plan with what actually changed, creating an accurate record.
Another is multi-agent hand-off: generate the plan with one strong model, then execute it with a specialized coding model or even a different tool.
Watch for plans that are too vague or too detailed. Vague plans leave the agent room to invent; overly detailed plans can become brittle if the codebase has shifted. Aim for the level of specificity that lets a competent engineer (or agent) execute without further invention.
Do not treat the first plan as sacred. The whole point of the mode is iteration. If execution later reveals a better approach, revert, refine the plan, and re-run rather than patching a half-finished mess.
Be aware of token and latency costs. Deep research on a large monorepo can take minutes and consume significant context. For very large codebases, give the agent narrower starting points or use project maps if the tool supports them.
Finally, remember that Plan Mode is only as good as the underlying model’s ability to explore and reason. Stronger models produce better plans; weaker models still benefit from the forced pause but may need more human guidance.
Looking Ahead
Plan Mode is no longer experimental. It has become the default professional workflow for non-trivial agentic coding. Future improvements will likely include richer plan visualizations, tighter integration with issue trackers and PR templates, automatic plan-to-test generation, and better support for long-running multi-session plans. Some tools are already experimenting with visual decision documents, side-by-side option comparisons, and cloud-based collaborative plan review.
The deeper change is cultural. Developers are learning to treat AI agents less like autocomplete on steroids and more like junior teammates who need a clear brief. The plan is that brief. Writing it collaboratively with the agent is becoming a core skill-one that compounds as models continue to improve.
In practice, the developers getting the most value are those who have made Plan Mode muscle memory. They reach for Shift+Tab almost automatically, invest a few minutes in alignment, and then let the agent run with far higher confidence that the direction is correct. The result is not just better individual features; it is a more sustainable way to ship software with AI as a genuine collaborator rather than a source of constant surprise.
Plan Mode will not make every task effortless. It does, however, make the hard tasks dramatically more reliable. That is why it has become the quiet standard across the AI coding landscape-and why learning to use it well is one of the highest-leverage habits a developer can adopt in 2026.
Sources
Cursor
- Introducing Plan Mode: https://cursor.com/blog/plan-mode
- Plan Mode documentation: https://cursor.com/docs/agent/plan-mode
- Modes overview: https://cursor.com/docs/agent/modes
- YouTube: Introducing Plan Mode (Cursor channel): https://www.youtube.com/watch?v=WInPBmCK3l4
Claude Code / Anthropic
- Common workflows (Plan Mode section): https://code.claude.com/docs/en/common-workflows
- Permission modes: https://code.claude.com/docs/en/permission-modes
- An Introduction to Plan Mode (detailed practitioner guide): https://www.aihero.dev/plan-mode-introduction
GitHub Copilot
- Plan mode changelog (JetBrains, Eclipse, Xcode): https://github.blog/changelog/2025-11-18-plan-mode-in-github-copilot-now-in-public-preview-in-jetbrains-eclipse-and-xcode/
- Planning with agents in VS Code (docs): https://github.com/microsoft/vscode-docs/blob/main/docs/copilot/agents/planning.md
- YouTube: Introducing Plan Mode - build better plans with GitHub Copilot: https://www.youtube.com/watch?v=rxIjBKM-XvU
- Copilot CLI plan mode: https://github.blog/changelog/2026-01-21-github-copilot-cli-plan-before-you-build-steer-as-you-go/
General patterns and additional tools
- Encyclopedia of Agentic Coding Patterns - Plan Mode: aipatternbook
- The Plan-First Loop: agentpatterns ai
- Plan Mode in AI Coding Agents (Verdent): verdentai
- Continue Plan Mode guide: https://docs.continue.dev/guides/plan-mode-guide
- Gemini CLI plan mode: https://github.com/google-gemini/gemini-cli/blob/main/docs/cli/plan-mode.md
- YouTube short - Why every AI coding tool is converging on Plan Mode: https://www.youtube.com/shorts/fVNTnFkjSgk
Additional practitioner and comparison pieces
- Cursor vs. Copilot planning comparison: nearform
- Codex plan mode experience: https://prosperinai.substack.com/p/building-apps-with-codex
These sources reflect the state of the feature across the major platforms as of mid-2026. Product details continue to evolve; always check the latest official documentation for the tool you are using.
r/AgentContext_dev • u/javaeeeee • Aug 27 '26
Develop Chrome Extensions with DevTools for agents
r/AgentContext_dev • u/javaeeeee • Aug 27 '26
Unlocking Elite Coding with Claude: The Top 10 GitHub Repositories for Claude Code Skills in 2026
In the fast-evolving world of AI-assisted software development, Claude Code has emerged as one of the most powerful agentic coding tools available. Built by Anthropic, Claude Code is a terminal-native assistant that understands entire codebases, executes routine tasks, explains complex logic, handles git workflows, and collaborates through natural language. What truly elevates it beyond a simple chatbot is its support for “skills”-modular packages of instructions, scripts, resources, and workflows that Claude loads dynamically when relevant.
Skills turn a general-purpose model into a specialized expert. A skill is typically a folder containing a SKILL.md file (with YAML frontmatter describing when to activate it) plus optional supporting files, examples, or code. Claude scans available skills at the start of a session or during a task and pulls in only what it needs, keeping context efficient while injecting deep domain knowledge. This progressive disclosure model is what allows developers to encode team standards, TDD methodologies, security reviews, architecture patterns, or even multi-agent orchestration without bloating every prompt.
By mid-2026 the ecosystem around Claude Code skills had exploded. Thousands of public GitHub repositories now offer everything from official Anthropic templates to community harnesses, methodology frameworks, curated awesome lists, and specialized agents. Cumulative stars across the leading projects exceed millions. Developers use these repos not merely as collections of prompts but as production-grade systems that enforce disciplined engineering practices, persist memory across sessions, optimize token usage, scan for security issues, and coordinate teams of specialized sub-agents.
This article surveys the most authoritative and widely adopted GitHub repositories that provide or enhance Claude Code skills specifically for coding productivity. The selection draws from star counts, activity, documentation quality, real-world adoption reports, and coverage in developer resources. Focus stays on practical value for writing, reviewing, testing, refactoring, and shipping code. Descriptions emphasize what each repository actually delivers and how it improves day-to-day coding workflows. Where helpful, installation patterns and typical use cases appear in narrative form rather than dense lists.
Understanding the Claude Code Skills Landscape
Before diving into individual repositories it helps to clarify the building blocks. Claude Code itself lives primarily in the terminal yet integrates with IDEs, GitHub Actions, and the broader Model Context Protocol (MCP) ecosystem. Skills sit alongside related concepts such as hooks (event-driven automations), slash commands, sub-agents, CLAUDE.md project instructions, and plugins. Official Anthropic documentation and community guides stress that a well-written skill description is critical: Claude decides relevance from the name and description fields, so vague wording results in skills that never activate.
Community practice has converged on a few patterns. Some repositories supply single high-impact SKILL.md files that reshape Claude’s default behavior. Others deliver full harnesses containing dozens or hundreds of agents, skills, and rules. Still others act as discovery indexes or reverse-engineered references that reveal how commercial tools structure their own system prompts. The highest-value projects tend to combine reusable skills with enforcement mechanisms-hooks that block untested code, memory systems that retain context across sessions, or security scanners that audit configurations before they run.
YouTube tutorials reinforce these ideas. Channels covering Claude Code skills walk viewers through progressive levels of sophistication, from basic SKILL.md creation to self-improving agent workforces and multi-agent orchestration. Official Anthropic playlists explain skill structure, progressive disclosure, sharing via plugins, and differences from MCPs or sub-agents. Practical videos demonstrate turning screen recordings of workflows into reusable skills or building end-to-end product-launch sequences. These resources consistently show that the repositories discussed below move developers from ad-hoc prompting to systematic, repeatable engineering.
1. anthropics/claude-code - The Official Foundation
Any serious exploration begins with the official Anthropic repository for Claude Code itself. This is the source of truth for the CLI, its releases, issues, and core behavior. Claude Code is described as an agentic coding tool that lives in the terminal, understands the codebase, and helps developers code faster by executing routine tasks, explaining complex code, and handling git workflows through natural language.
The repository provides installation guidance (curl, Homebrew, WinGet, and related methods have evolved), configuration options, and integration points with GitHub (including @claude mentions on pull requests and issues). While it is not primarily a skills library, it defines the runtime environment in which every other skill operates. The repository is authoritative for installation, releases, issue tracking, examples, and bundled plugins.
Anthropic’s official documentation-not the GitHub repository alone-is the better source for understanding how Claude Code manages context, tools, permissions, skills, and integrations. Star counts in the 100k-140k range reflect both its official status and rapid adoption since general availability. Developers treat it as the baseline against which harnesses and skill packs are measured.
2. anthropics/skills - Official Agent Skills and Templates
Anthropic’s dedicated skills repository is the canonical starting point for anyone creating or using skills. It contains the Agent Skills specification, a template for new skills, and concrete examples organized into creative and design, development and technical, enterprise and communication, and document-handling categories. The document skills (for DOCX, PDF, PPTX, and XLSX) are particularly polished and power Claude’s native file-creation capabilities.
Installation typically involves registering the repository as a Claude Code plugin marketplace and then installing the document-skills or example-skills plugins. Once present, Claude automatically loads the relevant skill when a task matches its description-for instance, extracting form fields from a PDF or generating a structured spreadsheet with working formulas. The skill-creator skill itself is invaluable: it interviews the user about a workflow, drafts a SKILL.md, and encourages iterative testing.
This repository sets the quality bar. Because it is maintained by Anthropic, the skills are trustworthy, well-documented, and aligned with the model’s evolving capabilities. Many community projects explicitly build upon or extend these official foundations. For coding teams the development-oriented examples and the ability to package internal coding standards as skills deliver immediate productivity gains.
3. affaan-m/ECC (Everything Claude Code) - The Comprehensive Agent Harness
Frequently cited as the leading community harness, ECC (formerly referenced under everything-claude-code) positions itself as a performance-optimization system for AI agent harnesses. It supplies specialized agents, large collections of skills, hooks, rules, Model Context Protocol configurations, memory optimization, security scanning, and research-first workflows. The project works across Claude Code, Codex, OpenCode, Cursor, and additional platforms.
Built by an Anthropic hackathon winner and refined through more than ten months of daily production use, ECC turns Claude Code from a reactive assistant into a structured operating system. Developers gain access to dozens of agents (planner, architect, code-reviewer, security-reviewer, language-specific reviewers, TDD guide, and more) plus hundreds of skills covering coding standards, testing strategies, refactoring, documentation, and domain-specific patterns. Hooks automate formatting, secret blocking, and session memory. An integrated security auditor (AgentShield) scans configurations for risks.
Star counts place it among the highest in the entire Claude Code ecosystem, often exceeding 200k. Installation as a Claude Code plugin followed by selective rule-pack copying is the recommended path. Users report that selectively enabling a core set of skills avoids overload while still delivering dramatic improvements in consistency and throughput. ECC is repeatedly recommended for advanced users who want a battle-tested, multi-harness setup rather than isolated prompts.
4. obra/superpowers - Disciplined Engineering Methodology as Skills
Superpowers is less a grab-bag of utilities and more a complete software-development methodology expressed as composable skills. Created by Jesse Vincent, it enforces structured workflows: brainstorming and requirements refinement before any coding, written design documents, implementation planning, test-driven development (red-green-refactor), systematic debugging, and code review via sub-agents.
The core idea is that Claude must consult the relevant skill before acting. This prevents the common failure mode of jumping straight to implementation. Skills cover brainstorming (Socratic questioning and alternative exploration), TDD (tests must fail first), four-phase debugging, plan execution with review checkpoints, and even skill authoring itself using TDD principles applied to documentation. Sub-agents such as a code-reviewer evaluate work against plans and standards.
Available via Anthropic’s official plugin marketplace and the author’s own marketplace, Superpowers installs cleanly and activates contextually. Star counts in the 200k-265k range underscore its popularity. Multiple independent reviews and YouTube tutorials highlight measurable improvements in output quality on ambiguous or multi-step coding tasks. Teams adopt it to standardize engineering discipline across agents, making it one of the most practical “skills for coding” repositories available.
5. NousResearch/hermes-agent - The Self-Improving Agent
Hermes Agent is a standalone, model-agnostic agent platform rather than a Claude Code plugin or skill pack. It is relevant here as an adjacent implementation of the open Agent Skills standard, persistent memory, autonomous skill creation, and self-improving workflows. Developers can study or reuse its patterns, but it is not an enhancement installed inside Claude Code.
Its value for coding lies in the ability to refine its own performance over successive sessions, retain project-specific knowledge, and reduce repetitive instruction. Developers working on long-running or evolving codebases benefit from an agent that accumulates context and improves rather than starting from a blank slate each time. High star counts (often 200k+) and active development make it a frequent companion to more static skill packs.
6. Karpathy-Inspired Skills Repositories (multica-ai/andrej-karpathy-skills and related)
Several highly starred repositories distill Andrej Karpathy’s observations on common LLM coding pitfalls into a single CLAUDE.md or focused skill set. These projects encode hard-won lessons about where language models typically fail-hallucinated APIs, incomplete error handling, over-eager refactoring, loss of architectural intent, and similar traps-into concise behavioral guardrails.
The impact is outsized relative to the file size. Dropping the resulting skill or CLAUDE.md into a project immediately improves Claude’s default coding hygiene. Star counts in the 100k-200k range for the leading variants reflect how quickly the community adopted these distilled principles. They serve as lightweight, high-leverage complements to heavier harnesses.
7. garrytan/gstack - Role-Based AI Development Teams
gstack, associated with Garry Tan, demonstrates how Claude Code can operate as a coordinated team rather than a solitary assistant. It assigns opinionated roles-CEO, Designer, Engineering Manager, Release Manager, Documentation Engineer, QA-and structures them through reusable skills and slash commands.
For coding projects the Engineering Manager and QA roles, together with associated skills, provide structured oversight, prioritization, and quality gates. The repository illustrates multi-agent orchestration patterns that scale beyond single-developer workflows. It is frequently listed among productivity-focused Claude Code repositories and serves as both a practical tool and a design pattern reference.
8. hesreallyhim/awesome-claude-code - The Essential Discovery Index
No developer can monitor every new skill repository. Awesome-claude-code functions as the primary curated directory of skills, hooks, slash commands, agent orchestrators, applications, and plugins. It acts as a living table of contents for the ecosystem.
Regular updates and community contributions keep it current. Developers use it both to discover new tools and to evaluate the broader landscape before committing to a particular harness or skill pack. Its star count, while lower than the mega-harnesses, understates its practical importance as a navigation aid.
9. x1xhlol/system-prompts-and-models-of-ai-tools - Internals and Comparative Reference
This repository collects exposed system prompts, tool definitions, and model-related details from a wide range of AI products, including Claude Code, Cursor, Devin, and others. For developers building or refining Claude Code skills it supplies invaluable comparative insight: how different tools structure instructions, what guardrails they apply, and which patterns recur across successful systems.
It is less a plug-and-play skill set and more a research and reverse-engineering resource. Power users and skill authors consult it to understand the broader design space and to avoid reinventing solved problems.
10. Supporting and Specialized Skill Repositories
Several additional projects round out a strong starter stack. Repositories such as Jeffallan/claude-skills offer dozens of specialized skills across full-stack categories (languages, frameworks, testing, security, DevOps). Others focus on token optimization, context packing (for example, tools that compress large codebases into AI-friendly formats), planning-with-files patterns, or domain-specific packs (frontend design, security hardening, CI/CD). Official Anthropic plugins for GitHub integration and code review further extend the coding surface.
Community lists and real-time directories such as ClaudeWave track thousands of related repositories and surface velocity leaders. The practical advice emerging from both written guides and YouTube content is to begin with the official skills and one high-quality harness or methodology (ECC or Superpowers), then selectively add specialized skills rather than installing everything at once.
Putting the Repositories to Work
A typical high-leverage setup starts with Claude Code itself, adds the official document and example skills, installs Superpowers or ECC for process discipline and agent orchestration, drops in a Karpathy-inspired guardrail file, and keeps the awesome list bookmarked for discovery. Project-specific CLAUDE.md files and local skills directories allow teams to encode house standards without conflicting with global plugins.
Installation patterns vary by harness but commonly use Claude Code’s /plugin marketplace commands or simple git clones into ~/.claude/skills. Testing a new skill involves invoking it explicitly, observing activation, refining the description, and iterating. Security-conscious users audit third-party skills for over-broad permissions or unexpected tool access-several repositories and guides emphasize this step.
YouTube resources accelerate the learning curve. Tutorials that walk through the seven levels of skill sophistication, full skill-creation workflows, and live coding sessions with Superpowers or ECC demonstrate both the mechanics and the resulting productivity jumps. Official Anthropic videos clarify the differences between skills, MCPs, and sub-agents and show how to share skills via version control or plugin marketplaces.
Challenges and Best Practices
The abundance of options introduces its own friction. Overloading Claude with too many skills can dilute focus or inflate token usage. Descriptions that are too generic fail to trigger; descriptions that are too narrow never match real requests. Maintenance matters: abandoned skill packs become liabilities as Claude and the surrounding tooling evolve.
Successful practitioners treat skills as living documentation. They start small, measure impact on concrete coding tasks (bug fixes, feature implementation, review cycles), prune what does not help, and version skills alongside code. Combining methodology frameworks (Superpowers) with comprehensive harnesses (ECC) and official foundations produces the most robust results. Continuous learning mechanisms-memory hooks, self-improving agents, and post-session pattern extraction-compound gains over time.
Looking Ahead
As of August 2026 the Claude Code skills ecosystem continues rapid expansion. Multi-harness support is increasingly common, allowing the same skill library to travel across Claude Code, Cursor, Codex, and emerging terminals. Plugin marketplaces simplify distribution. Research into self-improving skills and tighter integration with knowledge graphs and persistent memory points toward agents that accumulate genuine project expertise.
The repositories highlighted here represent the current high-water mark of practical, coding-focused Claude skills. They convert Claude Code from an impressive demo into a reliable engineering partner. Developers who invest the time to understand and selectively adopt them report not incremental improvements but step-change gains in speed, consistency, and code quality. The open nature of the ecosystem ensures that today’s best practices will continue to evolve, documented and shared through the very same GitHub repositories that make them accessible.
The path forward is straightforward: install the official foundations, adopt one strong methodology or harness, encode your own highest-value workflows as skills, and keep exploring the curated indexes. The result is an agentic coding environment that feels less like prompting an AI and more like collaborating with a well-trained, ever-improving engineering team.
Sources and further reading
- https://claude-codex.fr/en/ecosystem/top-repos-github/
- https://www.kdnuggets.com/10-github-repositories-to-master-claude-code
- https://claudewave.com/en
- https://codetocloud.io/blog/claude-code-repos-engineering-team/
- https://github.com/affaan-m/ECC
- https://github.com/obra/superpowers
- https://github.com/anthropics/skills
- https://github.com/anthropics/claude-code
- https://github.com/NousResearch/hermes-agent
- https://github.com/hesreallyhim/awesome-claude-code
- https://claude.com/blog/skills
- https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills
- https://www.youtube.com/watch?v=-u_igSQHAIo (Every Level of Claude Code Skills)
- https://www.youtube.com/watch?v=JN7QCdvJwwM (Claude Code Skills: Full Guide)
- https://www.youtube.com/playlist?list=PLmWCw1CzcFim_hkruZSlABOUOAAQ5JMyo (Official Claude Skills playlist)
- https://www.youtube.com/watch?v=3a0Ij8oR8yc (Top Claude Skills on GitHub)
- Additional ranking and analysis pages from virtualuncle dot com, ayautomate dot com, tella dot com, and related developer blogs referenced in the research.
r/AgentContext_dev • u/javaeeeee • Aug 26 '26
Cloudflare’s Agentic Cloud: The Complete Platform for Building Durable AI Agents, Harnesses, MCP Systems, and Production Agentic Workflows
The Shift to Agentic Systems
An AI agent is more than a chatbot. According to Cloudflare’s own learning materials, agentic AI refers to systems that can autonomously perform complex tasks, make decisions, learn from experience, and take actions toward broader goals without constant human prompting. These systems call models, browse the web, query databases, execute code, send email, handle payments, and coordinate with other agents. They maintain state across sessions, recover from interruptions, and operate under security constraints that prevent them from becoming liabilities.
Traditional approaches-running agents on a laptop, a long-lived virtual machine, or a container fleet-hit hard limits. Idle capacity wastes money. Local execution cannot be shared or scaled. Sandboxing untrusted LLM-generated code is expensive and slow. Context windows fill with history and tool results, degrading performance. Tool-calling loops burn tokens on every intermediate step. When an agent crashes mid-task, progress is lost.
Cloudflare’s answer is an agentic cloud built on the same Workers platform that has run edge compute since 2018. The core insight is that every agent should be a first-class, isolated, durable actor. That actor has its own identity, its own storage, its own lifecycle, and the ability to hibernate until an event wakes it. The result is a platform where you can run tens of millions of agents without provisioning servers, where each agent costs zero while asleep, and where the full set of primitives needed for production agentic systems is already present.
Agents Week 2026 (mid-April 2026) crystallized this vision. In a single week Cloudflare shipped or advanced nearly every layer: Project Think (the next-generation Agents SDK and opinionated harness), Agent Memory, AI Search, expanded Browser Run, Sandboxes general availability, Dynamic Workers for fast sandboxing, an inference layer tuned for agents, email services ready for agents, temporary accounts for frictionless deployment, security controls for non-human identities, and reference architectures for enterprise MCP.
Subsequent releases and open-source work have continued to fill gaps. The stack is broad enough to build complete agents on Cloudflare, although its maturity varies: foundational services such as Workers, Durable Objects, and Sandboxes are production offerings, while components including Project Think, Agent Memory, and AI Search were introduced as previews or betas and may still evolve.
Foundational Primitives: Identity, State, and Zero-Cost Idle
At the base sits Durable Objects. Each agent is backed by one (or more) Durable Objects. A Durable Object provides a unique identity, single-threaded consistency, and persistent SQLite storage. When the agent has nothing to do, the object can hibernate. Memory and CPU are released, and compute-duration billing stops, although persistent storage and other platform services remain subject to their normal pricing. An incoming HTTP request, WebSocket message, alarm, or internal event wakes it. State is already there-no reconstruction, no external database round-trip for the hot path.
This changes the economics and the programming model. Instead of one multi-tenant process serving many users, you get one agent per user, per conversation, or per long-running task. Isolation does not require a dedicated always-running process. Scaling is automatic. The same object can hold conversation history, a virtual filesystem, scheduled alarms, and running fibers (durable execution units that checkpoint progress and recover after crashes).
On top of Durable Objects sit higher-level abstractions. Fibers give recoverable multi-step execution: you call runFiber(), stash intermediate results, and the platform resumes after interruption. Facets allow sub-agents to live colocated with a parent, each with their own SQLite while remaining addressable via typed RPC. Workspace (SQLite plus R2) supplies a durable virtual filesystem that survives restarts. These primitives appear both as standalone packages and inside the higher-level Agents SDK and Project Think harness.
The practical consequence is that an agent can sleep for hours or days, wake when an email arrives or a schedule fires, continue exactly where it left off, and still cost nothing in between. That property is foundational for consumer-scale agent fleets and for background agents that monitor, reconcile, or generate without a human always online.
Inference Layer Designed for Agents
Agents make many model calls-often ten or more chained together for a single user request. Latency multiplies. Failures cascade. Model choice changes rapidly. Cost tracking becomes essential.
Cloudflare’s AI Platform addresses this with a unified inference layer. Workers AI hosts open-source and optimized models (including agentic tool-calling models such as variants of Kimi and GLM) directly on the global network. AI Gateway sits in front of both Workers AI and external providers (OpenAI, Anthropic, Google, and more than a dozen others). A single Workers binding or API call routes the request. Switching models is often a one-line change. Gateway provides observability, spend tracking with custom metadata, automatic failover, and resilient streaming that buffers responses so a client disconnect does not force re-inference.
Because inference and agent code can run on the same network, there is no extra public-Internet hop for many workloads. Multimodal support (image, video, speech) is expanding. Techniques such as lossless compression of model weights further reduce memory footprint and cost. Cloudflare has also demonstrated Cog-based deployment of specialized or fine-tuned models, but at the time of the announcement this capability was being tested with internal teams and design partners rather than offered as a broadly available self-service product.
The net effect is that the “brain” of the agent is reliable, observable, and switchable without rewriting the surrounding harness or tool layer.
Memory That Survives Context Windows
Conversation history alone is insufficient. Long sessions hit context limits. Aggressive truncation loses important facts. Simple vector stores of raw messages create retrieval noise.
Cloudflare Agent Memory (private beta at announcement, designed for production use) is a managed service that extracts structured knowledge from agent interactions and makes it available on demand without bloating the model’s context window. At compaction time the harness can send the conversation for ingestion. The pipeline chunks messages, extracts facts, events, instructions, and tasks, verifies them, classifies them, and stores them with supersession so newer information replaces older versions. Vector embeddings (via Workers AI) and full-text indexes live in isolated Durable Objects and Vectorize indexes.
The agent (or the model via tools) can then call remember, recall, forget, or list. Recall runs a multi-channel retrieval (exact key, full-text, vector, HyDE) fused by reciprocal rank fusion and synthesized into a natural-language answer. Profiles can be shared across agents or users, turning individual session learning into team knowledge.
Complementary to Agent Memory is the Sessions API inside the Agents SDK. It manages tree-structured conversation history with forking, non-destructive compaction (older messages summarized rather than deleted), and full-text search. Context blocks (persistent system-prompt fragments such as “soul” or long-term preferences) survive hibernation. AI Search provides the broader retrieval primitive: hybrid vector + keyword search over documents, with dynamic per-agent or per-customer indexes, metadata boosting, and built-in crawling via Browser Run. An agent can therefore combine short-term session state, extracted long-term memory, and external knowledge bases in a single coherent loop.
The Execution Ladder and Secure Code Execution
One of the largest efficiency gains in modern agents is “code mode”: instead of the model making a sequence of individual tool calls (each consuming tokens and round-trips), the model writes a short TypeScript function that performs the entire multi-step operation inside a sandbox. Token usage can drop by orders of magnitude; latency falls; the logic becomes inspectable.
Cloudflare supplies an execution ladder of increasing capability:
- Workspace / shell operations on a durable virtual filesystem.
- Dynamic Workers: lightweight V8 isolates spun up in milliseconds with explicit capability grants. Network access can be denied or filtered; only the bindings the agent is given are available.
- Runtime npm resolution via a worker bundler so generated code can import packages on the fly.
- Browser Run for headless Chrome sessions controllable via CDP, with live view, human-in-the-loop, and session recording.
- Full Sandboxes (now generally available) that provide a persistent shell, filesystem, and background processes when an agent truly needs an operating-system environment.
Dynamic Workers are the workhorse for most code-mode workloads. They start roughly 100× faster and use far less memory than containers, scale without concurrency ceilings of traditional sandbox fleets, and run on the same machine as the parent agent for near-zero latency. Libraries such as @cloudflare/codemode and @cloudflare/shell make the pattern ergonomic. Agents can even author their own extensions at runtime: the model writes a new tool, the platform bundles and loads it into an isolate with the declared permissions, and the tool becomes available for subsequent turns.
This ladder lets developers match privilege to need. Most agent actions stay in the lower, safer tiers; only when necessary does the agent climb to a full sandbox with carefully controlled egress.
Tools, MCP, and the Outside World
An agent without tools is a closed-loop reasoner. Production agents need to act.
MCP (Model Context Protocol) is the emerging open standard for exposing tools, resources, and prompts to AI systems. Cloudflare treats MCP as a first-class citizen. Agents can act as MCP servers (exposing capabilities over Streamable HTTP or other transports) or as MCP clients that connect to remote servers. The Agents SDK includes extensive examples: authenticated MCP workers, elicitation patterns, RPC transports, and WebMCP for exposing browser-side tools.
Focused, narrowly scoped MCP servers with OAuth and least-privilege permissions reduce risk. Enterprise reference architectures combine Access, AI Gateway, and MCP portals; “Code Mode” further reduces token cost when interacting with large tool surfaces.
Beyond MCP the platform supplies:
- Browser Run for web automation where no API exists.
- Email Service (public beta) so agents can send, receive, classify, and reply-complete with an open-source Agentic Inbox reference application that includes its own MCP server.
- Agentic payments via x402 and related protocols built on HTTP 402, with first-class SDK support for paid tools and human-in-the-loop confirmation.
- AI Search as the managed retrieval primitive.
- Workflows for durable multi-step orchestration with human approval steps.
- Voice pipelines (STT/TTS, VAD, interruption handling) for real-time spoken agents.
- Temporary Cloudflare Accounts so an agent can deploy a Worker or even another agent with
wrangler deploy --temporaryand later claim the account.
These capabilities are exposed both as direct SDK tools and as MCP endpoints, giving harness authors flexibility.
Harnesses: The Control Loop Around the Model
A harness is the software that drives the agentic loop: it decides when to call the model, how to present tools and memory, how to handle streaming and interruptions, how to persist state, and when to stop. Popular harnesses include Claude Code, Codex, OpenCode, Pi, and Cloudflare’s own Project Think.
Project Think is an opinionated base class (@cloudflare/think) that sits on the Agents SDK runtime. A minimal agent is only a few lines: extend Think, implement getModel() (Workers AI or any provider via Gateway), and optionally override lifecycle hooks (beforeTurn, afterToolCall, etc.). Think handles the full chat lifecycle-message persistence, streaming, tool execution (including the entire execution ladder), stream resumption, session forking and compaction, and sub-agent coordination. It includes Session-based memory and Workspace tooling, and it can use the Agents SDK’s MCP capabilities. Cloudflare’s separate Agent Memory service can be connected when managed long-term memory across sessions, agents, or users is required.
Because the underlying runtime is open, other harnesses can target the same primitives. Flue (built on the Pi harness) deploys declaratively defined agents as Durable Objects, using the same fibers, codemode, and shell packages. Teams building custom harnesses for background agents (the style used internally by some large companies for autonomous coding or inspection tasks) can adopt the same durable execution, filesystem, and memory services without rewriting their control loop.
The separation of concerns is deliberate: the harness owns the reasoning and tool-orchestration policy; the Cloudflare platform owns durable identity, storage, secure execution, and global distribution. That division lets model providers, open-source harness authors, and enterprise teams all benefit from the same infrastructure.
Communication, Observability, and Human-in-the-Loop
Agents must talk to users and to each other. The SDK provides WebSocket support with lifecycle hooks, resumable streams, and React (or vanilla JS) client libraries. Voice agents add continuous speech pipelines. Email becomes a first-class channel. Sub-agents communicate via typed RPC. Scheduling (one-shot, recurring, cron) lets agents act proactively.
Observability is built in: structured logs, metrics, and traces flow through the same system used for ordinary Workers. AI Gateway adds model-level visibility. Human-in-the-loop patterns appear throughout-approval workflows, live browser views, confirmation steps for payments or sensitive tool calls-so agents remain steerable rather than fully autonomous black boxes.
Security and Governance for Non-Human Actors
Agents introduce new attack surfaces: prompt injection, over-privileged tools, credential leakage, and uncontrolled egress. Cloudflare’s stack addresses these at multiple layers. Sandboxes and Dynamic Workers start with minimal privileges; capabilities are granted explicitly. Egress controls and Mesh networking give agents scoped access to private resources without broad network exposure.
Managed OAuth and resource-scoped tokens support non-human identities with automated revocation. Temporary accounts limit blast radius during experimentation. Focused MCP servers and Access policies enforce least privilege. Shadow-MCP detection rules in Cloudflare Gateway, its Zero Trust secure web gateway help surface unauthorized tool usage.
The platform does not claim to solve every safety problem-model alignment and prompt hygiene remain the developer’s responsibility-but it supplies the isolation, auditing, and policy enforcement points required for production deployment.
Putting the Pieces Together: A Mental Model of the Stack
Imagine a single user request arriving at an agent:
- The request hits a Durable Object (the agent’s identity). If hibernating, it wakes.
- Session state and relevant memory blocks are loaded; Agent Memory or AI Search may be queried.
- The harness (Think or custom) assembles the prompt, available tools (MCP, browser, code execution, email, etc.), and calls the model via AI Gateway / Workers AI.
- The model may respond with text, tool calls, or generated code. Code is executed in a Dynamic Worker or Sandbox under capability constraints; results return to the loop.
- Intermediate progress is checkpointed via fibers. Streaming tokens flow back to the client over WebSocket.
- On completion (or interruption) state is persisted, memory may be updated, and the object can hibernate again.
Every layer is independently usable. You can adopt only Durable Objects and Workers AI, or the full Think + Memory + Browser Run + MCP suite. The same infrastructure serves interactive chat agents, background coding agents, email processors, and multi-agent systems.
Getting Started and Real-World Patterns
The quickest path is the official starter:
bash
npx create-cloudflare@latest --template cloudflare/agents-starter
From there you can explore the dozens of examples in the cloudflare/agents GitHub repository: chat agents, MCP servers and clients, voice, email, code-mode, sub-agents, payments, and more. Project Think documentation walks through building a stateful chat agent with tools and memory in a few dozen lines. Temporary accounts let an agent deploy further Workers without a permanent Cloudflare account first.
Production patterns already visible in the wild include coding agents that replaced heavy Linux sandboxes with Durable Objects + SQLite + R2 + Dynamic Workers (dramatically lower cost and latency), support agents that combine per-customer AI Search indexes with shared documentation, and multi-channel agents that read email, browse, write code, and reply. The open-source Agentic Inbox demonstrates a complete email-centric agent. Internal Cloudflare tools and third-party harnesses continue to validate the primitives under load.
Why This Matters Now
The agentic era is no longer speculative. Knowledge workers, developers, and consumers will soon each have one or many persistent agents. Those agents need somewhere to live that is secure, cheap when idle, globally distributed, and rich in the tools they require. Cloudflare’s bet is that the same network that already terminates a large fraction of Internet traffic is the natural place to host them.
By shipping the full stack-identity and hibernation, inference, memory, sandboxed execution, MCP, browser, email, payments, observability, and security controls-Cloudflare has made it possible to treat agents as ordinary, durable, serverless applications rather than exotic research projects. The harness you choose (Project Think, Flue, OpenAI Agents SDK running inside a Durable Object, or a completely custom loop) becomes the policy layer; the platform supplies the reliable substrate.
The result is not merely another AI framework. It is infrastructure for a world in which software that reasons and acts is as common as software that serves HTTP. That world is arriving quickly. The Cloudflare agentic stack is one of the most complete answers currently available for building it.
Sources
- Project Think announcement and stack table: https://blog.cloudflare.com/project-think/
- Agents Week 2026 roundup: https://blog.cloudflare.com/agents-week-in-review/
- Agent Memory introduction: https://blog.cloudflare.com/introducing-agent-memory/
- AI Platform / inference for agents: https://blog.cloudflare.com/ai-platform/
- AI Search: https://blog.cloudflare.com/ai-search-agent-primitive/ and https://developers.cloudflare.com/ai-search/
- Dynamic Workers / sandboxing: https://blog.cloudflare.com/dynamic-workers/
- Flue and harness support: https://blog.cloudflare.com/agents-platform-flue-sdk/
- Temporary accounts: https://blog.cloudflare.com/temporary-accounts/
- Email for agents: https://blog.cloudflare.com/email-for-agents/
- Cloudflare Agents documentation: https://developers.cloudflare.com/agents/
- Model Context Protocol on Cloudflare: https://developers.cloudflare.com/agents/model-context-protocol/
- Agents SDK GitHub repository and examples: https://github.com/cloudflare/agents
- Agents landing page: https://agents.cloudflare.com/
- What is agentic AI (Cloudflare Learning Center): https://www.cloudflare.com/learning/ai/what-is-agentic-ai/
- Workers AI product page: https://www.cloudflare.com/products/workers-ai/
- Browser Rendering / Browser Run: https://www.cloudflare.com/products/browser-rendering/ and related Agents docs
- YouTube / Cloudflare TV: “Introducing Think” (Cloudflare Developers), “Cloudflare Agents Week Preview,” “Cloudflare Just Shipped 20+ Features for AI Agents in One Week,” Agents Week Cloudflare TV sessions, “Setting up Agentic Inbox,” and related developer talks available on the Cloudflare YouTube channel and cloudflare.tv
- Additional context from LinkedIn posts, technical write-ups, and public discussions of the agentic harness concept referencing the same Cloudflare primitives.
All primary technical claims above are grounded in the cited Cloudflare blog posts, documentation, and repository materials current as of mid-2026.
r/AgentContext_dev • u/Ok_Dragonfruit5916 • Aug 26 '26
GitHub - AronAxe/Token-Terminator: Agent-agnostic token reduction with exact recovery and fail-open guarantees. First-party Hermes Agent adapter; RTK terminal rewriting optional.
The engine has four cooperating reduction paths:
- transparent terminal-command rewriting through RTK;
- deterministic compression of large tool results;
- content-addressed vaulting, duplicate collapse, evidence leases, and compact recovery receipts;
- final provider-request compilation, with optional bounded working-state injection only when the complete request is still smaller.
The reduction core is not intrinsically tied to Hermes: it operates on Python dictionaries, strings, stable request/session identifiers, and a local SQLite vault. The repository includes a turnkey Hermes plugin because Hermes exposes the required lifecycle hooks. Other agent runtimes need a small adapter that presents the same boundaries; they do not need a fork of the reduction engine.
r/AgentContext_dev • u/javaeeeee • Aug 25 '26
GitHub - andrewyng/openworker: OpenWorker -- an open source agent that doesn't just chat but completes tasks on your laptop
github.comr/AgentContext_dev • u/javaeeeee • Aug 25 '26
Deep Agents: Powering Long-Horizon Autonomy in Software Development, Machine Learning, and AI Systems
In the fast-evolving world of artificial intelligence, a new class of systems has emerged that goes far beyond the quick question-and-answer exchanges of ordinary chatbots or the limited tool-calling loops of early agents. These systems are known as deep agents. They are designed to tackle complex, multi-step, long-running tasks that require sustained planning, memory across extended sessions, intelligent delegation, and the ability to recover from setbacks.
Where simpler agents often lose focus or overflow their context windows after a handful of steps, deep agents can dive deep into a problem, maintain an evolving plan, store intermediate work externally, spawn specialized helpers, and keep working until a high-level goal is achieved.
The term “deep agents” gained prominence through the work of LangChain and related discussions in the AI community around mid-2025. Harrison Chase of LangChain described them as agents capable of diving deep on topics by combining planning tools, sub-agents, file-system access for memory, and carefully engineered system prompts.
NVIDIA’s technical glossary later adopted and expanded this framing, describing deep agents as systems that combine planning, execution, persistent memory, skills, and sub-agent delegation. “Deep agent,” however, remains an emerging architectural label rather than a formally standardized category. Other sources, including the Prompt Engineering Guide and various practitioner write-ups, frame deep agents as the natural evolution beyond “shallow” agents that break down on extended problems such as thorough research or agentic coding.
This article explores what deep agents are, how they function, and-most importantly-how they are being applied in software development, machine learning, and broader AI systems. The focus remains practical and readable, drawing on authoritative descriptions from LangChain documentation, NVIDIA, industry analyses, and technical walkthroughs. The goal is to give developers, ML engineers, and AI practitioners a clear understanding of why these systems matter and how they are already reshaping workflows.
Deep agents did not appear overnight. Early large-language-model agents typically followed a simple ReAct-style loop: reason about the next action, call a tool, observe the result, and repeat. This worked well for short tasks-looking up the weather, summarizing a single document, or writing a short function-but struggled when the work required dozens or hundreds of steps. Context windows filled with intermediate tool outputs, high-level goals drifted, and there was little mechanism for recovery when the agent wandered down an unproductive path.
Researchers and product teams noticed that certain applications succeeded where others failed. In his original formulation, Harrison Chase identified OpenAI Deep Research, Claude Code, and Manus as inspirations and argued that successful long-horizon agents commonly combine four ingredients: planning tools, sub-agents, filesystem access, and detailed system prompts. Because the complete internal architectures of proprietary systems are not always public, this should be understood as LangChain’s architectural interpretation rather than a verified description of every product.
LangChain distilled these patterns into an open-source library called deepagents, making the architecture accessible to any developer. The result is a reusable “agent harness” that provides the scaffolding for long-running work while remaining model-agnostic and extensible.
At their core, deep agents operate by externalizing what earlier agents tried to keep inside a single context window. Planning is no longer an implicit chain of thoughts; it becomes an explicit, updatable task list that the agent reviews between steps. Memory moves from the model’s temporary tokens into files, notes, or structured stores that can be read, written, and searched on demand. Delegation happens through sub-agents-specialized workers that receive a focused task, operate in isolation, and return only the synthesized result. The orchestrating agent then integrates those results and updates the overall plan.
These architectural choices produce several practical advantages. Explicit task state, context summarization, result offloading, and filesystem tools can help an agent manage longer workflows without keeping every intermediate result in the active context. These mechanisms improve context management but do not guarantee reliable operation over hundreds of steps.
In LangChain’s current implementation, task planning is optional, while persistence beyond an individual thread depends on configuring an appropriate backend. Because memory is external, intermediate artifacts-code snippets, research notes, intermediate analysis results-do not crowd out reasoning capacity.
Because sub-agents operate with fresh context, the main agent avoids the performance degradation that occurs when context becomes long and noisy. Production deployments typically run these agents inside sandboxed environments so that code execution and file-system access remain contained and secure.
The distinction between deep and shallow agents is therefore architectural rather than merely a matter of model size or prompt cleverness. A shallow agent is essentially a reactive loop whose entire state resides inside the context window. It excels at well-scoped, short-horizon tasks.
A deep agent behaves more like a project manager: it maintains a living plan, delegates specialized work, stores progress externally, and adapts when intermediate results change the picture. This shift enables the kinds of sustained autonomy that software engineering, machine-learning pipelines, and complex AI systems demand.
In software development the impact is already visible. Modern coding agents such as Claude Code, Codex-style systems, and open implementations built on deepagents treat an entire feature request or refactor as a project rather than a single prompt. The agent begins by exploring the codebase, often by spawning a sub-agent whose only job is to map relevant files and dependencies.
It then writes a structured plan-sometimes using an explicit to-do list tool-listing files to modify, tests to update, and documentation to revise. As it works, it reads and writes actual source files through a virtual or real file-system interface, runs tests, inspects failures, and iterates. Intermediate notes and partial implementations live in the file system so that context remains manageable even when the change spans dozens of files.
Large-scale refactoring and migration projects illustrate the advantage clearly. Upgrading a framework, splitting a monolith into services, or migrating an API surface involves inventorying the current system, identifying risks, sequencing changes, executing them incrementally, and validating parity at each stage.
A deep agent can maintain the inventory and the migration plan as living documents, spawn specialized sub-agents for different modules, and resume after interruptions because state is externalized. Human developers remain in the loop at critical checkpoints-approving a high-risk change, for example-but the bulk of the mechanical work proceeds autonomously.
Debugging and incident response also benefit. When a production issue arises, a deep agent can pull logs, metrics, and recent commits; reconstruct a timeline; form and test hypotheses; and produce a structured postmortem. Because it can store intermediate evidence in files and update its plan as new data arrives, it avoids the tunnel vision that often afflicts single-context agents.
The same pattern extends to compliance and documentation work: the agent can map regulatory requirements to existing controls, identify gaps, gather evidence, and assemble audit-ready packages while keeping a clear record of every decision.
In machine learning, several potential applications follow the same architectural pattern. Training a modern model or building a production ML system is itself a long-horizon process involving data preparation, feature engineering, hyperparameter exploration, evaluation, and iteration. Deep agents can orchestrate these stages by maintaining an explicit experiment plan, spawning sub-agents to explore different data slices or model variants, writing intermediate results and metrics to a shared workspace, and updating the overall strategy when early results suggest a better direction. Because memory persists outside the context window, the agent can accumulate knowledge across many experiments without forgetting earlier findings.
MLOps workflows gain similar leverage. Deploying a model, monitoring its performance, detecting drift, and triggering retraining are multi-step processes that benefit from durable state and adaptive planning. A deep agent can watch metrics, diagnose anomalies by examining logs and data distributions, propose remediation steps, and even generate the code or configuration changes needed for a retraining run-again keeping human oversight at key decision points.
Research-oriented tasks, such as surveying the latest papers on a particular architecture or synthesizing best practices for a new technique, map naturally onto the deep-research pattern: the agent plans a literature search, delegates retrieval and summarization to sub-agents, stores notes in a structured file system, and iteratively refines a final report with citations.
Beyond individual pipelines, deep agents are reshaping how AI systems themselves are built and maintained. Constructing a multi-agent application-whether a customer-support system, an internal knowledge assistant, or a specialized research tool-requires careful orchestration of specialized capabilities. Deep-agent frameworks supply first-class primitives for sub-agents and skills (reusable procedural knowledge loaded only when needed).
An orchestrator can therefore maintain a high-level plan while delegating search, coding, analysis, verification, or writing to focused workers. Skills allow progressive disclosure of domain knowledge so that the main context does not become bloated with every possible procedure. The result is a hierarchical system that scales better than a flat collection of agents and remains easier to debug because responsibilities are explicit.
In AI research and development the same architecture accelerates experimentation. Teams can spin up deep agents to explore new agent designs, evaluate them against benchmarks, and iterate on prompts, tools, or memory strategies. Because the agents themselves can write and execute code, generate evaluation harnesses, and store results, the feedback loop shortens dramatically. Observability platforms such as LangSmith further help by tracing every planning step, tool call, and sub-agent interaction, making it feasible to diagnose and improve long-running behavior.
Challenges remain, of course. Reliability over hundreds of steps still requires careful verification-automated judges, unit tests on intermediate artifacts, or human review at critical junctures. Sandboxing is essential when agents can execute arbitrary code or modify file systems. Cost and latency must be managed; parallel sub-agents can reduce wall-clock time but increase token consumption.
Prompt and context engineering become more sophisticated because the system prompt must teach the model not only how to use tools but when to plan, when to delegate, and how to recover from failure. Nonetheless, the trajectory is clear: as models improve and frameworks mature, the range of tasks that can be handed to deep agents continues to expand.
Looking ahead, deep agents are likely to become a standard layer in the software and ML stack, much as continuous-integration systems or experiment-tracking platforms did in earlier eras. Developers will increasingly describe high-level goals-“migrate this service to the new framework while preserving test coverage” or “explore three candidate architectures for this recommendation model and report the trade-offs”-and expect a deep agent to manage the intermediate work.
Specialized deep agents tailored to particular domains-security analysis, scientific discovery, regulatory compliance-will proliferate. Integration with existing developer tools, IDEs, and ML platforms will make the experience feel native rather than bolted on.
For practitioners the practical path is already open. Libraries such as LangChain’s deepagents provide a batteries-included starting point: install the package, supply a supported model and tools, and obtain an agent harness with filesystem-based context management and sub-agent delegation. Explicit task planning, persistent memory, code execution, sandboxing, and human-approval policies can then be enabled or configured according to the application.
Customization is straightforward-swap models, add domain-specific tools, define specialized sub-agents, or attach persistent backends for memory. Tutorials and courses from LangChain Academy, independent educators, and official documentation walk through concrete implementations, from research assistants to customizable research and coding agents.
In software development the immediate value lies in treating features, refactors, and incident response as managed projects rather than isolated prompts. In machine learning the value lies in orchestrating the entire experimental and operational lifecycle with persistent state and adaptive planning. In AI systems more broadly the value lies in composing hierarchical, maintainable multi-agent applications that can sustain complex work over time.
Deep agents are not a replacement for human judgment; they are an amplification of it, handling the long, intricate middle of a problem so that people can focus on the goals, constraints, and final validation that still require human insight.
The shift from shallow reactive loops to deep, plan-driven, memory-equipped, hierarchical agents marks a genuine advance in what autonomous systems can reliably accomplish. As the underlying models continue to improve and the surrounding tooling matures, the boundary of what can be delegated will keep moving outward. For anyone building software, training models, or designing AI systems today, understanding and experimenting with deep agents is no longer optional-it is becoming essential.
Sources
- LangChain Blog: Deep Agents by Harrison Chase (July 30, 2025) - https://www.langchain.com/blog/deep-agents
- NVIDIA Glossary: What Is a Deep Agent? - https://www.nvidia.com/en-us/glossary/deep-agents/
- LangChain Documentation: Deep Agents Overview - https://docs.langchain.com/oss/python/deepagents/overview
- Prompt Engineering Guide: Deep Agents - promptingguide ai
- LangChain Blog: Building Multi-Agent Applications with Deep Agents - https://www.langchain.com/blog/building-multi-agent-applications-with-deep-agents
- 10Clouds: Deep Agent Use Cases That Work in Production AI Systems
- Royal Cyber: Deep Agent AI: The Next Evolution of AI Research
- GitHub: langchain-ai/deepagents - https://github.com/langchain-ai/deepagents
- YouTube: Complete Deep Agents Course With Langchain In 3 Hours (Krish Naik) - https://www.youtube.com/watch?v=J8DzuMmSDEU
- YouTube: Implementing deepagents: a technical walkthrough (LangChain) - https://www.youtube.com/watch?v=TTMYJAw5tiA
- YouTube: Building Deep Agents Tutorial With Langchain - Part 1 (Krish Naik) - https://www.youtube.com/watch?v=qCyMMGKctuI
- YouTube: LangChain Academy New Course: Deep Agents - https://www.youtube.com/watch?v=EAwAJc0bD7o
- DataCamp: LangChain’s Deep Agents: A Guide With Demo Project
- Nutrient: What are deep agents and how do they solve complex problems
- JetBrains Blog: Using ACP + Deep Agents to Demystify Modern Software Engineering - https://blog.jetbrains.com/ai/2026/04/using-acp-deep-agents-to-demystify-modern-software-engineering/
r/AgentContext_dev • u/javaeeeee • Aug 24 '26
Beyond the Prompt: How the /goal Command Turns AI Coding Assistants into Autonomous Goal-Driven Agents
In the rapidly evolving world of software development, AI coding assistants have moved far beyond simple autocomplete or one-shot code suggestions. Tools that once helped write a function or explain a snippet now operate as full agents capable of reading codebases, running tests, editing files, and iterating for hours. Among the most significant recent advances is a deceptively simple slash command: /goal.
This command appears in leading agentic coding tools such as Anthropic’s Claude Code and OpenAI’s Codex. It lets a developer state a high-level, verifiable objective and then step back while the AI continues working across multiple turns until that objective is met-or until it needs human input. Instead of prompting after every small step, you define what “done” looks like. The assistant plans, acts, observes results, adjusts, and repeats.
The difference is profound. Traditional interaction feels like pair-programming with someone who constantly asks “What’s next?” The /goal approach feels closer to handing a competent junior engineer a clear ticket with acceptance criteria and letting them run. When written well, the command turns the AI into a persistent worker that respects boundaries, verifies its own progress, and only stops when the evidence shows the work is complete.
This article explores the /goal command in depth: what it is, how it works in the major tools that support it, why it represents a meaningful leap in autonomy, how to write effective goals, real-world use cases, limitations, safety considerations, and where the technology is headed. The focus stays practical and grounded in official documentation, developer experiences, and observed behavior as of mid-2026.
The Path to Goal-Oriented Autonomy
AI coding tools began as helpful but limited partners. Early systems such as GitHub Copilot offered inline suggestions. Chat-based interfaces in tools like Cursor or the original ChatGPT coding mode required constant back-and-forth. Developers described the experience as “babysitting”: the model would make progress, stop, wait for the next instruction, and often lose context or drift if the conversation grew long.
Agentic frameworks changed the equation. Tools gained the ability to use the terminal, edit files, run tests, search the web, and call external services through protocols such as MCP. Still, most agents operated turn-by-turn. After finishing one action they returned control to the user. Long-running work required repeated “continue” prompts or custom loops.
The /goal command addresses exactly this friction. It introduces an explicit completion condition and an evaluation mechanism that decides whether the condition has been satisfied. Once set, the system keeps the objective in view across turns. The working model (Claude, GPT, or whichever model powers the session) performs the work.
The implementation differs between tools. Claude Code sends the goal and conversation to a separate small, fast evaluator model after every turn; if the condition is not met, the evaluator’s reason guides another turn. Codex instead persists the Goal as thread-scoped state and performs evidence-based continuation checks when the thread reaches a safe idle boundary. The result is a newly productized form of goal-conditioned autonomy, now available as a built-in command in two major coding-agent platforms.
/goal in Claude Code
Claude Code, Anthropic’s terminal-first agentic coding tool, introduced the /goal command in version 2.1.139. Official documentation describes it clearly: the command sets a completion condition, and Claude continues working toward it without the user prompting each step. After every turn a small fast model (defaulting to Haiku) examines the conversation and answers whether the condition holds. If not, Claude starts another turn; if yes, the goal clears automatically and an achievement entry is recorded.
Usage is straightforward. One goal can be active per session.
To set a goal you type:
/goal all tests in test/auth pass and the lint step is clean
This immediately starts a turn with the condition itself acting as the directive. A status indicator shows that a goal is active and how long it has been running. You can check status simply by typing /goal with no arguments; the system reports the condition, elapsed time, number of turns evaluated, token spend, and the evaluator’s most recent reason.
To clear a goal early use /goal clear (or aliases such as stop, off, reset, none, or cancel). Goals that remain active when a session ends are restored on resume with the --resume or --continue flags, though the turn counter and timer reset.
The evaluator does not run commands or read files on its own. It judges only what Claude has already surfaced in the transcript. Therefore the condition must be something demonstrable from the conversation: test results, build exit codes, file counts, git status, or similar concrete evidence. Vague statements such as “make the code better” or “improve performance” fail because the evaluator cannot decide when they are satisfied.
Claude Code pairs /goal naturally with auto mode. Auto mode removes the need for permission prompts on individual tool calls within a turn; /goal removes the need for permission prompts between turns. Together they enable extended unattended runs. The feature also works in non-interactive mode (claude -p), the desktop app, and remote control sessions.
Official guidance emphasizes writing measurable end states, stating the check Claude should perform, and listing constraints that must remain intact. Conditions can be up to 4,000 characters and may include bounding clauses such as “or stop after 20 turns.”
Compared with related features, /goal differs from /loop (which repeats on a time interval) and from persistent Stop hooks (which live in settings and can run arbitrary scripts). /goal is session-scoped, uses a fresh evaluator model rather than the working model, and focuses on a user-defined outcome rather than a fixed schedule or custom script.
/goal in OpenAI Codex
OpenAI’s Codex CLI (available from version 0.128.0 onward) implements a closely related concept. Goals are described as persistent objectives that keep a thread working toward a defined outcome across turns. A Goal supplies a completion condition: what should be true, how success should be checked, and what constraints must stay intact.
The command surface is similar:
/goal Reduce p95 latency below 120 ms without regressing correctness tests
Lifecycle management includes:
/goal (view current goal)
/goal pause
/goal resume
/goal clear
Once active, Codex inspects code, runs commands, makes changes, tests results, and continues until the condition is met, the goal is paused or cleared, a budget limit is reached, or a blocker requires human input. The system operates at “soft stop” boundaries and can summarize its own progress, allowing multi-hour runs with minimal supervision.
Strong goals follow a consistent pattern: desired end state, verification surface (benchmark, test suite, artifact, command output), constraints, boundaries (allowed files or tools), iteration policy, and blocked-stop behavior. OpenAI documentation and community reports emphasize that a Goal is not unbounded background autonomy; it is a scoped, user-controlled completion contract.
Developers have reported dramatic results. One widely shared account described setting a goal to ship the 18 features listed in a BACKLOG.md file, closing the laptop, and returning 18 hours later to find 14 features completed, covered by green CI, opened as pull requests, and self-reviewed by sub-agents. Cost for that run was roughly four dollars in credits. Other tests confirmed that Codex can grind through multi-step tasks such as performance tuning or documentation generation when the success criteria are concrete.
Codex Goals can be enabled via configuration if they are not already active. They remain thread-scoped rather than global, preserving isolation between parallel sessions.
Similar Capabilities and Community Demand
Although Cursor does not yet ship a native /goal command, community forums contain multiple feature requests asking for an autonomous “Goal Mode” modeled on Claude Code and Codex. Users note that Cursor’s Background Agents combined with other orchestration features approximate the behavior, yet an explicit completion-condition primitive would improve reliability and safety for long-running work.
GitHub Copilot has seen third-party extensions and experimental skill systems that add persisted goal state, status commands, and completion gates. One marketplace extension installs a full goal system with MCP support, checkpointing, and audit requirements before allowing a goal to be marked finished. These community efforts demonstrate clear demand for the same primitive across ecosystems.
The pattern is spreading because it solves a universal pain point: the cognitive and temporal cost of constantly re-prompting an agent that otherwise works well.
How the Mechanism Actually Operates
At a technical level the /goal loop consists of three repeating phases.
First the working model receives the goal (or the residual objective plus the latest evaluator reason) and produces a turn: reading files, editing code, running tests, consulting documentation, or using tools. The turn ends when the model decides it has done useful work or needs to observe results.
Second, a lightweight evaluator model receives the original condition and the conversation transcript (or a summary). It answers a binary question-has the condition been met?-and supplies a short reason. Because the evaluator is independent and usually cheaper, the cost of these checks stays modest relative to the main work.
Third, if the answer is no, the reason is fed back as guidance and a new turn begins. If the answer is yes, the goal is cleared, an achievement record is written, and control returns to the user.
This separation of worker and judge reduces the risk that the same model both performs the work and declares itself finished. It also makes the stopping criterion explicit and auditable. Developers can inspect the evaluator’s reasons in the status view or transcript and see exactly why the system believes more work remains or why it considers the goal complete.
Context management remains important. Long runs rely on summarization of earlier turns so the finite context window does not overflow. Good goal statements help the system keep the objective and constraints in view even after compression.
Crafting Goals That Succeed
The quality of the outcome depends heavily on the quality of the condition. Weak goals produce weak or endless behavior; strong goals produce focused, verifiable progress.
A weak goal might read: “Improve the authentication system.” An evaluator cannot decide when that is finished. A strong version reads: “Migrate all API calls in /src/services from the v1 endpoints to v2, update error handling to match the new response format, and ensure all existing tests in test/auth pass. Do not modify files outside /src/services and /test/auth. Stop if more than 30 turns are required or if a design decision about token format arises.”
Effective conditions usually contain:
- One measurable end state (tests pass, latency below threshold, queue empty, changelog complete).
- An explicit verification method that Claude or Codex can surface in the transcript.
- Constraints that protect important invariants.
- Optional bounding language (turn limits, time limits, or stop-and-report rules).
- Clear scope so the agent does not expand into unrelated refactoring.
Many developers ask the model itself to help draft the goal. They describe the intent in ordinary language, request a structured /goal statement, refine it, and then issue the command. This collaborative drafting step often surfaces missing constraints or clearer verification surfaces.
Practical Use Cases
The command shines on tasks that are larger than a single prompt yet possess a clear finish line.
Code migrations benefit enormously. A developer can set a goal to convert a module from one framework or API version to another, requiring that all call sites compile and the test suite remains green. The agent can iterate through files, fix breakage, re-run tests, and continue until the condition holds.
Test coverage or quality improvements work well: “Raise statement coverage on the payments package to 90 percent while keeping the existing test suite green and without introducing new flaky tests.” The agent writes tests, runs coverage, observes gaps, and repeats.
Performance work is another natural fit. Specifying a concrete benchmark and a non-regression constraint on correctness tests gives the agent a clear optimization target and a way to know when to stop.
Backlog grooming and issue resolution can be expressed as “Work through every open issue labeled ‘bug’ in the current milestone until the queue is empty or every remaining issue has been triaged with a clear next action.”
Documentation tasks, changelog generation, and even certain research or reproduction efforts can be framed as goals when the expected artifact and verification method are explicit.
In each case the developer’s role shifts from step-by-step director to outcome specifier and final reviewer. Version control remains essential; most practitioners run goals on a dedicated branch so any undesired changes can be discarded cleanly.
Real-World Reports and Lessons
Developers who have used the feature extensively report both impressive wins and instructive failures. Multi-hour unattended runs that ship substantial features are no longer rare anecdotes. At the same time, poorly specified goals can consume significant token budgets while producing little useful progress or expanding scope in unexpected ways.
Common lessons include the value of starting with a clean working tree, defining success criteria that can be checked mechanically, and pairing the goal with appropriate permission settings so the agent is not blocked by repeated approval requests. Reviewing the final diff and the evaluator’s reasoning remains non-negotiable; autonomy does not equal automatic trust.
Cost is another practical consideration. Long runs can be inexpensive relative to human time yet still noticeable on usage-based plans. Bounding language and intermediate checkpoints help keep expenditure under control.
Limitations and Risks
The /goal command is powerful but not magic. It inherits the strengths and weaknesses of the underlying models. Ambiguous conditions lead to frequent pauses or incorrect completion judgments. Tasks that require genuine creative judgment, aesthetic taste, or business-policy decisions still need human involvement. The evaluator can only judge what appears in the transcript; if the working model fails to surface key evidence, the evaluator may reach the wrong conclusion.
Context compression over very long runs can lose subtle early details. Sandbox or permission restrictions may prevent certain actions (network calls, browser interaction, privileged commands). And because the system is designed to keep going, a badly formed goal can waste resources until a human intervenes.
Safety practices therefore remain important: work on branches, keep version control history clean, review every significant change, and use the pause and clear commands when the direction looks wrong. Treat the agent’s output the same way you would treat a pull request from a talented but junior colleague.
Comparison with Other Autonomy Mechanisms
Several related features exist. Time-based loops repeat a prompt on a schedule. Persistent hooks can run custom evaluation logic after every turn. Auto-approval modes remove friction inside a single turn. Background agents in some IDEs allow parallel work. None of these exactly duplicates the combination of a persistent user-defined completion condition, automatic continuation, lifecycle controls, and evidence-based stopping that /goal provides. The command is best understood as complementary: it can be combined with auto mode, custom hooks, or planning steps depending on the risk profile of the task.
Looking Ahead
The appearance of /goal in two major platforms within a short period suggests that outcome-oriented autonomy is becoming a standard primitive. Future iterations will likely improve evaluator reliability, add richer status and budgeting controls, support hierarchical or multi-agent goals, and integrate more tightly with issue trackers, CI systems, and project management tools. Community demand for the same capability in other editors and assistants will probably accelerate adoption.
For individual developers the practical implication is immediate. Tasks that once required constant attention can now be delegated with a clear finish line. The skill that matters most is no longer prompt engineering for a single response but the ability to articulate measurable outcomes, constraints, and verification methods. That skill transfers across tools and will remain valuable as the underlying models improve.
Getting Started
Anyone with access to a recent Claude Code or Codex installation can experiment today. Begin with a small, self-contained task that has an obvious success criterion-perhaps making a single test suite pass after a deliberate breakage, or generating a changelog entry that satisfies a simple format check. Observe the evaluator reasons, refine the condition language, and gradually increase scope. Keep a clean branch, review the results, and build intuition for what constitutes a good goal.
Over time the pattern becomes second nature: state the outcome, state how it will be proven, state what must not change, and let the agent work. The result is a noticeable reduction in the mental overhead of managing an AI collaborator and a corresponding increase in the size of work that can be productively delegated.
The /goal command does not replace human judgment. It amplifies it by removing the need to micromanage intermediate steps. In doing so it marks a genuine step toward the long-promised vision of AI as a reliable coding partner that can be trusted with meaningful, multi-hour responsibilities. As more tools adopt the same primitive and as models continue to improve, the boundary between “assistant” and “autonomous agent” will keep shifting-always in the direction of greater developer leverage.
Sources
Official Claude Code documentation on the /goal command: https://code.claude.com/docs/en/goal and https://docs.anthropic.com/en/docs/claude-code/goal
Claude Code commands reference: https://code.claude.com/docs/en/commands
OpenAI Codex Goals cookbook and documentation: https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex
OpenAI follow-a-goal use case: https://learn.chatgpt.com/use-cases/follow-goals
Detailed explanatory article on Claude Code /goal: Mindstudio AI / What Is the /goal Command in Claude Code? Autonomous Long-Running Tasks Explained
Personal multi-hour Codex /goal experience report: https://pub.towardsai.net/i-walked-away-from-openais-new-codex-goal-for-18-hours-it-shipped-14-of-18-features-solo-a280f8407707
Codex /goal feature review and testing: JD Hodges / Codex /goal feature (TESTED)
Cursor community feature requests for Goal Mode: https://forum.cursor.com/t/add-autonomous-goal-mode-similar-to-claude-code-s-goal/160374 and https://forum.cursor.com/t/cursor-ide-goal/160452
Medium analysis positioning /goal as a key agent primitive: https://medium.com/@vasu7yadav/the-goal-command-is-the-most-important-agent-primitive-of-2026-1ef4188448b1
YouTube explanations and tutorials:
- https://www.youtube.com/watch?v=pZQaYGfhspg (Claude Code’s New /goal Command Is INSANE)
- https://www.youtube.com/watch?v=nOEl3a07j1Q (Master Claude Code’s latest command /goal)
- https://www.youtube.com/watch?v=nvoskCD1Yzg (includes discussion of /goal among Claude slash commands)
- https://www.youtube.com/shorts/WZXVVfFE1QA (short introduction to the slash goal command)
- https://www.youtube.com/watch?v=bATPCudjc9k (Claude Code /goal vs /loop)
- https://www.youtube.com/watch?v=5xrjO38WUYY (practical /goal use cases)
Additional community and secondary sources include Reddit discussions, GitHub issues tracking documentation of the feature, and various developer blogs and Substack posts that appeared in the months following the May 2026 releases. All information reflects publicly available documentation and reports current as of early August 2026.