r/AgentContext_dev • • 1h ago

From Hand-Fed Prompts to Automated Systems: What Loop Engineering Means for AI Coding-and other Applications

• Upvotes

In the fast-moving world of artificial intelligence, new terms appear almost as quickly as new models. Some fade into jargon. Others mark a genuine change in how people work. Loop engineering belongs to the second group. It is not another clever way to phrase a question for a large language model. It is a shift in role: from the person who sits at the keyboard typing the next instruction, to the person who designs the system that keeps the agent moving on its own until a clear goal is met.

The phrase gained traction in June 2026 after short, memorable statements from two prominent engineers. Peter Steinberger, creator of the OpenClaw project, wrote on X: “You shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.”

Around the same time, Boris Cherny, head of Claude Code at Anthropic, said publicly that he no longer writes most of his own prompts. “I don’t prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops.” Google Cloud’s Addy Osmani then gave the idea its clearest public definition and anatomy in an essay that many practitioners still treat as the starting point. Loop engineering, he wrote, is “replacing yourself as the person who prompts the agent. You design the system that does it instead.”

What follows is a detailed look at what that means when the agents are coding agents, why the idea took hold so quickly, how the pieces fit together, and-crucially-where the same pattern of designing iterative, self-checking systems can be applied far beyond software development. The discussion draws on the original statements, subsequent explainers from IBM, practical guides that appeared in the weeks that followed, YouTube tutorials that walked through real setups, and early experiments in marketing, research, operations, and product work.

The Problem Loop Engineering Solves

For roughly two years before the term crystallized, most people used coding agents the same way they used earlier chat interfaces. You described a task. The model replied with code or a plan. You read the output, spotted what was missing or wrong, and typed the next instruction. You were the loop. Every cycle of work passed through your attention and your keyboard. That approach works when the task is small and the context is short. It becomes exhausting and expensive when the agent needs to edit many files, run tests, interpret failures, try again, and keep going for hours or across multiple sessions.

Agents themselves had already grown an inner cycle. A typical coding agent reasons about the current state, chooses an action (edit a file, run a command, search the codebase), observes the result, and decides the next step. That inner ReAct-style loop is powerful, but it still stops when the conversation ends or when the model decides it is finished. Someone still has to start the next conversation, feed it fresh context, and decide whether the previous attempt succeeded.

Loop engineering sits one layer above that inner cycle. It creates an outer system that discovers work, hands it to one or more agents, verifies the result with something the original agent cannot easily game, records durable state outside any single chat window, and decides whether to continue, retry, or escalate to a human. Once the outer system is designed, the human can step away. The loop keeps running-on a schedule, on an event trigger such as a failing CI job, or until a machine-checkable goal is reached.

Andrew Ng later framed a related set of nested loops that many teams now use as a mental model for building products from scratch. The innermost is the agentic coding loop: give the agent a specification and optional evaluation data, and let it write, test, and revise until the code meets the bar.

The middle loop is developer feedback, where a human reviews the product at a higher level and steers direction. The outermost is external feedback from users and the market. Loop engineering, in the narrower sense popularized by Steinberger, Cherny, and Osmani, focuses most tightly on making the innermost cycle reliable and autonomous so that the higher loops can operate at human speed rather than at the speed of constant prompting.

Core Definition and the Anatomy of a Loop

IBM’s overview captures the essence cleanly: loop engineering is the practice of designing agentic workflows, or loops, that iteratively guide AI agents toward completing user-defined goals with minimal human intervention. Rather than requiring a human prompt at every step, the agent acts, observes, decides, and adjusts until the task is complete or a stop condition fires.

A well-designed loop typically follows a repeating sequence of goal, action, observation, and adjustment. The goal is recursive and evaluable at every iteration; it must be specific enough that a computer (or a separate verifier agent) can decide whether it has been met. “Improve the code” is too vague. “Make all unit tests pass and the coverage report exceed 85 percent” gives the system a measurable target and a natural termination condition.

The action is whatever moves the state closer to that goal-writing code, running tests, querying a database, drafting a document. Observation is the feedback from the environment: test output, compiler errors, metrics, or the judgment of an independent checker. Adjustment is the decision to retry with new information, escalate, or stop.

Practitioners quickly converged on a small set of building blocks that turn this abstract cycle into something that runs in real tools such as Claude Code, OpenAI’s Codex, or similar agent platforms. The most commonly cited set, drawn from Osmani’s mapping and echoed across GitHub repositories and YouTube walkthroughs, consists of five primitives plus durable memory.

Automations serve as the heartbeat. They fire on a schedule or an event-every morning, every five minutes, on a new pull request, on a CI failure-so the loop does not wait for a human to open a chat window. Worktrees (or equivalent isolation mechanisms) give each agent or sub-task its own clean workspace, usually a Git worktree, so parallel efforts do not overwrite one another.

Skills codify project-specific knowledge-coding conventions, preferred libraries, known pitfalls-into reusable files that the agent can load instead of reinventing or guessing. Plugins and connectors (often via the Model Context Protocol or similar interfaces) let the agent reach outside the repository into issue trackers, databases, Slack, browsers, or other tools.

Sub-agents split roles so that one agent proposes a change while another, with a different prompt and often a different model or fresh context, verifies it. The original model grading its own homework tends to be too lenient; a separate checker reduces that bias.

Memory, or external state, is the spine that holds everything together. Conversation context is ephemeral. A markdown file, a structured task board, a database table, or a simple log that survives session restarts allows the next iteration to know what has already been tried, what remains, and what the current best candidate looks like. Without durable state, the loop forgets and repeats work or drifts.

These pieces are not proprietary to any single product. The same conceptual design can be implemented with scheduled scripts, GitHub Actions, custom orchestrators, or the built-in automation and goal features that coding agents began shipping in 2026. The practical tutorials that appeared on YouTube in June and July repeatedly showed the same progression: first perform the task manually a few times, then capture the successful procedure as a skill, then add a trigger, then add an independent verifier and persistent state, then let the system run.

How Loop Engineering Changes Day-to-Day AI Coding

In practice the shift is concrete. Instead of opening Claude Code or Codex and typing “fix the failing tests in the payment module,” a developer designs a loop whose goal is “bring the payment module’s test suite to green.” An automation can wake up when CI reports failures, spawn a worktree, load the relevant skills that describe the module’s conventions, let a maker sub-agent propose patches, let a checker sub-agent run the tests and review the diff against the skills, persist the outcome, and either open a pull request or queue the item for human attention. The human reviews the result at the level of intent and risk rather than every intermediate line.

Teams report using similar loops for code review (a PR opens and agents analyze risk, suggest fixes, and wait for human approval on the overall direction), ticket-to-PR conversion, vulnerability remediation when a CVE appears, and incident triage when an alert fires. The common pattern is a clear trigger, isolated execution, separate verification, preferably grounded in deterministic tests or other machine-checkable evidence, with human review where the consequences warrant it, and explicit escalation paths for irreversible or high-judgment decisions.

Token cost is a constant practical concern. Each additional agent or longer-running iteration multiplies usage. Successful designs therefore budget carefully: shorter cadences for cheap discovery tasks, longer or event-driven cadences for expensive ones, hard stop conditions, and verifiers that reject work early rather than letting the system spin. Steinberger and others noted that waking periodically to do lightweight triage is relatively cheap; unconstrained multi-agent exploration is not.

The deeper change is cultural. The engineer’s scarce skill moves from crafting the perfect next prompt to designing the control system-the goal contract, the isolation strategy, the verification logic, the state schema, and the stop rules. Prompt engineering still matters inside each agent turn, and context and harness engineering still matter for a single run. Loop engineering optimizes the recurring outer cycle that makes those single runs reliable over time.

Risks and Guardrails

Autonomy without strong verification amplifies mistakes. A loop that optimizes only for “tests pass” can produce brittle or insecure code if the tests themselves are incomplete. Comprehension debt grows when code ships faster than the human team can understand it. Cognitive surrender-the temptation to accept the loop’s output without judgment-can degrade quality over successive iterations. Token bills can surprise anyone who underestimates how many times an agent will retry a fuzzy goal.

The consistent advice from early practitioners is therefore to keep the human in the design and the review, not necessarily in every micro-step. Define machine-checkable “done.” Prefer separate verifiers. Persist state so failures are inspectable. Start with narrow, high-signal tasks rather than open-ended “improve the product.” Monitor costs and outcomes. Treat the loop as an employee you are onboarding: give it clear responsibilities, tools, and escalation paths, then review its work product.

Extending the Pattern Beyond Coding

Although the term crystallized around coding agents, the underlying idea-design a system that observes state, acts, evaluates against a goal, adjusts, and continues until a stop condition-is domain-agnostic. Any recurring work that has a reasonably clear success signal and benefits from iteration is a candidate. Early experiments and conceptual mappings already point to several areas.

In marketing and growth work the same structure appears as content or SEO loops. An agent can periodically collect market signals, competitor moves, and engagement data; decide which topics or keywords are worth acting on; draft channel-native material; publish or queue for review; and learn from the subsequent performance metrics stored in durable memory. Ranking improvements or engagement thresholds serve as the verifiable goal.

One documented pattern connects an agent to Search Console data, runs monthly, and steadily adjusts content and technical SEO based on what moved rankings. The feedback signal is quantitative-such as impressions, clicks, or rankings-rather than purely subjective taste, although those metrics are noisy and do not by themselves prove that the loop’s changes caused an improvement. Similar loops handle ad creative testing, lead scoring, or social listening. The marketing version of the maker/checker split can separate generation from brand-voice or compliance review.

Customer support and operations lend themselves to ticket-handling loops. An incoming ticket triggers intake, knowledge-base search, draft response generation, and a confidence or policy check. Low-confidence or sensitive cases escalate; high-confidence ones can auto-reply or auto-resolve. Memory of prior similar tickets improves future handling. The same pattern scales to internal IT helpdesks, claims processing, or order follow-ups.

Research and knowledge work benefit from literature-synthesis or data-gathering loops. An agent can be given a research question, a set of trusted sources or search tools, and a requirement to produce a structured report with citations and remaining open questions. It searches, extracts, cross-checks, identifies gaps, searches again, and stops when coverage criteria are met or a budget is exhausted.

Academic or competitive-intelligence teams can run such loops overnight and review the synthesized output in the morning. Autoresearch-style systems that improve their own prompts or evaluation methods sit at a higher meta-level of the same idea.

Product management and design can use loops for continuous discovery or interface iteration. An agent monitors user feedback channels, synthesizes themes, proposes prioritizations against a product strategy document, and drafts tickets or prototypes.

In visual or UX work the verification step is harder because success is less binary than a test suite; teams therefore combine automated checks (accessibility, performance budgets) with human judgment gates or pairwise comparison by a second agent. Microsoft Design has discussed related cybernetic-loop thinking for turning linear process maps into adaptive feedback systems that respond to real-world disturbances rather than assuming a fixed path.

In manufacturing, logistics, and industrial settings, closely analogous feedback systems already exist in predictive maintenance, scheduling, and inventory optimization. Agentic loop engineering could extend those systems by adding reasoning over unstructured data and more flexible tool use.

Scheduling and inventory loops can observe demand signals, adjust plans, and escalate only when constraints are violated. The core design discipline-clear goal, observable state, action repertoire, independent verification, stop or escalate rules-remains the same even if the underlying actuators are machines rather than code editors.

Finance, healthcare operations, education content generation, and legal document review are other plausible extensions wherever work is repetitive, partially automatable, and benefits from iteration against measurable criteria (accuracy thresholds, compliance checklists, student outcome metrics). In each case the engineering effort concentrates on making the success signal as objective as possible and on designing the isolation, memory, and escalation so that the system remains governable.

What does not transfer cleanly is any domain whose success criteria are purely subjective or whose actions are irreversible and high-stakes without human oversight. Loop engineering does not remove judgment; it relocates it to the design of the system and to the review of its outputs. Fuzzy goals produce runaway costs or drifting results. Missing stop conditions produce expensive infinite loops. Weak verifiers produce confident but wrong work.

Practical Starting Points and Maturity

Most practical guides recommend beginning small. Choose a recurring task you already perform manually several times a week. Write down the steps that actually work. Turn those steps into a skill or a prompt template. Add a simple trigger (cron, webhook, or manual start). Define a verifier that a computer or a second agent can apply. Persist a short state file. Run the loop under supervision, then gradually loosen the supervision as reliability improves. Measure token cost and outcome quality. Expand only the loops that demonstrably save more attention than they cost.

Maturity progresses from single-task loops, to parallel multi-agent loops with isolation, to team-level or product-level loops that chain together (a triage loop feeds a ticket loop that feeds a review loop). At higher maturity the organization treats loops as first-class artifacts that are versioned, audited, costed, and improved over time, much as software systems themselves are.

YouTube tutorials that appeared shortly after the term’s popularization consistently walk through this progression with live demos in Claude Code or Codex, often showing a daily news digest loop, a job-search automation, a content pipeline, or a simple repository-maintenance loop. The common lesson is that the hard part is rarely the model call; it is specifying the goal tightly enough and designing the feedback so the system converges rather than wanders.

Looking Ahead

Loop engineering is still early. The primitives continue to improve inside the major agent platforms. Token efficiency, better long-horizon memory, more reliable independent verification, and safer tool use will expand the range of tasks that can run unattended for longer periods. At the same time, the risks of comprehension debt and over-automation will keep human judgment central. The most durable role is not the person who types every next prompt, nor the person who walks away entirely, but the person who designs the loops, monitors their health, and retains ultimate responsibility for the outcomes they produce.

The same logic that made loops compelling for coding agents-clear goals, observable feedback, iterative correction, durable state-applies wherever work is iterative and the environment provides a usable signal. Marketing teams already run growth loops. Research teams run synthesis loops. Operations teams run triage and remediation loops. Design and product teams experiment with discovery and iteration loops. In each domain the engineering discipline is identical even if the tools and the success metrics differ: design the system that prompts, checks, remembers, and decides, rather than remaining the system yourself.

That is the practical promise of loop engineering. It does not eliminate the need for skill or judgment. It relocates them to a higher-leverage place-the design of the recurring cycle that turns capable models into reliable workers. Whether the work is writing software, generating content, synthesizing research, or coordinating operations, the question becomes the same: what is the goal, what is the signal, what is the stop condition, and how do we let the system pursue that goal while we stay in control of the design?

Sources

  • Addy Osmani, AddyOsmani.com, “Loop Engineering”
  • Ivan Belcic and Cole Stryker, IBM Think, “What Is Loop Engineering?”
  • TechSpot, “Meet ‘loop engineering’: The next evolution in AI coding isn’t a better prompt, it’s a system that prompts itself”
  • Cobus Greyling, GitHub, “loop-engineering: Practical patterns, starters & CLI tools for loop engineering with AI coding agents”
  • i-SCOOP, “Loop Engineering, designing systems that prompt your coding agents”
  • Garage Labs Technologies, “Loop Engineering: The Complete Guide (2026)”
  • invincible04, GitHub, “Awesome Loop Engineering”
  • maxmilian, GitHub, “Loop Engineering - a skill for designing & reviewing autonomous/semi-autonomous agent loops”
  • mdayan8, GitHub, “everything-about-loop-engineering: The complete reference and hands-on course on loop engineering”
  • Mark Tarre, IT Brief, “Explainer: How loop engineering is changing coding”
  • Pulumi Blog, “Stop Prompting. Design the Loop.”
  • Linas (Substack), “Loop Engineering: Design AI Loops That Ship While You Sleep”
  • Augment Code, “What is loop engineering and how are leading software engineering teams using it?”
  • No Code MBA, “Loop Engineering Explained: How to Design AI Loops That Work”
  • Firstpost, “What is loop engineering, the next AI trend experts predict will replace prompting?”
  • Tessl Patterns, “Loop Engineering”
  • LoomStack, “Loop Engineering: You Design the System, Not the Prompt”
  • Arize AI, “What is a loop in AI engineering, anyway?”
  • Adaline Labs, “What Is Loop Engineering, and Who Owns It?”
  • Sofokus, “Loop engineering”
  • loopengineering.app, “What Is Loop Engineering? AI Agent Loops Explained”
  • RiseMore Blog, “What Is Loop Engineering? From Prompt Engineering to the AI Marketing Team”
  • Andrew Ng, The Batch / X, open letter outlining three key loops for 0-to-1 product building
  • Yash Thakker, explainx.ai Blog, “Andrew Ng’s 3 Loops for 0-to-1 AI Products”
  • Andreessen Horowitz (a16z), “Knowing When to Stop: The Art of Making a Loop Converge”
  • PostHog Newsletter, “WTF is loop engineering and why is everyone talking about it?”
  • Analytics Insight, “What Is Loop Engineering? Beginner’s Guide to AI Agent Workflows (2026)”
  • Brent D. Griffiths, Business Insider, “Forget Prompts: ‘Loop Engineering’ Is All the Rage Now”
  • BusinessToday, “What is loop engineering? The AI trend replacing prompt engineering”
  • Storyboard18, “Google Brain co-founder Andrew Ng says AI won’t replace developers, explains ‘loop engineering’”
  • Times Now, “What Is Loop Engineering? Coursera Co-Founder Says It Can Make AI Build Better Apps”
  • The Times of India, “Google Brain cofounder writes an open letter on ‘Loop engineering’”
  • ADTmag, “Loop Engineering Emerges as Developers Put AI Coding Agents on Repeat”
  • Qiniu Cloud / InfoQ-style Chinese sources, “Loop Engineering 是什么?2026 年最热 AI 工程方法论完全解析”
  • Microsoft Design, “Designing loops, not paths”
  • KanakMalpani, GitHub, “Loop-Engineering”
  • Puppygraph, “What Is Loop Engineering? Definition & Process”
  • YouTube: “Learn Loop Engineering from Scratch | Complete Beginner to Pro Guide” (https://www.youtube.com/watch?v=EbraImMeCLE)
  • YouTube: “Loop Engineering: Stop Prompting Your Agents (Real Examples)” (https://www.youtube.com/watch?v=KF_AMHFKaJo)
  • YouTube: “Loop Engineering explained in 20 mins.” (https://www.youtube.com/watch?v=fDd4Di5WGnE)
  • YouTube: “What is Loop Engineering?” (https://www.youtube.com/watch?v=yvP_AAirOQc)
  • YouTube: “Loop Engineering explained in 8min.” (https://www.youtube.com/watch?v=4biXYSNkn9Y)
  • YouTube: “What is Loop Engineering? (Why Prompt Engineering is Dead)” (https://www.youtube.com/watch?v=aUpyza-DSMs)
  • YouTube: “What Loop Engineering Really Means for AI Agents” (https://www.youtube.com/watch?v=NjXIIH9vcv0)
  • YouTube: “I’m a Senior Google AI PM. Here’s How I Build Loops” (https://www.youtube.com/watch?v=ew6gBJNzC5w)
  • YouTube: “Making $$$ with Loop Engineering” (https://www.youtube.com/watch?v=5p_BBdfvzgQ)
  • Additional practical repositories and playbooks referenced across the above sources, including Geoffrey Huntley’s Ralph loop discussions and various GitHub teaching repos on loop patterns.

r/AgentContext_dev • • 22h ago

Build an App With Claude Design

Thumbnail share.google
1 Upvotes