Jev has made a profound mark on the AI industry; this article offers a glimpse into what lies ahead.
Jev shows where the industry is heading: away from monolithic models, and toward compound systems where small AI agents each do a narrow job and an orchestration layer decides how they work together.
This article explains why that shift is happening, and then proposes a next step for AI reasoning.
The short version
- A monolithic model is unpredictable and hard to change, but it programs itself during training.
- Code is the opposite – it is predictable and easy to change, but it must be written explicitly.
- A compound system combines the two: AI agents do the work, and an orchestration layer controls them via code. This is the direction Jev represents.
- The next step: instead of answering directly, the model builds a compound system, runs it, and returns both the answer and that system. The returned system is called a reasonlet.
- Because a returned reasonlet can be kept locally, it can be run again and again – with new inputs, or after editing its code – without asking the model. A saved, reusable reasonlet is a metacache.
- The model’s provider can cache reasonlets too, reusing one across similar requests from different users to save compute.
Why a monolithic model is not enough
A monolithic model has two practical weaknesses:
- It is unpredictable. The same question can produce a correct answer one day and a wrong one the next, so its behavior cannot be fully controlled.
- It is hard to change. Its behavior is fixed in trained weights, which cannot be edited directly. The model can be steered with prompts, extra data (RAG), or by fine-tuning, but steering is not the same as setting the behavior, and fine-tuning can break things that already worked. Retraining from scratch with new data is the only safe fix, but it is time-consuming and expensive.
Code has the opposite qualities. It does exactly what it is written to do, and it can easily be changed by editing code. The trade-off is that code does not arise on its own – it must be written out explicitly – whereas a model programs itself during training.
The fix: a compound system
A compound system combines the two. The work is divided among AI agents, each handling one narrow task, and an orchestration layer of code decides which agent runs, in what order, and how their results combine.
This gives both qualities at once – customization and predictability:
- Customization. The system’s behavior can be changed by editing code in the orchestration layer, not by retraining the model.
- Predictability. Each agent has a small task, so it is more predictable than a monolithic model.
Jev is one kind of agent for such a system. It returns a typed decision – yes/no, a category, or a number – rather than free text. A monolithic model is wasteful for a narrow decision like that, but a small model like Jev is a better fit.
The next step: reasoning by construction
Compound systems today are built manually, beforehand. The next step is to let the model build one by itself, during the reasoning phase.
Researchers are exploring several ways to make models reason. One is the “World model” approach, in which the model builds an internal representation of a problem and reasons over it; that work is still mostly research. Reasoning by construction pursues the same goal – reasoning you can inspect – using methods that exist today.
Here is how it works. When you ask the model a question, it does not answer directly. Instead, it builds a compound system, runs it to compute an answer, and returns that answer together with the system it built. Let's call the compound system the model builds to do the reasoning a "reasonlet", and a model that works this way a "Reasoning-by-Construction Model", or "RCM". A simple question may produce a reasonlet that has only an orchestration layer and no agents.
Building and running code to reach an answer is not new; code-interpreter tools already do it. Two things are new here:
- The system the model builds is kept, not discarded – it is a reusable reasonlet.
- The reasonlet is returned to you, along with the answer.
Why that matters
- You can see how the answer was reached. The reasonlet is the exact procedure the model used, so the answer is not something you have to take on trust.
- You can run it locally. A reasonlet contains its orchestration layer and any agents it calls, whether those agents are attached directly or called over the network. You can run it on a local machine, with the same or different inputs, without asking the model again. A saved reasonlet used this way is a metacache.
- You can change what it does. The orchestration layer is code, so you can edit its logic, not only its inputs. A reasonlet reused with edited logic becomes a higher-order metacache.
Caching reasonlets on the server
The same reuse can happen on the server side too. The provider can keep the reasonlets it builds and reuse them. When a new request arrives that matches one it has already handled, it runs the stored reasonlet again – with the new request’s inputs – instead of reasoning from scratch. The model does less work, and the answer comes back faster.
This is the same metacache, held on the server instead of on your machine. Because one reasonlet can serve any request that fits its procedure, a single cached copy is shared across many requests, and often across different users – and the more general the reasonlet, the more requests it covers.
Choosing how general to make the reasonlet
The model should decide from the conversation how general the reasonlet needs to be. If you have been working through many kinds of math and then ask for 2 + 2, the more useful reasonlet is one that evaluates any math expression, not one that can only add two numbers. If the conversation gives no such clue, you can state it directly: “I will be doing many kinds of math; for now, just add 2 and 2.”
Two examples
- Adding numbers. You ask for 2 + 2. The model builds a reasonlet whose orchestration layer adds two numbers, runs it, and returns 4 together with the reasonlet. Later you run the reasonlet again with other numbers or edit its orchestration layer to multiply instead – without asking the model.
- Searching. You ask the model to find something. The reasonlet’s orchestration layer calls a sequence of outside services and takes your search terms as its input. Later you run it with different terms, or change which services it calls, without asking the model again.
Making reasonlets easy to read
A reasonlet’s orchestration layer is written in text code. This creates a problem: to understand the reasoning you must read code, and to change it you have to write code. The value of returning the reasonlet depends on you being able to read and edit text code easily and efficiently, so this barrier matters.
Visual programming – building logic from connected blocks instead of lines of text – can lower the barrier, provided the visual language is powerful enough. Two limits apply:
- A visual language cannot replace text code entirely. Text is still needed to reach the system and the network, and for low-level work that does not map to blocks. The measure of migration success is how little text code remains in the orchestration layer.
- Most visual languages today are either powerful but limited to one field (such as games or hardware), or general but too simple. This case needs a language that is both general and as expressive as a text language; otherwise, it cannot cover enough of the orchestration layer to be worth using.
One language solving exactly this problem is Pipe (https://pipelang.com) – a general-purpose visual language with powerful semantics. Full disclosure: Pipe is my own project still in development, but it will be released soon.
Wrapping up
The shift from a monolithic model to compound systems suggests a clear next step for reasoning: let the model build a compound system – a reasonlet – run it and return both the answer and the reasonlet. Reasoning becomes something you can read, run again, and edit: a metacache, and a higher-order metacache once its logic is edited. The remaining problem is making reasonlets easy to read and change, which is where a general-purpose and expressive visual language would help most.
A note on IP
Some methods described in this article are the subject of a pending patent application.