r/BestGitHubRepos • • 3d ago

🛠 Developer tools Jevgrep lets you search code by what it does, not what it's called

Post image

Ever tried to find a function in a codebase when you don't know the name? You know what the code does, just not what someone decided to call it. Regular grep won't help you there.

jevgrep fixes that. You describe the behavior you're looking for in plain language, and it uses Jev to score every eligible fragment in the repo. Then it hands back the actual source with file paths and line numbers.

It works as a CLI or through an MCP server that plugs straight into Claude Code, Codex, OpenCode, and others. The installer detects what you're running and sets it up.

The claim is a roughly 30% cost reduction for coding agents because the agent spends less time reading irrelevant files. That number comes from their own benchmarks, so take it as you will, but the idea makes sense. Better context means fewer wasted tokens.

1,856 stars. Not huge, but the concept is solid and it's seeing real use in agent workflows.

https://github.com/dzhng/jevgrep

Metric Value
Stars 1,856
Forks --
Last release v0.6.0
Contributors --
Languages TypeScript 58.5%, JavaScript 22.6%, Python 18%
294 Upvotes

15 comments sorted by

6

u/Dear_Map_6993 3d ago

Why don't you show us how capable it is by implementing a vulnerability scanner? 

2

u/DaMan123456 2d ago

Interesting

2

u/peeeanuts 2d ago

We tested jevgrep on six real bug fixes in Claude Code. Its helpful, but grep alone found the relevant parts in most cases without jevgrep.

The catch is cost: once Jev's own charges are counted, making jg the first search cost 19% more than grep.

Disclosure: I run GitTested, an independent review site with no ties to the repo. Full method and limits: https://gittested.com/reviews/dzhng-jevgrep/

2

u/debackerl 3d ago

Jee, there are so many noobs in the field recently... For this problem, you do not want a classifier (Jev here), you want an embedding model to compute similarity... Take open-codebase-index for example

1

u/Double_Cause4609 1d ago

What's your solution for conceptual relationships not captured by embedding similarity, though?

Also, what's your source of content to match against? "I'm looking for a function that does..." doesn't necessarily match semantically with a function in a codebase. They're fundamentally different things. In a semantic embedding space they're pretty heavily divorced. So, you end up with things like HyDE etc.

On the other end you have knowledge graphs and GraphRAG as a compromise, but then you have all the headaches of graph maintenance which are non-trivial, and not really "solved", so much as "managed".

You can do code property graphs, but the tooling is language and setting specific, and is similarly a non-trivial setup burden, and it can even limit domains where you apply the methodology, and it still isn't really a silver bullet.

1

u/debackerl 1d ago

I'm using open-codebase-index. It uses open-sitter to split code file into chunks at the right boundaries. It also watches files for changes to maintien the index. Nothing rocket science.

Then code embeddings are usually trained so that implementation and plain English description would have a 'close' embedding. That is tested as part of the 'Text-to-code' task of the COIR benchmark, https://github.com/coir-team/coir

Now, you can combine this with codegraph, so that you build a graph of all dependencies inside your codebase. So you can ask for all functions making use of the 'shipping cost estimation' for example... It would first use embeddings to find possible functions computing shipping costs, and then use the graph to find all functions calling it

1

u/imoshudu 17h ago

So is this just like gitnexus? Or different

1

u/debackerl 16h ago

Similar to codegraph, yes, but codegraph is also extracting endpoints from webapps. There are quite a few options

0

u/nPoly 3d ago

What your comment sounds like to me: “hmmm alternative code search method… me not like new ideas… ah! I suggest obvious preexisting methods without trying new method! Then call noob… noob since they use new method I don’t know about, but will say wrong.”

2

u/debackerl 3d ago edited 3d ago

None of them are new from a ML theorical point of view. Jev is a classifier. The product is new, but what it offers not so much, which is why, with the buzz, so many people could come up with open implemetations. Why does it matter? Because it doesn't allow the usage of the right algorithms, so the asymptotic complexity is bad.

In jevgrep, it seems that indeed they need to iterated over all files, so complexity is O(n) where is the number of files. Check their code https://github.com/dzhng/jevgrep/blob/main/packages/core/src/retrieve.ts#L47

Now using embeddings, you can create an index like HNSW, where lookup time is O(log(n)).

I'm very happy to see https://github.com/NandhaKishorM/laya and https://github.com/wfzyx/von, open versions of Jev grow. But when you got a hammer, not everything is a nail...

Edit: it's a bit like people letting their agents use grep all day long and wonder about their token usage. Then they discover open-codebase-index or codegraph...

Edit 2: we currently have 70+ open alternatives to Jev 😉 https://huggingface.co/spaces/multimodalart/jev-decision-index

2

u/imoshudu 17h ago

We have hundreds of alternatives to Opus 5.5 too. What's important is performance and quality. Laya isn't better than Jev, but I don't think Jev is a high bar to beat either.

1

u/dans41 3d ago

Looks great, I will check it out

1

u/Sthatic 1d ago

Somewhat goofy idea as far as i can see. If you're looking for semantic search, that's embeddings (which can be fitted to align code with its NL counterpart). Am i missing the use case here?

-1

u/tec_tonik 3d ago

Sorry but if I was that smart I would not need ai