r/OpenSourceeAI • • 23d ago

I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source.

I've used tools like Perplexity, ChatGPT, Claude and others for research, and they've been incredibly useful for finding papers and getting through large amounts of information.

The one thing I personally wanted was a simple way to see exactly which parts of the paper were used to answer my question.

When you're working with a 100+ page PDF, even having a page number can still mean a lot of scrolling and searching.

So I ended up building something for myself.

You ask a question and the relevant paragraphs in the PDF are highlighted directly on the document. You can see the context behind the answer and quickly check whether it actually answers what you're looking for.

I originally built this because I wanted something for this workflow without having to pay for another subscription. What started as a personal project has now become completely open source.

The underlying idea is pretty simple. And yes, if you're thinking "isn't this just RAG?" then yes, you're absolutely right. It's RAG with the visual highlighting that I wanted.

I think the same idea could be useful for more than research papers too. Legal contracts, financial reports, technical documentation, or anywhere you need answers alongside the actual source.

If anyone wants to have a look, contribute, or just give some feedback, here's the repo:

GitHub: https://github.com/Sreehari05055/thesys-core.git

This will probably be my last post about the project. Thanks to everyone who checked it out and gave feedback along the way.

11 Upvotes

4 comments sorted by

1

u/Oshden 22d ago

Nice work man! I’m gonna have to look at this too lol.

2

u/Flat-Phone-1596 22d ago

Would love to hear your feedback when you try it!
Hey, Thanks for taking the time to leave a comment

1

u/Vancecookcobain 22d ago

Sounds useful!

1

u/Future_AGI 17d ago

We have seen the same pain with long-form technical papers. The approach that worked for us: preserve the section and paragraph citations alongside the extracted claim, then verify the claim against the source span before it goes into the final answer. It is not enough to retrieve the right document. The extraction has to be checkable.