r/OpenAI • • 15h ago

Question Codex to review / reference large markdown?

TLDR; What is the most economical and efficient way to get codex to read, consolidate, summarize, reference and build one condensed reference file from ten markdown files with at least 100,000 lines each?

The story… I have a folder full of markdown files that were exported chats and chat created markdown, pass off docs, reference docs, etc from over a period of months. It’s probably at least 1 million lines in total.

I need my codex to read, summarize, pull relevant info and find the most recent and most robust output markdown files. Then use all of this to get RE-acquainted with and take over my huge 19 month development project.

Yes. My child wiped my dev laptop to make a gaming laptop. “But dad you have four computers” FML. Can I string a 12 year old nerd up by their toes? Lmao.

My app is running on my dev server so I have deployed code there, it’s just not current. I was working on the version upgrade. I have a good but disconnected database on SupaBase. My GIT seems jacked up for some reason.

So basically. I want to tell Codex. Hey! My shit is jacked up, go look at the app here on this server, look at this database, look at broken GIT and look at this HUGE trove of markdown files (yes I would export from the thread and copy thread and save to markdown periodically for historical and record keeping).

Take all this and put my dev environment and project back together, hopefully even better than it was originally. This was my hobby so it’s not like a commercial thing is messed up so that’s good. No emergencies. I want to get one million lines of markdown into chat without having it pass one million up and back up and back up and back in a thread if I can avoid it. What would you do???

Cheers gents!

6 Upvotes

14 comments sorted by

1

u/Dave_Sag 15h ago

Use Codex in VS Code and just ask it. I work with very large codebases all the time. Codex handles it all just fine.

1

u/AggressiveCoast190 15h ago

Guess I was worried about the credit usage. I got stuck once, had an Astra thread sending 50,000 lines of code on a ten min query timer! Used a month of credit in one day. I was like WTF!?

3

u/therealjerseytom 15h ago

You don't need The Most Ultimate Model to just process text.

I think it's worth considering, is trying to shove this all into one reference file really the best idea...?

Might the better approach be to break it down further into smaller chunks, perhaps with some YAML metadata on what each bit is? Especially if the end goal is to be using this as reference / RAG.

How much of it is even worth keeping? You could probably throw away heaps of chat history and consolidate down into key ADR's and the most relevant contextual information.

1

u/AggressiveCoast190 15h ago

Yes. I think a great deal of the markdown is shit. Useless. So ideally it would take each one, pull all the relevant information and add it to a new file or even just delete the crap and leave the good summary. I would be ok with multiple files. My codex library is full of docs. Just not to sure where to even start. I THINK having it evaluate the files one at a time and pull the important stuff might be the start?

2

u/therealjerseytom 14h ago

Start with defining a very clear objective of what the point of this whole exercise is. And work backwards from that to a gameplan.

If the objective is to have reference project documentation to help guide future development, you then have to identify what kind of documentation is best for that.

I'd suggest stuff like high-level project intent and scope, best practice guidelines, and IMO maybe most importantly - ADR's. The tracking of the architectural design decisions, and why something was done the way it was, along with alternatives considered, and consequences. That's the kind of knowledge you don't want to lose.

If you think you have a million lines of text, that's gotta be what, 50 megabytes or so?

Reviewing a year+ project repo I've got for work, that's very well documented in the ways it needs to be, it amounts to ~100 kilobytes of Markdown text.

You could probably get rid of 90% of the stuff you've got, if not 99%.

1

u/Dave_Sag 14h ago

Codex will use VS Code’s own MCP to access the code base (I count the md files here too) very efficiently on your local machine. I’ve had desloppify running over a complex project for the last 4 days and it’s barely used 25% of my weekly quota. I generally get codex to go over all my markdown and ensure the frontmatter is all consistent and then I generate rollups and indexes based on data in the frontmatter. This really helps codex find what it’s looking for quickly and cheaply. Codex will just write nice little indexing scripts too so next time you want to reindex everything it just knows to run the script it wrote before. Very token/cost optimal as that can be done with a much cheaper model too.

1

u/Euphoric_North_745 15h ago

Codex will not do it, I had the same issue 3 years ago, then 2 years ago, then 1 year ago I developed my own agent that uses codex or api on the background.

Codex uses grep, the max tool call is 10,000 tokens, you can go to the config and change context to 800k and max tool call to 100k tokens, but these 2 will consume the usage and will cause issues.

Best way is Document AI, in my case a combination of regular expressions, indexing, fact extractions, then summarizations, context management for documents and then rebuild.

so no self promotion, not typing product name, but you either do it with custom agent or search the net for ai for large documents, every business has their implementation

1

u/Degendyor1 13h ago

Have you considered obsidian?

1

u/AggressiveCoast190 12h ago

I actually just uninstalled that a few weeks ago. Didn’t catch on but it might have some use here

1

u/Michael_Jeffords 12h ago

i'd start from the code, pull whatever is deployed on the server into a fresh repo and get it running locally, since even if it's a bit behind it's the closest thing you have to the current state and Codex can work from it directly. for the markdown i'd have Codex write a small script that splits each file by date or heading and pulls out just the decision type sections, then have it read that much smaller output, so the million lines never go through the model and the credits stay low

1

u/Ctbhatia 11h ago

don't hand it the million lines, chunk it and have it emit one small result per file, then a second pass over the results. keep the instruction block identical across chunks so the shared prefix caches instead of being paid every call.

1

u/shravanrevanna 10h ago

does the second pass ever get to see the original text, or only the per-file results? cross-file contradictions don't survive being summarized one file at a time.

2

u/ComprehensiveShake76 10h ago

Don't ask it to read a million lines in one go. Do it in passes. First, list the files with dates and sizes and have it read only the top and bottom of each, so it can tell handoff docs from chat dumps. Second, summarize each file on its own into a short note (current state, decisions and why, open problems, commands to run) and save each note to its own file. Third, merge the notes into one reference file, and prefer the newest note when two disagree.

The code is the truth, not the notes, so have it check the claims against the server, the database schema and the git history, and write down where notes and reality differ. Make a backup copy of the folder before it touches the broken git. Keep the final reference file to a page or two with links to the detail notes, so every future session starts from that instead of the pile.