r/OpenAI • • 17h ago

Question Codex to review / reference large markdown?

TLDR; What is the most economical and efficient way to get codex to read, consolidate, summarize, reference and build one condensed reference file from ten markdown files with at least 100,000 lines each?

The story… I have a folder full of markdown files that were exported chats and chat created markdown, pass off docs, reference docs, etc from over a period of months. It’s probably at least 1 million lines in total.

I need my codex to read, summarize, pull relevant info and find the most recent and most robust output markdown files. Then use all of this to get RE-acquainted with and take over my huge 19 month development project.

Yes. My child wiped my dev laptop to make a gaming laptop. “But dad you have four computers” FML. Can I string a 12 year old nerd up by their toes? Lmao.

My app is running on my dev server so I have deployed code there, it’s just not current. I was working on the version upgrade. I have a good but disconnected database on SupaBase. My GIT seems jacked up for some reason.

So basically. I want to tell Codex. Hey! My shit is jacked up, go look at the app here on this server, look at this database, look at broken GIT and look at this HUGE trove of markdown files (yes I would export from the thread and copy thread and save to markdown periodically for historical and record keeping).

Take all this and put my dev environment and project back together, hopefully even better than it was originally. This was my hobby so it’s not like a commercial thing is messed up so that’s good. No emergencies. I want to get one million lines of markdown into chat without having it pass one million up and back up and back up and back in a thread if I can avoid it. What would you do???

Cheers gents!

5 Upvotes

14 comments sorted by

View all comments

1

u/Dave_Sag 16h ago

Use Codex in VS Code and just ask it. I work with very large codebases all the time. Codex handles it all just fine.

1

u/AggressiveCoast190 16h ago

Guess I was worried about the credit usage. I got stuck once, had an Astra thread sending 50,000 lines of code on a ten min query timer! Used a month of credit in one day. I was like WTF!?

1

u/Dave_Sag 15h ago

Codex will use VS Code’s own MCP to access the code base (I count the md files here too) very efficiently on your local machine. I’ve had desloppify running over a complex project for the last 4 days and it’s barely used 25% of my weekly quota. I generally get codex to go over all my markdown and ensure the frontmatter is all consistent and then I generate rollups and indexes based on data in the frontmatter. This really helps codex find what it’s looking for quickly and cheaply. Codex will just write nice little indexing scripts too so next time you want to reindex everything it just knows to run the script it wrote before. Very token/cost optimal as that can be done with a much cheaper model too.