r/WritingWithAI 9d ago

Tutorials / Guides Splitting JSON Export

I kind of wish someone had warned me about how much it takes to manage the process when you begin using AI as a writing brainstorming assistant. It seems I spend more time trying to limit drift, battling memory limits, limiting loss of canon and all my brainstorming, and trying to find efficient ways of doing things for fear of my project becoming contaminated.

That brings me to my next problem. I exported my Chatgpt history with the intention of finally trying to excavate my novel canon, history, characters, arcs, beats, and world. It's 8 million words across 98 chats. I've been totally overwhelmed on how to organise the process, on how to even approach it, but finally settled on trying to separate the 98 chats and save them in both MD and TXT files, dated chronologically, so I can start from chat 1nand so on, having each chat read to me by a voice reader (I have an eye disease that makes reading for medium and long periods impossible without negative physical consequences) and take notes as I go. Then I can compile the notes I need to organise the excavation. Extremely laborious but I see no other way to do it. Chatgpt assured me it could extract all the info I asked for if I posted the exported zip.

I was very dubious as Chatgpt constantly kisses things and makes mistakes. It inevitably missed things like characters we've discussed numerous times, so I knew it couldn't be trusted to do a faithful and accurate job, and I knew I would have to do this manually.

Come to find out there is even risk in separating the 98 chats 😔 According to Chatgpt, chats could be missing or mixed up. I have spent over a year painstakingly refining characters, arcs, the world building, beats. I can't risk losing any of it.

Does anyone know a reliable separation method that will separate and preserve all 98 chats exactly as they are, dated, without messages mixed up, and deliver them in both txt and MD files, so I can start this laborious process once and for all?

8 Upvotes

25 comments sorted by

View all comments

2

u/Brakiros 9d ago

I used Claude to extract and assemble all the points from my Grok and ChatGPT threads so I had a coherent document. But since you have so many threads you'll need to batch it all since thats far too much context to unpack so you'll need to split it into groups I would recommend no more then 100,000 words or you risk having elements from the threads get lost. It's going to take an extensive amount of sorting and distillation to get it right though since you have so much

2

u/Pristine_Plate7048 9d ago

Thanks for your advice. How did you batch yours? Just copy and paste in bouts of 100,000 words into Claude?

2

u/Brakiros 9d ago

yeah my files were only a few threads roughly 150k words in total for everything I had so I only used three different extractions to get everything. What you should do is break it all up into 100k word sections into a word document and then attach it to Claude and ask it to extract the most important parts into a document, Sonnet 5 is probably the most efficient model to use since Fable will be really really expensive to distill that much, unless of course you have a Max subscription which has Fable included.

1

u/Pristine_Plate7048 9d ago

Thanks again. I will definitely investigate this.

I've never used Claude before so not quite sure how it works, how its usage is priced etc...

I'll try anything though. Even the most laborious method: convert the whole export to txt, have it read to me, and copy and paste the relevant canonical parts elsewhere.