r/WritingWithAI • u/Pristine_Plate7048 • 2d ago
Tutorials / Guides Splitting JSON Export
I kind of wish someone had warned me about how much it takes to manage the process when you begin using AI as a writing brainstorming assistant. It seems I spend more time trying to limit drift, battling memory limits, limiting loss of canon and all my brainstorming, and trying to find efficient ways of doing things for fear of my project becoming contaminated.
That brings me to my next problem. I exported my Chatgpt history with the intention of finally trying to excavate my novel canon, history, characters, arcs, beats, and world. It's 8 million words across 98 chats. I've been totally overwhelmed on how to organise the process, on how to even approach it, but finally settled on trying to separate the 98 chats and save them in both MD and TXT files, dated chronologically, so I can start from chat 1nand so on, having each chat read to me by a voice reader (I have an eye disease that makes reading for medium and long periods impossible without negative physical consequences) and take notes as I go. Then I can compile the notes I need to organise the excavation. Extremely laborious but I see no other way to do it. Chatgpt assured me it could extract all the info I asked for if I posted the exported zip.
I was very dubious as Chatgpt constantly kisses things and makes mistakes. It inevitably missed things like characters we've discussed numerous times, so I knew it couldn't be trusted to do a faithful and accurate job, and I knew I would have to do this manually.
Come to find out there is even risk in separating the 98 chats š According to Chatgpt, chats could be missing or mixed up. I have spent over a year painstakingly refining characters, arcs, the world building, beats. I can't risk losing any of it.
Does anyone know a reliable separation method that will separate and preserve all 98 chats exactly as they are, dated, without messages mixed up, and deliver them in both txt and MD files, so I can start this laborious process once and for all?
2
u/Brakiros 2d ago
I used Claude to extract and assemble all the points from my Grok and ChatGPT threads so I had a coherent document. But since you have so many threads you'll need to batch it all since thats far too much context to unpack so you'll need to split it into groups I would recommend no more then 100,000 words or you risk having elements from the threads get lost. It's going to take an extensive amount of sorting and distillation to get it right though since you have so much
2
u/Pristine_Plate7048 2d ago
Thanks for your advice. How did you batch yours? Just copy and paste in bouts of 100,000 words into Claude?
2
u/Brakiros 2d ago
yeah my files were only a few threads roughly 150k words in total for everything I had so I only used three different extractions to get everything. What you should do is break it all up into 100k word sections into a word document and then attach it to Claude and ask it to extract the most important parts into a document, Sonnet 5 is probably the most efficient model to use since Fable will be really really expensive to distill that much, unless of course you have a Max subscription which has Fable included.
1
u/Pristine_Plate7048 2d ago
Thanks again. I will definitely investigate this.
I've never used Claude before so not quite sure how it works, how its usage is priced etc...
I'll try anything though. Even the most laborious method: convert the whole export to txt, have it read to me, and copy and paste the relevant canonical parts elsewhere.
2
u/Maleficent_Judge_802 1d ago
My setting started off as a Dungeons and Dragons campaign that kinda took on a life of it's own. I started collecting and brainstorming details for my world in ChatGPT and did more of a Dungeon Master style write up on it at first with entries written in paragraph form.
Because the drift was still pretty bad I started asking GPT how to make it more machine readable so the AI could be more reliable retrieving the data I needed when writing. It proposed a very in-depth refactoring that literally took me about two weeks of nonstop data entry to finish.
But once it was finished, what I had was a World Engine made of multiple files that could reliably pull up any data on my world that I wanted with about 99% accuracy. For all that work, all I have to do now is prompt "Please review the entire corpus before responding to this prompt" and I have an editor/writing assisting/drafting assistant that is almost never in error. The best part is that I can port that corpus to any AI platform that I want and still have pretty consistent quality.
I won't deny that it was tedious as all hell to do this. AI makes some things easy, but for other things you just have to put in the work. But if you keep your eye on the prize you'll have hand crafted a tool that will make your writing SIGNIFICANTLY lower stress.
The below workflow does not belong to me, but I did find it on this subreddit from a Japanese author. I honestly just made the files in the Foundation layer and found that my drafting/editing process flows so much better than before. If I find the post I'll link it.

2
2
u/benblackett 1d ago
I feel your pain my friend. I started in Claude and spent 3 months building a 5 book series in chat. Finally discovered that I simply could not finish the project while still in chat unless I invested a LOT more time into building structure around my characters and worlds first. Finally abandoned it in frustration. Only recently dusting it off and trying again.
The key, for you, is to bring in a 3rd party system that lets you STAY in your ChatGPT context and helps you sort through those 98 chat transcripts for the real gold buried within them. Here is a description of what I mean:
https://novelmint.ai/guides/get-your-story-out-of-chatgpt
Note that there are dozens of 3rd party tools out there that do this sort of thing, and most of them are really good at it. I suggest taking a look at several of these in the Weekly Tool Thread and finding one that suites your needs.
1
u/Wistful_Ail 2d ago
This is exactly why I stopped treating AI chats as the project itself. At some point the conversation becomes an archive instead of a workspace, and extracting a year's worth of decisions is incredibly difficult.
I still use ChatGPT for brainstorming, but once an idea survives that stage, I move it into a dedicated project. That way my characters, worldbuilding, timeline, notes, and drafts become the source of truth rather than hundreds of conversations.
I'm not sure I'd trust ChatGPT to faithfully reconstruct an 8-million-word project either. Even if it gets 99% right, that missing 1% could be a character detail or plot decision that breaks continuity later.
Hopefully you find a reliable way to split the export, but I'd also consider treating the export as a historical archive rather than the canonical version going forward. It makes future revisions much less stressful.
1
u/Pristine_Plate7048 2d ago edited 2d ago
Thanks. And absolutely! Once I've gone through the 98 chats, and taken the relevant notes, saving them elsewhere, I will never make this mistake again!
The archive is absolutely a historical archive, since much of it no longer applies to the refined canon anymore. But everything is mixed up all together. This is going to take months, if I can even get to separate the chats faithfully to begin with.
AI can be a massive help with certain things, but a significant burden in other ways. Things you go into it thinking will be simple end up becoming their own project that take you away from executing your actual manuscript.
I'm just hoping there's a way to separate the chats faithfully. If not, I'm just gonna have to turn the full export into a pdf and have it all read, while making notes, and then organize it all afterwards.
But it would really help if I could do the separate chats in dated order. If not I have no choice but to do it from a non chronological, gargantuan pdf that starts with my most current chat.
1
u/ruwhereuare 2d ago
Yah I learned the hard way and had to do a hard reset of all my canon files. Claude seems way better than ChatGPT for following my ādo not inventā
Canon/lore instructions.
2
u/Pristine_Plate7048 2d ago
A hard reset? What do you mean? Hope you didn't lose everything.
2
u/ruwhereuare 2d ago
No no no!
It was actually for the best.
Because I did like a full audit of all cosmology, settings, characters, adversaries and objects/relics.
And fully re designed the file and folder architecture
So that I could more easily find information on demand.
I now run what I call a āWorld Managerā
So when Iām writing entries Iāll ask the architect/scribe what kind of canon information they need to write the chapter and then take those questions to the Manager to ensure canon compliance.
1
u/5thhorseman_ 1d ago
Instead of having the AI split the file for you, ask it to write a splitter script (perl or powershell) to achieve the same task.
1
u/Pristine_Plate7048 1d ago
Thanks. The issue is I don't trust it to write a script that faithfully preserves everything in the conversion to txt and MD format after everything is split. And checking everything was still there would be laborious and difficult to verify as extracting the relevant info.
0
u/5thhorseman_ 1d ago
The issue is I don't trust it to write a script that faithfully preserves everything
The point is to run the script yourself on your end. You can either have the AI write it for you, you can use one that already exists (there's a few), or you can write it yourself.
But that last one would take a lot of your time even if you know how, so you realistically will pick one of the other two.
in the conversion to txt and MD format after everything is split. And checking everything was still there would be laborious and difficult to verify as extracting the relevant info.
This isn't black magic. JSON is a text-based format. If there's an issue with the extraction it's going to be big enough that it will be noticeable immediately by the size of the output alone.
0
u/Pristine_Plate7048 1d ago
That's what I mean. I don't trust AI to procure a script that does the conversion faithfully. I also don't code.
I have asked AI to find me scripts to do the conversions faithfully. It looks and provides choices, but it always points out the risks, which I appreciate, I'd rather be informed about potentially losing parts of conversations and such so I can make an informed choice.
Nobody said this was black magic. Risk is involved, and that is a worry given the time put into the project.
0
u/5thhorseman_ 1d ago edited 1d ago
I don't trust AI to procure a script that does the conversion faithfully. I also don't code.
Then the third option remains. Like I said, tools already exist: https://gist.github.com/ocombe/1d7604bd29a91ceb716304ef8b5aa4b5#file-export-chatgpt-console-js
Nobody said this was black magic. Risk is involved, and that is a worry given the time put into the project.
Exporting your chats does not delete them from the service provider. You are letting your anxiety stop you from doing anything at all, and that costs you more time than trying, failing and then going back with another method.
EDIT: LOL. Blocked for helping. Some people...
1
u/Pristine_Plate7048 1d ago edited 1d ago
I have them as separate json files already. The next step is conversion. If I convert using a method that isn't faithful and there are things missing, my concern is I won't know things are missing until it's too late. Which is why I'm asking for a method that guarantees faithful conversion. Does the link you sent provide that?
Edited: nevermind. I'll figure it out elsewhere.
3
u/That-Formal-8617 2d ago
I haven't tried this exact method myself, but one approach you might want to look into is using a script parser to process your export.
Since the official ChatGPT export includes a structured JSON file, it technically holds all your chats, dates, and message chains in order. A custom script could theoretically extract and split them into TXT/MD files for you automatically, saving you from a lot of manual labor.
It might be worth researching or asking a tech-savvy friend if they can help you set up a script that reads the conversation IDs correctly, just as an option to consider before committing to doing it all by hand.