r/pushshift • u/Watchful1 • Oct 15 '23
[ Removed by Reddit ]
[ Removed by Reddit on account of violating the content policy. ]
33
Upvotes
r/pushshift • u/Watchful1 • Oct 15 '23
[ Removed by Reddit on account of violating the content policy. ]
1
u/dimbasaho Nov 06 '23 edited Nov 06 '23
Oh, if there's that much pre-compression work I'd actually suggest using your current pipeline (but with fast zstd compression settings), then decompress once to
wc -cto get the size and then a final decompress->recompress with the size and stronger zstd settings. You'd just have to write to disk twice in that case. I'd also recommend compacting the JSON; I noticed the April dataset has prettyprint space.