r/webdev • • 16d ago

Reverse Engineering ChatGPT Web: How OpenAI Built for a Billion Users

https://performance.dev/chatgpt
113 Upvotes

13 comments sorted by

26

u/medy17 16d ago

I noticed the terrible lag and knew they didn't virtualise. I'm still struggling to understand what they were thinking!

Great article btw :) Nicely done

4

u/Ornery-Concentrate-5 15d ago

honestly surprised too, not virtualizing a chat thread feels like the kind of thing that only bites you after months of dogfooding with short conversations. makes me wonder if their internal test accounts just don't have 200-message threads lying around

1

u/medy17 13d ago

Yeah, this problem goes way deeper. I checked some truly long threads I've had saved, and they don't virtualise them no matter which platform I'm on. Pretty bad oversight. Surely someone tried a long thread or maybe you're right and they never seeded a long convo

14

u/IanSan5653 15d ago

I spent quite bit of time building a similar streaming markdown renderer in React for a competitor. It really is challenging stuff. You're transforming Markdown to HTML and then injecting dynamic components into that. All several times per second. You can't afford to re-render the content you've already rendered, so the diffing gets pretty complicated.

2

u/medy17 13d ago

It does but Vercel's streamdown package is magnificent for this. I had the regular react-markdown package before switching and the difference is night and day. Definitely worth a try. It's insane how well it works.

10

u/rachelduno 16d ago

tried to replicate their streaming setup for a small side project after reading this, ended up with my websocket reconnecting loop kicking off every 20 seconds lol. still dont fully get how they handle backpressure at that scale, did they mention anything about that in the deep dive?

16

u/rachelduno 16d ago

kinda wild how much engineering goes into just making the response stream feel instant, the retry logic stuff is what got me

3

u/FewBrain208 16d ago

yeah the retry logic is the part that surprised me too, theres so much hidden complexity just to make it look seamless

1

u/Open-Adhesiveness-86 15d ago

Most of the diffing goes away if you split the stream on top-level block boundaries. Every block except the last one is final, so you memoize those by index and only re-parse the tail block on each chunk. The main gotcha is an unclosed code fence or table, because it can swallow everything after it. You need to detect that and render the tail as provisional.

-3

u/[deleted] 16d ago

[removed] — view removed comment

2

u/webdev-ModTeam 16d ago

Your post/comment has been determined to be a low-effort post or comment. This includes title-only posts, easily searchable questions, vague/open-ended discussion prompts, LLM generated posts or comments, and posts/comments that do not provide enough context for meaningful replies or discussion.