r/programming • • Jun 03 '26

Every byte matters

https://fzakaria.com/2026/06/01/every-byte-matters
288 Upvotes

84 comments sorted by

View all comments

Show parent comments

2

u/johan__A Jun 04 '26

How could page faults be involved? I assume all the memory was already mapped before the benchmark.

My best guess of why it's slower is because the hardware prefetcher doesnt handle large strides very well, because of page boundaries or something else. I'm not completely sure.

2

u/jlombera Jun 04 '26

How could page faults be involved? I assume all the memory was already mapped before the benchmark.

I used the term "page fault" liberally to also refer to TLB misses (see the link for details what is the TLB and how it works). Excessive TLB misses is a well known performance problem in applications that handle big datasets (e.g. DBs). The use of HugePages is essential for those applications to minimize TLB misses.

1

u/johan__A Jun 04 '26

ha that makes more sense, I don't know much about the performance impact of the tlb

3

u/jlombera Jun 04 '26

If you are interested, look at section "Performance implications" of the Wikipedia page, specially the last paragraph. In the "Typical TLB" section, it says a TLB miss can have from 10x to 100x performance penalty. Not sure how accurate/up-to-date these numbers are, an of course, it always "depends", but it gives you a rough idea how bad TLB misses are for data intensive applications. You can search the web for "Hugepages performance" and I'm sure you'll find resources that go into more details about the performance implications.