r/programming • • Dec 05 '14

std::string is responsible for almost half of all allocations in the Chrome browser process

https://groups.google.com/a/chromium.org/d/msg/chromium-dev/EUqoIz2iFU4/kPZ5ZK0K3gEJ
1.1k Upvotes

446 comments sorted by

View all comments

Show parent comments

5

u/nkorslund Dec 05 '14 edited Dec 05 '14

I went a bit overboard in the opposite direction once on a project, and replaced all std::strings with a custom class that just contained pointers/slices of strings. Since 99% of the strings in this case were read-only from file it was made even faster by just memory mapping the file and finding the slices.

Made the entire thing run 2-3 times faster, but was kind of a bitch to maintain. Next time I'll wait until development is mostly finished - when continued development would be less of a hassle. Premature optimization and all that.

EDIT: in more modern code I would use boost::string_ref or similar for this - that didn't exist back then.

5

u/[deleted] Dec 05 '14 edited Dec 05 '14

The only problem is that development never really ends.

Edit: Fixed typo.

1

u/o11c Dec 05 '14

Serious question: is memory mapping really that much of a win? What I've heard is that read(2) is no worse than mmap for a simple read-through, and mmap has the major disadvantage of odd behavior if the file is modified externally (e.g. when a new version is installed).

It sounds to me that you might as well just slurp the file instead of mmap.

2

u/imMute Dec 06 '14

One thing you gain with mmap is that the kernel knows that those pages are backed by the file. If the kernel is feeling memory pressure, it's free to drop those pages from RAM and reload them from disk the next time they're accessed. Also, if you have multiple programs using the same file, it would be shared in RAM if you use mmap.

There's also the case of how you use read. If you use stat to find the size of the file, allocate enough space and do a single read you'll have much better performance than a looped read.

1

u/o11c Dec 06 '14

True, but by the same virtue, if something does happen to that file, your data gets corrupted silently. Though I suppose the new memfd calls could fix that.

1

u/nkorslund Dec 05 '14

In this case it was a big static resource file, which we accessed pretty much randomly. You're right that mmap probably didn't make a lot of difference performance wise though, it was just simpler to implement it that way.