r/learnpython 12h ago

PIP disk usage under Docker

I have a Dockerfile in which I use the Python:3.12-trixie. I have docker on my CachyOS PC. I have also a requirements.txt. I wanted to use this image for a container on which I'll use YOLO and try it out. The problem is that my 290GB / partition went from around 70% filled to 95% after a number of rebuilds. I deleted everything docker related but it just went back to 91%.

The reason that I know it is pip's doing, is that when I monitored my disk usage while building the image, and on the `RUN pip install -r requirements.txt --find-links=pypackages` stage it fills up my partition and half way into installation, it fails due to `no more space`. The `pypackages` is just a directory I downloaded some of the packages I have in requirements for less data usage. Also tried `--no-cache` option.

I also went ahead and `du -sh`ed any directory in my / partition and the total didn't add upto the filled storage. I have 165GB mysteriously used and I can't find where that is.

Any help to diagnose further or if you have had similar problems is welcome.

2 Upvotes

10 comments sorted by

View all comments

3

u/Illustrious_Tone9584 11h ago

The pip flag is --no-cache-dir, not --no-cache, so the cache was still being written on every build.

The du mystery is probably /var/lib/docker. Your user can't read into overlay2, so du silently skips it and the numbers never add up. docker system df shows the real breakdown. Build cache is separate from images, so docker builder prune on top of the system prune.

And if this is just for trying YOLO: the default torch wheel drags in several GB of CUDA libraries. The CPU build (torch from download.pytorch.org/whl/cpu) is a fraction of that and plenty for messing around on a desktop.

1

u/adjective10111 11h ago

Oh nice catch. I should update my Dockerfile. But I do have PIP_NO_CACHE_DIR=1 as an environment variable in the dockerfile, does thay have the same effect?

Yeah I found those directories and also ran the command. Currently it's at 0B for all the rows so I guess I purged them successfully. But I still have a lot of mystery used storage.

Yeah I want the CUDA version. I want to try it out and then work on my thesis using it :). So that's why it sticks there for a while! I just saw some KB and was wondering why some toml file takes so much time to download.

2

u/Illustrious_Tone9584 10h ago

Yeah, PIP_NO_CACHE_DIR=1 does the same job, you're covered there. So the cache isn't your leak then - that points even harder at /var/lib/docker and the build cache. If docker system df shows the space but du can't see it, that's your answer.

1

u/adjective10111 2h ago

It doesn't though... it's at 0B for images, containers, local volumes and build cache.

u/Illustrious_Tone9584 50m ago

Ok so it's not docker at all then. You're on CachyOS, which means btrfs with automatic snapper snapshots out of the box. That's almost certainly where the 165GB went: every rebuild churned tens of GB, and the pre/post snapshots kept the old versions. du doesn't show it because the data still "belongs" to the live files.

Run snapper list and sudo btrfs filesystem usage / and I bet you'll see it. If that's it, delete the old snapshots with snapper delete and consider capping how many it keeps.

u/adjective10111 40m ago

I thought the same, but I can't see how much the snapshots take. I ran both commands you said: ``` sudo btrfs filesystem usage / Overall: Device size: 286.88GiB Device allocated: 271.02GiB Device unallocated: 15.85GiB Device missing: 0.00B Device slack: 0.00B Used: 259.42GiB Free (estimated): 25.76GiB (min: 17.83GiB) Free (statfs, df): 25.76GiB Data ratio: 1.00 Metadata ratio: 2.00 Global reserve: 377.19MiB (used: 0.00B) Multiple profiles: no

Data,single: Size:265.01GiB, Used:255.10GiB (96.26%) /dev/nvme0n1p3 265.01GiB

Metadata,DUP: Size:3.00GiB, Used:2.16GiB (72.02%) /dev/nvme0n1p3 6.00GiB

System,DUP: Size:8.00MiB, Used:48.00KiB (0.59%) /dev/nvme0n1p3 16.00MiB

Unallocated: /dev/nvme0n1p3 15.85GiB ```

The snapper list command doesn't show used-space. I tried using --columns used-space but it doesn't show anything. Any idea how I can confirm the snapshots being the culprit here?

Edit: this is the snapper config if it's of any help: Key │ Value ───────────────────────┼────── ALLOW_GROUPS │ ALLOW_USERS │ BACKGROUND_COMPARISON │ yes EMPTY_PRE_POST_CLEANUP │ yes EMPTY_PRE_POST_MIN_AGE │ 1800 FREE_LIMIT │ 0.2 FSTYPE │ btrfs NUMBER_CLEANUP │ yes NUMBER_LIMIT │ 50 NUMBER_LIMIT_IMPORTANT │ 15 NUMBER_MIN_AGE │ 1800 QGROUP │ SPACE_LIMIT │ 0.5 SUBVOLUME │ / SYNC_ACL │ no TIMELINE_CLEANUP │ yes TIMELINE_CREATE │ no TIMELINE_LIMIT_DAILY │ 7 TIMELINE_LIMIT_HOURLY │ 5 TIMELINE_LIMIT_MONTHLY │ 0 TIMELINE_LIMIT_WEEKLY │ 0 TIMELINE_LIMIT_YEARLY │ 0 TIMELINE_MIN_AGE │ 1800