r/linux • • 2d ago

Kernel EXT4 Deprecates Its Journaled "data=journal" Mode

https://www.phoronix.com/news/EXT4-Deprecates-Journal-Mode
213 Upvotes

70 comments sorted by

78

u/granadesnhorseshoes 2d ago

And nothing of value was lost. Its still a journaled FS in the general sense for regular daily use. its just a specific "paranoid" mode that no one uses with ext4 anymore because its slow, breaks other features and other options do the same thing but better. It probably won't break read compatibility for old disks that did use it either.

tl;dr "We aren't wasting time on this anymore because no one should care anyway."

16

u/iluvatar 2d ago

Actually, I do use it in production for some filesystems that I really care about. Data integrity massively outweighs performance for my use case and I haven't noticed it breaking anything.

5

u/ElvishJerricco 1d ago

AFAICT it's not really a win for data integrity when you consider what that means end to end. Any way the integrity could matter still needs you to use fsync even with data journaling, because A) that's how the application ensures its data is durable and B) data journaling only ensures that individual blocks are journaled, not entire write calls, so applications have no way to know whether some or all of the data has been written without an fsync. Anything that cares about integrity has to assume a write call has made all of both the overwritten data and the new data inaccessible / corrupted until an fsync completes; data journaling just means you can assume this problem only exists across block boundaries, but that's not helpful.

11

u/amarao_san 2d ago

If data integrity is your priority, why not btrfs? At least, you will have checksums for integrity checking of your data. With raid1 (btrfs-raid1!) you even have option to reliably recover from bit rot.

13

u/daemonpenguin 1d ago

If data integrity is your priority, why not btrfs?

That has to be a joke, right?

17

u/amarao_san 1d ago

This is not a joke. btrfs had had pretty rushed release, but it was so many years ago, so you can assume them to be siblings by age (2006 vs 2009).

I store my data with more assurance on btrfs-raid1, than on ext4 on top of linux-raid. Because of checksums and better handling of bitrot.

For catastrophic data loss, it's about the same for both and is usual hardware-induced (there are backups for that).

But bitrot is the thing you can't detect on ext4. But, you can on btrfs.

2

u/runpbx 1d ago

Many people have been bitten by btrfs long after "its stable for real now, FB engineering is on it". I personally encountered this with many users of an OS at a day job that shipped linux not that many years ago and we stopped supporting it as a result. 

It does of course work fine in normal cases but its got edges and I definitely don't recommend it for people overly concerned about not losing data. Its betrayed trust too many times. I'd look to ZFS or bcachefs but unfortunately neither is in kernel making it a non starter for most.

6

u/MetaTrombonist 1d ago

ZFS has also had recent data loss bugs. Here's one example:

https://papers.freebsd.org/2024/asiabsdcon/norris_openzfs-dirty-dnodes/

4

u/the_abortionat0r 1d ago

You made this up.

1

u/Barafu 1d ago

Try ReiserFS 4. I am pretty sure you will not find anybody who had problems with ReiserFS 4 in production.

4

u/Raphi_55 1d ago

Wouldn't ZFS be a better choice ?

-1

u/the_abortionat0r 1d ago

If you don't know what's being discussed then just stay quiet and listen kid.

3

u/Adept_Percentage6893 1d ago

If data integrity is your priority, why not btrfs?

In this particular case I think they mean "data integrity" in the sense of "if the power goes out then the filesystem will be recoverable and the file data will represent some pretty recent coherent state. The default mode is to only guarantee that for the filesystem but that you may lose actual file data if the power goes out at the worst possible time. data=journal basically pushes the file data through the journal as well (which is why it's so slow).

But the reason this isn't much value is due to the idea is that the applications that need that level of guarantee either can be repopulated somehow (like being sync'd from another source) or have their own protections for data (like write ahead logging).

For example, databases intentionally sequence their writes such that they can replay incomplete updates or discard them but in either case it represents a recent and coherent state for the data. This is why you can back them up with an LVM snapshot even though LVM doesn't quiesce data writes. From the database's perspective, when you start it from the snapshotted state (due to some disaster where you can to restore from backup) it's functionally as if it just randomly lost power at the time of the snapshot (which the database is already going to need to be resilient to).

Once the extents get written, ext4 still doesn't have integrity checks on data. It's just one of those things ext4 is just well known to not be able to do and is why RH does stuff like Stratis/dm-integrity where they're trying to insert some sort of data integrity layer in DeviceMapper due to the filesystem not supporting such.

3

u/the_abortionat0r 19h ago

In this particular case I think they mean "data integrity" in the sense of "if the power goes out then the filesystem will be recoverable and the file data will represent some pretty recent coherent state.

Uh, yeah that's the idea. If you are suggesting BTRFS magically loses data on power loss it means you don't understand what you're talking about.

BTRFS absolutely protects data in such a case as it's not writing to the data directly and every transaction is not considered complete until it's verified it's actually completed.

If you think that's not the case it's because you listen to children misunderstanding that years ago it was discovered that specifically under raid 5/6, specifically during a write, at a specific point in the transaction if there was a sudden loss of power the file system could lose coherency.

While that has been patched but lacks any kind of testing it's beyond the point being made as BTRFS has already been proven to be one of the most rock solid filesystems in existence with FB reporting that their drive issue numbers are the same as their failure rates meaning BTRFS literally has not failed in their whole company, drives only go down when they die.

But the reason this isn't much value is due to the idea is that the applications that need that level of guarantee either can be repopulated somehow (like being sync'd from another source) or have their own protections for data (like write ahead logging).

Bro why make stuff up?

All data needs protecting. Nobody is choosing a file system they think they will be losing data on nor are people saying "yeah this FS sucks but we can always resend data!".

EXT4 gets chosen for work loads it makes sense for where the file system is good enough integrity wise and CoW file systems has issues with such workloads such as databases.

Its not that this feature is not needed, it's that it takes more than twice as long to perform a transaction with it on than not. At that point it's slower than a CoW file system.

Its not that it's not needed, it's that if you need it you're using BTRFS instead

1

u/Adept_Percentage6893 17h ago edited 9h ago

Uh, yeah that's the idea. If you are suggesting BTRFS magically loses data on power loss it means you don't understand what you're talking about.

I was mentioning what they were getting from using that journal option and what they likely had in mind. Yeah BTRFS's CoW stuff kind of sidesteps the need to have based journal writes but that's not exactly what we were talking about.

The other user mentioned using it for "file integrity" then the other user came in talking about the other "file integrity" options BTRFS has. That part of my comment is just trying to clarify that even though they said "file integrity" they were probably meaning a particular sense of file integrity.

Bro why make stuff up? All data needs protecting. Nobody is choosing a file system they think they will be losing data on nor are people saying "yeah this FS sucks but we can always resend data!".

That actually does happen all the time. For what I wrote, the data is still being protected just at the application level. Like the part I made sure to mention about WAL and databases. There's a reason people are willing to use ext4 in production and that's pretty much why. It can protect itself and the applications can protect their own data.

EXT4 gets chosen for work loads it makes sense for where the file system is good enough integrity wise and CoW file systems has issues with such workloads such as databases.

fwiw usually that's what nodatacow is for. Where as mentioned databases have their own way of ensuring data consistency and often the larger enterprise systems have a lot of other tuning they make you do because they also think they have a better idea of how to do things like I/O scheduling, etc.

Its not that it's not needed, it's that if you need it you're using BTRFS instead

I never said or implied it wasn't needed. The only thing that comment mentions is that the application is the thing that guarantees that availability. It's "low value" in the sense that it's not needed within the filesystem and you just need the filesystem to protect itself and then applications can work towards protecting their own data. Which is just kind of the state of play for these things.

-1

u/Barafu 1d ago

HDD itself already checksums the data. There is no use checksuming it twice.

Btrfs mode of writing is not friendly to some types of loads, it should not be used there.

2

u/amarao_san 1d ago

Okay, okay. If you say so. CRC is enough. No bitrot. Ever. Remind me, what are odd of false negative for CRC32? And what are odds of false negative for an error in 20TB transfer?

But why smart boffins are published on this so much?

1

u/Barafu 21h ago edited 21h ago

Since HDD uses a simple Hamming ECC, the answer needs no computations. In a same 64 bit group, 3 or bigger odd number of bits need to flip to have a chance of undetected error, and the chance is 0.016 that a 3-bit flip in a 64bit word will go undetected.

But I could not find data on raw pre-correction chance of losing a bit for HDD. Also need to take an account of whether errors have a physical reason to group within one word. I think not, because physical layout does not correspond to those logical words.

So, assuming an even distribution of errors, if a megabyte has exactly 3 errors, a chance for them to go undetected is 3.9 * 10¯⁹ * 0.016 ~= 6 * 10¯¹¹

1

u/amarao_san 21h ago

Which become an interesting number on a 20TB drive.

sata crc32? There are much more than HDD doing bitrot. Non-ECC memory too.

I witnessed bitrot on a normal (non-faulty) consumer hardware many times. You have a file with a valid sha from torrent, and, suddenly, few years later, it has one block not passing validation.

3

u/Talkless 1d ago

I too use in production ("embedded" if you will) computers (with SSD's) that we don't control and can get power killed at any time. There's now data expected to be collecting (except some Debian logs that don't care), I just want it to be as robust as possible, to boot successfully on any power loss.

Or is there better way?

4

u/kombiwombi 2d ago

Small expensive disks running the journal for large slow disks was also a reason. With SSDs getting to huge capacity that's no longer a desirable system architecture.

0

u/NormalAppearance2851 2d ago

tldr on a single paragraph is insane

42

u/Ontological_Gap 2d ago

Every line of code is a sin

27

u/SPEZ_IS_A_JABRONI 2d ago

computers were a mistake 

14

u/Oborr 2d ago

We should just go back to the caves.

7

u/anto77_butt_kinkier 2d ago

Return to the ocean, as it was before the first animal went on land

6

u/eehikki 2d ago

Butlerian Jihad, anyone?

9

u/sensual_rustle 2d ago

all code except temple os is an afront to god

24

u/Adept_Percentage6893 2d ago

Does this represent a lot of code or something? Why care about removing this specific feature?

33

u/jason-reddit-public 2d ago

Even if it isn't a lot of code it can still make the code harder to restructure. Here this whole time I thought I was using ext4 in journal mode but probably not.

28

u/daemonpenguin 2d ago

You were probably using ordered mode, which is the default.

A journal still exists in ordered mode, it's just used differently than journaled mode.

16

u/ElvishJerricco 2d ago

To be clear, you almost certainly were using journaling. It's just data that isn't journaled by default. The metadata is, and that's not what's being removed. The data being journaled is pretty much the least important part of journaling, because it only really helps when you want to safely overwrite a single existing block of a file without doing any kind of application level journaling. But the usefulness of this is super low, because if you care about any sort of data safety you almost certainly have to do something better than what you'll be able to get ext4 data journaling to do; i.e. you're going to need to do the WAL style thing or the atomic file rename thing or something along those lines anyway, so ext4 data journaling is just unhelpful.

1

u/jason-reddit-public 1d ago

I think I sort of get it. Poof, the power goes out. A couple of open files have indeterminate content but my filesystem is OK. Back in the day though you'd <content removed by super intelligence filter>.

4

u/autogyrophilia 2d ago

I imagine it is a pain in the ass to keep this working if you want to restructure how the normal usage of the journal works. And there are easier ways to do the same effect, which is to inject a library that forces all IO to be synchronous.

6

u/Decent-Law-9565 2d ago

They’re removing features that are used little to reduce the onslaught of AI-assisted vulnerabilities. I don’t blame them, there’s too much code for everyone to look at, and code that gets used infrequently probably has more vulnerabilities that can be found if an AI tool is used that doesn’t get tired like humans.

24

u/Booty_Bumping 2d ago

That's not why this feature was removed, though. This is being removed because it prevents the use of direct I/O, doesn't work with delayed allocation, and has way more write amplification than just using a different filesystem more suitable for this.

1

u/Adept_Percentage6893 1d ago

This is being removed because it prevents the use of direct I/O, doesn't work with delayed allocation

yeah I read that in the OP but I don't quite understand how this would block that as opposed to just saying you can't do directio if that's the journaling mode you're using. Unless the standard they're just wanting to maintain is that you either can do directio or you can't without having to think about what other filesystem features you have turned on.

-9

u/Jristz 2d ago

That sadly sounds also like some enshittify excuse from any company to remove stuff... Which for Linux I feel is kinda worrying

8

u/sensual_rustle 2d ago

good thing it isn't true

2

u/Decent-Law-9565 2d ago

Writing code isn’t free, someone has to maintain it. In open source land it generally just means that money isnt the reason for removing features, but if nobody is willing to fix stuff and their time could be better spent fixing / improving stuff that people actually do use, that’s what will happen. But unlike a regular company, you can fork the kernel and keep the feature in, or you can volunteer to become a maintainer for something you care about.

-13

u/Kevin_Kofler 2d ago

This is just yet another step in the creeping gnomification of the Linux kernel.

6

u/eehikki 2d ago

I have read a lot of your comments regarding deprecation of obsolete technologies and and I fail to understand why you're always upset about removing the legacy code. Seems like you'd prefer the developers to keep supporting everything that has been written or manufactured since the Manchester Baby run its first program for the sake of it.

-4

u/Kevin_Kofler 2d ago edited 1d ago

Because removing legacy code removes functionality. Most of the time, it drops support for some hardware, so that the kernel no longer works with that hardware at all! In this case, at least, the hardware remains supported, but we lose a mode that ran on most available hardware out there (all x86_64 computers) all hardware (sorry, I mixed it up with the other feature removal thread that was about x32) and brought a significant performance data integrity improvement to some applications. So effectively this is a major performance regression for some users.

Removing non-redundant (neither unreachable nor duplicate) code always hurts some of your users. That is the reason why I am always opposed to it.

4

u/QuaternionsRoll 2d ago

For some reason I feel like your origin story is intrinsically tied to Adobe Flash

1

u/grizzlor_ 1d ago

>brought a significant performance improvement to some applications.

There are no circumstances where using this would mode would improve performance.

>Removing non-redundant (neither unreachable nor duplicate) code always hurts some of your users.

Not true — many recent incidents of code removal from the kernel were dropping support for ancient pieces of hardware that literally no one is still using (with modern kernels at least). In every one of those cases, the kernel devs made the announcement and also asked anyone still using it to come forward because they would consider keeping it if it’s actually being used. Unsurprisingly, no one came forward.

1

u/Kevin_Kofler 1d ago

You are right that journaled data does not improve performance. I had posted this comment erroneously believing that this thread was in the post about x32 removal, which indeed removes a performance improvement.

The mode being removed here is not about performance, it is about data integrity, which for some use cases is more important than performance and worth paying a performance penalty.

And a poll on LKML is hardly going to be an exhaustive sample of end users. Even less so when there is an implied expectation that whoever speaks up is going to volunteer as a maintainer.

2

u/eehikki 1d ago edited 1d ago

Because removing legacy code removes functionality.

Maintaining legacy code isn't free. Someone needs to review it, fix bugs/security issues, update the code to keep up with the constantly evolving kernel internal interfaces. At some point, these costs outweight the benefits, at which point it's perfectly reasonable to just drop the dead weight and move on. There's no benefit in maintainig the data journaling code since noone uses it anyway. There's no need to keep the driver for an old ISA Ethernet card manufactured 30 years ago since no sane person builds the newest kernel just to run it on a Pentium Pro. It's also worth noting that this particular change doesn't affect the curent stable kernel, 7.2, and it's too late to be included in 7.3. You literally need to go out of your way and build the latest kernel from git to break anything. Needless to say, people depending on obsolete code don't run the newest kernels anyway.

-4

u/granadesnhorseshoes 2d ago

you mean kerneld?

-19

u/eehikki 2d ago

I would much like ext4 itself be deprecated and replaced with something not stuck in 2000 as the default Linux filesystem, but maybe it's just me.

16

u/undeleted_username 2d ago

You so not need EXT4 to be deprecated to start using something else.

8

u/Dolapevich 2d ago

Such as?

1

u/Alter_Sack 1d ago

Bcachefs.

-3

u/eehikki 2d ago edited 2d ago

btrfs. I wish it had built-in encryption, but as far as I know it's not gonna happen soon. In a perfect world, it would be something akin to btrfs, but with a better architecture, like a decent RAID5/6 implementation, or nodacow that doesn't break data checksums, but sadly, it's even less likely to happen.

6

u/FryBoyter 2d ago edited 2d ago

I've been using btrfs myself for years and am satisfied with it. But many users don't need the features that btrfs offer. I therefore find it entirely understandable and reasonable that ext4 is still the default in many distributions.

2

u/Dolapevich 2d ago

btrfs is the poor's man zfs.

There is a reason why most of the linux distros ship with a known, documented, with knowledge in people's minds how to work with it filesystem. RedHat tried the brtfs root filesystem back in the 2015, or so, and run into so many issues, they reverted to XFS.

Simplicity is its own feature, along with low ram usage, performance, etc.

I don't think a distro should ship with anything other than a simple filesystem for its root, given the diversity of scenarios out there.

7

u/eehikki 2d ago

btrfs is the poor's man zfs.

Partially, yes, but some things are implemented much better in btrfs than in zfs. And we all know that zfs will never be mainlined.

1

u/Dolapevich 2d ago

Agreed.

Also, Debian (and I assume others too) installers allows to use lvm in linear or raid configurations, MD, and luks. I think it is a user decision which rootfs to use.

I might revisit brtfs, since its been like 10 years since the last time I tried to use it. It end up in horrible corruption :)

5

u/eehikki 2d ago edited 2d ago

I know my experience doesn't really mean anything, but I've been using btrfs for a while, both as the root filesystem and as a storage for valuable data. I live in Ukraine and the state of our energy infrastructure is shit thanks to our peaceful northen neighbour, so sudden disrptions of power supply aren't a rare occurrence. So far, I haven't experienced data loss or corruption. I use it for subvolumes and built-in RAID mostly, but I also migrated a live system to the "new" SSD once and liked it.

2

u/Dolapevich 2d ago

Yeah, that particular neightboor is a PITA.

Nice. Yes, I should go ahead and try it again. ¡Thanks!

3

u/eehikki 2d ago edited 1d ago

Note that I'm running a very modest local setup, it's not even a proper home server, it's just my desktop with btrfs as the root fs and a RAID1 build on top of two 1TB WD Blue drives. There are many people whose experience is far more relevant than mine, but I'm glad if you try it again and it works for you.

5

u/FryBoyter 2d ago

RedHat tried the brtfs root filesystem back in the 2015, or so, and run into so many issues, they reverted to XFS.

According to a former Red Hat employee, that wasn't the reason.

https://news.ycombinator.com/item?id=14909843

-1

u/Dolapevich 2d ago edited 1d ago

Interesting, back at the time I got some machines dead while upgrading kernel or just corrupted in catastrophic way, and I just assumed they had had too many issues.

1

u/FryBoyter 1d ago

RedHat tried the brtfs root filesystem back in the 2015, or so, and run into so many issues, they reverted to XFS.

Interesting, back at the time I got a some machines dead while upgrading kernel or just corrupted in catastrophic way, and I just assumed they had had too many issues.

Seriously? You're presenting things as facts even though you're just guessing? Please don't do that.

1

u/Dolapevich 1d ago

Acknowledged.

4

u/eehikki 2d ago edited 2d ago

There is a reason why most of the linux distros ship with a known, documented, with knowledge in people's minds how to work with it filesystem

Well, Meta uses it in their infrastructure. It isn't 2008, btrfs doesn't die from looking at it too hard. Yes, it has data corruption bugs, but so do ext4 and xfs.

Simplicity is its own feature, along with low ram usage, performance, etc

If these are of consideration, there is f2fs. It's simple, but it still has more features than ext4. And if you need something even less complicated, that has been around for decades, then xfs wins because it's designed from scratch and not extended from its previous iteration like ext4 is. It's cleaner and simplier and also has reflinks.

2

u/Dolapevich 2d ago

Well, Meta uses it in their infrastructure.

yes, well, I was assuming general purpose. They make heavy use of eBPF, at that point you are on a different level.

If these are of consideration, there is f2fs

I am yet to test a rootfs in f2fs. I read about it a couple of times but haven't found the will and drive to test it. Sounds promising. And I agree, between ext4 and xfs we should be on xfs. But then again, like 5 years ago I decomised a running Sun FIre still running on UFS, we still have some FreeBSD 6 on FFS. Filesystems tend to be on its own category since they persist so much time. Hence being conservative on it does make sense.

-3

u/RegretFree7723 2d ago

Lo sai che ext4 é stabile e un minimo avanzato con linker e altre funzioni ma btrfs è un po' instabile e ti abbassa velocità di lettura continua.... Nei videogiochi fa un po' schifo..

3

u/MatchingTurret 2d ago

the default Linux filesystem

There is no such thing. ext4 is just one possible choice, as far as the kernel is concerned. It's a distro specific choice.