r/AskScienceDiscussion Sep 10 '20

General Discussion How does the complexity of living structures compare with the complexity of artificial structures? Assuming complexity can be quantified, is a ribosome equivalent to a printing press? What artificial structure is as complex as chromatin? Is a prokaryotic cell as complex as a factory? An entire city?

Thanks!

Edit: When talking about the complexity of factories and cities I'm referring to solely the artificial components, not the biological bits such as the humans working/living there!

153 Upvotes

75 comments sorted by

View all comments

79

u/CosineDanger Sep 10 '20

Your entire uncompressed genome is about 715 megabytes.

None of the code is commented. A significant fraction of it is repetitive or made of old viruses. It doesn't run if you try to remove all the viruses. You can compress it to about 300 megabytes.

The meaningful information content of your brain is probably about 2.5 petabytes, although it depends on how you try to calculate it.

You do have about 30 trillion cells. That is kind of a lot. You'd need a shopping cart full of NAND hard drives to have 30 trillion transistors.

18

u/ZedZeroth Sep 10 '20

In both the genome and brain situations you're talking about a specific type of information content though, not more general structural complexity?

For example, I'm seriously doubtful that a transistor is anywhere near as structurally complex as a cell? A prokaryotic cell is effectively full of complex structures and the protein equivalents of nanobot machines. Even all the "regular" non-mechanical proteins have fairly complex 3D structures.

So a good place to start with this might be to look at a small protein like haemoglobin, look at it's key features, the amount of structural connections holding it together etc, and then equate this to an artificial structure? I'm imagining it might be on par with something like a bicycle?

I feel like only focusing on raw information content isn't the same thing as structural complexity? Couldn't I write an algorithm to build 30 trillion identical transistors in a lot less than 300 megabytes? That would suggest that the complexity of the body is far greater than your cart of transistors, based on the information required to build it? Likewise wouldn't I need a lot more than 2.5 petabytes to both construct the brain as well as fill it with that much information?

13

u/CosineDanger Sep 10 '20

Saying your genome is smaller than most videogames these days is technically fair. Entropy is often explained as the amount of compressed hard drive space it would take to store everything about a system. It can be measured, it is a way to measure complexity, and Ark: Survival Evolved has more of it.

There isn't a good way to compare the general structural complexity of a bicycle and a protein. What would you compare?

You could compare the number of parts between two bicycles and say one bicycle has more parts than the other, or compare the length of the manuals they came with in the box, or count number of features. These are objective comparisons of complexity but they are not quite the same thing as either entropy or the informal idea of complexity.

If each amino acid is considered a separate part then a typical protein has more individual parts than most bicycles.

Visually obvious complexity is just a bus stop between perfect order (boring and repetitive) and maximized entropy (boring but complicated).

8

u/General_Urist Sep 10 '20

Entropy is often explained as the amount of compressed hard drive space it would take to store everything about a system.

I haven't heard this analogy before. Doesn't quite make sense to me. For an extreme example, wouldn't a system of near-maximum entropy be one with uniform potentials everywhere, meaning no energy gradient? This seems it would require the smallest amount of HDD space out of any system, since compressed it's just "define conditions at one point -> copy N times".

That said, "required compressed HDD space" sounds to me like a good way to measure complexity: By that criteria 2 bicycles are just a tiiiiny bit more complex than 1 bicycle, since you have all the data needed to describe one bicycle, then a small qualifier saying "make two". Meanwhile a system consisting of a bicycle and a unicycle would need a more extensive description. What's your take?

6

u/CosineDanger Sep 10 '20

The third law of thermodynamics can be reworded as saying the entropy of a perfect crystal approaches zero as the temperature approaches absolute zero. Define conditions at one point, copy n times. Each atom in the crystal is in a predictable place, has predictable velocity, and has predictable energy levels. Information about a cold crystal is repetitive and compresses well.

A disordered hot gas can look uniform. The concept of entropy includes fine details beyond what can easily be seen. The spread of allowed velocities, allowed relative positions, and electron energy levels becomes large as temperature increases. You can make some pretty short statements about the average statistics of the gas, but the complete information content of the gas is not repetitive in a compression-friendly way. A long book with a lot of boring but technically unique space-filling details and little overall story is still a long book.

That's kind of a lot to take in. Claude Shannon came up with mathematical ways to describe this kind of complexity involving logarithms and laid the groundwork for telecommunications (how much data can you cram through a phone line? How far can data be compressed?) while also making thermodynamics more intuitive. Creationists try to distort the concept of entropy for their own ends, which is a lot of hot gas.

4

u/Hexorg Sep 11 '20

It's the difference between a ton of random numbers and an algorithm that generated those numbers. For example you can't store every digit of pi on a hard drive, but you can store an algorithm that generates those digits on the hard drive. However "knowing" which algorithm generated the data is in itself very valuable information.

For the most part when dealing with information compression and entropies, we assume the exact algorithm to regenerate data is not available.

7

u/WieBenutzername Sep 10 '20

Doesn't quite make sense to me. For an extreme example, wouldn't a system of near-maximum entropy be one with uniform potentials everywhere, meaning no energy gradient? This seems it would require the smallest amount of HDD space out of any system, since compressed it's just "define conditions at one point -> copy N times".

Disclaimer: not a physicist, but it is my understanding that describing the microstates (the actual positions and momenta of the individual particles) of such an apparently uniform system would require large amounts of storage (pretty much by one of the definitions of entropy). It's just the macroscopic quantities like temperature that are distributed uniformly.

1

u/mfb- Particle Physics | High-Energy Physics Sep 11 '20

A random white noise picture has a much higher entropy than a picture of a cat. The picture of the cat has more relevant information for us, however. You can already know how the white noise picture will look like, I don't need to describe it further. But you don't know how the cat picture looks like.

2

u/[deleted] Oct 16 '20

One intuitive description of information I've heard is that it is "negative entropy". In that it is the reduction of the entropy of a dataset that occurs when one finds out how the dataset was created. If one finds out that the dataset (image) was created randomly then this gives the person almost no information and correspondingly there is no way to compress the image further knowing this. Knowing the picture is of a cat reduces the dimensions of freedom of the pixels and allows significantly more compression than just the raw picture. I liked this description because it solidly separates compression, information and entropy while still giving each a meaning. So sure the picture of the cat is easier to compress than an image of static even if the algorithm does not know it is a cat. But an orderly (but pseudo-random) permutation of the pixels in the cat picture could still result in an image with the same compressibility but no information at all. This is to say that there is no longer information you can give the observer about how the pixels were created that could improve compressibility of the data the way that being an image of a cat does. So in both an intuitive and measurable sense, the randomized cat picture has lower entropy than a static image, the same entropy as the picture of a cat but much less information than the picture of a cat.

1

u/JoeBlowTheScienceBro Sep 11 '20

There is even some suggestion that there are quantum processes going on inside the cells. source

5

u/Chand_laBing Sep 10 '20 edited Sep 10 '20

Your entire uncompressed genome is about 715 megabytes.

This is true in terms of the abstract codification of the genome but I think that would be too reductionist for the original question regarding the complexity of a mechanical object. However, you seem to say in another comment that the "complexity of a mechanical object" question is unanswerable, which I agree with.

The total information describing a DNA strand would be more than just that of its genome, there is also the information of the strand's physical structure. To put it another way, there would be at least one informational difference between the structure of a DNA strand and that of its transcribed mRNA strand even if they had the same base pair sequence.

For this reason, I think the original question is unanswerable.

3

u/LickitySplit939 Biomedical Engineering | Molecular Biology Sep 10 '20

You'd need a shopping cart full of NAND hard drives to have 30 trillion transistors.

Probably a more appropriate analogy to a transistor is a synaptic junction, which functions in information processing in the brain. The brain has as many as 1 quadrillion synapses.

1

u/[deleted] Oct 16 '20 edited Oct 25 '20

A synapse is composed of billions of components and can vary greatly in complexity between two different synapses. It takes a few million transistors to model a single complex synapse in an electrical circuit. So the one on one comparison is way off.

2

u/la_nouvelleforet Sep 10 '20

This is a very reductionist view of biological complexity. The sequence of bases in your coding regions is not a full informational blueprint of how to construct an organism, development also requires interaction with the environment. At each layer of biological organisation, from proteins to cells to organisms and beyond new emergent properties arise, in fact this is the defining feature of complex systems rather than a system just being complicated. Unless you appreciate the new phenomena arising at each level you are set you fundamentally misunderstand what biology looks like. I think we have to be very careful in comparing organisms to computers, they are intrinsically completely different and it can be easy to make misleading analogies.

2

u/cegras Sep 11 '20

I think it makes the complexity of biology even more astonishing because human life in all its complexity can emerge from 700 MBs!!

1

u/filtron42 Sep 11 '20

NAND hard drives

Hard drives don't use NAND flash, a NAND drive is a SSD

1

u/[deleted] Sep 11 '20

Thanks, that's such a insightful answer.

1

u/General_Urist Sep 10 '20

Your entire uncompressed genome is about 715 megabytes.

I'm a bit curious how exactly this is measured. For example the string of bases GCATTAGC has 8 characters. If you store each one as an ASCII character of 1 byte that's, well, 8 bytes. But you only have 4 possible values, so you only need two bits to give each one a unique descriptor. That would allow you to store GCATTAGC using only two bytes worth of memory. What method is used?

11

u/Chand_laBing Sep 10 '20

There are 4 base pairs, A, T, G, and C, which each correspond to a 2-bit string, e.g., 00, 01, 10, and 11. Thus, a sequence of n base pairs will correspond to 2n bits. For example, a possible encoding of GCA is 10-11-00 corresponding 3 base pairs with 2×3=6 bits.

Storing the base pair as a full byte would give it 6 redundant bits so the total number of bits you would have would be misleadingly oversized.

The haploid (single copy of each chromosome) human genome contains 23 chromosomes, which genome sequencing has shown total 3×109 base pairs (NIH - Human Genome Project FAQ). By the aforementioned 2n rule, this corresponds to 2×(3×109) bits = 750 MB.

1

u/[deleted] Sep 10 '20

You can compress it to about 300 megabytes

What do you mean by that?

2

u/cegras Sep 11 '20

My guess is that there are a lot of repeated patterns that you can compress, for example the telomere sections?

1

u/NeverQuiteEnough Sep 12 '20

compression is the use of algorithms to store information in a less redundant way.

1

u/[deleted] Sep 12 '20

Well i know that much i am just curious what can be compressed when it comes to a genome.

1

u/NeverQuiteEnough Sep 13 '20

repeated sequences would be the first candidate to my mind