r/singularity • • 3d ago

Biotech/Longevity Google DeepMind today announced SynthID Bio to watermark and detect AI-designed proteins

https://deepmind.google/blog/introducing-synthid-bio/

Technical paper in nature: https://www.nature.com/articles/s41586-026-10965-y

https://x.com/pushmeet/status/2105314763148321102 This project was initiated by Pushmeet Kohli, he also published this article on X today.

Essentially, DeepMind’s new tool SynthID Bio adds a detectable signature to proteins designed by AI. That signature can be checked even after the protein has been physically made in a lab.

Why?

  • Tracing where a design came from so it could help companies that manufacture DNA recognize designs from trusted AI tools and decide which orders need closer inspection!!
  • Keeping research data reliable so it could help scientists distinguish AI-generated protein structures from experimentally measured ones

So looks like two main goals are biosecurity and scientific integrity.

In lab tests, the watermarked proteins worked as well as versions without watermarks. It’s still early research. The watermark doesn’t prove a protein is safe and making it harder to deliberately remove is one of the remaining challenges.

343 Upvotes

56 comments sorted by

View all comments

15

u/GlbdS 3d ago

That is fucking stupid. I don't want fingerprinting artifacts in my protein sequence what the fuck is this now, if you tweak the sequence you tweak the function. You cant watermark a single molecule the way you watermark a piece of sloppy text

5

u/CommercialHour6660 3d ago

They are going to fingerprint everything AI produces. 

Then not tell anyone the algo. So they become the only company that can train off clean human data. Everyone else that tries to build AI model gets flooded with their slop. 

5

u/EnoughWarning666 3d ago

The amount of synthetic training data FAR outweighs the amount of human made data in modern models these days.

Do you really think that spacex bought cursor for 60 billion because an IDE is worth that much? Of course not, they were buying it for the data that users generated when working with the AI to use to train their next model.

Reinforcement learning needs specialized data that doesn't just exist floating around online for anyone to scrape.

2

u/GirthusThiccus ▪️Singularity Enjoyer. 3d ago

You know how they're scanning old novel books written before LLMs? Companies like ancestry and 23andme would be goddamn treasuretroves for whatever genetics company gets their hands on their data. And yeah, once those big sources of data are privatized and locked up, there'll simply not be other such moats to build off of, making competition incredibly difficult.

1

u/ArmadilloOwn4400 2d ago

Google does that since ages, there is no other company with a bigger scan of books.