Google figures out how to watermark AI-designed proteins
Ars Technica · LC · trust 46/100

That’ll leave a mark Google figures out how to watermark AI-designed proteins Intended to help with biosecurity, it works with a popular AI protein design tool.
27 Credit: ALIOUI Mohammed Elamine Credit: ALIOUI Mohammed Elamine Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav AI-based tools seem to be causing security threats on a nearly daily basis, in part because we’ve been slow to recognize potential threats. One area where we seem to be ahead of the game, however, is in biosecurity.
As with most things biological, utility goes hand in hand with threats. We’ve developed increasingly sophisticated tools for designing proteins and seen some major successes, such as AI-designed enzymes that can digest plastics or block venom proteins . But these same tools could be used to make toxins or alter the behavior of viral proteins.
And the software we use to identify DNA sequences that encode potentially threatening proteins doesn’t pick out AI-designed proteins , since nobody has characterized them well enough to know that they’re threats. Nearly a year after that risk was flagged, it still wasn’t clear what anyone could do about it.
On Wednesday, the DeepMind team at Google published a research paper offering a potential solution: protein watermarking. The system creates a watermark on protein sequences themselves without compromising the protein’s function. This allows new proteins designed by trusted researchers to be identified, opening everything else up to closer scrutiny.
The work was based on Google’s SynthID tech , which can add a subtle watermark to AI-generated digital material. The watermark influences the probability of certain choices the AI makes, and that bias ends up systematically distributed throughout the product, whether it’s text or images. Because you can’t identify the watermark without knowing how it was encoded, it’s impossible to remove. And because it’s distributed throughout the image, it can survive basic exporting, resizing, and so on.
It’s pretty easy to see how this can work with subtle differences in things like the colors of a photo. It’s a whole lot harder to see how you can do it with a protein.
Proteins are composed of only 20 amino acids, any of which could be essential for structural integrity or catalytic activity. While some of these amino acids are chemically similar (like leucine and isoleucine), others have opposite charges. Many proteins have significant regions where limited changes to their amino acid sequence are tolerable and other areas where even a slight deviation from the existing sequence inactivates the protein.
Proteins are also small. While images often contain millions of pixels, proteins containing 500 amino acids are fairly large. That’s a lot less raw material to hide any sort of signal in.
So it wasn’t clear the SynthID tech would work; it might be unable to hide sufficient signal in a typical protein, or, if it crammed in enough information to create a functional watermark, the resulting proteins might be inactive. The only way to find out was to try it.
To better understand how the system works, it helps to know a bit about protein chemistry. Amino acids have a constant section primarily made of two carbon atoms linked to a nitrogen. A protein is made by linking up a series of these constant sections to form a long chain called a backbone. Each amino acid also has what is called a side chain hanging off it. These can range in complexity from a single hydrogen atom to large ring structures; the side chains can be basic hydrocarbons, acids, bases, and more.
The side chains determine how the backbone folds up in three-dimensional space. In a soluble protein, all of the pure hydrocarbon side chains end up packed into the center, while the acids, bases, and hydroxyl-containing side chains face the water. Once the protein folds into its final form, the backbone will describe the protein’s overall shape.
The Google team started with one of the most popular AI protein design tools, ProteinMPNN . (This software was developed by the Baker Lab, and David Baker was honored with the same Nobel Prize that was shared with the head of DeepMind.)
ProteinMPNN works in a two-stage process. First, a separate tool describes a backbone configuration that is appropriate for the design. Next, ProteinMPNN works its way down the backbone, placing side chains one amino acid at a time. Each amino acid is chosen based on its ability to fit into the shape defined by the backbone, interact with neighboring amino acids, and fit any other constraints defined by the experiment. (Those constraints can include things like forming catalytic pockets or interacting with another protein.)
A variant of Google’s SynthID, called SynthIDBio, steps in during this process. It uses a key (similar to a cryptographic key) and the identity of the previously chosen amino acids to suggest a new one. ProteinMPNN then determines whether the amino acid suggested by SynthID works from the perspective of forming a functional protein. If it doesn’t, it rejects it. If not, it moves on. Put differently, as the system works through the backbone one amino acid at a time, it only incorporates watermark amino acids when they’re consistent with a functional protein.
One way to think about this is that, when the system comes across a location where a set of chemically related amino acids will all work (like leucine/isoleucine/valine or serine/threonine), it will use one that’s consistent with the watermark when possible. Another way to look at the process, suggested by one of the people involved in developing the system, is that it searches through the space occupied by functional proteins for the subset that happens to have a sufficient number of watermark amino acids.
As a result, the watermark is randomly distributed across the entire length of the protein, and detecting one isn’t a…
Read the original at Ars Technica →
Open in TruthVane →