How Physics, AI, and Supercomputers Reveal How Proteins Work | Biology
Kathrine
Thank you for joining me today. To start, would you mind introducing yourself, your field and the main questions your research group studies?
Professor Carnevale
Of course. So first of all, thank you very much, Kathrine. So excited to be here with you.
So my name is Vincenzo Carnevale. I am an Associate Professor in the Department of Biology at Temple University. So my original training was in theoretical physics, but today I would describe myself more as a computational biophysicist.
So my research group studies how the physical interactions among atoms and molecules gives rise to biological functions. And proteins are the major focus of our research. So we're interested in many diverse questions, such as how does a protein's amino acid sequence determines the three-dimensional structure?
So how does that structure then in turn move and perform a function? How do mutations alter these properties? And over much longer timescales, then we are interested in how these proteins evolve while preserving this function or sometimes even adapting and transforming, changing their function.
And we do this using several approaches. So we use molecular dynamic simulations. We use techniques from statistical physics, bioinformatics, and more and more we use machine learning and generative artificial intelligence.
So a major part of our work concerns membrane proteins and in particular ion channels, but we also study protein evolution, protein design, processes called self-assembly processes. And in other system, in various system, and, you know, the common theme is usually how relatively simple interactions produce complex collective behavior.
Kathrine
So, yeah. Next, what originally inspired you to pursue using physics and computation to study biological systems?
Professor Carnevale
Yeah, I mean, so personally, I mean, I was initially very much attracted to physics as a student because, you know, it seeks relatively simple general principles to explain complicated phenomena. And then during my training, you know, I became increasingly interested in biology, exactly because biological systems are extraordinarily complex. However, after all, so they have to obey to the law of physics.
So proteins, you know, still obey thermodynamics, electrostatics, mechanics, they can be described using statistical physics. So what I was really fascinated at some point during my career was by the possibility of using this general principle to understand biological organization. And so even though proteins may contain thousands of atoms and, you know, they can look at, you know, extremely and hopeless, hopelessly complex systems, ultimately each atom has to obey, you know, the rules of physics.
And most importantly, and even more interesting than that is the fact that when you look at these complex systems, so their collective behavior, so the fact that they are a bunch of atoms, but collectively these atoms produce very much specific and somewhat simple behavior. So ultimately, for instance, going back to the example of membrane proteins and ion channels, for instance, an ion channel ultimately, you know, even though it's made out of many, many, many, many atoms, it turns out to have a behavior that is not so dissimilar from one of our, let's say, of the transistors that we use in electronic devices. So ultimately, you know, the behavior turns out to be simpler than, you know, the complexity of the system itself.
And so this is the hallmark of, you know, complex systems that may have, you know, the so-called emergent properties. And therefore, this behavior that can be modeled using relatively simple equations and, you know, relatively simple mathematics. Now, in between those two, right, so there is, you know, an entire complexity due to the fact that the system is big, there are many atoms and there is an emergent property that seems to be obeying to relatively simple rules.
So in between this and that, there is, you know, a gap. So how can we reconcile these two worlds, right? So these atoms that move and, you know, we somewhat, we find challenging to follow the motion of all these atoms, but up to this, let's say, emerging behavior, we see something that is on-off and, you know, the response to an electrostatic field, as in the case of the ion channels.
Well, to bridge this gap, then it's the entire, let's say, that's where, you know, our research really is focused on. And so computation is providing this bridge. So using computers, using supercomputers to just tame this complexity and match or bridge the microscopic physical rules to this biological behavior.
And this is done through simulations, so-called molecular dynamic simulations. So in practice, we sort of built into the computer the laws of physics, and then we observe processes, processes that typically are too small, too fast, too rare, too difficult to isolate, to be observed experimentally. And so ultimately, this is, you know, the ultimate fascination for me was this.
How can I use physics to bridge this entire and very wide gap that we have between these microscopic things that we know how to describe and instead this emergent biological scale that is still so fascinating and mysterious to some extent.
Kathrine
Absolutely. So next, can you walk us through one specific research project that your research group did where you studied protein evolution or the effects of these mutations, starting from the original research question all the way to the final conclusion?
Professor Carnevale
Sure. So one recent example involves a protein, an enzyme actually. So this is called UCHL1.
And it's a protein particularly abundant in neurons. So the normal protein, so the one that, you know, each of us has, we have in the neurons, has a rather unusual property. Its backbone forms a deep molecular knot.
So the protein is a polypeptide, can be seen as a little thread that folds into itself. But this one in folding forms a knot. It's not very common in proteins.
And we've been studying this protein because we were interested in this somewhat peculiar mathematical property, the topology of the backbone, the so-called backbone of the protein. And in studying that, so we studied not only the naturally occurring one, but also, you know, the mutation, a mutant protein in which an isoleucine is mutated into methionine. So this mutation has been associated with neurological disease.
And we wanted to understand why and if topology had anything to do with that. And, you know, at a first glance, so this, the change from isoleucine to methionine is somewhat small. So they're both hydrophobic amino acids.
And if you look at the overall folded structure of the mutant and so-called wild type, so the one occurring in, you know, the largest part of the population, well, the two structures look very much the same. However, you know, a mutation may alter some other properties, right? Not only just the structure.
So proteins are not really static objects, though they're rather dynamic. And the mutation can, in principle, we thought, alter some other aspects, not really how the protein look like on average, rather how the protein moves. And here, you know, we deployed, you know, all the most advanced techniques in terms of molecular dynamics, simulations, and use of supercomputers.
And what we realized is that, so this mutation was affecting the stability of this very peculiar topological state. So it was affecting the stability of this knot. And in particular, the mutation was promoting the so-called unknotting of the protein.
And in doing so, there was an effect downstream, we realized, which is the fact that when these proteins as a whole, in the neuron, so this enzyme is present in so-called bimolecular condensates. So it phase separates, so it forms relatively large aggregates that are reversible. So the protein can be in solution or in this so-called aggregated state.
So now when the protein is in the normal native knotted state, so this process of aggregating and disaggregating in a reversible way occurs without problems. On the other hand, when the knot is lost, so the aggregated state instead forms entangled states across chains. So these chains now become so much disordered that they form entangled aggregates that cannot be reversibly dissolved.
So there is now an important question that we are trying to address and it is whether or not, so these proteins evolved in order to have this molecular knot. And in doing so, being protected against the formation of these tangled states that could be, we think, could be one of the reasons for neurotoxicity. So in neurodegenerative disorders, at least in some of them, these mutations can exacerbate the formation of aggregates and amyloids because of these insoluble aggregates.
So ultimately, I mean, so this is one of the cases in which several scales are at play, in which simulations and molecular evolution and artificial intelligence, they all conspire, right? So to sort of getting from the microscopic interactions to these general principles that govern a cellular process.
Kathrine
So next, as you mentioned before, your work combines molecular simulations, statistical physics, generative models, and even deep learning. So how do you decide what method to use for a certain problem?
Professor Carnevale
Sure, that's a great question. So of course, we begin with the scientific question, not the method. And different methods provide different kinds of information.
So there is no single technique that is best for every problem. So when we need an anatomically detailed description of a protein, perhaps, for instance, to understand in an ion channel how an ion binds or how water molecules go through a channel, then in this case, we very often use molecular dynamic simulations. So as I was mentioning before, so this amounts to sort of encoding or coding, really, the laws of physics into a computer, and then describe the system with the maximum possible level of details and just follow the time evolution of this complex molecular system.
So these simulations then become physically interpretable. So they tell us a lot, but they can become computationally prohibitively expensive, and especially when we are trying to capture very, very slow processes or rare events. And so in these other cases, often, you know, we may use instead approaches from statistical physics because these are instead, let's say, more useful to understand general principles and typically the collective behavior.
For instance, so this is the typical scenario when we want to understand how interactions among amino acids constrain and shape protein evolution. So in addition to that, of course, and in the last few years, more and more important, more and more importance have gained machine learning and generative models. Because now, whenever, you know, we have a problem in which the search space, the space of solutions we're looking for becomes astronomically large, enormous, then we cannot necessarily simulate because, as I said, you know, perhaps the search space is too large and statistical physics may or may not be useful because perhaps we're looking for a denier resolution at a larger, let's say, degree of chemical accuracy. And then in these cases, so generative models can learn patterns, so can just learn the patterns either from, you know, experimental data or from our simulations and just, and just in this way, just extrapolate and learn from these patterns and give us new solutions. So ultimately, these generative models can propose us, for instance, molecular structures that are very likely to be plausible from the point of view of both physics and biology.
And so overall, I mean, the most productive approach is to combine all these methods. So we can generate a hypothesis with machine learning, identify candidates and then use physical simulations, therefore these molecular dynamic simulations, to test and be more quantitative. And then so we can learn general rules, as I was saying, from the collective behavior and use statistical physics to describe this behavior and understand along the way molecular evolution, because all of this is not necessarily the result of pure randomness, right?
So it's the result of natural selection.
Kathrine
Absolutely. And so next, how can those computation predictions about protein structures, mutations or functions eventually be tested experimentally?
Professor Carnevale
Yeah, that's another good question. So, you know, a good computational prediction should always be lead to an experimentable, testable statement. So otherwise, I mean, it would be a little bit empty as an activity.
So ideally, I mean, so what we aim at is typically something that could be resolved, tested or disproved by an experimental collaborator, for example. So in a simulation, we may predict how the result of a mutation destabilizes the folded state or a particular conformational state. For instance, it can alter the opening of an ion channel or it can strengthen the interaction between two proteins.
It can change the affinity with a binding partner. And so an experimental collaborator then can produce the mutant protein and then test this prediction through biochemical measurements or structural techniques, electrophysiology, microscopy, whatever is appropriate to measure the phenotype of this mutated protein. So in the specific area of protein engineering, then we can use computation to propose new sequences.
So we can, so to speak, invent proteins that nature hasn't discovered yet. And then we can predict whether or not they will fold into the desired structure that we designed them to fold into or to bind to a particular target. And then, you know, in this case, the experimental testing is done by synthesizing this DNA and then have bacteria synthesize the protein.
And then after purification, this can be experimentally tested. And usually, I mean, so it's not only that, you know, a prediction gets tested, but, you know, experiments are, per se, extremely informative. And so they give us extremely valuable lessons.
Typically, this is more like a cycle. And after a round of experiments, we know what else and we should, how we should, let's say, refine our models. And also what kind of question we should be asking to our theoretical models.
Kathrine
Yeah, absolutely. And the next step of this process would be, in what way can understanding these mutations contribute to areas such as drug development, disease research, or protein engineering?
Professor Carnevale
Yeah, understanding mutations is, of course, you know, a direct connection to disease whenever these mutations are naturally occurring mutations. And so if a disease has, you know, a genetic component, then study the behavior of the protein carrying that particular mutation gives us direct access to the molecular underpinnings of that disease. So, for instance, you know, a mutation typically may shift, you know, some very subtle and delicate equilibrium between two conformational states, for instance, in an ion channel, an open and closed state or for an enzyme, the inactive versus the active state.
And so this is all useful information because it can be translated in a rather direct way, I wouldn't say in a simple way, but certainly directly into a strategy, into a therapeutic strategy. So when we have this level of understanding of the molecular process that is going wrong, then we can think about a drug molecule that may restore the, let's say, the physiological, let's say, process. And so in drug discovery, in drug development, then we use, we make every use of this kind of structure, structural and dynamic information.
Because, for instance, in this case, and my group is involved in several of these drug discovery projects. So we can, once we understand the normal functioning of biomolecules and we understand how a mutation affects this functioning, then we can start thinking about small molecules, small organic molecules that binding to the mutant can partially restore this, affect this equilibrium between these conformational states, for instance. And so we can use these principles over and over again.
So having the ability to make predictions about the affinity of a drug molecule for a specific target and understanding how binding will affect downstream the behavior of the protein is really the necessary input for the medicinal chemist to explore the chemical space and propose drug candidates. So in the end, so this sits really in the early phases of drug discovery, when really the concept of the drug itself is developed. And when the first, let's say, when the chemical properties of the drug are being identified.
Kathrine
And finally, what is a realistic project that you would recommend to high school students so that they could explore your field?
Professor Carnevale
Oh, yeah, a very realistic first step would be certainly to learn some basic Python. So the programming language that has become the universal language. So I would definitely suggest a small project in Python involving some protein structure.
So there is this protein data bank, which is a publicly accessible data bank in which all the biomolecular structure in general, so protein-based structures have been deposited. And one can download little files containing the coordinates of all the atoms of these proteins and visualize them using programs called Pymol or Chimera. And so I would suggest, first of all, starting from there.
So starting from downloading these proteins, using, learning how to use these programs, and then start to write simple Python programs to analyze some simple property of these molecular structures, for instance, by looking at the different amino acids and counting them, locating mutations, measuring distances. And the very first step, perhaps, would be to compare closely related protein sequences and protein structure, for instance, from different species and identify what is the conserved core in the structure and what is instead that has changed. One would realize, for instance, that, you know, the same protein in two different species that are closely related, for instance, human and chimpanzees are practically the same.
And already when we go to mouse, we start to notice some differences. So some amino acids will be different, the structure will have subtle variations. And going back even further in terms of time and meaning that comparing with a more distantly related species, for instance, at the extreme comparing a similar protein contained, for instance, in bacteria would reveal that along evolution, I mean, nature has explored a lot the chemical space and that's changed a lot.
And yet, you know, a good chunk of the physical properties of these biomolecules are kept constant. So one would be surprised of seeing how billions of years didn't change the overall structure, even though many, many of the amino acids are changed. So analyzing this, analyzing variations, I think would be an interesting