A robust pipeline for characterising multiple conformations of proteins in solution
Lead Research Organisation:
University of Sheffield
Department Name: School of Biosciences
Abstract
Proteins do most of the work in biological systems. They form the enzymes that catalyse all the reactions that occur, with remarkable selectivity and rate enhancement; they transmit signals; they form most of the structural components of the cell; they form the immune system which defends us against attack. In order to do this, they need to interact with a wide range of other molecules, for example enzyme substrates, signalling molecules, and other proteins in molecular complexes. When they do this, they change their shapes in relatively minor but functionally important ways. For example, a loop of polypeptide at the edge of the protein may change its orientation in order to bind better, or to shield an enzyme substrate from water. A particularly important example is that as far as we can tell, every enzyme changes its structure slightly when it binds to its substrate: this is often called the induced fit (or conformational selection) model. The significance of this example is that if we want to find inhibitors or allosteric modulators of enzymes, it is often the case that the most selective such molecules (and thus the best pharmaceutical drug targets) turn out to be those that bind not to the "relaxed" conformation of the enzyme but to the activated state.
The problem is that it is often not easy to identify what these states look like. The normal ways for determining the structure of a protein are X-ray crystallography, or recently the AI program AlphaFold, which is designed to predict structures that are very close to the crystal structures, and does it very well. Both of these will typically find the relaxed state, but are less good at identifying activated states. This means that for many proteins, we know what the relaxed state looks like, but we have almost no idea how many other alternative states there are, or what they look like. This proposal uses an alternative method, namely NMR. NMR is the little brother of crystallography: it is slower and typically less accurate. However, it has two big advantages: it operates in solution (not in the crystal); and what you see in NMR is what is there, as an average. So if we can tease out the constituent structures from the averaged NMR data, we can work out what conformations are present, and therefore start to target them.
This proposal sets out to do this. The novelty in the proposal comes from two directions. First, we avoid a lot of the most tedious and time-consuming parts of NMR structure calculation by starting with the AlphaFold prediction. Second, we propose a different way of comparing structures to experimental data, which gets round most of the technical difficulties in previous methods, and is called the NOE R factor in this proposal. It is a better method than current approaches because it provides much stronger discrimination between right and wrong structures.
We have selected four proteins for study, which are all known to have alternative structures of different kinds. We propose to develop a semi-automatic pipeline for analysis, with minimal user intervention, which produces a detailed list of the structures present in solution. The most exciting aspect of the proposal is that there is currently no protein for which we can confidently say that we know what conformations it can adopt, or even how many different structures exist. This proposal will generate this information for four proteins, and thus for the first time allow us to start making some general rules about what conformations may be present for any protein, and in what proportions. We expect that this will encourage the design of novel inhibitors and allosteric effectors that bind to activated states.
The problem is that it is often not easy to identify what these states look like. The normal ways for determining the structure of a protein are X-ray crystallography, or recently the AI program AlphaFold, which is designed to predict structures that are very close to the crystal structures, and does it very well. Both of these will typically find the relaxed state, but are less good at identifying activated states. This means that for many proteins, we know what the relaxed state looks like, but we have almost no idea how many other alternative states there are, or what they look like. This proposal uses an alternative method, namely NMR. NMR is the little brother of crystallography: it is slower and typically less accurate. However, it has two big advantages: it operates in solution (not in the crystal); and what you see in NMR is what is there, as an average. So if we can tease out the constituent structures from the averaged NMR data, we can work out what conformations are present, and therefore start to target them.
This proposal sets out to do this. The novelty in the proposal comes from two directions. First, we avoid a lot of the most tedious and time-consuming parts of NMR structure calculation by starting with the AlphaFold prediction. Second, we propose a different way of comparing structures to experimental data, which gets round most of the technical difficulties in previous methods, and is called the NOE R factor in this proposal. It is a better method than current approaches because it provides much stronger discrimination between right and wrong structures.
We have selected four proteins for study, which are all known to have alternative structures of different kinds. We propose to develop a semi-automatic pipeline for analysis, with minimal user intervention, which produces a detailed list of the structures present in solution. The most exciting aspect of the proposal is that there is currently no protein for which we can confidently say that we know what conformations it can adopt, or even how many different structures exist. This proposal will generate this information for four proteins, and thus for the first time allow us to start making some general rules about what conformations may be present for any protein, and in what proportions. We expect that this will encourage the design of novel inhibitors and allosteric effectors that bind to activated states.
Technical Summary
This ambitious proposal aims for the first time to determine the minimal set of conformations required to represent the structure in solution, for 4 proteins. We note that this is completely different from typical NMR ensembles, which contains a set of structures each of which is an approximation to the single best average structure. We will do this by NMR. We start from the AlphaFold structure and validate it (in itself a very useful goal). We then generate a large number of alternative conformations, from which we select the minimal group. This is done using an NOE R factor, which compares experimental NOEs to those calculated from structures. R factors have been proposed before, but this is different because it uses NOEs detected on 13C, which has the crucial advantage of having no diagonal and therefore having a much smaller dynamic range than conventional NOESY, as well as better resolution and more uniform lineshapes, and is therefore better for quantitative analysis. We add conformations one at a time, each one giving the biggest improvement in R factor, up to a maximum of 20, and then refine the populations. This will indicate how many conformations are present in solution and how different they are from each other - information not known for any protein so far. We will then validate the structures to check they are correct and better than conventional ensembles, using a range of standard measures - which will also serve to evaluate the measures and improve our validation tools. We will construct a pipeline to streamline the process and make it available. We will also check if adding RDCs and hydrogen bond restraints improves the conformations. An understanding of the set of alternative conformations actually populated in solution will permit the design of ligands that bind to these states, which are likely to be good inhibitors or allosteric modulators.
Organisations
People |
ORCID iD |
| Michael Williamson (Principal Investigator) |