Subscribe for the Newsletter

Mobile Navigation

Designing Drugs Against Moving Proteins: Drug Discovery Enters the Era of Protein Conformational Ensembles

Designing Drugs Against Moving Proteins: Drug Discovery Enters the Era of Protein Conformational Ensembles

Jul 6, 2026PAO-06-26-PA-17

Key Takeaways

  • Protein conformational ensembles provide a more complete view of drug targets than single static structures by capturing multiple interconverting states.

  • Generative AI and molecular dynamics can help identify cryptic binding pockets and transient conformations that may support new drug-discovery strategies.

  • State-selective inhibitors and allosteric modulators can exploit specific protein conformations to improve selectivity and control target function.

  • Protein dynamics can explain some drug-resistance mutations that do not directly alter the binding site.

  • Integrating cryo-EM, NMR, molecular dynamics, and generative modeling will be essential for validating dynamic protein states and translating them into drug candidates.

From the Protein Structure Problem to the Protein Motion Problem

The ability to infer a protein’s three-dimensional structure from its amino acid sequence has changed the scale at which structural hypotheses can enter drug discovery. Deep-learning systems can now produce highly accurate models for many proteins that previously lacked experimentally determined structures, giving discovery teams a starting point for target assessment, binding-site analysis, mutational interpretation, and ligand design.1

That achievement does not make proteins static. A predicted or experimentally determined structure generally captures one configuration of a molecule that can bend, rotate, open, close, partially unfold, and reorganize as it performs its biological functions. Even when a structural model accurately represents a highly populated state, it may not show every conformation relevant to ligand binding, catalysis, regulation, or disease. Proteins often occupy distributions of interconverting states, some common and readily observed, others rare and transient. The latter can still determine whether a drug binds, whether an allosteric signal propagates, or whether a resistance mutation disrupts an inhibitor’s mechanism.2,3

This distinction has practical consequences. A single structure may display a well-defined active site, but it may also conceal a binding pocket that opens only briefly. It may show an inactive state while omitting a sparsely populated active conformation or depict an unbound protein whose geometry changes substantially in the presence of a ligand. A mutation located far from a binding site may appear innocuous in a static comparison but alter the stability or accessibility of the state that a drug requires.

Structure prediction asks what a protein can look like. Ensemble prediction seeks to determine which conformations it visits, how their populations are distributed, and, in more complete models, how the protein moves between them. The growing interest in conformational ensembles reflects a recognition that some of the most important drug-discovery questions concern not only molecular shape, but molecular motion.

What Makes a Conformational Ensemble Drug-Relevant?

A conformational ensemble is a distribution of structural states available to a protein under a defined set of conditions. The differences among those states can be subtle, such as side-chain rotations or local loop movements, or extensive, such as domain opening, activation-loop rearrangement, partial unfolding, or reorganization of an intrinsically disordered region. The relevant ensemble may also change with temperature, pH, posttranslational modification, membrane environment, binding partners, or ligand occupancy.3,4

Three properties of an ensemble are key to drug discovery. The first is structural diversity: which conformations are accessible at all. The second is population: how frequently each state occurs. The third is kinetics: how rapidly the protein moves between states and which pathways connect them. A set of plausible structures may describe diversity without accurately representing populations or transition rates. That distinction becomes important when a drug binds a rare state, when access to a pocket depends on a slow transition, or when a mutation changes the balance among states without introducing a visible structural defect.

The energy-landscape model offers a useful way to understand these relationships. Relatively stable conformations occupy favorable regions of the landscape and therefore appear more often. Less favorable states may be sparsely populated, but they are not necessarily unimportant. A ligand can preferentially bind one of those states and stabilize it, shifting the distribution of the entire ensemble. A mutation can produce a similar redistribution by favoring one basin, destabilizing another, or changing the energetic barrier between them.

Ligand binding can involve both conformational selection and induced fit. In conformational selection, a compound recognizes a state that already exists within the unbound ensemble. In induced fit, interactions formed during binding promote additional structural rearrangement. These mechanisms are not mutually exclusive, and some binding processes include an initial selection step followed by further ligand-driven change.5

The drug-discovery value of an ensemble therefore depends on more than generating numerous structural variants. The key questions are whether those variants correspond to physically and biologically relevant states, whether their relative prevalence is credible, and whether the model captures the conformational changes that control ligand recognition or function.

From Molecular Dynamics to Generative Ensemble Models

Molecular dynamics (MD) has long provided a computational route from static structures to time-dependent behavior. By simulating atomic motion under a physical force field, MD can produce trajectories that reveal local fluctuations, domain movements, transient interactions, pocket-opening events, and ligand-entry pathways. It can also help explain how mutations alter flexibility or communication between distant regions of a protein.3,5

However, this value comes with substantial limitations. Conventional simulations can be computationally expensive, particularly when a biologically important event occurs on a timescale much longer than the simulated trajectory. Rare transitions may require the system to cross high free-energy barriers, leaving important states unexplored. The resulting ensemble also depends on the force field, starting structure, solvent model, and other simulation choices. Enhanced-sampling and adaptive-sampling methods can improve coverage, but they add methodological complexity and do not eliminate the need for validation.2,3

Generative modeling offers a different route. Rather than calculating every intermediate timestep in a trajectory, a generative model learns a statistical distribution from training data and then draws new conformations from that learned distribution. The training information may include MD trajectories, experimental structures, sequence data, physical constraints, or combinations of these inputs. In principle, this approach can produce many structurally independent samples without repeatedly simulating the full path between them.

Early work demonstrated direct generation of physically realistic ensembles from coarse-grained simulations of intrinsically disordered peptides, establishing that machine learning could approximate conformational distributions rather than predict only a single structure.2 Subsequent studies explored whether learned representations could transfer to sequences and conformations beyond those used directly in training and whether models could generate atomistic ensembles conditioned on variables like temperature.4,6

More recent systems have expanded the ambition of the approach. BioEmu was developed to generate equilibrium-like ensembles for folded proteins at scale and was reported to reproduce motions associated with cryptic-pocket formation, local unfolding, and domain rearrangements.7 Such capabilities could allow researchers to generate candidate states rapidly, identify unusual conformations, and then direct more expensive simulations or experiments toward the most promising hypotheses.

These systems remain emerging tools rather than general solutions to protein dynamics. Their performance can depend on the protein class, training distribution, structural resolution, thermodynamic conditions, ligand state, and available validation data. A generated ensemble may look physically plausible while misrepresenting rare states, relative populations, or transition kinetics. The strongest near-term role for generative models may therefore be to expand conformational search and prioritize questions, while physics-based calculations and experiments determine which predictions deserve confidence.

Revealing Cryptic Pockets and Ligand-Induced States

One of the clearest applications of ensemble thinking is the search for binding sites that are absent from commonly observed structures. A protein may appear to lack a suitable small molecule pocket when assessed in a ground-state or unbound configuration, yet structural fluctuations can create cavities that open only intermittently. These cryptic pockets may expose interaction surfaces that are unavailable in the dominant structure and provide opportunities to modulate targets that otherwise appear difficult to drug.3,8

MD can reveal pocket-opening events by allowing the protein to move away from its starting configuration. The challenge is that the relevant transition may be rare, and sufficiently broad sampling across many targets can become expensive. PocketMiner illustrates a hybrid approach in which information derived from simulations is used to train a graph neural network to identify residues likely to participate in cryptic-pocket formation from a single input structure. In its reported test set, the method identified experimentally confirmed pockets and operated much faster than the simulation-based comparison methods used in the study.8

The implication is not that machine learning removes the need for dynamics. Instead, expensive dynamic information can be converted into a scalable screening tool. A model can scan many structures for regions that merit attention, after which focused simulations, biophysical studies, and medicinal chemistry can test whether a predicted pocket actually opens and binds useful ligands.

Ligand-specific structural change presents a related challenge. A pocket may not exist in the unbound ensemble at sufficient population to appear in a conventional structure, but interactions with a ligand may promote or stabilize its formation. DynamicBind models protein and ligand conformations together rather than treating the receptor as rigid. In a reported case involving SETD2, the system recovered a ligand-associated pocket that was blocked in the starting AlphaFold model.9

These examples illustrate different routes to the same discovery problem. One approach predicts where fluctuations may create an opening. Another predicts how a particular ligand may reshape the protein. Generative ensemble models may add a broader inventory of candidate states before a compound is selected. None of these methods, however, establishes druggability on its own. Detecting a cavity does not prove that it can bind a ligand with sufficient potency, selectivity, pharmacological effect, or developability. The value lies in revealing structural opportunities that a single snapshot would never place in front of the discovery team.

Designing State-Selective and Allosteric Modulators

Different conformations of the same protein can present distinct pocket geometries, interaction networks, electrostatic environments, and solvent exposure. A compound may therefore bind much more strongly to one state than to another. This creates the possibility of designing drugs that control function by stabilizing a selected member of the ensemble.

State-selective inhibitors can trap a protein in an inactive configuration, prevent access to an active state, or distinguish among closely related proteins whose dominant structures appear similar. Their activity depends not only on the local contacts formed in a binding site, but also on whether the target can adopt and maintain the conformation the compound recognizes.

The binding of imatinib to Abl provides a detailed example. The inhibitor initially recognizes an autoinhibitory DFG-out conformation, after which a larger rearrangement of the activation loop completes the bound complex. The process therefore includes recognition of a pre-existing state followed by additional structural change.5 From a medicinal chemistry perspective, the binding site cannot be separated from the pathway that makes it accessible.

Allosteric modulators extend this logic beyond the primary active or orthosteric site. By binding at a distant location, an allosteric ligand can favor an inactive state, disrupt domain coupling, alter the accessibility of a catalytic region, or reshape a separate binding pocket. The effect arises from redistribution of the conformational ensemble rather than from direct competition at the principal functional site.

Cryptic and allosteric pockets may overlap, but the concepts describe different properties. A cryptic pocket is defined by its conditional structural accessibility. An allosteric site is defined by its functional coupling to another region of the protein. A transient cavity may serve as an allosteric site, but not every transient pocket regulates distant activity, and not every allosteric site is hidden in the dominant structure.

The Abl system also illustrates how allostery could intersect with resistance. Binding at the myristoyl site stabilized a region destabilized by certain resistance mutations, suggesting that an allosteric ligand can, in some circumstances, compensate for an unfavorable conformational shift.5 This does not establish a universal strategy for reversing resistance, but it demonstrates how an ensemble-level mechanism can reveal intervention points beyond the original drug-binding site.

An ensemble-centered design program would therefore optimize more than affinity for a static pocket. It would ask which state a compound recognizes, which state it stabilizes, how it changes the broader population distribution, and whether that redistribution produces the desired pharmacology.

Conformationally Selective Antibodies and Dynamic Epitopes

The relevance of conformational ensembles extends beyond small molecules. Antibodies recognize three-dimensional epitopes whose geometry and accessibility can change as a protein moves. A surface that is buried in the dominant state may become exposed in a rare conformation, while a flexible region may present several distinct antigenic shapes.

Conformation-selective antibodies can exploit these differences. They may bind preferentially to one state, stabilize a low-population conformation, or report whether a small molecule has produced a desired structural change. They can also serve as structural chaperones, holding dynamic proteins in configurations that are easier to characterize experimentally.

Conformation-locking antibodies developed against KRAS demonstrate this principle. These antibodies distinguished and stabilized rare or inhibitor-associated states, including an open conformation linked to covalent inhibitor binding in KRAS G12C.10 Their value was not limited to recognizing the target. By locking specific conformations, they provided tools for studying inhibitor mechanisms and characterizing states that would otherwise remain difficult to observe.

Generative ensemble models could eventually support this type of discovery by identifying state-specific epitopes, transiently exposed surfaces, or rare conformations worth stabilizing. Experimental antibody selection could then test whether those states are accessible and whether their stabilization changes target behavior.

A clear distinction remains necessary between an antibody used as a mechanistic probe and a clinically viable therapeutic. Selective recognition of a rare conformation does not by itself establish favorable potency, tissue penetration, safety, manufacturability, or pharmacokinetics. Even so, conformationally selective antibodies show that dynamic states can become practical discovery reagents rather than remaining abstract features of an energy landscape.

When Resistance Rewrites the Conformational Landscape

Resistance is often explained through direct structural changes: a mutation introduces a steric clash, removes a hydrogen bond, alters electrostatics, or reduces complementarity within the binding site. Those mechanisms remain important, but they do not account for every loss of inhibitor activity.

Some mutations act by changing the stability, accessibility, or lifetime of a drug-compatible state. Others alter the pathway required to reach that state or transmit effects from a distant region through a dynamic network. In these cases, the resistant protein may retain a structurally recognizable binding site while visiting the relevant conformation less often or paying a greater energetic cost to adopt it.

The Abl–imatinib system provides a strong example. Several resistance-associated mutations examined in the study were located outside the immediate drug-contact region. They destabilized part of the C-lobe involved in the large conformational transition required for imatinib binding, making the final inhibitor-compatible state less favorable. The same mutations had smaller effects on dasatinib, which binds through different conformational requirements.5 Resistance in this case was selective for the structural pathway used by one inhibitor rather than evidence that the target had become universally undruggable.

Studies of human immunodeficiency virus type 1 (HIV-1) protease similarly show that resistance mutations can alter the dynamics of an inhibitor-bound enzyme without producing only obvious local changes in a static structure.11 When multiple mutations accumulate, their combined effects may redistribute behavior across the protein rather than create one discrete defect.

Coupling MD with machine learning offers one way to interpret such distributed mechanisms. In analyses of darunavir potency across HIV-1 protease variants, simulation-derived features from multiple regions of the enzyme were linked with experimentally measured losses in inhibitor activity.12 The approach does not prove that current models can forecast every clinically important mutation, but it shows how dynamic information can expose relationships that a binding-site comparison may miss.

A similar integrative strategy has been used to investigate the H172Y mutation in severe acute respiratory syndrome coronavirus 2 main protease, combining simulation with measurements of protein stability, catalytic activity, and inhibitor susceptibility.13 Such studies reinforce the need to connect computational ensembles with experimental phenotypes.

Ensemble analysis may therefore contribute to resistance research in three ways: by explaining why a known mutation reduces activity, by identifying alternative states that remain targetable, and by revealing allosteric regions that could bypass or compensate for the disrupted pathway. It should not be presented as a universal theory of resistance or as a mature system for predicting every future variant. Its more defensible role is to illuminate mechanisms that static structures leave unresolved.

Integrating Cryo-EM, NMR, Molecular Dynamics, and Generative Modeling

No single method can capture every structure in an ensemble, determine its exact population, and measure the kinetics connecting all states. Each technology observes a different portion of the problem.

X-ray crystallography can provide high-resolution atomic structures but often favors well-ordered conformations compatible with crystallization. Cryogenic electron microscopy (cryo-EM) can separate multiple structural classes in suitable macromolecular systems and reveal large-scale heterogeneity. Nuclear magnetic resonance (NMR) spectroscopy can probe proteins in solution and provide information about interactions and conformational exchange across a range of timescales. MD supplies atomically detailed trajectories and mechanistic hypotheses but remains limited by sampling and force-field accuracy. Generative models can expand structural exploration rapidly, although their outputs reflect the strengths and biases of their training data.3

The most credible ensemble workflows will combine these approaches. Experimental measurements can constrain a computational ensemble, adjust the relative weighting of candidate states, or reject conformations inconsistent with observed data. Simulations can connect experimentally resolved structures, propose intermediate states, and suggest mechanisms for transitions that cannot be observed directly. Generative models can produce a broad set of hypotheses, helping researchers decide where to invest more intensive simulation or experimental effort.

A practical discovery workflow might begin with all available experimental structures, predicted models, ligand complexes, sequence variants, and biochemical information. MD, enhanced sampling, or generative modeling could then produce an initial set of states. Cryo-EM, NMR, hydrogen–deuterium exchange, mutational analysis, binding kinetics, or other biophysical measurements could test and refine that set. Researchers could search the resulting ensemble for cryptic pockets, dynamic epitopes, allosteric networks, and mutation-sensitive conformations before screening compounds across a selected group of states.

The process would remain iterative. New ligands would become experimental probes of the ensemble, while observed binding modes and functional responses would update the computational model. A compound that stabilizes an unexpected state could reveal a new mechanism, and a failed prediction could expose weaknesses in the assumed landscape.

This framework is a long way from being a standardized industrial process today. Its importance lies in shifting validation from a final checkpoint to a continuous exchange between modeling and experiment. Ensemble generation becomes most useful when every predicted state carries an evidentiary context: why the model produced it, what data support it, and which experiment could disprove it.

From Structural Snapshots to Dynamic Drug Design

Ensemble-based discovery still faces substantial scientific and practical barriers. Generating diverse structures is easier than proving that their populations are accurate. A model may reproduce broad conformational variation while misrepresenting transition kinetics or rare states. Training data may overrepresent certain folds, motions, or experimental conditions. Apo ensembles may fail to capture ligand-induced rearrangements, and models generated under one temperature, pH, membrane context, or cofactor state may not transfer cleanly to another.

Resolution also matters. Coarse-grained models can reveal global movements but omit local interactions needed for medicinal chemistry. Atomistic models provide more detail but introduce greater computational cost and more opportunities for force-field error. Rare states may be particularly difficult to validate because their low population makes them challenging to isolate experimentally.

Drug-discovery interpretation adds another layer of uncertainty. A predicted pocket may open but remain too shallow, flexible, hydrated, or nonspecific to support a useful ligand. A compound may bind a modeled state without producing the expected functional response. Conformational selectivity may improve potency while creating liabilities elsewhere in the target’s biological network. Prospective evidence that ensemble models improve real discovery decisions will be more persuasive than retrospective agreement with known structures or mechanisms.

Generative systems are therefore unlikely to replace MD, cryo-EM, NMR, crystallography, or established biophysical methods. Their strongest role may be connective: expanding the set of conformations that can be considered, integrating evidence from multiple sources, prioritizing rare states for validation, and making dynamic reasoning more scalable.

The field cannot yet produce a complete and quantitatively exact conformational landscape for any arbitrary target on demand. That standard may not be necessary for meaningful progress. A partial ensemble that is experimentally grounded and pharmacologically informative could reveal a transient pocket, distinguish the binding requirements of two inhibitors, explain a distal resistance mutation, or expose an epitope available only in a disease-relevant state.

The first structural biology revolution made protein structures broadly accessible. The next may make it practical to design against the changing populations those structures represent.

References

1. Jumper, John, et al. Highly Accurate Protein Structure Prediction with AlphaFold.” Nature. 596: 583–589 (2021).

2. Janson, Giacomo, et al.Direct Generation of Protein Conformational Ensembles via Machine Learning.Nature Communications. 14: 774 (2023).

3. Wei, Haixin, and J Andrew McCammon.Structure and Dynamics in Drug Discovery.” npj Drug Discovery. 1: 1 (2024).

4. Janson, Giacomo, Alexander Jussupow, and Michael Feig. “Deep Generative Modeling of Temperature-Dependent Structural Ensembles of Proteins.” Communications Chemistry. 8: 354 (2025). https://doi.org/10.1038/s42004-025-01737-2

5. Ayaz, Pelin, et al.Structural Mechanism of a Drug-Binding Process Involving a Large Conformational Change of the Protein Target.” Nature Communications. 14: 1885 (2023).

6. Janson, Giacomo, and Michael Feig.Transferable Deep Generative Modeling of Intrinsically Disordered Protein Conformations.” PLOS Computational Biology. 20: e1012144 (2024).

7. Lewis, Sarah, et al.Scalable Emulation of Protein Equilibrium Ensembles with Generative Deep Learning.Science. 389: eadv9817 (2025).

8. Meller, Artur, et al.Predicting Locations of Cryptic Pockets from Single Protein Structures Using the PocketMiner Graph Neural Network.” Nature Communications. 14: 1177 (2023).

9. Lu, Wei, et al. DynamicBind: Predicting Ligand-Specific Protein–Ligand Complex Structure with a Deep Equivariant Generative Model.” Nature Communications. 15: 1071 (2024).

10. Davies, Christopher W, et al.Conformation-Locking Antibodies for the Discovery and Characterization of KRAS Inhibitors.” Nature Biotechnology. 40: 1325–1333 (2022).

11. Cai, Yufeng, et al. Drug Resistance Mutations Alter Dynamics of Inhibitor-Bound HIV-1 Protease.” Journal of Chemical Theory and Computation. 10: 3438–3448 (2014).

12. Leidner, Florian, Nese Kurt Yilmaz, and Celia A Schiffer.Deciphering Complex Mechanisms of Resistance and Loss of Potency through Coupled Molecular Dynamics and Machine Learning.” Journal of Chemical Theory and Computation. 17: 2054–2064 (2021).

13. Clayton, Joseph, et al.Integrative Approach to Dissect the Drug Resistance Mechanism of the H172Y Mutation of SARS-CoV-2 Main Protease.Journal of Chemical Information and Modeling. 63: 3521–3533 (2023).

Nice Insight is the market research division of That's Nice LLC, the leading marketing agency serving life sciences.
Subscribe for the newsletter
© 2026 PHARMA'S ALMANAC. All rights reserved.