Key Takeaways:
Advances in genetic engineering, x-ray crystallography techniques, and AI/ML-based computational modeling have led to great strides in therapeutic protein engineering.
In addition to traditional random and rational design approaches for the improvement of existing proteins, de novo design to generate completely novel proteins is increasingly common.
Advances in synthetic chemistry, synthetic biology, and high-throughput screening capabilities are also enabling more rapid design and development of increasingly complex and novel proteins,
Despite the amazing progress achieved to date, protein design continues to have limitations that will only be addressed with further advances in computational modeling, development of high-quality data sets, standardization, and the use of appropriate benchmarks.
Evolution of Protein Engineering
The field of protein engineering has been undergoing advances over the last several decades.1 Directed evolution approaches were first introduced in the late 1960s and early 1970s. Use of molecular biology and the polymerase chain reaction (PCR) for DNA amplification was also first applied in the 1970s. Non-natural evolution approaches were then introduced in the 1980s to early 2000s, including gene mutagenesis and saturation mutagenesis (focused randomization). Error-prone polymerase chain reaction (epPCR) technology was used to generate combinatorial libraries of antibodies in the early 1990s. DNA shuffling is another form of random mutagenesis that was introduced in the 1990s.
These random methods suffer from the need for extensive physical screening of large numbers of protein variants. To overcome this limitation, newer approaches leveraging rational protein design have emerged.
Fundamental Approaches to Protein Engineering
Both genetic engineering and X-ray crystallography techniques have enabled significant advances in protein engieneering.1 More recently, computer modeling leveraging advanced statistics and more recently artificial intelligence (AI) and machine learning (ML) have enable rational approaches to protein design. A whole host of strategies have been applied to the development of monoclonal antibodies, bispecific antibodies, novel fusion proteins, and other protein-based therapies.2
In addition to random and rational design, hybrid approaches that combine elements of both strategies are also widely used and offer specific advantages and disadvantages for therapeutic protein design.1,2 Advances in high-throughput screening and other lab automation technologies have been crucial to enabling the rapid analysis of large libraries of potential new variants. In addition, with recent advances in computational capabilities, it is now possible to not only improve the properties of existing proteins, but design entirely new proteins.
Directed evolution involves the generation of random mutations in genes that encode known proteins to generate large numbers of variants that are screened to identify those with improved properties. This approach mimics natural selection and is a random process. EP-PCR is a key technique for generating those random sequences. The advantage of this approach is that no prior knowledge of the protein structure or mechanism of action is required. This strategy can also uncover unexpected functionality that might not be identified computationally.
The rational method typically involves site-directed mutagenesis or creation of point mutations at specific locations in a coding sequence based on extensive knowledge of the structure and function (mechanism of action) of the existing protein. While this approach has a greater likelihood of resulting in proteins with improved properties, it is still often difficult to make accurate predictions using computational methods, such as molecular dynamics (MD) simulations and computational mutagenesis, and requires access to atomic- and molecular-level insights that is often not available.
Semirational (hybrid) protein design combines elements of both directed evolution and rational design methods to better guide variant generation and thus typically results in smaller libraries of higher-quality protein candidates. Often, structural and mechanistic information is used to generate targeted mutational libraries that are then subjected to directed evolution. Alternatively, data generated on variants produced via directed evolution can be used to identify structural motifs of importance, the knowledge of which is used to guide rational design efforts. These approaches are particularly useful for the engineering of highly complex proteins.
All of these methods are intended to engineer proteins with improved performance with regard to stability, specificity, and/or biological activity (or catalytic activity in the case of industrial enzymes). De novo design approaches generate new proteins by combining amino acids in specific sequences known to have different structural and functional properties. Advanced AI and ML-based models are key to achieving success using this approach. AI/ML is also being leveraged to improve hybrid approaches.
Graduating to Autonomous Protein Engineering
Autonomous platforms for protein engineering are helping to reduce the labor-intensive aspects of protein engineering and accelerate the development of novel protein designs.1,3 These systems combine robotics, self-driving laboratories, and automated learning algorithms. AI/ML approaches accelerate the design and learn phases of the design–build–test–learn cycle by predicting optimal mutations from sequence data and evaluating analytical results. High-throughput screening technologies combined with automated liquid handling and sample preparation systems and robotic workstations greatly speed the build and test phases. Together, they enable “closed-loop” optimization and accelerate protein engineering even using random mutagenesis methods.4 Some achieve protein expression in bacterial or other cell-based methods, while others use cell-free systems.
Synthetic Chemistry Advances
Advances in synthetic chemistry and structural elucidation technologies have been crucial to the structural successful generation of large libraries of novel protein variants.5 Highly specific chemical transformations are necessary to achieve the atom-level modifications required to produce customized proteins., including side-chain alterations, sequence refinements, and backbone edits.
Techniques have, for example, been developed to introduce non-natural disulfide bridges to increate thermal stability. Macrocyclization, bicyclization, stapling, and controlled oligomerization, which are considered structure-guided covalent crosslinking strategies, have also been introduced to enhance protein stability and achieve desired folding behaviors. Incorporation of unnatural amino acids is another strategy for engineering proteins with new properties.
All of these modifications must be achieved without impacting the structures, native folding behaviors, and functions essential to the therapeutic performance of the proteins being optimized. As a result, the new methods, such as solid-phase peptide synthesis (SPPS), chemo selective ligation, and late-stage bioconjugation, are designed to be highly precise. The AminoX platform uses a synthetic biology approach to generate functional intermediates that can transfer non-standard amino acids to the ribosome in a cell-free process, enabling their incorporation into novel protein structures.6
Big Impacts of AI and ML
Some of the greatest leaps in protein engineering have occurred since the emergence of more advanced natural language and image-processing systems leveraging AI and ML have been applied to protein structure prediction. The 2024 Nobel Prize in Chemistry went to scientists advancing AI-driven de novo protein design and protein structure prediction, specifically the developers of RosettaFold diffusion (RF-diffusion) and AlphaFold2.7 AlphaFold 3 released in 2024 also predicts protein interactions with DNA, RNA, ligands, and antibodies.2
Advanced systems incorporate many types of data sources, including “protein sequences, structures, dynamics derived from molecular-dynamics simulations, functional annotations from resources such as the Gene Ontology, experimental data, and even natural-language descriptions.”8 Specific advances highlighted by some include “generative modeling of sequences, backbone structure, and atoms; tailoring general versions of such models to design proteins with specific properties; modeling for extraction of protein representations and scoring candidate protein sequences; and developing techniques for library design, including synthesis-aware approaches.”9
These computational models use vast amounts of biological data and complex statistical models to identify patterns not visible to humans.2 AI-based protein design approaches generally fall into one of two categories: those that leverage large language models and those that use diffusion algorithms often used for image generation.9 ProtGPT2 is a generative AI model trained to read and understand amino acid sequences as a language and then generate new ones. The company Profluent is developing its own large language model to generate functional protein sequences. ProteinMPNN and ThermoMPNN generate amino acid sequences that fold into the shapes of existing proteins using message-passing neural networks.1 CobaFold comprises one algorithm for creating structures, another for designing sequences, and a third for exploring protein folding in one platform to generate optimal engineered proteins.10
AI and ML algorithms are also being used to rapidly screen large sets of protein variants to determine which have optimal properties (e.g., stability, solubility, binding affinity) for a specific application.2 As with all AI/ML systems, the level of success achieved depends not only on the quality of the algorithm but the quality of the data used to train them.
Innovations keep coming, too. Just in May 2026, an upgraded ESM protein atlas generated using the improved a “protein language” model ESMFold2 was introduced.11 It goes well beyond both its previous version and the AlphaFold databases with more than one billion predicted protein structures and billions more protein sequences.
Companies using these advanced technologies include Cradle Bio, Archon Biosciences, Levitate Bio, and Cyrus Biotechnology.7 Others, such as Tamarind Bio, AI-DT, and OpenProtein.AI, offer applications that enable non-specialists to use advanced protein engineering software tools. The latter provides a web interface for uploading data and using various open-source protein engineering tools, including the firm’s Protein Evolutionary Transformer (PoET) language model.12
De Novo Protein Design
De novo protein design involves the generation of completely new proteins, often ones not possible to form via natural evolutionary processes.4 Advanced AI/ML-based computational tools are used to design proteins containing non-natural amino acids and with the ability to fold into new structures with varying functions. The algorithms are used to predict how specific amino acid sequences will behave and to construct more stable, less immunogenic proteins with highly specific features, including customized binding pockets and unique catalytic activities, among others.
The approach is particularly attractive for the development of novel protein therapeutics, such as highly specific tumor-targeting capabilities, as well as others classes of molecules including synthetic cytokines, targeted protein degraders, and immune cell engagers. It also provides the ability to easily incorporate many different nonstandard amino acids, which without the predictive capabilities supporting de novo design would be extremely challenging and require extensive time and effort to identify optimal solutions.6 OzempicⓇ/WegovyⓇ is one example of a blockbuster drug designed in this manner. It contains non-natural amino acids that enable the drug to remain stable and circulate in the blood stream for longer periods.
Moving to Function-Driven Protein Engineering
The advances in computational modeling are driving an important shift in protein engineering away from structure-focused approaches to strategies driven by desired protein function.13 De novo protein design using AI/ML algorithms is making this shift possible. No longer is extensive structural and mechanistic data for a starting protein required. Current platforms leveraging generative algorithms can optimize specific protein activities and functions and then generate novel sequences (including nonstandard amino acids) to achieve them, leading to new proteins with new structures. LigandMPNN, for instance, enables ligand-binding protein design, evaluating interactions of protein sequences with many different types of ligands, from small molecules and nucleic acids to metals.
Increasingly, protein engineering is being achieved using multiple types of models, such as large language, graph networks, and molecular dynamics, often in combination with high-throughput closed-loop iterative synthetic biology-based systems.13 As an example, ESM-3 simultaneously uses protein sequence, 3D structure, and functional generation models. Advanced systems also enable analysis of dynamic, conformational data and not just static structural information. The latter is crucial as much key protein functions are related to conformational changes. VibeGen is a newer AI model that uses specified motion patterns (e.g., flexing, vibrating, shifting shapes) to generate new protein structures that exhibit them.15 A random sequence of amino acids is refined using a seance generator and a motion predictor, with iterations between the two taking place until the algorithm cornages on a seance that exhibits the desired motion.
Improving Protein Therapeutics
One of the biggest challenges in the engineering of therapeutic proteins is achieving the right combination of performance properties, as often improving one feature can negative impact others. For instance, improving stability can have detrimental impact on activity, particularly for complex proteins such as bispecific antibodies and antibody–drug congratulates (ADCs).3
One approach combines hybrid engineering strategies (e.g., directed evolution with computational modeling), expression system optimization, new formulation strategies, and improved purification processes to minimize protein folding and aggregation, degradation, impurity formation, and other issues. Incorporation of nonnatural amino acids must, meanwhile, been done quite thoughtfully to avoid potential undesired immune responses. Often deimmunization and humanization are achieved using computational tools designed to identify and alter potential problematic regions.3
Today, both large and small biopharma companies and many academic groups are leveraging both traditional and advanced computational methods for protein engineering to accelerate the design and development of not just improved but novel protein therapeutics.
In one case, researchers The University of Texas at Dallas developed ProSSpeC (protease substrate specificity calculator), an ML-based model for predicting the behavior of specific protease enzymes and designing new synthetic protease that could potentially serve as mire effective targeted therapies.15 Researchers at Queensland University of Technology are design synthetic “smart” proteins that switch between active and inactive states via small structural changes under very specific conditions within cells, making them attractive biosensoirs.16
Archon Biosciences is designing “antibody cages” (AbCs), a new class of antibodies that are bound to AI-designed protein sequences that cause the antibodies to fold in a highly specific manner and exhibit improved activity and specifciity.7 ESMFold2 is being used to design new antibodies that bind strongly to other proteins known to be involved in cancer and immunological disease pathways.10 Boehringer Ingelheim is using OpenProtein’s platform to engineer proteins for the treatment of cancer and autoimmune diseases.12 Cyrus Biotechnology is focused on improving the performance of existing proteins using computational modeling and deep mutagenesis.7
Many Opportunities for Further Advances
The progress made in protein engineering is clearly stunning. Even so, there are still limitations to current technology and many opportunities for achieving further advances.3,13 Protein engineering is highly complex and involves large numbers of variables. Even the most advanced AI/ML algorithms in use today generate novel protein that suffer from issues ranging from stability to manufacturability to immunogenicity.
More knowledge and understanding of human biology and disease pathways are needed to improve current predictive models and overcome many of these challenges. Limitations of current models with subunit assembly and interactions for complex proteins must be overcome as well. More advanced computing technologies, such as quantum computing, have potential to design proteins that are simultaneously optimized across more variable. Such approaches could better simulate protein folding and interactions, thereby including essential dynamic, rather than just static, information. Emerging ensemble predictors such as ensemble/clustered AlphaFold-style predictors, flow- or flow-matching diffusion models, MD-guided AF2 variants, and multi-state design frameworks are showing promise as well.
Advances in more traditional computational modeling and directed evolution techniques will also lead to improved protein design, particularly when combined with synthetic biology and increasingly advanced high-throughput screening capabilities and autonomous labor tarries. Application of state-of-the-art protein engineering to next-generation, complex modalities including multipiece, ADCs, and protein-based nanoparticle therapeutics will also accelerate the development of these important new drugs.
Development of more high-quality data sets will also be key to improving the performance of protein design software. Finally, avoiding selective reporting of only successful results, use of current and appropriate benchmarks, standardization, and sharing of the rationale behind the operation of advanced modeling platforms to enable greater interpretability and validation of results will help build trust and acceptance.13
References
1. Shi, Jinghao, et al. “Recent Advances on Protein Engineering for Improved Stability.” BioDesign Research. 7: 100005 (2025).
2. Bose, Priyom. “Insights Into Protein Engineering: Methods and Applications.” The Scientist. 29 Oct. 2024.
3. Naderiyan, Zahra, and Alireza Shoari. “Protein Engineering Paving the Way for Next-Generation Therapies in Cancer.” International Journal of Translational Medicine. 5: 28 (2025).
4. Weigmann, Konstantin FG, Uwe T Bornscheuer, and Mark Doerr. “Advances and Critical Evaluation of Autonomous Protein Engineering: Towards Transparent, Accessible, and Reproducible Platforms.” Current Opinion in Biotechnology. 97: 103395 (2026).
5. Nithun, Raj V, et al. “Advancing Protein Engineering via Organic Chemistry.” Communications Chemistry. 9: 161 (2026).
6. Kuru, Erkin, et al. “AminoX: Making Better Protein Drugs, Quicker and Cheaper.” Wyss Institute at Harvard University.
7. Seydel, Caroline. “Designer Proteins Take Shape.” Genetic Engineering & Biotechnology News. 45: 10 (2025).
8. Yang, Yuhang, et al. “Recent Advances in Multimodal-Driven Protein AI.” Biophysics Reviews. 7: 011309 (2026).
9. Listgarten, Jennifer, and Hanlun Jiang. “How Artificial Intelligence Is Reengineering Protein Engineering.” Science. 392: 159–166 (2026).
10. Howes, Laura. “Generative AI Is Dreaming Up New Proteins.” Chemical & Engineering News. 101: 20–25 (2023).
11. Callaway, Ewen, and Miryam Naddaf. “Move Over, AlphaFold: Open-Source Model Predicts Shape of 1 Billion Proteins.” Nature. 27 May 2026.
12. Winn, Zach. “Bringing AI-Driven Protein-Design Tools to Biologists Everywhere.” MIT News. 17 Apr. 2026.
13. Jin, Shuming, et al. “Breaking Evolution’s Ceiling: AI-Powered Protein Engineering.” Catalysts. 15: 842 (2025).
14. Martinovich, Stephanie. “MIT Engineers Design Proteins by Their Motion, Not Just Their Shape.” MIT News. 26 Mar. 2026.
15. Horner, Kim. “A Protein Engineering Method May Lead to More Exact Cancer Treatments.” Phys.org. 20 Apr. 2026.
16. Dhar, Payal. “Synthetic Smart Proteins That Function as Biological Switches.” Chemical & Engineering News. April 22, 2026.












