Foundation models (FMs) are redefining biological design by moving beyond static prediction toward intelligent creation. The latest systems no longer just infer protein structures—they integrate text, sequence, and experimental context to reason across biological scales. By embedding design intent directly into generative logic, they merge understanding and invention into a single computational framework. In part 1, we explore how multimodal and controllable architectures are enabling AI to design therapeutics tuned not only for function but for manufacturability and experimental realism.
Introduction: From Prediction to Creation
The breakthroughs of AlphaFold2 and RoseTTAFold in 2021–2022 represented a turning point in computational biology. For the first time, machine learning systems achieved near-atomic precision in predicting how proteins fold from their amino acid sequences, an achievement that solved one of biology’s most intractable challenges. However, even as these models transformed structure prediction, they left a critical gap unaddressed: they were diagnostic, not creative. They could describe how a protein should fold under ideal conditions, but they could not design new sequences that would reliably fold, express, and remain stable under the complex realities of bioprocessing. In practice, the leap from knowing a structure to producing a viable therapeutic remained immense.
A new generation of artificial intelligence is now emerging to close that gap. These systems, known as foundation models (FMs), move beyond single-task networks toward architectures capable of learning general biological principles from massive, diverse data sets. Initially adapted from natural language processing (NLP), early biological FMs, such as ESM and ProtT5, treated amino acid sequences as a kind of biological text, learning the “grammar” of evolution through self-supervised pretraining on millions of sequences. From there, the field rapidly expanded toward multimodal models that integrate multiple streams of information, including sequence, structure, and experimental metadata.1,2 By training across modalities, FMs begin to capture not only how molecules look, but also how they behave, linking three-dimensional conformation to thermodynamic stability, binding specificity, and even phenotypic outcomes.
The conceptual advance mirrors developments in general-purpose artificial intelligence (AI). Just as large language models learn to generate coherent text by understanding statistical relationships among words, biological FMs learn to generalize across the language of biomolecules. They translate among modalities — sequence, structure, chemical environment, and cellular phenotype — without explicit supervision, developing representations that can be fine-tuned for nearly any downstream task.3 This capacity for cross-domain reasoning makes them uniquely suited to the challenges of therapeutic design, where data are fragmented across molecular, biophysical, and process contexts.
What distinguishes the current wave of FMs from their predecessors is not merely scale but intent. Instead of predicting structure as an end in itself, these models are increasingly designed to generate novel, functional molecules — proteins, enzymes, RNA constructs, and multimolecular assemblies — while embedding constraints relevant to manufacturability and clinical translation. The objective is no longer to infer the most likely fold for a sequence but to design a sequence that folds correctly and performs reliably under industrial conditions. Emerging work in generative design models, such as those that integrate mechanical property prediction and process-compatible folding dynamics, explicitly incorporate variables like solubility, aggregation propensity, and thermal stability into their optimization loops.4,5 In doing so, they begin to bridge the long-standing divide between computational discovery and bioprocess engineering.
This evolution marks the beginning of a broader convergence. Biological FMs operate at the intersection of three formerly separate domains: molecular design, experimental validation, and manufacturing scalability. Their generative capabilities enable rapid exploration of sequence space, while their multimodal understanding ties in process-relevant parameters, such as expression yields, formulation stability, and purification compatibility, that define whether a theoretical design can become a practical therapeutic. In this sense, foundation models are not just tools for discovery but engines of integration across the continuum from molecular invention to production readiness.6
The implications FMs present for the biopharmaceutical industry are profound. By embedding process-aware constraints into the earliest stages of design, companies could dramatically reduce the time and cost required to bring biologics from concept to clinic. Instead of discovering first and optimizing later, the design process itself becomes optimized for manufacturability from the start. The next generation of FMs, trained on both biological and process data, may thus give rise to AI systems that not only invent biomolecules but also anticipate how they will perform during expression, purification, and formulation. What began as a revolution in structure prediction is rapidly transforming into a revolution in creation, where design, function, and production are learned as part of the same biological language.
The Rise of Biological Foundation Models
The conceptual journey from early protein language models to today’s multimodal FMs unfolded rapidly but followed a clear technological logic. Each stage in this evolution expanded the scope of what AI could represent about biology — from sequences, to structures, to multimodal contexts — and in doing so, redefined the relationship between computation and experiment.
Early Sequence-Based Language Models (2019–2022)
The earliest biological language models borrowed directly from the architecture of natural language processing. Models such as ESM, ProtBERT, TAPE, and ProtT5 treated amino acid sequences as if they were sentences in a biological language, applying masked-token prediction and other transformer-based training objectives to vast corpora of natural proteins. These models learned to reconstruct missing residues based on contextual information from surrounding amino acids, analogous to how the NLP model BERT learns to fill in missing words. Though trained only on raw sequences, they discovered implicit rules governing secondary structure formation, domain boundaries, and conserved motifs.1,7
These models became powerful, general-purpose representations of protein function and evolution. Without explicit supervision, they captured the statistical signatures of secondary structure, hydrophobic patterning, and coevolutionary dependencies that shape folding stability. These sequence embeddings could then be used for “zero-shot” tasks, such as predicting thermostability, variant effects, or enzymatic function, without additional training. This capability hinted that deep learning could internalize aspects of protein biophysics directly from sequence distributions, effectively learning the grammar of evolution itself.
Structure-Aware and Inverse Folding Models
The next phase extended these models beyond language-like understanding to geometric reasoning. Structure-aware systems, such as ESM-IF, ProteinMPNN, and RoseTTAFold Diffusion, trained jointly on paired sequence–structure data, allowing them to connect amino acid patterns with three-dimensional folding outcomes. This shift was transformative: rather than modeling proteins as linear strings of tokens, the networks learned to encode physical relationships, including bond lengths, torsion angles, and spatial proximities, that govern how a sequence folds into a stable conformation.7
Equally important was the inversion of the prediction task. Traditional models forecast structure from sequence; inverse folding models, by contrast, generate sequences predicted to adopt a given backbone. This reorientation reframed structure prediction as a design problem. By learning which residues best stabilize a target geometry, these inverse models began to function as generative engines for creating new proteins. AlphaDesign further formalized this idea, using AlphaFold itself as an oracle in a closed feedback loop to iteratively generate and evaluate candidate sequences.8 Together, these advances marked a turning point: AI could now not only describe nature’s proteins but also propose new ones that had never existed before.
The Foundation Model Leap
As data volumes and computational resources grew, researchers began training models that generalized across multiple biological modalities. The defining principle of this new class was scale — both in data and architecture. Performance followed empirical scaling laws, improving predictably with the size of datasets, model parameters, and computational throughput.2 With sufficient capacity, a single pretrained model could be adapted to many downstream tasks, each requiring only minimal fine-tuning.
These foundation models unified multiple learning paradigms — masked modeling, contrastive learning, and structure-conditioned decoding — within a single framework. Instead of training bespoke models for folding, annotation, or design, one model could perform all three. This multitask capability emerged naturally from self-supervised pretraining across diverse inputs, enabling the model to encode a shared latent representation of biological space.6 In practice, the same pretrained model could be fine-tuned for predicting the impact of mutations, annotating enzyme function, optimizing binding sites, or generating entirely new sequences. This universality positioned FMs as the biological equivalent of large language models: flexible backbones that learn the rules of the system itself rather than a single task within it.
Challenges Emerging with Scale
Scale, however, introduced new challenges. The same data sets that empowered foundation models also introduced bias: well-characterized protein families and organisms dominate public sequence repositories, leading to overrepresentation of familiar motifs and underrepresentation of disordered or rare structures.1 Models trained on such data can generate sequences that appear novel but actually reflect these statistical biases, producing “hallucinated” functions that do not correspond to biophysically plausible behavior.
Interpretability also remains a pressing concern. While FMs can predict or generate sequences with high statistical confidence, the biological rationale behind their outputs often remains opaque. Mechanistic transparency — the ability to link learned features to specific physicochemical phenomena — lags behind predictive performance. This lack of interpretability complicates their use in regulated environments, where explainability is increasingly viewed as a prerequisite for trust.
The field’s rapid growth has also created benchmarking fatigue. As dozens of new architectures emerge, comparing performance across data sets and tasks has become inconsistent. Initiatives such as ProteinBench aim to address this fragmentation by providing standardized evaluation tasks, metrics, and datasets for protein folding, variant effect prediction, and generative design.9 These frameworks help ensure that model improvements reflect genuine advances rather than overfitting to specific leaderboards or data splits.
Taken together, the rise of biological FM represents more than just a scaling milestone; they represent a conceptual shift toward a unified framework for representing biological information. From sequence-only learning to geometry-aware reasoning and multimodal generalization, FMs now encapsulate much of what defines molecular biology itself: the coupling of structure, function, and context. The challenge ahead is to channel this capability toward real-world outcomes by designing molecules that not only exist in silico but also perform predictably in manufacturing and therapy.
Expanding Beyond Proteins: RNA, DNA, and Cross-Modal Biology
The early breakthroughs of foundation models centered almost entirely on proteins, where abundant sequence data and established benchmarks accelerated progress. Yet proteins are only one part of the molecular continuum that defines biological systems. In cells, nucleic acids serve not just as blueprints for proteins but as dynamic, structured molecules with their own regulatory and catalytic functions. Recognizing this, researchers are now extending the foundation model paradigm beyond proteins to encompass RNA, DNA, and even cross-modal architectures that learn shared representations across all molecular classes. The goal is nothing less than a universal biological model capable of reasoning across the entire central dogma, from transcription to translation and beyond.
RNA as the Next Frontier
The surge of interest in RNA therapeutics following the success of mRNA vaccines transformed RNA biology from a niche field into a central pillar of drug development. Messenger RNA, once viewed simply as a transient intermediary, is now understood as a tunable therapeutic modality whose efficacy is governed by fine details of sequence structure, folding kinetics, and translational efficiency. The challenge lies in designing mRNA molecules that are simultaneously stable, efficiently translated, and manufacturable at scale.
Traditional computational tools could predict secondary structure or codon usage, but they lacked the ability to capture the full interplay between sequence, structure, and expression. RNA-focused foundation models are beginning to close that gap. By training on large corpora of RNA sequences, paired with experimental data on folding, degradation, and translation, these models learn the implicit grammar of RNA biology: the coupling between sequence features, three-dimensional conformations, and regulatory performance. They are used to optimize untranslated regions (UTRs), reduce immunogenicity, and design sequences with enhanced in vivo stability and expression yield.10,11
Key Examples
Among the most promising developments is RiboDiffusion, a diffusion-based generative framework for RNA inverse folding.10 Rather than predicting structure from sequence, it begins with a desired three-dimensional RNA backbone and generates compatible sequences that fold to match it. This allows the model to design RNAs with defined tertiary architectures while preserving sequence diversity, a critical factor for balancing structural fidelity and translational efficiency. By learning to operate in both sequence and structure spaces, RiboDiffusion achieves superior recovery rates and design diversity compared with conventional inverse folding algorithms.
Helix-mRNA reflects a complementary advance focused specifically on therapeutic mRNA design.11 Unlike models that handle short or cropped sequences, Helix-mRNA employs a hybrid attention and state-space architecture capable of processing full-length transcripts, including both coding and untranslated regions. It predicts translation rates, secondary structure stability, and likely degradation pathways in a single integrated framework. This allows simultaneous optimization for manufacturability attributes, such as sequence integrity during in vitro transcription, and for biological performance in vivo. Together, these models move RNA design beyond piecemeal optimization toward holistic sequence engineering informed by the full complexity of RNA biology.
Unified Biological Modeling
The next step in this progression is integration: building models that understand the relationships among DNA, RNA, and protein as parts of a shared biological language. One approach introduces a unified tokenization scheme that allows a single foundation model to represent nucleic acid and protein sequences within a common embedding space.12 By aligning the vocabularies of different molecular types, the model can perform cross-modal transfer learning: for example, predicting protein-level effects of DNA mutations or inferring RNA folding changes from genomic variants. This integration hints at a biological analogue of multimodal AI systems in vision and language, where shared representations enable seamless translation between distinct data types.
A related conceptual framework envisions foundation models that encode “biological meaning” across molecular hierarchies.13 Instead of treating DNA, RNA, and proteins as separate inputs, the model learns to represent the entire information flow — from gene sequence to expressed structure to phenotypic effect — as a continuous computational process. Such architectures could, in principle, capture emergent dependencies that span multiple layers of regulation, from transcriptional control to post-translational modification.
Implications
The expansion of FMs into nucleic acid space carries far-reaching implications for both research and therapeutics. For the first time, it is becoming possible to model the full transcription–translation pipeline within a single representational framework, predicting how sequence changes at the DNA or mRNA level propagate through protein folding and cellular expression. This unification supports a new class of design workflows in which nucleotide sequences can be optimized not only for their own properties but also for the behavior of the proteins they encode.
In practical terms, this convergence enables rational co-design of RNA–protein systems, from ribonucleoprotein complexes and RNA-guided enzymes to hybrid gene therapies that combine coding and regulatory elements. By capturing shared structure–function principles across modalities, cross-domain FMs could predict how altering an mRNA’s UTR might change the folding dynamics or stability of the encoded protein, or how a synonymous DNA variant might impact both transcriptional efficiency and downstream manufacturability.
Ultimately, the move toward universal, cross-modal biological architectures marks a decisive step toward treating molecular biology as a single interconnected design space. As RNA foundation models mature and multimodal training becomes routine, researchers will gain the ability to navigate that space seamlessly, designing DNA, RNA, and proteins not as isolated entities but as coordinated components of therapeutic systems. The implications for drug discovery, synthetic biology, and biomanufacturing are profound: foundation models are evolving from specialized predictors into unified engines for understanding and engineering life’s molecular code.10–13
Part 2 examines how process-aware models are embedding manufacturability, stability, and scalability into molecular design and how emerging benchmarks are redefining what it means to validate AI in biomanufacturing.
References
1. Bjerregaard, Andreas, et al. “Foundation models of protein sequences: A brief overview.” Current Opinion in Structural Biology. 91: 103004 (2025).
2. Guo, Fei, et al. “Foundation models in bioinformatics.” National Science Review. 12: nwaf028 (2025).
3. Li, Michelle M and Marrinka Zitnik. “Context Matters for Foundation Models in Biology.” Kempner Institute. 16 Aug. 2024.
4. Ni, Bo, David L Kaplan, and Markus J Buehler. “ForceGen: End-to-end de novo protein generation based on nonlinear mechanical unfolding responses using a language diffusion model.” Science Advances. 7 Feb. 2024.
5. Rettie, Stephen A, et al. “Accurate de novo design of high-affinity protein-binding macrocycles using deep learning.” Nature Chemical Biology. 20 Jun. 2025.
6. Si, Yunda, et al. “Foundation models in molecular biology.” Biophys. Rep. 10: 135–151 (2024).
7. Yang, Soojung, et al. “Probing the Embedding Space of Protein Foundation Models through Intrinsic Dimension Analysis.” 38th Conference on Neural Information Processing Systems. 2024.
8. Jendrusch, Micahel A, et al. “AlphaDesign: a de novo protein design framework based on AlphaFold.” Mol. Syst. Biol. 21: 1166–1189 (2025).
9. Ye, Fei, et al. “ProteinBench: A Holistic Evaluation of Protein Foundation Models.” Bytedance Research. Accessed 16 Oct. 2025.
10. Huang, Han, et al. “RiboDiffusion: Tertiary Structure-based RNA Inverse Folding with Generative Diffusion Models.” Bioinformatics. 2024.
11. Wood, Matthew. “Helix-mRNA: A Hybrid Foundation Model For Full Sequence mRNA Therapeutics.” Arxiv. 11 Mar. 2025.
12. He, Yong, et al. “Generalized biological foundation model with unified nucleic acid and protein language.” Nature Machine Intelligence. 7: 942–953 (2025).
13. Kalfon, Jeremie, Laura Cantini, and Gabriel Peyre. “Towards foundation models that learn across biological scales.” bioRxiv. 18 May 2025.












