Key Takeaways
Microphysiological systems (MPS) and organ-on-chip technologies are moving from experimental demonstrations toward validated tools designed to support specific drug-development and regulatory decisions.
The FDA’s emerging framework for new approach methodologies emphasizes context of use, human biological relevance, technical characterization, and fit-for-purpose performance rather than universal validation of a technology.
Cross-platform benchmarking, reproducibility, automation, standardized metadata, and controlled cellular inputs will be essential for translating MPS from specialized academic systems into scalable drug-development platforms.
Organ-on-chip systems are likely to supplement conventional evidence before replacing animal studies, with adoption progressing from additional evidence to greater decision-making influence and, eventually, substitution in well-supported applications.
Animal study replacement is expected to proceed application by application as individual MPS demonstrate that they can provide reliable human-relevant information for clearly defined development questions.
The Question Is Shifting
Microphysiological systems (MPS) are moving into a different phase of development. These systems combine human cells, often in three-dimensional configurations, with engineered environments that can incorporate fluid flow, mechanical forces, tissue interfaces, and interactions among multiple cell types. They have been explored for disease modeling, drug development, and safety assessment, with the broader goal of reproducing selected aspects of human physiology in experimentally accessible systems.1
Animal models can provide information that cannot yet be reproduced comprehensively in vitro, but species differences can also limit their ability to predict human responses. Human-cell-based systems may offer particular advantages when the biological process being studied differs meaningfully between humans and the experimental species.2
The regulatory environment has also changed substantially. New approach methodologies (NAMs), a broader category that includes advanced in vitro systems such as MPS along with computational and other nonanimal approaches, are now the subject of an explicit validation framework from the U.S. Food and Drug Administration (FDA). The agency encourages the submission of NAM data when these methods can improve the predictivity of nonclinical studies and recognizes that a fit-for-purpose in vitro NAM may, in appropriate circumstances, replace an animal study.2
What makes the current moment different is that scientific maturation is increasingly being matched by regulatory frameworks, cross-platform validation efforts, quality initiatives, and formal standards development. At the same time, routine use still lags behind that infrastructure. Organ-on-chip and MPS methods were rare among NAM studies submitted to the FDA during a 15-year period examined in a recent analysis.3 A 2025 technology assessment by the U.S. Government Accountability Office (GAO) likewise concluded that organ-on-chip technologies were not yet capable of broadly replacing animal testing and identified insufficient benchmarking and validation, inconsistent access to high-quality human cells, limited data sharing, and regulatory uncertainty among the barriers to wider adoption.4
The practical challenge is determining what evidence transforms an MPS from an impressive experimental model into a drug-development tool that can support a development or regulatory conclusion with enough confidence to reduce, supplement, or replace information that would otherwise come from an animal study.
Replacement is unlikely to occur through a single determination that organ chips are validated and animal testing is no longer necessary. It is more likely to happen incrementally, as individual systems demonstrate that they can answer particular questions reliably enough for defined uses.1 That makes validation, reproducibility, benchmarking, transferability, and regulatory qualification the real story of MPS adoption.
Defining What Replacement Actually Means
The complexity of an MPS can be visually compelling. A system may incorporate multiple cell types, extracellular matrix, fluid flow, mechanical stimulation, sensors, and complex analytical endpoints. But increasing the number of biological features does not necessarily increase its value for drug development.
The more relevant question is whether those features improve its ability to answer the question for which it was designed. MPS can be developed around applications including toxicity, barrier function, inflammatory responses, drug penetration, and patient-specific therapeutic responses. The appropriate model is therefore not necessarily the one that most closely resembles an entire organ, but the one that reproduces the biology needed for its intended application with sufficient reliability.
This principle is central to the FDA’s concept of context of use (CoU), which defines the purpose for which a NAM is intended to be used in drug development. Validation is consequently tied to an application rather than to the technology in the abstract.2
A toxicity-focused MPS, for example, does not need to recreate every function of the organ it represents. It needs to capture the biological processes that matter for the toxicity question well enough to produce a dependable answer.
That narrower scope also places limits around what a successful result means. Evidence that a model performs well for one endpoint cannot automatically be extended to long-term systemic toxicity, multi-organ interactions, chronic effects, complex immune responses, or other whole-body phenomena that remain particularly difficult to reproduce with current MPS.1
The practical goal is not to reproduce an animal in miniature. It is to determine when a human-based experimental system can provide the answer that an animal study was being used to obtain.
Validation Begins With the Decision
Draft guidance released by the FDA in March 2026 provides a useful framework for evaluating whether a NAM is ready to support that kind of use. It identifies four central components of validation: context of use, human biological relevance, technical characterization, and fit-for-purpose performance.2
In practical terms, those components translate into four questions. What development or regulatory conclusion is the model intended to support? Does it contain the human biology necessary to address that question? Can the system generate reliable and reproducible results under defined conditions? And does its performance justify relying on those results for the intended purpose?
Human biological relevance involves more than simply incorporating human cells. The selected cell types, tissue architecture, culture conditions, device characteristics, exposure conditions, and measured endpoints all have to correspond meaningfully to the biological process being investigated.
Technical characterization addresses the engineered system surrounding that biology. Organ chips can be influenced by device materials, geometry, fluid flow, extracellular matrix, shear forces, media composition, and interactions between the test article and the device itself. These variables have to be understood well enough to distinguish a biological response from an artifact of the platform.
Fit-for-purpose performance then connects the assay to its intended use. When an established comparator exists, the FDA recommends demonstrating that the NAM characterizes the relevant risk at least as well as the established approach, using performance measures appropriate to the application. The agency also recommends identifying limitations that could affect reliability, reproducibility, or interpretation.2
This framework leaves room for several forms of adoption. A NAM may supplement existing evidence, provide information where conventional approaches are inadequate, or ultimately replace an established study. Complementary use may be an important stage in building confidence because it allows developers and regulators to compare MPS results with existing evidence before deciding whether a conventional study can be removed.
Validation is also distinct from formal qualification. The FDA can consider appropriately validated NAM data within an individual development program without requiring the method to have completed a broader qualification process. Qualification establishes that a drug-development tool can be relied upon for a specified interpretation and application beyond a single program. This allows useful MPS applications to begin generating regulatory experience while broader qualification efforts continue.
Prediction Has to Be Proven Against Something
A biologically convincing response inside an organ chip is not enough to establish predictive value. If an MPS is expected to replace or materially influence established evidence, its performance has to be tested against meaningful reference outcomes.
That requires suitable compounds, relevant human data, defined endpoints, and performance criteria that reveal not only when the model succeeds but when it fails. Benchmarking against trusted human reference data and compounds with known clinical effects is therefore central to qualification-oriented MPS development.
Drug-induced liver injury (DILI) provides one of the clearest examples of how that progression can work. In an evaluation published in Communications Medicine in 2022, a human Liver-Chip was tested with 27 drugs with known toxic or nontoxic outcomes across 870 chips. Under the study’s qualification framework, the system achieved 87% sensitivity and 100% specificity for the tested compounds.5
The significance of that study lies partly in the form of the evidence. The platform was not simply shown to reproduce liver-like behavior. Its predictions were compared with known outcomes and expressed through quantitative performance measures. That creates a more useful basis for determining how much confidence developers should place in the result.
A newer effort extends this logic across multiple platforms. Eight commercially available liver MPS are being evaluated through a harmonized program developed in collaboration with the FDA’s Center for Drug Evaluation and Research (CDER). The study uses blinded compounds, shared endpoints, and independent data analysis to assess performance across systems.6
Its CoU is deliberately narrow: retrospective assessment of elevated liver signals observed during early clinical trials. The participating systems are not being asked to prove that liver MPS can replace all animal toxicology. They are being tested against a defined clinical problem for which performance can be evaluated systematically.
That illustrates how qualification-oriented evidence may develop. A promising platform can first demonstrate predictive performance against known outcomes. Cross-platform evaluations can then ask whether comparable approaches produce sufficiently consistent and interpretable results under harmonized conditions. If the evidence supports the intended application, developers and regulators have a clearer basis for deciding when the method can influence or replace an existing source of information.
Cross-platform work also addresses a persistent problem in emerging technologies. When individual developers use different compounds, endpoints, protocols, and performance criteria, it becomes difficult to determine whether differences among results reflect differences among the technologies or simply differences in how they were tested. Harmonized and blinded evaluation moves the field closer to establishing expectations that can apply across a class of systems rather than only to one laboratory or commercial platform.6
The path from demonstration to routine use therefore depends on establishing meaningful reference outcomes, challenging the method with relevant cases, measuring predictive performance, and making its limitations visible.
Reproducibility Is Also an Engineering Problem
Strong predictive performance in one study is not enough if another laboratory cannot reproduce it.
MPS bring biological and technical variability into unusually close contact. Biological variability can be valuable, particularly when systems are intended to capture differences among donors or patients. Technical variability can produce superficially similar differences for reasons unrelated to human biology. Distinguishing one from the other is essential if MPS are expected to support conclusions about safety, efficacy, or patient heterogeneity.7
The challenge extends across the biological material, the engineered device, and the experimental workflow. Cell source, donor characteristics, cell lots, media, device fabrication, flow conditions, drug preparation, sampling, imaging, analytical processing, and operator technique can all influence results. Experimental metadata may also need to capture events such as unexpected flow changes, bubbles, or differences in organoid size that could help explain an anomalous response.
This is where the transition from academic prototype to industrial platform becomes particularly important. A specialized laboratory may successfully operate a complex system because investigators have accumulated tacit knowledge about how it behaves. Routine drug development requires the method to work beyond the laboratory that created it. Results eventually have to remain interpretable across operators, batches, experiments, and, where appropriate, different sites.
That does not mean eliminating variability altogether. If cells from two donors respond differently because of genuine biological differences, the variation may be precisely what the model is intended to reveal. The objective is to control and characterize technical variability well enough that meaningful biological variation can be recognized as biology rather than dismissed as experimental noise.
For MPS, reproducibility is therefore partly an engineering and process-control challenge. A useful model has to include not only the right biology but a sufficiently controlled method for producing, operating, measuring, and documenting that biology.
An Assay Has to Travel
Transferability provides a useful way to connect automation, standardization, data management, and quality systems.
An assay does not become a scalable drug-development platform simply because the physical device can be manufactured repeatedly. The biological inputs, operating procedure, analytical methods, controls, and data generated by the system all have to travel with it.
Manual operations create obvious vulnerabilities. Chip preparation, cell seeding, dosing, sampling, imaging, and analysis can introduce operator-dependent variation, particularly when small procedural differences affect cell state or system performance. Selective automation of fragile or influential steps can therefore improve more than throughput. It can reduce dependence on individual technique and make execution more consistent across users and laboratories.
Standardization has to extend to what is measured. A proposed framework for MPS evaluation distinguishes quantitative features controlled through system design from physiological characteristics that emerge from cellular function, providing a way to assess how measured features relate to human biology and predictive value in a given application.8
The data generated also need enough structure to allow experiments to be reconstructed and compared. International standards work now under development addresses requirements for MPS data and metadata, including experimental design, device- and assay-level information, data provenance, versioning, and connections with computational models.9 A separate committee draft addresses quality control, documentation, and critical quality attributes for cellular components used in MPS experiments.10
These remain draft standards rather than finalized requirements, but their scope reflects the practical demands of industrialization. A laboratory receiving an established MPS method needs to know not only what device to use but which cells and materials are acceptable, how the assay should be operated, which variables must be controlled, what metadata have to accompany the result, and how deviations should be interpreted.
The broader NAM ecosystem is beginning to build infrastructure around the same needs. The NIH Validation and Qualification Network is intended to support standardized reporting and common data elements, conformity assessments, quality-management approaches, and preparation of NAM data packages for regulatory validation and qualification.11
Meeting these requirements will involve more than MPS developers alone. Platform developers need manufacturable, reproducible devices. Cell and tissue-model suppliers influence the consistency of biological starting materials. Automation providers can reduce operator dependence. Analytical and computational technologies can help convert complex experimental outputs into interpretable evidence. Contract research organizations can make validated methods accessible to drug developers that do not plan to establish every capability internally.
The common requirement is transferability. If the method cannot leave the hands of its inventors without losing reliability, its role in routine development will remain limited.
The Regulatory Pathway Exists, but Experience Is Still Sparse
The regulatory issue surrounding organ-on-chip technologies is increasingly less about whether the FDA will accept MPS data in principle and more about how developers can generate evidence that supports a defined use.
Multiple pathways now exist. The Innovative Science and Technology Approaches for New Drugs program explicitly identifies tissue chips and MPS intended to answer safety or efficacy questions as examples of novel drug-development tools that may be considered for qualification.12 At the same time, the FDA’s current NAM framework allows validated methods to contribute to individual drug programs without requiring formal qualification first.2
The National Institutes of Health (NIH) and the FDA are also supporting a more deliberate qualification infrastructure. Four Translational Centers for Microphysiological Systems have been established to promote wider use of tissue-chip technologies in drug discovery and development, with the stated goal of advancing systems with defined CoUs toward qualification as drug-development tools by the FDA.13
What remains scarce is accumulated regulatory experience. Despite increasing attention to NAMs, organ-on-chip and MPS studies appeared only rarely in the CDER submissions evaluated over the 15-year period analyzed, while in silico approaches and more conventional in vitro methods were much more common.3
That gap matters. A formal regulatory route can establish how a technology should be evaluated, but confidence grows when reviewers and drug developers see validated methods used repeatedly across compounds and programs. Early applications may therefore supplement conventional evidence before they replace it. As performance becomes better characterized and the same method is encountered across more development programs, the basis for relying on it alone in a particular setting can become stronger.
The infrastructure is increasingly in place. The remaining task is to build a body of reproducible, well-benchmarked evidence.
From Validated Tool to Development Decision
Even a validated and regulatorily acceptable MPS will not become routine unless it improves drug-development decisions.
In many settings, adoption may begin by adding human-relevant evidence to a package that still includes animal studies. If that information repeatedly proves useful, it may begin to carry more weight in candidate selection, interpretation of safety signals, or other development choices. Only after sufficient confidence has accumulated may a conventional study become unnecessary. The progression may therefore move from supplementation, to influence, to substitution rather than directly from experimental demonstration to animal replacement.
Cost, accessibility, scalability, specialized expertise, and compatibility with existing workflows will still affect how readily that progression occurs. The value of an MPS ultimately depends on whether the information it generates justifies the resources required and changes what a development team can decide with confidence.
Narrowly defined applications may therefore advance faster than broad claims of model superiority. Where the relevant biology is sufficiently understood, strong human reference data are available, reproducibility can be demonstrated, and the consequences of an incorrect prediction can be assessed, developers have a clearer basis for determining whether an MPS is ready to carry greater evidentiary weight.
The practical threshold is reached when scientific, technical, operational, and regulatory confidence is sufficient for the information to alter the next step in a development program.
Replacement Will Happen Piece by Piece
Current MPS cannot reproduce every function for which whole animals are used. Long-duration physiology, systems-level interactions, chronic toxicity, complex immune responses, and coordinated effects across multiple organs remain particularly difficult to model comprehensively. Inconsistent access to high-quality human cells, incomplete validation, limited data sharing, and other adoption barriers also remain unresolved.
Those limitations argue against treating animal replacement as a single technological finish line. They do not require waiting until an MPS can reproduce an entire organism.
A more realistic transition will proceed through narrowly defined uses in which MPS first complement established evidence, then influence development decisions more directly, and eventually replace conventional studies where the accumulated evidence supports doing so.
The liver examples show how narrow those initial steps can be. One system has demonstrated predictive performance against known DILI outcomes, while a broader cross-platform effort is evaluating multiple liver MPS for the specific problem of interpreting elevated liver signals observed in early clinical trials. Neither result establishes that liver MPS can answer every toxicological question. Both help define where reliance on these systems may become justified.
The decisive transition will come when an organ chip is judged less by how convincingly it resembles an organ and more by whether developers and regulators can trust the answer it produces for the question in front of them. Animal studies will not disappear because organ chips become more elaborate. They will be displaced where specific human-based systems demonstrate, with evidence, that they can provide the information a development decision actually requires.
References
1. “5 perspectives on translating microphysiological systems from concept to application.” Nature Communications. 17: 9148 (2026).
2. General Considerations for the Use of New Approach Methodologies in Drug Development. Draft Guidance for Industry. U.S. Food and Drug Administration. 18 Mar. 2026.
3. Dao, Tyna, Nakissa Sadrieh. “A CDER perspective: Landscape of New Approach Methodologies (NAMs) submitted in drug development programs.” Regulatory Toxicology and Pharmacology. 165: 106007 (2026).
4. “Human Organ-On-A-Chip: Technologies Offer Benefits Over Animal Testing but Challenges Limit Wider Adoption.” U.S. Government Accountability Office. GAO-25-107335. 21 May 2025.
5. Ewart, Lorna, et al. “Performance assessment and economic analysis of a human Liver-Chip for predictive toxicology.” Communications Medicine. 2: 154 (2022).
6. LaFollette, Megan R, et al. “Introduction to the 3RsC's cross-platform microphysiological systems evaluation for drug-induced liver injury in collaboration with FDA-CDER.” Regulatory Toxicology and Pharmacology. 172: 106189 (2026).
7. Miedel, Mark T, et al. “Validation of microphysiological systems for interpreting patient heterogeneity requires robust reproducibility analytics and experimental metadata.” Cell Reports Methods. 5: 101028 (2025).
8. Nahon, Dennis M, et al. “Standardizing designed and emergent quantitative features in microphysiological systems.” Nature Biomedical Engineering. 8: 941–962 (2024).
9. “Microphysiological systems and Organ-on-Chip — Digital twins and computational modelling.” International Organization for Standardization. ISO/CD 25591, Committee Draft, 2026.
10. “Microphysiological systems and Organ-on-Chip — Quality control and documentation for cellular components.” International Organization for Standardization. ISO/CD 25530, Committee Draft, 2026.
11. National Institutes of Health Common Fund. “Complement-ARIE Validation and Qualification Network (VQN).” 12 Feb. 2026.
12. “Innovative Science and Technology Approaches for New Drugs (ISTAND) Program.” U.S. Food and Drug Administration. Accessed 28 Aug. 2026.
13. “Tissue Chip Projects & Initiatives.” National Center for Advancing Translational Sciences. 16 Jan. 2026.












