Tempus AI has published results showing that PRISM2, a multimodal artificial intelligence foundation model for digital pathology, matched or exceeded the performance of specialised clinical-grade products across selected prostate, breast and lymph-node cancer-detection tasks. The study, published in Nature Medicine and highlighted by Tempus AI on August 4, 2026, evaluated whether one large model trained through pathology images and clinical dialogue could support diagnostic, biomarker and prognostic applications without requiring an entirely separate artificial intelligence system for every task.
PRISM2 was trained on approximately 2.3 million whole-slide images and 14 million diagnostic question-and-answer pairs generated from nearly 700,000 pathology reports. The model combines visual information from routine haematoxylin and eosin-stained tissue slides with language derived from real clinical diagnoses, allowing researchers to ask questions about a case and obtain probability-based responses.
The model demonstrated particularly strong results in cancer detection, achieving performance comparable with Paige Prostate and Paige Breast and exceeding Paige’s breast lymph-node product on the respective evaluation datasets. Researchers also reported that PRISM2 representations supported cancer subtyping, tumour grading, biomarker prediction and estimates of patient survival or recurrence risk.
Those findings are commercially relevant for Tempus AI because the company acquired digital pathology specialist Paige in August 2025 and is integrating pathology artificial intelligence into a broader precision-oncology platform spanning genomic testing, clinical data and drug-development services. Tempus AI generated $289.3 million in diagnostics revenue during the second quarter of 2026 and has also agreed to acquire molecular residual disease specialist Personalis, giving the company several complementary technologies across cancer diagnosis, treatment selection and recurrence monitoring.
The study does not mean that PRISM2 can immediately be used to diagnose patients. The publicly released model is intended for non-commercial and non-clinical research, and prospective studies, workflow validation and an appropriate regulatory authorisation would be required before any PRISM2-based diagnostic product could independently influence patient care in the United States.
What did PRISM2 achieve in the Nature Medicine digital pathology study?
PRISM2 was developed to overcome a major limitation of earlier pathology foundation models. Many existing systems analyse small image tiles extracted from enormous digital slides, but they require an additional task-specific model to aggregate those fragments and answer a clinically relevant question about the complete patient case.
The new model instead creates representations at the whole-slide level. One or more digital pathology slides are divided into tiles and processed using Virchow2, Paige’s underlying image foundation model. The resulting information is aggregated into a slide-level representation and connected to a four-billion-parameter language model trained to interpret pathology-oriented questions and diagnostic statements.
PRISM2 generates two principal types of representations. Its diagnostic embedding is designed for tasks such as determining whether cancer is present or identifying a tumour subtype. Its base embedding retains more general information that can be adapted for applications not explicitly represented in diagnostic reports, including biomarker and survival prediction.
In pan-cancer detection involving 16 tissue origins, the PRISM2 diagnostic embedding achieved an area under the receiver operating characteristic curve of 0.967. The original PRISM model achieved 0.947, while TITAN produced 0.931. Performance declined only modestly to 0.957 across a subset of seven rare cancer categories, suggesting that the model retained useful discrimination beyond the most common tumour types.
PRISM2 also performed strongly on external public datasets evaluating breast, lung and kidney cancer subtyping, lymph-node tumour detection and staging, prostate grading and colorectal cancer grading. The study authors said the evaluation data had not been used during PRISM2 or Virchow2 training, an important safeguard against directly testing a model on images it had already encountered.

How does clinical dialogue help PRISM2 understand an entire pathology slide?
The central innovation is the use of clinical dialogue as a training signal. Rather than asking the model only to match an image with a report, the researchers transformed pathology reports into several forms of structured conversation, including open-ended questions, yes-or-no questions, multiple-choice questions and report-generation tasks.
The training process used the diagnostic information already recorded by pathologists to create approximately 14 million question-and-answer pairs. Examples could ask whether invasive cancer was present, whether lymphovascular invasion could be identified or which tumour subtype best matched the tissue appearance.
This structure gives the model a more explicit connection between tissue morphology and the reasoning reflected in clinical reports. A conventional image model may learn that two groups of slides look different without understanding the diagnostic language associated with those differences. PRISM2 attempts to connect visual patterns with concepts pathologists use when classifying disease.
The model is not being proposed as an autonomous conversational chatbot for pathology. The study authors explained that dialogue was primarily used to produce clinically grounded slide representations and probability-based answers that could later support specialised applications.
Potential applications could include prioritising suspicious slides for specialist review, suggesting additional immunohistochemistry or sequencing tests, pre-populating structured pathology-report fields and providing decision support to general pathologists handling unusual cases. These remain proposed uses rather than authorised clinical functions of PRISM2 itself.
Why did comparisons with Paige’s clinical-grade cancer products matter?
The researchers tested PRISM2 against Paige Prostate, Paige Breast and Paige BLN, which are specialised products trained for defined cancer-detection tasks. Paige Prostate received United States Food and Drug Administration De Novo authorisation as a class II device intended to assist pathologists in detecting prostate cancer on digitised tissue slides.
Using prompt-based yes-or-no questions without task-specific retraining, PRISM2 matched the performance of Paige Prostate and Paige Breast and exceeded Paige BLN on the corresponding testing datasets. Earlier slide-level foundation models, including PRISM and TITAN, did not consistently reach the performance of the clinical-grade products under the same prompt-based evaluation.
This is important because specialised diagnostic artificial intelligence traditionally requires a curated dataset, model-development programme and regulatory strategy for each tissue type and clinical question. A foundation model capable of supporting several tasks could reduce the time and data required to develop subsequent applications.
The result should nevertheless be interpreted carefully. PRISM2 was evaluated retrospectively on established datasets, not through a prospective study in which pathologists used the model during routine diagnosis. Matching the performance of an authorised product in a research experiment does not give the foundation model the same regulatory status or prove that it will produce equivalent results across hospitals, scanners and patient populations.
The comparison also used the product-testing datasets associated with the specialised Paige applications. Independent studies conducted by health systems without developer involvement would provide stronger evidence of generalisability and resistance to institution-specific bias.
Can PRISM2 predict cancer biomarkers without replacing genomic testing?
The study evaluated whether information within routine haematoxylin and eosin slides could predict molecular biomarkers that frequently require genomic or immunohistochemical testing. PRISM2’s base embeddings performed at least as well as competing foundation-model representations across the selected biomarker tasks.
On the Memorial Sloan Kettering Cancer Center biomarker datasets, PRISM2 recorded an average area under the curve of approximately 0.854, compared with 0.846 for the closest competing model. On overlapping tasks drawn from The Cancer Genome Atlas, PRISM2 produced an average of 0.784, narrowly ahead of TITAN at 0.781.
These results suggest that tissue morphology contains signals associated with genetic or molecular alterations. A future product could potentially use artificial intelligence to identify patients most likely to carry a biomarker, accelerate triage or determine which specimens should receive confirmatory testing first.
The performance does not support replacing validated molecular diagnostics. Even a strong area-under-the-curve result can leave clinically significant false-positive and false-negative cases, while treatment eligibility may depend on precise confirmation of a mutation, amplification or protein-expression threshold.
The more realistic near-term use is an artificial intelligence screening or prioritisation layer. A model might rapidly analyse every available slide, flag likely biomarker-positive cases and reduce the chance that an appropriate genomic test is overlooked. Confirmatory testing would remain necessary when a biomarker determines access to a targeted therapy.
What did PRISM2 reveal about cancer survival and recurrence prediction?
Researchers created an additional survival-focused version of PRISM2 using data covering more than 225,000 cases and nearly 100,000 patients. They evaluated whether the resulting embeddings could predict colorectal cancer recurrence-free survival and disease-specific survival across multiple cancers.
For colorectal cancer recurrence-free survival, the fine-tuned PRISM2 model produced a concordance index of 0.809, compared with 0.773 for a specialist survival model trained from the beginning on the same dataset. The result suggests that broad pathology pretraining can provide useful information before a model is adapted to a narrower prognostic task.
This could be valuable for developing tests that estimate recurrence risk or identify patients who may require more intensive treatment and surveillance. Pathology slides are routinely generated during cancer diagnosis, meaning an artificial intelligence prognostic tool could potentially extract additional information without requiring another biopsy.
However, prognostic models require a particularly high evidentiary standard. Performance must be validated across treatment eras, demographic groups, disease stages and institutions because outcomes are influenced by therapy selection, surgical quality, follow-up practices and access to care, not tissue morphology alone.
A model predicting poorer survival could also create unintended consequences if clinicians interpret the output as a fixed biological destiny. Any future prognostic application would need clear labelling, transparent calibration and evidence that using the prediction improves clinical decisions rather than merely classifying risk.
Could PRISM2 automatically complete pathology reports during routine diagnosis?
The researchers explored whether PRISM2 could populate fields in a College of American Pathologists-style report for invasive breast carcinoma biopsies. The model answered questions concerning histological type, ductal carcinoma in situ, lymphovascular invasion, necrosis, nuclear grade and other reporting elements.
Performance varied substantially by field. For histological type across the represented categories, the model achieved an adjusted mean recall of 0.519, rising to 0.731 after probability calibration. When the question was narrowed to distinguishing invasive ductal carcinoma from invasive lobular carcinoma, performance was considerably stronger at 0.933.
The variation illustrates both the promise and the weakness of general-purpose pathology models. PRISM2 can recognise several clinically meaningful concepts, but it may struggle when rare subtypes, ambiguous morphology or poorly calibrated probability estimates are involved.
A report-completion tool could reduce administrative burden by proposing preliminary entries for pathologist review. It should not silently generate a final diagnosis or transfer uncertain predictions into the medical record without human verification.
The most useful deployment may resemble an intelligent drafting assistant. The model could prepare a structured report, highlight uncertain fields and direct the pathologist to the slide regions that most influenced its conclusion. The pathologist would remain responsible for reviewing the tissue and signing the diagnosis.
Why is PRISM2 not yet an FDA-authorised diagnostic medical device?
PRISM2 is a research model rather than a marketed diagnostic product. Tempus AI has made the model weights available through Hugging Face for non-commercial and non-clinical research, but the publication and open release do not constitute Food and Drug Administration clearance or approval.
A clinical product based on PRISM2 would require a defined intended use. The developer would need to specify the type of cancer, specimen, scanner, user, clinical workflow and decision the software is intended to support.
Regulators would then expect evidence covering analytical performance, clinical validation, cybersecurity, human factors, failure modes and performance across relevant demographic and technical subgroups. The Food and Drug Administration has said that artificial intelligence systems combining pathology, patient information and other data create new evaluation questions involving data harmonisation, missing information and generalisability.
Foundation models create an additional challenge because they can support many downstream applications. Authorising a prostate cancer detector does not automatically validate the same underlying model for breast cancer, biomarker prediction or survival forecasting.
The developer may therefore need to commercialise PRISM2 through several regulated products with narrower claims. Each application could use the foundation model as a common technological base while undergoing validation for its specific clinical purpose.
How could Tempus AI commercialise PRISM2 across oncology and drug development?
Tempus AI is positioned to use PRISM2 through both diagnostic products and services for pharmaceutical companies. The company already owns Paige’s digital pathology technology and has launched Paige Predict, a biomarker-prediction offering intended to help researchers identify molecular signals from pathology images.
PRISM2 could accelerate the development of tissue-based companion diagnostics, patient-selection models and retrospective analyses of clinical-trial specimens. Pharmaceutical companies frequently hold large archives of pathology slides but may not have corresponding genomic data for every patient.
A flexible foundation model could help researchers search those archives for morphological patterns associated with drug response, resistance or adverse outcomes. The strongest findings could then be tested through sequencing, prospective studies or clinical-trial enrolment strategies.
Tempus AI has already reported delivering an oncology foundation model to AstraZeneca and signed approximately $200 million in new data and applications licences during the second quarter of 2026. That commercial activity indicates that foundation models are becoming part of Tempus AI’s services for biopharmaceutical customers rather than remaining an academic exercise.
The proposed Personalis acquisition would extend the platform into molecular residual disease monitoring. In principle, Tempus AI could connect pathology-based diagnosis and biomarker prediction with genomic profiling, treatment data and blood-based recurrence monitoring across a patient’s cancer journey.
What data, bias and validation risks could prevent PRISM2 from reaching clinics?
Scale does not eliminate bias. PRISM2’s training dataset is extremely large, but performance can still be influenced by the institutions, scanners, tissue-processing techniques, reporting styles and patient populations represented in the source material.
Pathology foundation models can learn non-biological signals such as staining patterns, slide preparation artefacts or institutional conventions. A model may appear accurate during internal testing while performing less consistently when moved to a laboratory using different equipment or serving a different population.
The dialogue-training process also depends on historical pathology reports. Those reports contain valuable expert knowledge, but they may include inconsistent terminology, incomplete descriptions and errors. Automatically converting reports into millions of question-and-answer pairs could reproduce those weaknesses at scale.
Interpretability remains another concern. Attention maps can show which tissue regions influenced a model, but they do not always explain why a particular biological conclusion was reached. Clinicians need to recognise when an artificial intelligence system is uncertain or operating outside the population on which it was validated.
Independent prospective validation will therefore matter more than another retrospective benchmark. The decisive evidence would come from studies showing that pathologists using a PRISM2-based product diagnose cases more accurately or efficiently without increasing clinically significant errors.
Expert view: PRISM2 narrows the gap between foundation models and usable pathology tools
PRISM2 represents a meaningful advance because it performs several clinically relevant tasks through one slide-level architecture. Its ability to match specialised cancer-detection products without task-specific training suggests that general-purpose pathology models are becoming more than reusable image encoders.
The study’s strongest feature is the combination of scale and clinical grounding. Training on millions of whole-slide images gives the model visual breadth, while pathology-report dialogue connects those images to the language and concepts used during diagnosis.
Its limitations are equally important. PRISM2 has not been tested as an autonomous diagnostic system, its results remain retrospective and some report-completion tasks required substantial calibration. Biomarker and survival predictions are promising research signals rather than replacements for established tests.
For Tempus AI, the study provides technological validation for its Paige acquisition and supports a wider strategy linking digital pathology with genomic diagnostics and oncology data. The commercial opportunity could extend beyond selling individual algorithms to providing a foundation on which hospitals and drug developers build multiple cancer applications.
The next stage will determine whether PRISM2 can move from an impressive benchmark to a dependable medical product. That transition will require prospective evidence, independent validation, carefully limited clinical claims and regulatory scrutiny.
PRISM2 has shown that one model can extract diagnostic, molecular and prognostic signals from routine pathology slides. It has not yet shown that one model should be trusted to make all those decisions in real patients. That distinction will define the next phase of digital pathology.
