Research analysis · Drug discovery

Two hundred six patients in, tumour-immune organoids still cannot choose your immunotherapy

The first prospectively registered systematic review of patient-derived tumour-immune organoids as predictors of checkpoint-inhibitor response has arrived, and it is a model of the genre: 23 studies, 206 patients, a pooled sensitivity of 0.70, and a headline specificity of 0.97 that rests on exactly one false positive. Every included study carried high risk of bias. The authors' own conclusion is that these assays should not yet determine whether a patient gets immunotherapy, and the numbers say they are right.

Source: Patient-derived tumour-immune organoids as functional biomarkers of checkpoint-inhibitor response: a systematic review and exploratory meta-analysis, medRxiv, 2026. Primary source. Read: full text retrieved from medRxiv on the run date; all pooled figures, counts and eligibility rules checked against the retrieved text.

What the work claims

This is a synthesis, and a disciplined one, so it should be read as the state of an evidence base rather than as a new result. Tan and colleagues asked a precisely clinical question: can a three-dimensional, patient-derived tumour model that retains or rebuilds an immune compartment predict whether a specific patient will respond to immune checkpoint blockade? They searched five sources from January 2018 through 5 August 2026, preregistered the protocol on PROSPERO, and applied a prespecified rule that only peer-reviewed full reports with at least five paired patients and extractable two-by-two data would enter quantitative synthesis. Of 3,344 records, 23 independent studies survived, covering 206 deduplicated patients with paired ex vivo and clinical observations; the median study had 3 patients, with an interquartile range of 2 to 9.51.

The headline numbers, all from the protocol-concordant primary set of five reports and 102 patients: pooled sensitivity 0.70 (95% credible interval 0.48 to 0.89), model-implied specificity 0.97 (0.88 to 1.00), with crude proportions of 35/52 and 49/50 respectively. A broader post hoc pooling of all 20 studies with classifiable patients (n=154; 54 true positives, 1 false positive, 18 false negatives, 81 true negatives) gave sensitivity 0.83 and specificity 0.98, which the authors explicitly de-emphasize as evidence-boundary analysis. Their conclusion: biological and translational promise, feasibility and early clinical association, but not clinical validity, not clinical utility, and no basis for using the assays to add or withhold immunotherapy.

How it works

The review's real contribution is a taxonomy of three non-equivalent ways to build a tumour-immune organoid, each sacrificing different biology. Endogenous platforms (7 studies, 53 patients) keep the tumour's resident immune and stromal cells in short-window organotypic cultures, preserving tumour-experienced immune states at the cost of limited culture duration. Reconstructed platforms (11 studies, 143 patients) expand tumour epithelium and add the patient's own lymph-node cells, PBMCs, tumour-infiltrating lymphocytes or effusion lymphocytes, gaining control and serial sampling while remodelling the native immune state through expansion and activation. Engineered systems (4 studies, 9 patients) impose controlled geometry, matrix and perfusion on chips and assembloids, adding platform-specific variables with the control. A fourth study combined an immune organoid with a humanized xenograft.

Two implementation findings matter for anyone who builds or buys these assays. First, assay success is a selection event, and it is inconsistently reported: at least one organotypic method succeeded in 22 of 31 melanoma samples; 28 of 30 hepatocellular carcinoma cultures met a day-7 viability gate; 98 of 328 lung-cancer organoids were established in one cohort and 95 of 111 in another; 35 of 36 assembloids succeeded in a fourth; and only 5 of 13 specimens yielded interpretable immunotumoroids in a fifth. Every accuracy estimate downstream is conditional on surviving that funnel. Second, decision thresholds were heterogeneous and largely circular: residual ATP below 50%, 20% or 30% viability reductions, a growth-inhibition index above 0.245, a 31.93% ROC-derived cutoff and a cohort-median organoid-killing-index split, almost all derived and evaluated in the same small cohort, none externally calibrated. The bivariate random-effects model standard in diagnostic meta-analysis could not even be fit on five studies with one false positive, which is itself a statement about the evidence density.

Where a skeptic should push

The load-bearing number is the specificity, and it is nearly hollow. A point estimate of 0.97 with a credible floor of 0.88 sounds like near-perfect rule-out performance, but it is identified by a single false positive across the entire primary set. With one event, the posterior for the false-positive rate is whatever your prior shape says it is; the authors say so plainly and label the estimate model-implied and weakly identified. A sponsor quoting this specificity to a payor or regulator would be quoting a number the source paper itself declines to stand behind. Sensitivity of 0.70 is sturdier but carries its own clinical weight: in the intended use, deciding whether to add a checkpoint inhibitor to an otherwise reasonable regimen, a sensitivity of 0.70 means roughly three in ten responders would be read as negative. The review quantifies the stakes: 17 of the 52 patients with clinical benefit were ex vivo negative, so a negative assay result, acted on, would have withheld durable benefit from a third of benefitters, while a false-positive add decision exposes patients to immune toxicity, delay and cost.

Layer on the design problems the authors catalogue: retrospective or selected sampling, thresholds tuned in the evaluation cohort, no blinding, one-class evidence where responder-only cohorts inform only sensitivity, combination regimens that prevent attributing an ex vivo effect to checkpoint blockade itself, and reference outcomes that mix RECIST response, durable stable disease and postoperative recurrence. All 23 studies were judged high overall risk of bias and certainty was very low. One-sentence stress test of the whole enterprise: every pooled figure here is an association measured on the patients whose tumours grew well enough to test, interpreted with a cutoff chosen on the same patients. That is a development program, not a validation study.

The evidence bar for immuno-organoid tests

For organoid-based drug discovery this review does something the field needs badly: it converts a marketing question (do these models work?) into an engineering specification. The authors' list of what a decisive validation study requires is, in effect, a requirements document for the next generation of organoid biomarker programs: enroll consecutive patients before assay success is known; lock biopsy number, sampling sites, transport time, culture conditions, immune source, drug concentration, exposure and readout; blind laboratory and clinical assessors; test the exact planned regimen; keep every failure and indeterminate result in the denominator; and compare head-to-head against PD-L1, MSI/MMR, tumour mutational burden and clinician choice, with outcomes that include turnaround, decision change, toxicity, cost and equity. Any organoid company that cannot put a checkmark against most of that list is selling development-stage science, and the review gives payors and trial designers the vocabulary to say so.

The opportunity is that the specification is achievable, and the first groups to meet it will own the clinical category. A common minimum data model spanning specimen provenance, matrix lot, immune-cell source and ratio, activation conditions, exposure, raw readout, decision rule and every failed case, followed by ring trials across laboratories, is exactly the kind of unglamorous infrastructure that converts a cottage craft into an industry; the review's finding that establishment success varies from 5 of 13 to 35 of 36 across protocols shows how much room there is to compete on reliability. The threat is the mirror image: the base rate of organoid-positive anecdotes plus a pooled sensitivity of 0.70 is more than enough to sustain premature clinical claims, and each overclaimed case report raises the evidential bar for everyone else while inviting a regulatory crackdown on the whole functional-precision-medicine category. There is also a quieter strategic point. The intended use the authors defend is adjunctive, helping decide whether to add an ICI, not rule one out; a platform positioned as rule-out would need a false-negative rate this evidence base cannot currently support, and positioning it that way would be the fastest route from promise to harm.

The bottom line

Established by this review: tumour-immune organoids are biologically plausible, technically feasible across at least three architectures, and associated with clinical response at a pooled sensitivity around 0.70 in a small, biased evidence base. Not established: clinical validity, clinical utility, or incremental value over existing biomarkers and clinician judgment; no comparative prospective study showed that organoid-guided addition of an ICI improves any outcome. The review's own verdict, which the numbers fully support, is that these assays should remain investigational. What would confirm the field is a prospective, locked-threshold, failure-inclusive multicentre study with blinded assessment and biomarker comparators; what would break it is a well-run version of that study showing no decision-by-assay interaction, or showing that apparent accuracy evaporates once establishment failures and threshold arbitrariness are priced in. Until then, the honest caption for every tumour-immune organoid accuracy figure is the one this review supplies: feasibility and early association, priced accordingly.

Frequently asked questions

What is a tumour-immune organoid?

A three-dimensional culture built from a patient's tumour that also contains an immune compartment, either preserved from the original tissue, rebuilt from the patient's own immune cells, or engineered on a chip. Checkpoint inhibitors act through immune cells, so epithelial-only organoids cannot test them.

What did the review pool?

Five peer-reviewed reports with at least five paired patients each, 102 patients total. Pooled sensitivity was 0.70 (95% credible interval 0.48 to 0.89) and model-implied specificity was 0.97 (0.88 to 1.00), against cell counts of 35 true positives, 1 false positive, 17 false negatives and 49 true negatives.

Why is the specificity estimate fragile?

Because exactly one false positive occurred across the entire primary set. With a single event, the estimate is dominated by the statistical model's prior rather than the data, which is why the authors call it model-implied and weakly identified, and why quoting it as established performance is not defensible.

What does a sensitivity of 0.70 mean for patients?

In the intended use of deciding whether to add a checkpoint inhibitor, about 30% of true responders would test negative. In the primary evidence, 17 of 52 patients with clinical benefit were ex vivo negative, so acting on negative results would have denied benefit to a substantial fraction of responders.

How often do these assays succeed at all?

Success rates varied widely and were often incompletely reported: 22 of 31 melanoma samples, 28 of 30 hepatocellular carcinoma cultures, 98 of 328 lung-cancer organoids in one cohort, 95 of 111 in another, 35 of 36 assembloids, and 5 of 13 immunotumoroid specimens. All accuracy figures are conditional on passing these gates.

What would make these assays clinically usable?

By the review's specification: prospective consecutive enrollment, locked thresholds set before evaluation, blinded assessment, exact regimen matching, failure-inclusive denominators, external calibration, and direct comparison with established biomarkers and clinician choice in a randomized or otherwise decisive design.

Should a negative organoid result stop a patient getting immunotherapy?

No. The review's conclusion is explicit that these assays should not yet determine whether immunotherapy is added, and the pooled sensitivity of 0.70 with a third of clinical benefitters testing ex vivo negative shows exactly why.

References

  1. Tan C, Wang B, He S, Gong Y, Zhang L, Wang H, Tang Q, Li X, Xiong G, Zhou L, Li X. Patient-derived tumour-immune organoids as functional biomarkers of checkpoint-inhibitor response: a systematic review and exploratory meta-analysis. medRxiv. 2026. doi:10.64898/2026.08.17.26360042. https://www.medrxiv.org/content/10.64898/2026.08.17.26360042. Accessed 2026-09-10.