Comparative efficacy trials have historically served as visible clinical confirmation within biosimilar development, but their incremental scientific value has become increasingly contested as analytical characterization, functional testing, pharmacokinetic comparison, pharmacodynamic biomarkers, and immunogenicity assessment have become more sensitive. This critical evidence review evaluates whether comparative efficacy trials still resolve uncertainties that materially affect biosimilar development or regulatory decision-making. A transparent, question-led search and appraisal approach was used to prioritize literature addressing the totality of evidence, analytical similarity, reference-product variability, endpoint sensitivity, pharmacokinetic and immunogenicity uncertainty, and the consequences of retaining or waiving comparative efficacy testing. The central tension is that patient-level efficacy endpoints appear clinically direct but may be less capable than upstream methods of detecting or attributing small product-related differences. Disease heterogeneity, background treatment, saturated dose-response relationships, broad equivalence margins, endpoint measurement error, and reference-product variability can further reduce discrimination. Nevertheless, manufacturing-related immune risk, inadequately characterized functional differences, or uncertainty that cannot be resolved through analytical or clinical-pharmacology evidence may still justify targeted comparative clinical testing. The reviewed evidence supports neither automatic retention nor universal elimination of comparative efficacy trials. Their value depends on whether a clinically plausible and product-related uncertainty remains, whether the selected clinical model can detect its expected consequence, and whether the result could alter a development or regulatory conclusion. Comparative efficacy testing should therefore be treated as a conditional uncertainty-resolution tool rather than a routine final step. This judgement remains product-, mechanism-, assay-, population-, and regulatory-context dependent.