For optimal use, please visit numiqo on your desktop PC!

Actual outcomes
Test values or scores
Compare by group (optional)

How-to

Compare ROC Curves

Author: Dr. Mathias Jesussek
Updated:

ROC curves are paired when the tests were measured on the same participants and independent when the curves come from separate participant samples. The distinction determines how the uncertainty of the AUC difference must be calculated.

Open paired example Open independent example Open grouped example

Select the correct comparison design

DesignData layoutAnalysis
Several tests on the same participants One actual outcome and several test values or scores Paired ROC curves and paired DeLong comparisons
Tests from separate participant samples Separate actual outcome–test pairs Independent ROC curves and unpaired DeLong comparisons
One test compared across groups One outcome, one test and one grouping variable Independent group curves and unpaired DeLong comparisons

Paired ROC curves

Select one actual outcome and all tests that were recorded on the same participant rows. numiqo plots the curves together and compares every pair with the paired DeLong method. The covariance between AUC estimates is retained; treating these observations as independent would discard information about their shared participants.

Each individual AUC uses the valid outcome–test pairs available for that test. A paired comparison uses rows on which the outcome and both tests in that comparison are valid. Therefore, when missingness differs between tests, the comparison’s AUC difference can differ from subtracting the two AUC estimates displayed in the individual-results table.

Independent ROC curves from separate samples

Select the same number of actual outcomes and test values or scores. numiqo pairs them by their left-to-right column order and displays the mapping before the results. Each pair may have its own positive category and high-or-low direction. Because participants are not shared across curves, comparisons use the unpaired DeLong variance.

Compare performance by a grouping variable

For tidy data such as group, disease, marker, select disease as the actual outcome, marker as the test value or score and group under Compare by group. One curve is calculated for each group containing at least one valid positive and one valid negative observation. Rows with missing grouping values are reported separately. Degenerate groups are suppressed rather than displayed as valid ROC analyses.

What the comparison table reports

  • AUC difference for each pair of curves
  • Standard error and 95% confidence interval for the difference
  • Two-sided p-value for the null hypothesis of equal AUCs
  • Paired or unpaired design used for the comparison
  • Holm-adjusted p-value when more than one pairwise hypothesis is tested

Do not infer an AUC difference merely because one individual AUC is significant and the other is not, or because two individual confidence intervals appear to overlap. The direct comparison accounts for the uncertainty of the difference and, for paired data, the covariance between AUC estimates.

Multiple comparisons

With three or more curves, several pairwise hypotheses are tested. numiqo reports the raw p-value and applies Holm’s sequential adjustment across the displayed comparisons. The adjustment controls the family-wise error rate while retaining the ordering and effect estimates needed for interpretation. Report the AUC differences and confidence intervals, not only whether an adjusted p-value crosses a threshold.

Confidence bands and interpretation

Optional shaded bands are pointwise 95% percentile bootstrap intervals for sensitivity at fixed specificity values, based on 2,000 stratified resamples per curve. They show uncertainty within each curve but are not a simultaneous region and are not a substitute for the DeLong test of an AUC difference.

Related ROC guides

References

Cite numiqo: numiqo Team (2026). numiqo: Online Statistics Calculator. numiqo e.U. Graz, Austria. URL https://numiqo.com