Confusion Matrix Formulas
All 22 formulas in one table, with plain-English meanings and a worked example.
Free · No sign-up · Runs in your browser
Open the calculator Loads TP 45 · FN 5 · FP 27 · TN 423 Jump to the cheat sheetThe Complete Formula Reference
The Four Building Blocks
Every formula uses four counts: TP (true positive), TN (true negative), FP (false positive) and FN (false negative). New to them? Start with the walkthrough.
All 22 Formulas
| Metric | Formula | Plain English |
|---|---|---|
| Core Metrics | ||
| Accuracy | (TP+TN) / Total | Share of all predictions that were correct |
| Sensitivity (Recall, TPR) | TP / (TP+FN) | Of actual positives, how many were caught |
| Specificity (TNR) | TN / (TN+FP) | Of actual negatives, how many were cleared |
| Precision (PPV) | TP / (TP+FP) | Of positive calls, how many were right |
| Negative Predictive Value | TN / (TN+FN) | Of negative calls, how many were right |
| Prevalence | (TP+FN) / Total | How common the positive class actually is |
| Composite Scores | ||
| F1-Score | 2·P·R / (P+R) | Harmonic mean of precision and recall |
| Balanced Accuracy | (Sens+Spec) / 2 | Honest accuracy when classes are imbalanced |
| Matthews Correlation (MCC) | (TP·TN−FP·FN) / √(...) | Correlation between predictions and truth, −1 to 1 |
| Cohen's Kappa | (p₀−pₑ) / (1−pₑ) | Agreement with the truth beyond chance |
| Youden's J | Sens+Spec−1 | 0 is useless, 1 is flawless |
| Markedness | PPV+NPV−1 | How trustworthy both predictions are, together |
| Jaccard Index (IoU, CSI) | TP / (TP+FP+FN) | Overlap of predicted and actual positives, ignoring TN |
| Fowlkes–Mallows | √(P·R) | Geometric mean of precision and recall |
| Rate Complements | ||
| False Positive Rate | FP / (TN+FP) | False-alarm rate on actual negatives |
| False Negative Rate | FN / (TP+FN) | Miss rate on actual positives |
| Error Rate | (FP+FN) / Total | 1 − Accuracy |
| False Discovery Rate | FP / (TP+FP) | 1 − Precision |
| False Omission Rate | FN / (TN+FN) | 1 − NPV |
| Diagnostic Ratios | ||
| Positive Likelihood Ratio | TPR / FPR | How much a positive result shifts the odds up |
| Negative Likelihood Ratio | FNR / TNR | How much a negative result shifts the odds down |
| Diagnostic Odds Ratio | LR+ / LR− | Overall discriminative power in one number |
Worked Example: A Diagnostic Test
A test screens 500 patients, 50 of whom have the condition. It catches 45 (TP=45, FN=5) and clears 423 of the 450 healthy ones (TN=423), with 27 false alarms (FP=27).
| Test positive | Test negative | |
|---|---|---|
| Actually has it | TP = 45 | FN = 5 |
| Actually doesn't | FP = 27 | TN = 423 |
| Metric | Calculation | Result |
|---|---|---|
| Accuracy | (45+423) / 500 | 93.6% |
| Sensitivity | 45 / 50 | 90% |
| Specificity | 423 / 450 | 94% |
| Precision | 45 / 72 | 62.5% |
| NPV | 423 / 428 | 98.83% |
| F1-Score | 2(.625)(.90) / (.625+.90) | 73.77% |
| MCC | (TP·TN−FP·FN) / √(...) | 0.72 |
| Positive Likelihood Ratio | 90% / 6% | 15.0 |
Accuracy (93.6%) and sensitivity (90%) look good, but precision is 62.5%: over a third of positives are false alarms. That is low prevalence (10%) at work, and accuracy alone hides it. See also likelihood ratios, or try your own counts.
Frequently Asked Questions
Why isn't there just one confusion matrix formula?
Because errors cost different things. A spam filter that lets spam through and one that blocks real mail are both wrong, but not equally. Each formula isolates one failure mode.
What's the difference between a rate and a ratio in this table?
A rate (accuracy, sensitivity, FPR) is a proportion from 0 to 1. A ratio (LR+, DOR) divides one rate by another and has no upper bound. LR+ = 15 means a positive result is 15 times likelier from a true case than from a false alarm.
Which formula should I use to compare two models?
On imbalanced data, Balanced Accuracy or MCC, since both use all four cells. If one error costs more, use the metric that captures it: precision, recall or a weighted F-beta.
Do these formulas work for multi-class problems?
Not directly; they assume a 2×2 matrix. For 3+ classes, compute each one per class (one vs. rest), then macro-average (classes equal) or weighted-average (by class size). The calculator's Multi-Class tab does this.