Screening Test Validity for FMGE: Sensitivity, Specificity, and the Two Numbers Examiners Actually Test

By Dr. Utsav Bhattacherjee, MBBS, MBA · 9 September 2026 · 11 min read
Every sensitivity/specificity question is really the same question wearing a different clinical scenario: fill in a 2x2 table, then read off the ratio the question is actually asking for. The four terms sound similar and are easy to confuse under time pressure — the fix isn’t memorizing more definitions, it’s memorizing the table.
The 2x2 Table, Fixed in Memory
Every screening test validity question reduces to four numbers, arranged the same way every time:
| Disease Present | Disease Absent | |
|---|---|---|
| Test Positive | a (True Positive) | b (False Positive) |
| Test Negative | c (False Negative) | d (True Negative) |
Once this table is filled in from a vignette, every other value is a simple ratio built from these four cells. The skill being tested is filling the table correctly from a word problem, not the arithmetic that follows — which is why it is worth practising Community Medicine previous year questions until filling the table is automatic.
Sensitivity and Specificity: Fixed Properties of the Test
Sensitivity is a/(a+c) — of everyone who actually has the disease, what proportion did the test correctly identify as positive. It answers: how good is this test at catching real cases?
Specificity is d/(b+d) — of everyone who doesn’t have the disease, what proportion did the test correctly identify as negative. It answers: how good is this test at correctly clearing healthy people?
The property worth understanding, not just memorizing: sensitivity and specificity are fixed characteristics of the test itself. They don’t change based on how common or rare the disease is in the population being tested — a test with 95% sensitivity has 95% sensitivity whether you’re screening a high-risk population or a low-risk one, because sensitivity is calculated entirely from among people who already have the disease, and specificity entirely from among people who don’t.
PPV and NPV: The Numbers That Actually Depend on Prevalence
Positive predictive value (PPV) is a/(a+b) — of everyone who tested positive, what proportion actually has the disease. This is the number a patient with a positive result actually cares about: given my positive test, what are the real odds I’m sick?
Negative predictive value (NPV) is d/(c+d) — of everyone who tested negative, what proportion is actually disease-free.
Here’s the detail that separates a fast, confident answer from a guess: PPV and NPV change with disease prevalence, even when sensitivity and specificity stay exactly the same. In a population where the disease is rare, even a highly specific test will generate a meaningful number of false positives relative to the small number of true positives — so PPV drops, sometimes dramatically, even though the test’s own sensitivity and specificity haven’t moved at all. In a high-prevalence population, the same test’s PPV rises, again without the test itself changing in any way.
Why This Prevalence Effect Is the Single Most-Tested Concept Here
A classic exam pattern: the same test is applied to two different populations — say, a general screening population versus a high-risk clinic population — and the question asks how PPV changes between them. The trap is assuming sensitivity or specificity must have changed too, when in fact neither did; only the prevalence in the tested population changed, and that alone is enough to move PPV and NPV substantially.
A concrete way to hold this intuition: imagine a test with excellent 99% sensitivity and 99% specificity, applied to a population where only 1 in 10,000 people actually has the disease. Even at that accuracy, the sheer number of disease-free people being tested means the absolute count of false positives can end up rivaling or exceeding the count of true positives — so PPV can be surprisingly low despite the test being highly accurate on paper. This is exactly why population-wide screening for rare conditions is approached cautiously: a "highly accurate" test can still generate more false alarms than real detections when the underlying disease is rare enough.
Putting the Four Together
| Term | Formula | Changes with prevalence? |
|---|---|---|
| Sensitivity | a/(a+c) | No — fixed property of the test |
| Specificity | d/(b+d) | No — fixed property of the test |
| PPV | a/(a+b) | Yes — rises with prevalence |
| NPV | d/(c+d) | Yes — falls with prevalence |
A Worked Example, With Real Numbers
Formulas stay abstract until you’ve filled in a table once. Say a screening test is used on 1,000 patients in a population where 100 actually have the disease:
| Disease Present (100) | Disease Absent (900) | |
|---|---|---|
| Test Positive | 90 | 90 |
| Test Negative | 10 | 810 |
From this table: Sensitivity = 90/100 = 90%. Specificity = 810/900 = 90%. PPV = 90/(90+90) = 50%. NPV = 810/(810+10) = 98.8%.
Notice what happened: even with a solid 90% on both sensitivity and specificity, PPV lands at only 50% — a coin flip — because the disease is uncommon enough in this population (10%) that the test’s false positives (90) end up equal in number to its true positives (90). Run the same 90%/90% test on a population where 50% actually have the disease instead of 10%, and PPV rises sharply without sensitivity or specificity moving at all. This is the prevalence effect from above, made concrete with actual numbers instead of just stated as a rule.
Likelihood Ratios: The Same Information, Framed Differently
Positive likelihood ratio (LR+) is sensitivity divided by (1 − specificity) — it tells you how much more likely a positive result is in someone with the disease compared to someone without it. An LR+ well above 1 (conventionally, above 5 or 10 is considered strong) means a positive result meaningfully raises the probability of disease.
Negative likelihood ratio (LR−) is (1 − sensitivity) divided by specificity — it tells you how much a negative result should lower your suspicion of disease. An LR− well below 1 (below 0.1 or 0.2 is considered strong) means a negative result is genuinely reassuring.
The advantage likelihood ratios have over PPV and NPV directly: likelihood ratios, like sensitivity and specificity, don’t change with disease prevalence — they’re calculated purely from the test’s own performance characteristics, which makes them useful for combining with a patient’s individual pre-test probability (based on their specific clinical picture) to arrive at a post-test probability, rather than relying on a population-wide prevalence figure that may not match any individual patient in front of you.
Overall Accuracy: The Summary Measure Worth Knowing, and Its Limit
Overall accuracy (also called efficiency) is (a+d)/(a+b+c+d) — the proportion of all test results, positive and negative combined, that were correct. It’s intuitive but has a real limitation worth flagging: in the worked example above, overall accuracy would be (90+810)/1000 = 90%, which sounds reassuring but says nothing about the 50% PPV sitting underneath it. A single accuracy figure can mask a real, clinically important weakness in one specific direction — which is exactly why sensitivity, specificity, PPV, and NPV are reported separately rather than being collapsed into one overall number.
How to Solve These Questions in the Exam
1. Identify what’s being asked — is the question framed from the test’s perspective (sensitivity/specificity) or the patient’s perspective, given a result (PPV/NPV)?
2. Fill in the 2x2 table from whatever numbers the vignette gives — this is usually the actual bottleneck, not the formula itself.
3. Check whether prevalence is changing between two scenarios in the same question — if it is, the answer almost certainly hinges on PPV or NPV moving, not sensitivity or specificity.
A Note on Screening Test Selection
Screening sits inside the wider biostatistics and epidemiology block — see how to prepare Community Medicine for FMGE for where it fits, and every Community Medicine question this site holds to test it. This framework also explains a related, frequently tested principle: highly sensitive tests are preferred for initial screening, because a false negative on a screening test means a real case gets missed entirely and never reaches confirmatory testing — a costly failure mode for a screening programme. Highly specific tests are preferred for confirmation, once a positive screening result needs to be verified, because a false positive at the confirmatory stage means unnecessarily alarming and treating someone who was never actually sick. This is why many screening programmes are structured as a highly sensitive first-pass test followed by a highly specific confirmatory test, rather than relying on one test to do both jobs well.