Diagnostic testing

· Q 1. I wish to test the accuracy of my pain prediction test in cancer patients with the help of data I have available for a representative sample of patients.  I have heard of the notions of sensitivity and specificity but I don’t really understand them. How can I find out more?

A. Have a look at slides 42-43 of the presentation slides Comparing two or more groups and measures and tests of association.

Loader Loading...
EAD Logo Taking too long?

Reload Reload document
| Open Open in new tab

This will allow you to practice relevant calculations using a worked example!

· Q 2. I have also heard of the Positive and Negative Predictive Value. How do they compare with sensitivity and specificity when interpreting diagnostic tests?

A. Have a careful look at slides 42 – 45 of the presentation slides Comparing two or more groups and measures and tests of association (see solution to Q. 1,above). This resource provides illustrations showing how to calculate all four of these statistics and why sensitivity and specificity may not be enough in assessing the quality of your diagnostic test.

N.B. Please consult the StatsforMedics WordPress page CROSS-TABULATING FREQUENCIES OR PERCENTAGES to discover how the measures discussed under Q.’s 1 and 2 on that page can be conveniently generated using SPSS.

· Q 3. I wish to calculate the 95% confidence intervals for sensitivity, specificity, positive and negative predictive values. I was going to use the exact binomial method (using the F-distribution). However, I then read that the score confidence interval by Agresti and Coull (1998) is even better than the exact binomial, even for the smallest sample sizes (as the exact confidence intervals tend to be conservative (and as such, comparatively wide)). Do you have any recommendations?

A. Two of the better methods are the Clopper-Pearson (1934)  “exact” and Wilson (1927) “score” method. The C-P method is more conservative in the above sense than the Wilson method. The first block of the embedded Excel template below calculates the Wilson confidence interval (CI) for a proportion. To replace the existing frequencies with your own and calculate the corresponding CI limits, click on the link ‘Download for personal editing’. Once you have downloaded the template, make sure that you click on the button ‘Enable Editing’ at the top.

TIPS ON HOW TO USE THE EXCEL TEMPLATE
You will see the expression ’81 out of 263′. You need to replace the ’81’ and ‘263’ by the numerator and denominator, respectively of the fraction which you used to calculate your sensitivity, specificity, NPV or PPV value. By way of illustration, for the example provided under Q. 1, the solution ‘Sensitivity = a/(a+c)=997/(997+6)=0.99’ is provided. This involves a numerator of 997 and a denominator of 1003 (or ‘997’ out of ‘1003’).  From the Excel  template provided, using the Wilson score method and rounding to 3 decimal places, the corresponding 95% CI is  (0.987, 0.997).

Having gone through the above worked example, you should now know which values to insert in the spreadsheet and verify your own work on the basis of this illustration.

 Access the original paper which presents the Wilson score method

*Thanks is due to Professor Robert G Newcombe, Professor Emeritus of Statistics, Cardiff University for providing the above Excel template.

· Q 4. I would like advice on comparing sensitivity measures across different diagnostic tests – ESR and CRP – by way of assessing the relative quality of these tests. I would also like to repeat this approach for other measures of diagnostic quality, such as specificity, NPV and PPV. Can you recommend some suitable literature on appropriate methodologies?

Case A. Patients who receive ESR belong to an independent group to those who receive CRP
In this case, could use Fisher’s Exact test of the chi-square test of association to perform your comparisons. You will want to perform your choice from these two tests using summary data. For help with this, refer to the section Handling Summary Data in the solution to Q. 2 on the StatsforMedics page HYPOTHESIS TESTS FOR CATEGORICAL DATA. For an example of what data are required, have a look at the section Comparing Two Sensitivities in the document Tests for Two Independent Sensitivities.

Case B. The same patients are receiving the two interventions
In terms of hypothesis testing, when considering sensitivity or specificity I would recommend the McNemar test – not the chi-square test of association – for this purpose, and this is backed up by the literature. For  example,  please consider the letter and the full text of the articles in the reference list for this letter available under the title McNemar Test Is Preferred for Comparison of Diagnostic Techniques.  I would particularly recommend the reference Matchmaking and McNemar in the Comparison of Diagnostic Modalities, as the illustrations are so clear and analogies can be conveniently formed with your own data.  Thanks is due to the author, Dr Andrew J Dwyer, for kindly providing a copy of this paper for use at this site.

The paper by Dwyer also provides useful information on confidence intervals for differences in sensitivity. While the illustrations apply to sensitivity comparisons, the methodologies can very easily be adapted to cover comparisons for other measures. Within the above recommended textbook, you should also consider the corresponding advice provided on the use of Cohen’s effect size index.

Tip: If you are still in the process of collecting your own data, it is highly advisable for you to practise techniques using the data from the above paper by Dwyer. As well as using SPSS to perform the McNemar test, you could set up a template in Excel for calculating the relevant confidence intervals in Excel by using the available arithmetic functions. When your data are ready, you will then be in a much better position to complete the relevant work efficiently. 

Additional note. When pondering the choice of the McNemar test, please bear in mind that the data leading to the sensitivity measures for the different diagnostic tests are likely to be correlated (and therefore not independent). This is because these data arise from the same patients and therefore there are underlying patient characteristics which may prevent the diagnostic test results from being independent for each patient. This correlation needs to be taken into consideration when the standard error of measurement  (SEM) is estimated for the differences in sensitivities. Having the wrong hypothesis test can lead to an under-estimation of this SEM and in turn, to the p-value being too low. Ultimately, this can lead to the wrong conclusions for your study!

NPV and PPV need to be dealt with differently, as now you must commence with the results of the diagnostic tests and if one considers PPV, for example, you need to compare across ESR and CRP the proportion of diagnosed positives that are indeed true positives, and of course, the diagnosed positives are not necessarily the same for ESR and CRP. You are therefore no longer comparing like with like. Thus, a reasonable case could be defended for using the chi-square test of association (or Fisher’s Exact test if needed instead) for comparing each of NPV and PPV across ESR and CRP. Yes, there is some degree of overlap in the diagnosed positives (or negatives) for each of ESR and CRP but there is not such a strong case for using the McNemar test. This is implicit from Dwyer’s paper.

· Q 5. I would like to have a measure of how much a positive diagnostic test result (e.g. serum ferritin of around 60 mmol/l) improves the accuracy of the pre-test odds of having a particular disease (in this case, anaemia) so as to see if the test is worthwhile. Can you recommend a suitable statistic?

A. Yes; you should consider the Likelihood Ratio for a positive result (LR+). Please note that there is a similar statistic (LR-) for assessing the test in relation to negative test results. More details can be found under Likelihood Ratios. Please be aware of the handy resource Confidence Interval Calculator for obtaining the corresponding 95% CIs. The relevant sheet from the Excel template carries the title two-level likelihood ratios. If you hover over the cells of the table for entering your frequencies, you will be advised what is needed in each case.

· Q 6. I am about to conduct a study involving the evaluation of a survey procedure for diagnosing pneumonia in developing countries. This is to involve assessing the sensitivity and specificity of the survey mechanism. At this stage, I am interested in performing a sample size calculation for my study in order to ensure that sensitivity and specificity can be estimated to a particular degree of accuracy (defined in terms of the width of their confidence intervals). Can you recommend how to proceed?

A. Please think very carefully before deciding on a sample size calculation. A precise formula is used to determine the minimum sample size needed and therefore you will require to be rather accurate about the estimates you enter into the formula.
To see exactly what is involved, please refer to the article
Statistical Methdology: I. Incorporating the Prevalence of Disease into the Sample Size Calculation for Sensitivity and Specificity by Dr Nancy M Fenn Buderer (formerly of St Vinent Medical Centre, Toledo, Ohio).
For your convenience, here is a reminder (as an adaptation from the above article) of the specifications which you must provide:

Specifications

You should specify:

i) the maximum clinically acceptable width of the confidence interval (CI). Call it W.*
*[W is actually half the width of the confidence interval, as it is the value you would usually express in the form ‘± W’ for a two-sided confidence interval]
ii)  an estimate for the prevalence of disease in the target population: call it P;
iii) a value for the expected sensitivity of the new diagnostic test:  call it SN;

and

iv) a value for the expected specificity of the new diagnostic test: call it SP.

If you have difficulties with iii) and iv), note the recommendation under point 3. on p. 897.
(For purposes of calculations, W, P, SN and SP are expressed as numbers between 0 and 1, rather than as percentages.)

· Q 7. I have been advised to use an ROC curve to assess whether the TIMI score is a good predictor of an adverse outcome for patients presenting with suspected cardiac chest pain. However, to be honest, I hadn’t previously heard of an ROC curve and don’t really understand why it is useful. Where can I find a gentle introduction to ROC curves?

A. First note that ROC (Receiver Operating Characteristic Curves) apply specifically to the case where your response variable has two outcomes (usually ‘adverse’ or ‘normal’). If your response variable has more than two outcomes, you may wish to merge two or more of these outcomes to make your data amenable for using with an ROC curve. Also, do not attempt to understand ROC curves until you are sure that you understand the notions of sensitivity, specificity and positive and negative predictive value (see Q 1. and Q 2., above).
Now let’s progress with some fundamental points about ROC curves which should address your immediate concerns.

· Q 8. I understand from the above resource that it is possible to categorize the quality of the TIMI and GRACE scores as predictors of adverse outcome according to the area under the ROC curve. However, is there a special test which I can implement which would allow me to test for the significance of the difference in the areas under two ROC curves as a means of measuring the relative qualities of these predictive scoring systems?

A. Yes, some good research has been done in this area. The article The Magnificent ROC Curve presents the results of this research at a gentle pace. You will in turn need to refer to Table 1 from the original paper by Hanley and McNeil, which is a look-up table for a correlation coefficient used in determining the test statistic for your hypothesis test.  Alongside the test statistic, z, you should also report a corresponding p-value to represent statistical signficance.  Here is a handy calculator to help you:

P Value Calculator.

· Q 9. Which University supported package would you recommended for generating one or more ROC curves?

A. SPSS. The most up-to-date version of IBM SPSS currently supported by the University of Edinburgh (Version 22.0) provides two very useful real-life examples by way of explaining how to generate the relevant output and interpret it. You are also directed to the relevant datasets, hivassay.sav and bankloan.sav, which come with the package.

These spreadsheets can be readily accessed via the University’s desktop managed PC’s by means of the menu File within SPSS. Just follow the path

***UPDATE***

File –> Open –> Data and where the resultant dialogue box specifies ‘Look in:’, follow the path App Virt(Q:) –> SPSS 19 x64 –> Samples –> English.

To access the tutorial itself, simply click on worked examples on creating ROC curves using SPSS.

. Q. 10. I have heard that ROC slopes can be used to assess the relative qualities of diagnostic tests for diabetes miellitus. What is this all about? (My diagnostic test has two possible outcomes – postive or negative – and the experimental measurements are on a continous scale.)

A. For measurement (that is, continuous rather than categorical) data there are two slopes of interest here.  These are to assess which of two or more tests will best help us to rule in or rule out disease in a given patient. The values of the slopes are defined as likelihood ratios. However, these ratios can also be defined in terms of the sensitivity and specificity of a particular test – notions with which you should already be familiar.
It is recommended that in the first instance, you read the paper Slopes of a Receiver Operating Characteristic Curve and Likelihood Ratios for a Diagnostic Test so as to get a feel for the mathematics involved.
So as to ensure that you are not left out at sea, however, you should then proceed to the resource on Likelihood Ratios provided under the solution to Q. 5, above.  You should then be able to make some informed judgments regarding the usefulness of your diagnostic tests based on the results that you obtained.

CC BY-NC-ND 4.0 Diagnostic testing by Margaret MacDougall is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

The WordPress site for supporting undergraduate medical student learning in statistics for short research projects

  • If you are visiting StatsforMedics for the first time, welcome!
  • Please take time to visit the page SCOPE OF SITE (see menu bar, below) for advice on how to make best use of the site and how to contact me.
  • University of Edinburgh undergraduate medical students: feel free to contact me if you need further assistance with your *curricular* activities.