· Q. 1. A new technique for the measurement of exophthalmos has been developed using digital photography. I wish to compare measurements taken for the same subjects using this technique with those obtained using the standard Hertel technique. Also, for any one technique, I would like to assess inter-observer reliability with respect to my own measurements and that of a more experienced doctor. Can you recommend techniques for representing the agreement in measurements across a) two methods or b) two observers for a sample of different patients?
A. In such circumstances, it is not enough to compare summary measures, such as means or medians. The Bland-Altman method should provide some valuable insight. however, into levels of agreement. Please take a careful read of Bland and Altman’s 1986 Lancet paper to assess, for example, how changes in level of agreement can be taken into account as the magnitude of the exophthalmos changes. This paper is provided below as follows:
Statistical Methods for Assessing Agreement Between Two Methods of Clinical Measurement
In relation to part b) of the current question, you may also wish to consider Q. 4 and the accompanying solution, below,
· Q. 2 Having read pp. 5 to 6 of this article, I realise that I need to log-transform my differences data for one of the following reasons: a) my difference data are not Normal and I wish to endeavour to use a log-transformation to Normalize the data or b) there is evidence of systematic bias in the difference in my Bland-Altman plot in the sense explained in the above paper. The details of how to proceed with the log-transformed data are provided on pp. 5 – 6 of the paper but I am not clear how to conveniently log-transform the data. Can you offer some tips?
A. The resource How to do log transformation in SPSS/PASW includes instructions on how to carry out this process using SPSS. If you originally carried out tests of Normality on the untransformed data and discovered that your data were not Normal, you should now repeat the same tests on your log-transformed data.
NB. The procedure described in the tutorial of adding 1 to your original variable before taking logs is only illustrated as a means of avoiding taking the log of zero. You may not have any zeros in your data, in which case you don’t need to include the constant. If you have negative values in your data, you will need to replace ‘1’ by a number which, when added to all of the untransformed values, ensures that they are all positive. (Remember, the log function is only defined for positive numbers!)
· Q. 3. Can I perform a sample size calculation for use prior to implementing the Bland-Altman method.
A. Yes, you can; but check the assumptions concerning your data for running the Bland-Altman test before proceeding. These can be found in the above paper by Bland and Altman. For advice on the relevant sample size calculation, please refer to
How can I decide the sample size for a study of agreement between two methods of measurement?
· Q. 4. Could I adapt the Bland-Altman method to compare results for two raters using the same method? Also, if yes, should I use the coefficient of repeatability or coefficient of variation (assuming these measures are different)?
A. Regarding your first question, yes, that would be fine. Personally, I have a preference for the intra-class correlation coefficient rather than the Bland-Altman method for assessing inter-rater agreement, partly because it is a natural substitute for the Pearson Correlation for measuring agreement and it provides a sense of the extent of agreement on a scale from 0 to 1 which appeals to common sense. It is also easy to test this measure for statistical significance. On the other hand, I would like to reassure you that you that it isn’t wrong to adapt the Bland-Altman method for your purposes. If you opt for this approach, I suggest that a coefficient of repeatability would be a useful complement. The coefficient of repeatability is a measurement provided in the same units as your measurements. Precisely, it can be defined as 1.96 times the standard deviation of the differences between the repeated measures. This measure has a natural alliance with the Bland-Altman plot, as it represents the distance from the line for the mean difference to any of the limit lines (see the MedCalc resource Bland-Altman plot). As such, assuming the difference between measurements are Normally distributed, it helps you estimate the limits within which most of the mean difference for the repeated measures could lie with 95% certainty. You should then ask whether you think these limits are satisfactory for the measurements to be clinically useful. Is the possible error too high? The question of too high depends on the clinical context. How high would you wish your average difference in measurements across observers (as an absolute value) to go? Are the limits too wide from a clinical perspective? These are the kinds of questions you should ask. Please note that the coefficient of variation is the standard deviation of measurements calculated as a proportion of the mean and as such, differs from the coefficient of repeatability. If you wish to explore the intra-class correlation coefficient, I would recommend looking at sections 11.8.2 and 11.8.3 of Kinnear, P.R. & Gray, C.D. (2006) PASW Statistics 17 Made Simple, Psychology Press, Hove and New York. The steps and explanation on these pages relating to obtaining the intra-class correlation coefficient are rather user-friendly.
If you are registered with the University of Edinburgh, you can check the loan status of hard copies of the above books via the University’s library discovery system, DiscoverEd.
Also, for more comprehensive advice and details on use of the intra-class correlation coefficient, please refer to the MedStats WordPress page STATISTICAL INDICES FOR MEASURING AGREEMENT AND CONSISTENCY BETWEEN GROUPS OF CATEGORICAL AND MEASUREMENT DATA.
. Q 5. I would like to estimate the precision of the limits of agreement in my Bland-Altman plot by means of confidence intervals (CIs) using the equations that are recommended in the original Bland-Altman article. Assuming, I wish to use 95% CIs, can you recommend a quick way for obtaining the t-statistic used in these equations?
A. You can try the t-distribution generator provided under the indexed item Distributions at VassarStats: Website for Statistical Computation. Use the button reload to include your degrees of freedom (sample size – 1). Bear in mind that since you wish an upper and a lower limit, you are looking for t-statistic corresponding to a two-tailed p-value of 0.05. To check that you understand how to use the calculator, make sure that you would obtain the value 2.12 if you were carrying out the work for the example provided in the Bland-Altman paper
.
The Bland-Altman Method by Margaret MacDougall is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.