Medical Statistics

Jonas Ranstam, BSc, PhD, is a medical statistician and former professor of medical statistics at Lund University, Sweden. He has published more than 300 scientific articles in the fields of experimental and observational medical research and been cited over 30,000 times. He is a member of the World's Top 2% Scientists Network and was in 2016 recognised by STAT as "The World's Top Peer Reviewer".

Conditional and marginal models

Statistical analyses of related data need to take the relations into account to avoid misleading results. Two main types of statistical models are used, conditional and marginal models (1). A conditional model can be fitted within the framework of generalized linear mixed models (GLMM) and a marginal model using generalized estimating equations (GEE). Analyses based on conditional and marginal models give answers to different questions. While a conditional model can be used to estimate the ou...
Read post

Fixed effects, random effects, and mixed models

While classical statistical methods are based on an assumption of independent observations, many currently used statistical models include observations that are related instead of independent, such as repeated measurements from the same patient and patients randomised at the same centre. Fixed effects models, random effects models, and mixed models provide different ways to deal with independent and related, and a mix of independent and related observations (1). Fixed effects estimate a populat...
Read post

Efficacy, effectiveness, and efficiency

The concepts of efficacy, effectiveness, and efficiency are frequently misunderstood. They all pertain to the outcome of a medical intervention, but it is crucial to comprehend their differences accurately. Efficacy refers to whether an intervention produces more benefit than harm when tested under tightly controlled, ideal circumstances, such as strict inclusion/exclusion criteria, close monitoring, and high adherence to the treatment protocol. This effect is typically investigated in a classi...
Read post

Blinding the statistician?

Blinding randomised patients and doctors by masking treatment, where this is possible, is an established approach in clinical trials, aiming to prevent bias. It is not uncommon that published trial reports state that the statistician was also blinded, not only during the planning of the final analysis, which has been routine for a long time, but also during the analysis. Professional clinical trials units seem to have variable approaches to the blinding of trial statisticians (1). You might arg...
Read post

Frequentists and Bayesians

Followers of today's two main traditions of statistical inference are known as frequentist and Bayesian. Ronald Fisher attempted during the 1930s to develop a third school called fiducial inference (1), but this was broadly considered controversial and do not play a major role today. The fundamental difference between frequentists and Bayesians is that they define probability in different ways. For a frequentist, a probability is an objective measure, a long-run relative frequency. For example,...
Read post

Parameters and estimands

A search in PubMed shows that the use of the statistical term estimand has increased markedly during the past ten years. Yet, the term remains unfamiliar to many readers. In statistics, a parameter is a numerical characteristic of a population, probability distribution, or statistical model. Examples include a mean, proportion, variance, regression coefficient, hazard ratio, or risk. Because the population is rarely observed in its entirety, such quantities are usually inferred from sample data...
Read post

Aleatoric and epistemic uncertainty

Statistical inference is used to evaluate sampling uncertainty in medical research. The two most commonly used uncertainty measures are confidence intervals and p-values. However, it is often useful to distinguish between the uncertainty resulting from random variation (aleatoric uncertainty) and the uncertainty caused by incomplete knowledge (epistemic uncertainty). A small p-value indicates disagreement between observed data and a tested null hypothesis, but it is not in itself a direct measu...
Read post

Confirmatory trials and their interpretation

Unfortunately, the findings of confirmatory trials are often misinterpreted. There is no guarantee that a hypothesis is true just because it passes a statistically significant test. The significance level, typically 5%, represents nothing more than the likelihood of a false positive result. Hence, systematic reviews and meta-analyses, which lessen the uncertainty by integrating the findings from multiple trials of the same endpoint, play a significant role in the pursuit of truth. Conversely, c...
Read post

Exploratory studies, confirmatory trials, and Bonferroni correction

Medical research is primarily performed using samples of humans, laboratory animals, or cells, but the studied phenomena are rarely limited to what can be observed in samples of these. On the contrary, the aim is almost always to learn about the population from which the sample was drawn. However, this leads to generalisation problems. Sampling variability makes the results from sample studies uncertain, and the consequences of non-random sampling may induce an uncertain amount of bias. The un...
Read post

Incidence

The term 'incidence' may seem straightforward, but it can refer to three distinct metrics: the total number of new cases, cumulative incidence (risk), and incidence density. The calculations of these metrics can vary in difficulty (1). As demonstrated by Havers-Borgersen et al. (2) mistakes are often published, and numerous errors probably remain undetected as a result of unclear methodological descriptions. 1. Incident numbers In the most basic form, incidence begins with the number of new ca...
Read post

Quartiles, range, and interquartile range

Misuse of statistical terminology is very common in medical research reports. The misuse not only indicate methodological ignorance, it also threatens the consistency of the statistical terminology. For example, Nahoui et al. (1) state that in their sample of patients, those "in 3rd and 4th quartiles of median PES [esophageal pressure] had increased mortality risk compared to 1st quartile". Given that only three quartiles exist, this statement is remarkable. A quartile is defined (2) like this:...
Read post

Nonparametric data

Using the correct terminology helps to ensure that the same words are used for the same concepts, which is crucial for a clear communication and for avoiding misunderstandings. The term 'nonparametric data' appears often in the statistics section of research reports. For example, Dugan et al. (1) state that the "data from patients in the two groups were compared using Mann-Whitney U tests for nonparametric data." However, statistical tests are used to evaluate sampling uncertainty, to test a hy...
Read post

A very brief history of statistics in medicine

Medical science is one of the youngest sciences, at least as we define science today. From 1665, when the first scientific journals were established, to the mid-20th century, when modern medical research emerged, medical research publications were primarily descriptive (case reports) or presenting subjective comments (expert opinions). It lasted until the mid-20th century, until medical research focused on empirical evidence collected from samples of patients and evaluating the uncertainty of th...
Read post

Effect Measures: RR, HR, and OR

While the risk (or incidence) in absolute numbers is an important measure from an individual and a public health perspective and for planning health care resources, the biological effect of a beneficial or harmful exposure is always measured and medically interpreted in terms of relative risk. With some prospective study designs, the relative risk (RR) can be measured directly from risks or indirectly from incidence density rates. Other statistical methods produce other effect measures, such a...
Read post

Univariate, multivariate, univariable, and multivariable

The terms univariate, multivariate, univariable, and multivariable often appear in scientific medical publications. For example, Kanbaş et al., claim that they have evaluated factors prognostic for cervical cancer using multivariate Cox regression. However, Cox regression is a semi-parametric technique that they use to evaluate how a single response variable (survival time) is associated to one or more explanatory (potentially prognostic) variables. That is not a multivariate analysis. Univar...
Read post