Meta-Analysis and Forest Plots
How Many Studies Become One Estimate

One Pooled Answer From Several Studies
Say five different trials test whether the same drug lowers the risk of a stroke. One finds a strong benefit, two find a modest benefit, one finds almost nothing, and one finds a slight increase in risk. Which one is right? The honest answer is usually none of them individually. Each trial is a noisy estimate of the same underlying truth, and meta-analysis is the formal way of combining them into a single, more precise estimate rather than picking a favorite or averaging them by eye.
The tool used to both perform and communicate that combination is the forest plot. It looks intimidating the first time you see one, rows of lines and boxes stacked above a diamond, but every piece of it has a specific job, and once you know what each piece represents, reading one becomes mechanical.
Reading a Forest Plot
Each row of a forest plot is one study. The square (or circle) on that row marks the study's point estimate, the effect size that trial actually measured, whether that is a risk ratio, an odds ratio, or a mean difference. The horizontal line through the square is that study's confidence interval, showing the range of effect sizes the data are consistent with. A short line means a precise estimate; a long line means a lot of uncertainty.
The size of the square itself is not decorative. Larger squares represent studies that carry more weight in the pooled result, and smaller squares represent studies that carry less. At the bottom of the plot sits a diamond, the pooled estimate across all studies combined. The width of the diamond is the confidence interval around that pooled estimate, and it is almost always narrower than any single study's interval, because combining data reduces random error.
A worked example makes this concrete. Imagine four trials of a blood pressure drug's effect on stroke risk, reported as risk ratios: Trial 1 finds a risk ratio of 0.82 with a wide confidence interval, 0.60 to 1.10, because it only enrolled 200 patients. Trial 2, with 1,500 patients, finds a risk ratio of 0.75 with a tight interval, 0.68 to 0.83. Trial 3 finds 0.90 with a moderate interval, 0.70 to 1.15. Trial 4 finds 0.78 with another tight interval, 0.65 to 0.94. None of the individual confidence intervals is dramatically different from the others; they overlap. Pooled together, the diamond lands around a risk ratio of 0.79 with a confidence interval of 0.72 to 0.87, a narrower, more confident estimate than any single trial produced on its own.
Why Studies Get Weighted Differently
Trial 2 and Trial 4 pulled the pooled estimate harder than Trial 1 and Trial 3 did, and that was not arbitrary. The standard approach is inverse-variance weighting: a study's weight in the pooled estimate is proportional to the inverse of its variance, which in practice means more precise studies (larger sample sizes, narrower confidence intervals) get more say, and smaller, noisier studies get less. This is the same logic behind why a national poll of ten thousand people is trusted more than a street survey of twenty, applied to clinical trials instead of polls.
Fixed Effects vs. Random Effects Models
There are two common ways to do the pooling, and they rest on different assumptions about what the included studies are actually estimating.
A fixed effects model assumes every study is estimating the exact same true effect, and any differences between study results are purely due to random sampling error. Under that assumption, studies are weighted almost entirely by their sample size and precision, since the only source of disagreement being modeled is chance.
A random effects model assumes the true effect itself varies across studies, perhaps because the populations, dosages, or follow-up periods differed, and each study is estimating its own version of a related but not identical effect. This model adds an extra layer of variance to account for that between-study variation, which typically widens the pooled confidence interval compared to a fixed effects analysis of the same data. In practice, random effects models are the more common and more conservative default, especially once there is any reason to suspect the studies were not all measuring precisely the same thing.
Reasons to Look Twice
A forest plot can still mislead even when the arithmetic is correct. Two issues are worth knowing before trusting a pooled diamond at face value.
The first is heterogeneity, when the individual studies disagree more than chance alone would explain. This often shows up visually as confidence intervals that barely overlap or point estimates scattered widely across the plot, and is formally quantified with statistics like I-squared. High heterogeneity is a sign that averaging the studies together may be combining apples with oranges, and that the pooled estimate deserves less confidence than its narrow diamond suggests.
The second is publication bias: studies that found a clear, positive effect are more likely to get published than studies that found nothing, which means the pool of available trials for any given meta-analysis can be skewed toward favorable results before the pooling even begins. This is sometimes called the file drawer problem, since the unflattering studies tend to sit in a drawer rather than a journal. A funnel plot, a separate chart entirely, is the usual way analysts check for this kind of asymmetry.
A forest plot's diamond is only as trustworthy as the studies feeding into it, pooling precision cannot fix disagreement or a biased sample of trials.
Key Takeaways
Meta-analysis combines several studies into one pooled estimate rather than relying on any single trial.
In a forest plot, each row is a study: the square marks its point estimate, the line marks its confidence interval, and the diamond at the bottom marks the pooled result.
Studies are weighted by precision, not by vote; larger, more precise studies pull the pooled estimate harder through inverse-variance weighting.
Fixed effects models assume one true effect across studies; random effects models allow the true effect to vary and produce wider, more conservative pooled intervals.
High heterogeneity means the included studies disagree more than chance explains, and should lower confidence in the pooled estimate.
Publication bias can skew which studies are even available to pool, since positive results are more likely to get published than negative ones.
Up Next:
Next, we will shift from pooling separate studies to looking within a single study over time, with Kaplan-Meier curves and how censoring works when some patients leave a study before the outcome of interest ever happens.