Research Pitfalls That Can Make Evidence Look Stronger Than It Is

Research can be carefully performed and still leave important questions unanswered. Study design, participant selection, outcome measurement, missing data, analysis, and reporting decisions all influence what the findings can reasonably support.

For DNP students, critical appraisal requires moving beyond labeling a study as simply “good” or “bad.” The more useful questions are what conclusions the methods support, how much confidence the findings deserve, and how much weight the study should carry when making a practice decision.

Start With What the Study Design Can Actually Answer

Research design places limits on the conclusions that can be drawn from the results. Randomization can strengthen causal inference about an intervention by reducing systematic differences between comparison groups, provided that allocation and other methodological safeguards are successfully implemented. Observational studies can answer important questions about associations and patterns in clinical populations, but confounding and selection bias can provide alternative explanations for observed relationships (Higgins et al., 2024; Vandenbroucke et al., 2007).

Problems arise when the conclusion becomes stronger than the design. A cross-sectional study may identify an association between burnout and turnover intention, for example, but measuring both variables at the same point makes temporal order difficult to establish (Vandenbroucke et al., 2007). The study alone may not determine whether burnout contributed to turnover intention, turnover concerns contributed to burnout, or another factor influenced both. Causal language therefore deserves greater scrutiny when the methods cannot adequately distinguish a proposed effect from competing explanations.

Pay Attention to Who Entered the Study and Who Did Not

A large sample is not necessarily representative of the population where the findings will be applied. Recruitment methods determine who had an opportunity to participate, while eligibility criteria determine who was intentionally excluded. Those decisions can affect both selection bias and how broadly the findings can be applied (Vandenbroucke et al., 2007).

Consider a study of a complex outpatient intervention that excludes patients with cognitive impairment, unstable housing, limited English proficiency, serious comorbidities, or difficulty attending follow-up visits. Those exclusions may be justified by the research question, but they also narrow the population represented by the results.

Loss to follow-up can create another source of bias. Participants with missing outcomes may differ systematically from those who remain, particularly when the reason for missingness is related to the outcome or differs between study groups (Higgins et al., 2024). Participant flow therefore provides important context for the final sample. How many people were assessed, excluded, enrolled, lost, and analyzed can substantially change how the reported result should be interpreted.

Ask Whether the Outcome Measures What You Think It Measures

Research conclusions depend on how outcomes are defined and measured. A measure may be collected consistently while answering a narrower question than the one a reader assumes the study answered. An educational intervention might improve knowledge immediately after training, but that finding does not establish that clinicians changed their practice or that patient outcomes improved. A referral initiative might increase the number of referrals entered without demonstrating that patients successfully connected with the service.

Outcome measurement can also introduce bias when methods differ between groups or when knowledge of the intervention can influence assessment. Cochrane treats measurement of the outcome as a distinct risk-of-bias domain in randomized trials and considers whether outcome assessment could differ between groups or be influenced by knowledge of the intervention received (Higgins et al., 2024). The conclusion should remain tied to the variable the investigators actually measured rather than a more clinically meaningful outcome that was never evaluated.

Missing Data Can Change the Result

Missing data deserve more attention than the percentage of participants who completed follow-up. Suppose an intervention begins with 300 participants and reports outcomes for 215. Interpretation depends partly on who was lost, why their data are missing, whether missingness differed between groups, and whether the reason for missingness could be related to the outcome.

A follow-up intervention provides a useful example. Patients who are easiest to contact may also have more stable housing, reliable phone access, transportation, or greater ability to navigate healthcare. If people who are harder to reach disproportionately disappear from the analysis, the findings may describe the patients for whom the intervention was easiest to deliver better than those facing greater barriers. Cochrane cautions against judging missing-data bias from the proportion missing alone because risk also depends on why outcomes are missing and whether missingness may depend on the true value of the outcome (Higgins et al., 2024).

Look Beyond Statistical Significance

A statistically significant result does not establish that an effect is large enough to matter clinically. The effect estimate and its confidence interval provide additional information about the magnitude and precision of the result. A small effect can reach statistical significance in a sufficiently large sample, while a potentially important effect may remain uncertain when its confidence interval is wide. Critical appraisal therefore considers what the estimate suggests, how precise it is, and whether the range of effects compatible with the data would meaningfully change a clinical decision.

Confidence intervals are particularly useful when a result does not meet a conventional threshold for statistical significance. A nonsignificant result does not automatically demonstrate that two interventions are equivalent or that no clinically important difference exists. The interval may still include effects large enough to matter in practice. CONSORT 2025 recommends reporting effect estimates with measures of precision, such as 95% confidence intervals, rather than relying on P values alone (Hopewell et al., 2025).

Look for Results That May Have Been Selected After the Data Were Known

A study can measure several outcomes, evaluate them at multiple time points, and analyze the data in different ways. That flexibility becomes concerning when the reported result is selected because it is more favorable than other available results. Trial registration, protocols, and statistical analysis plans can help readers determine what investigators intended to measure and analyze before the results were known. CONSORT 2025 calls for transparent reporting of trial registration, access to protocols and statistical analysis plans, prespecified outcomes, and changes made after a trial began (Hopewell et al., 2025).

A primary outcome that changes between registration and publication deserves attention, as does a paper that emphasizes a favorable secondary result after the prespecified primary outcome was less convincing. A secondary or post hoc finding may still be useful, but readers should be able to distinguish what was planned from what was investigated after the data became visible. Cochrane similarly considers whether a reported result may have been selected from multiple measurements or analyses on the basis of the findings (Higgins et al., 2024). Selective reporting can make an evidence base appear more favorable even when the underlying data have not been fabricated.

Strong Methods, Clear Reporting, and Applicability Are Different Questions

A methodologically strong study may still be a poor match for the clinical setting where its findings are being considered. An intervention tested in a highly resourced academic center may depend on staffing, technology, specialist availability, patient characteristics, or follow-up capacity that differs substantially from another setting. Participant characteristics also influence whether findings can reasonably be generalized beyond the population studied (Vandenbroucke et al., 2007).

Reporting quality is another separate issue. CONSORT and STROBE are reporting guidelines intended to improve the completeness and transparency of research reports (Hopewell et al., 2025; von Elm et al., 2007). Complete reporting makes methodological strengths and weaknesses easier to identify, but following a reporting checklist does not itself establish that the underlying study was well designed or conducted. STROBE explicitly states that its recommendations concern reporting and are not intended to function as an instrument for evaluating the quality of observational research (von Elm et al., 2007).

A clearly reported limitation therefore remains a limitation. Transparency allows readers to identify where confidence should be reduced, methodological appraisal addresses how much those limitations threaten the findings, and applicability asks whether otherwise credible evidence fits the population and setting where it might be used.

Read for the Link Between the Question and the Conclusion

Critical appraisal becomes easier when a study is viewed as a sequence of methodological decisions. The research question leads to a design, participants are selected, outcomes are measured, data are analyzed, and those results are used to support a conclusion. Weakness at any point can narrow what that conclusion reasonably means.

The purpose of appraisal is not to find a flaw large enough to dismiss every study. Nearly all research has limitations, and different limitations carry different consequences. A useful appraisal identifies which weaknesses threaten the conclusion being considered, examines the magnitude and precision of the findings, and adjusts the weight given to the evidence accordingly.

A statistically significant result does not settle that judgment. The central question is whether the methods and results support the conclusion strongly enough to deserve the confidence being placed in it.

References

Higgins, J. P. T., Savović, J., Page, M. J., Elbers, R. G., & Sterne, J. A. C. (2024). Assessing risk of bias in a randomized trial. In J. P. T. Higgins, J. Thomas, J. Chandler, M. Cumpston, T. Li, M. J. Page, & V. A. Welch (Eds.), Cochrane handbook for systematic reviews of interventions (Version 6.5). Cochrane. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-08

Hopewell, S., Chan, A.-W., Collins, G. S., Hróbjartsson, A., Moher, D., Schulz, K. F., Tunn, R., Aggarwal, R., Berkwits, M., Berlin, J. A., Bhandari, N., Butcher, N. J., Campbell, M. K., Chidebe, R. C. W., Elbourne, D., Farmer, A., Fergusson, D. A., Golub, R. M., Goodman, S. N., . . . Boutron, I. (2025). CONSORT 2025 statement: Updated guideline for reporting randomised trials. BMJ, 389, e081123. https://doi.org/10.1136/bmj-2024-081123

Vandenbroucke, J. P., von Elm, E., Altman, D. G., Gøtzsche, P. C., Mulrow, C. D., Pocock, S. J., Poole, C., Schlesselman, J. J., & Egger, M. (2007). Strengthening the reporting of observational studies in epidemiology (STROBE): Explanation and elaboration. PLoS Medicine, 4(10), e297. https://doi.org/10.1371/journal.pmed.0040297

von Elm, E., Altman, D. G., Egger, M., Pocock, S. J., Gøtzsche, P. C., & Vandenbroucke, J. P. (2007). The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. PLoS Medicine, 4(10), e296. https://doi.org/10.1371/journal.pmed.0040296

Discover more from YourDNP.com

Subscribe now to keep reading and get access to the full archive.

Continue reading