Develop the skills to critically appraise scientific claims by dissecting sample sizes, placebo controls, statistical versus clinical significance, publication bias, p-hacking, and the limitations of peer review. These competencies will equip you to evaluate health news headlines, pharmaceutical marketing, and policy arguments that invoke "studies show" as their authority.
Most scientific claims reach you not as raw experiments but as studies boiled down to a headline. Judging them means knowing what to look for. How the study was designed. How big and representative the sample was. Effect size versus mere statistical significance. Conflicts of interest, and whether the result has actually replicated. This exercise drills the checklist that separates a confident reader of research from a credulous one.
This exercise trains you to appraise scientific claims by examining sample size, placebo controls, publication bias, and the gap between statistical and clinical significance. You practice reading past a headline finding to the design that produced it, which is where most overstated claims become visible.
Background
The 2000s were humbling for research credibility. When teams tried to replicate 100 well-known psychology studies, only a third to a half held up. Similar efforts in medicine and economics landed just as hard. The lesson isn't that research is worthless. It's that any single finding's credibility depends on details the headline almost always leaves out.
A few failure modes recur. P-hacking runs many analyses and reports only the ones that 'worked,' quietly inflating a 5% false-positive rate. HARKing invents the hypothesis after seeing the results, then presents it as a prediction. And publication bias means null results stay in the drawer, tilting the whole literature toward positive findings. Ask about replication, sample size, and conflicts before you believe a study. See Scientific Thinking.
Questions
0 of 6 answered
Question 1
A supplement brand's Instagram ad states: "Clinically proven! In a study at a leading university, participants who took NeuroFocus showed a 22% improvement in sustained attention scores after just two weeks (n = 16, p = 0.04)." You find the paper and discover it had no placebo control, used an unvalidated attention measure created by the company, and was funded entirely by the manufacturer. What is the most critical problem?
Question 2
A news headline reads: "New study proves acupuncture reduces chronic low back pain by 35% compared to no treatment (n = 240, p < 0.01)." The study compared 12 weeks of acupuncture (n = 120) to a waiting-list control (n = 120) who received no additional care. It did not include a sham acupuncture group. Why does the choice of control group matter so critically here?
Question 3
A pharmaceutical company tests 20 different molecular compounds against major depressive disorder in separate randomized controlled trials, each with adequate sample sizes. Nineteen trials show no significant benefit over placebo, but one compound shows a statistically significant improvement (p = 0.03, d = 0.35). The company publishes only the positive trial in a high-impact journal and files the drug for regulatory approval. What phenomenon does this represent, and why is it dangerous?
Question 4
A health news headline reads: "Researchers discover statistically significant link between daily coffee consumption and reduced colorectal cancer risk (p = 0.02, OR = 0.97, 95% CI: 0.94-0.99, n = 480,000)." Your statistically literate colleague says the finding is "significant but clinically meaningless." What do they mean?
Question 5
A paper published in a peer-reviewed nutrition journal claims that a common artificial sweetener causes gut microbiome disruption in humans based on a 4-week study of 36 participants. A parent in an online forum writes: "If experts reviewed and approved it for a top journal, that settles it. I am eliminating all artificial sweeteners from my family's diet immediately." Why is this reasoning incomplete?
Question 6
A psychology lab runs an experiment on whether background music improves creative problem-solving. They measure creativity using four different scales, test participants at three time points, and analyze results for the full sample and six demographic subgroups. After running all these analyses, they find one statistically significant result: women aged 18-25 scored higher on the Alternate Uses Task at the 30-minute mark (p = 0.04). They report this as their primary finding. What methodological problem does this represent?