TORTINI

For your delectation and delight, desultory dicta on the law of delicts.

Acetaminophen – Bradford Hill Confounded Part 06

October 1st, 2026

“The credulous man is father to the liar and the cheat… .”[1]

On the appeal from the MDL court’s exclusion of plaintiffs’ expert witnesses’ causation opinions, Judge Calebresi’s acknowledged both the complete uncertainty of any causal claim,[2] as well as the “significant debate in the relevant scientific communities.”[3] Even these acknowledgements were understatements. While citing Daubert, the court ignored the general acceptance prong of that opinion. The significant debate and acknowledged uncertainty of any causal claim were vivid demonstrations that there was no general acceptance of the causal conclusion. Although the court cited the amended Rule 702,[4] it ignored subsection (b), which requires the trail court to ask whether the “the testimony is based on sufficient facts or data.” Baccarelli answered the sufficiency question in the negative, in September 2025, but Judge Calebresi ignored the answer in front of him.

The appellate court chose to frame the issue on appeal as how closely may a district court examine an expert witness’s conclusions when he has deployed accepted methodologies. This framing of the issue obscured the actual issue in dispute, which was whether the challenged expert witness had used an accepted methodology in an unacceptable, unreliable, and invalid way. The district court, following the mandate of Rule 702, had inquired not only into the methodology of inferring causation (navigation guide review and Bradford Hill analysis), but the application of those methodologies.  The Second Circuit abandoned the critical analysis of witnesses’ actual application of methodologies in favor of a credulous approach.

The appellate court did not provide an exhaustive review of the district court’s evaluation of supposed errors of interpretation. Judge Calebresi completely ignored the application of the so-called navigation guide review, which allowed him to skip over the most obvious duplicity in the plaintiffs’ expert witness reports. The appellate court asserted that the MDL court had arrived at definitive interpretations of certain studies when “reasonable scientists” could have interpreted the studies differently.[5] Among the many aspects of study interpretation missed by the appellate court were the hypothesis-generating nature of some studies, the speculative interpretations offered in the absence of fully supportive data, and the inability of the extant studies to support a conclusion of causality. The appellate court was willing, if not eager, to indulge the plaintiffs’ expert witnesses’ rank speculation in their attempts to support their litigation claims. The trial court had not adjudicated the claims, as the appellate court charged, but simply identified the logical and factual lacunae that Dr. Baccarelli and other witnesses had tried to traverse with suggestions of “plausibility” and “possibility.”

A brief review of how the appellate court parsed some of the so-called Bradford Hill factors will illustrate the superficiality of the appellate review. On the subject of dose-response, one of Hill’s factors, the Circuit court faulted the MDL trial court for its treatment of Baccarelli’s evaluation of whether a dose-response relationship existed to support his analysis. The trial court correctly noted that dose-response was a key issue in the studies relied upon by Baccarelli, and that none of the studies recorded actual doses ingested by women in their pregnancies.[6] The Circuit court then proceeded to fault the trial court for not citing scientific authority for the need for precise dosage data from the study participants,[7] when many of the studies relied upon did not even have imprecise dosage data.

The Circuit court claimed, without evidence, that exact dosages would rarely be available in observational studies. This claim by the appellate court is clearly wrong for prescription medications, and even for over-the counter medications when dosage regimens are prescribed or when data are collected concurrently in a prospective study.

The studies cited by Baccarelli attempted to use proxy data to estimate dosage, and their authors acknowledged the significant limitations on their analyses. This body of evidence hardly constituted robust data to claim that the exposure-gradient consideration had been “satisfied.” In an effort to sanitize Baccarelli’s overclaiming, the Circuit court pointed to Baccarelli’s citation of Alemany 2021 as “recognizing that multiple prior studies found a dose-response relationship.”[8]

Alemany and colleagues never addressed dose-response in their meta-analysis study.[9] The Alemany paper acknowledges that it presented no dose-response data, but noted that “previous studies have shown dose-response effects” for symptoms, not diagnoses. The Alemany was a meta-analysis of six studies, but these authors cited only three studies with any dose-response evidence. The first paper cited by Alemany involved attention deficit and hyperactivity; it did not study autism at all.[10] The second paper cited was also a paper on symptoms; and so it could not possibly have addressed dose-response with autism itself.[11] This second paper stated unequivocally that the authors “were unable to evaluate the effects of dosage because of mothers’ difficulties in recalling the dose taken.”[12] To the extent that this second paper tried to analyze its symptom data by “persistent” or “sporadic” use, its results across multiple symptoms and both sexes were inconsistent, with no attempt to correct for multiple testing.[13]

The third study cited by Alemany, which was advanced by Baccarelli in support of his claim that the dose-response consideration was satisfied, also presented doubtful data on dose. The authors noted that their dose-response analyses were based upon “counting total weeks of use,” because “more than 80% of the interviewed women were unable to recall” exact dosages.[14] The clear implication was that even “total weeks” of use was likely inaccurate, but the Circuit court failed to explain why this paper’s estimates of total weeks use were credible estimates. Instead the Circuit faulted the MDL court for insisting that the dose-response data were inadequate when scientists had obviously attempted to construct dose-response analyses. What the Circuit missed is that scientists will sometimes use less than adequate data in preliminary or provisional analyses, such as for dose-response, but that is not the same as overclaiming that Bradford Hill’s dose-response consideration was satisfied.

Weak Associations

Baccarelli contended in his report that strength of association was an important consideration, which he had found to be satisfied in his analysis. The trial court, in turn, found Baccarelli’s claim to be relying upon a “strong” association to be invalid. Most of the studies identified by Baccarelli as supporting his opinion showed risk ratios below 2.0, and legal authority within the Circuit consistently labeled such associations as “weak.”[15]

The Circuit court engaged in considerable second guessing of the district court, but its opinion suggests that the appellate judges did not understand the difference between strength of association and the presence or absence of statistical significance. With respect to what constitutes a “strong” association, the Circuit credulously embraced Baccarelli’s self-serving explanation that there was “no general rule for how large an association needs to be” to satisfy this criterion.[16]

Of course, Sir Austin characterized strength as a consideration, not a criterion. The Circuit might well have let the matter go, but it improvidently chose to dig in an attempt to resurrect Baccarelli’s opinion. As part of this effort, the Circuit cited a dubious case from the Ninth Circuit, which embraced risk ratios less than two as strong, without any scientific support.[17]

The Circuit purported to discern that Baccarelli had shown that epidemiologic practice supported the satisfaction of the strength consideration with a risk ratio less than 2.0.[18] Here the court confused instances in which the scientific community concluded that causality existed with low risk ratios with the community’s agreement that the association was strong. Nothing of the sort was going on; rather, there are instances of scientists’ finding causal associations in the presence of weak, not strong, associations when the additional supporting evidence is robust, consistent, and free of threats to validity. Residential exposure to radon is generally accepted as a cause of lung cancer among nonsmokers, although the risk ratio for exposed persons is about 1.2. The magnitude of the association is weak, but other considerations dominate to lead to a conclusion of causality.

The Circuit also blessed Baccarelli’s pure speculation that the magnitude of association as shown in the available studies might have been greater if studies had not relied on maternal self reports. Baccarelli advanced this speculation in the absence of evidence, and in the presence of well-known recall biases that affect mothers whose children are afflicted with the condition under investigation.[19] 

The Circuit court’s exposition suggests further that the appellate judges do not even understand what an association is. Their opinion uncritically accepted Baccarelli’s assertion that one study[20] found a positive association in a subgroup, when the reported risk ratio was 1.03 (for first trimester exposure to acetaminophen and ASD).[21] The appellate opinion failed to present the confidence interval around the reported risk ratio of 1.03, but the cited paper provided that the 95 percent interval as spanning ratios of 0.82 to 1.29. This interval represents an approximate p-value of 0.80, which is clearly nowhere close to statistical significance.

The probability of obtaining an exact point estimate risk ratio of 1.00 in an observational study, when the true risk ratio is 1.00, is zero. Even if we were to specify a very narrow interval around 1.0, say 0.99 to 1.01, the probability of obtaining a result in that interval would be very close to zero. And the probability of finding a result outside that narrow interval would be close to 100 percent. The actual, reported point estimate of 1.03 is statistically indistinguishable from no association at all, and completely consistent with random error.

Professional judgment of what constitutes a “strong” association is not entirely subjective, and the Circuit court’s gullible acceptance of Baccarelli’s assertion was unwarranted. Back in 1980, Breslow and Day, two respected cancer researchers, noted in a publication of the International Agency for Research on Cancer, that “[r]elative risks of less than 2.0 may readily reflect some unperceived bias or confounding factor, those over 5.0 are unlikely to do so.”[22] The following year, in 1981, two other eminent epidemiologists, Sir Richard Doll, and Sir Richard Peto, expressed a similarly skeptical view about risk ratios less than two, in assessing the causality of associations:

“when relative risk lies between 1 and 2 … problems of interpretation may become acute, and it may be extremely difficult to disentangle the various contributions of biased information, confounding of two or more factors, and cause and effect.”[23]

In 1990, two epidemiologists confronted the issue of what constitutes a “strong” association in the context of studying lung cancer and exposure to environmental tobacco smoke (ETS). They stated the widely shared view that

“An association is generally considered weak if the odds ratio is under 3.0 and particularly when it is under 2.0, as is the case in the relationship of ETS and lung cancer. If the observed relative risk is small, it is important to determine whether the effect could be due to biased selection of subjects, confounding, biased reporting, or anomalies of particular subgroups.”[24]

A few years later, the journal Science published an influential essay on the limits of epidemiology. The author quoted Marcia Angell, former editor of the New England Journal of Medicine, as stating that “[a]s a general rule of thumb, we are looking for a relative risk of 3 or more [before accepting a paper for publication], particularly if it is biologically implausible or if it’s a brand new finding.”[25]

The usefulness of studies with risk ratios under two has been consistently questioned as weak, and likely to be driven by confounding or bias.[26] One textbook gave a more discerning evaluation of “weak,” in the differing contexts of cohort and case-control studies:

“Even after attempts to minimise selection and information biases and after control for known potential confounding factors, bias often remains. These biases can easily account for small associations. As a result, weak associations (which dominate in published studies) must be viewed with circumspection and humility. Weak associations, defined as relative risks between 0.5 and 2.0, in a cohort study can readily be accounted for by residual bias (Fig. 7.2). Because case-control studies are more susceptible to bias than are cohort studies, the bar must be set higher. ln case-control studies, weak associations can be viewed as odds ratios between 0.33 and 3.0 (Fig. 7.3). Results that full within these zones may be due to bias. Results that full outside these bounds in either direction may deserve attention.”[27]

Conclusion

Judge Calebresi and his co-panelists demonstrated remarkable credulity and poor scholarship in their attempt to resuscitate the opinion testimony of Dr. Baccarelli. Despite their acknowledgement that the causal issue at hand was controversial and quite uncertain, they focused on Baccarelli’s reliance on scientific studies, as though such reliance, without critical evaluation can stand under Rule 702. Perhaps even more discouraging, the judges of the Second Circuit were confronted with Baccarelli’s duplicity in proffering a causal opinion in litigation, while admitting that a causal conclusion was beyond the data at hand in his professional publication and in his communication to both the profession and the public, two years after submitting his expert witness report in the acetaminophen MDL. On September 11, 2026, the defendants filed a brief in support of their motion to reconsider the panel’s decision. The Second Circuit now has the opportunity to fix the problems that they created.


[1] William Kingdon Clifford, The Ethics of Belief, in Leslie Stephen & Frederick Pollock, eds., 2 LECTURES AND ESSAYS 177, 186 (1879).

[2] Slip op. at 35.

[3] Slip op. at 5.

[4] Slip op. at 24-25.

[5] Slip op. at 40.

[6] 707 F. Supp. 3d at 350.

[7] Slip op. at 35.

[8] Slip op. at 35.

[9] Silvia Alemany, et al., Prenatal and postnatal exposure to acetaminophen in relation to autism spectrum and attention-deficit and hyperactivity symptoms in childhood: Meta-analysis in six European population-based cohorts, 36 EUROPEAN J. EPIDEMIOL. 993, 1000 (2021).

[10] Zeyan Liew, et al., Acetaminophen use during pregnancy, behavioral problems, and hyperkinetic disorders, 168 JAMA PEDIATRICS 313 (2014).

[11] Avella-Garcia, et al., Acetaminophen use in pregnancy and neurodevelopment: attention function and autism spectrum symptoms, 45 INTERNAT’L J. EPIDEMIOL. 1987 (2016).

[12] Id. at 1994.

[13] See id. at 1992, Table 3.

[14] Zeyan Liew, et al., Maternal use of acetaminophen during pregnancy and risk of autism spectrum disorders in childhood: A Danish national birth cohort study, 9 AUTISM RES. 951, 956 (2016).

[15] See, e.g., In re Mirena IUS Levonorgestrel-Related Prods. Liab. Litig. (No. II), 341 F. Supp. 3d 213, 243 (S.D.N.Y. 2018), aff’d, 982 F.3d 113 (2d Cir. 2020) (addressing risk ratios of 3.90 and 7.69); Daniels-Feasel v. Forest Pharms, Inc., No. 17-cv-4188, 2021 WL 4037820, at *8 (S.D.N.Y. Sept. 3, 2021), aff’d, No. 22-146, 2023 WL 4837521 (2d Cir. 2023) (risk ratio of 2.2).

[16] Slip op. at 38 (citing App’x 1901).

[17] See id. (citing Hardeman v. Monsanto Co., 997 F.3d 941, 966 (9th Cir. 2021), for the proposition that “a hardline increase in a risk statistic, or even an adjusted odds ratio above 2.0, is [not] necessary for finding a strong association”), abrogated on other grounds sub nom. Monsanto Co. v. Durnell, U.S. Supreme Court, No. 24-1068, 2026 WL 1825691 (Jun. 25, 2026).

[18] Slip op. at 38.

[19] Slip op. at 38.

[20] Zeyan Liew, Beate Ritz, Jasveer Virk & Jørn Olsen, Maternal Use of Acetaminophen during Pregnancy and Risk of Autism Spectrum Disorders in Childhood: A Danish National Birth Cohort Study, 9 AUTISM RESEARCH 951, 955 (2016) (Table 2, Hazard Ratios (HR) for Autistic Spectrum Disorders and Infantile Autism in Children According to Maternal Acetaminophen Use during Pregnancy, first trimester 1.03 (95% C.I., 0.82–1.29).

[21] Slip op. at 48-49.

[22] Norman E. Breslow & Nicholas E. Day, STATISTICAL METHODS IN CANCER RESEARCH, VOL. I – THE ANALYSIS OF CASE-CONTROL STUDIES at 36 (IARC Sci. Publ. No. 32, 1980).

[23] Richard Doll & Richard Peto, THE CAUSES OF CANCER 1219 (1981).

[24] Ernst L. Wynder & Geoffrey C. Kabat, Environmental Tobacco Smoke and Lung Cancer: A Critical Assessment, in H. Kasuga, ed., INDOOR AIR QUALITY (1990)

[25] Gary Taubes, Epidemiology Faces Its Limits, 269 SCIENCE164, 168 (Jul. 14, 1995). 

[26] R. Bonita, R. Beaglehole & T. Kjellström, BASIC EPIDEMIOLOGY 93 (W.H.O. 2d ed. 2006) (“A strong association between possible cause and effect, as measured by the size of the risk ratio (relative risk), is more likely to be causal than is a weak association, which could be influenced by confounding or bias. Relative risks greater than 2 can be considered strong.”); Brian L. Strom, Basic Principles of Clinical Epidemiology Relevant to Pharmacoepidemiologic Studies, chap. 3, in Brian L. Strom, Stephen E. Kimmel & Sean Hennessy, eds., PHARMACOEPIDEMIOLOGY 48 (6th ed. 2020) (“Conventionally, epidemiologists consider an association with a relative risk of less than 2.0 a weak association.”); David A. Freedman & Philip B. Stark, The Swine Flu Vaccine and Guillain-Barré Syndrome: A Case Study in Relative Risk and Specific Causation, 64 LAW & CONTEMP. PROBS. 49, 61 (2001) (“If the relative risk is near 2.0, problems of bias and confounding in the underlying epidemiologic studies may be serious, perhaps intractable.”).

[27] Kenneth F. Schulz & David A. Grimes, ESSENTIAL CONCEPTS in CLINICAL RESEARCH: RANDOMISED CONTROLLED TRIALS AND OBSERVATIONAL EPIDEMIOLOGY at 75 (2d ed. 2019) (internal citations omitted). Previously these authors had stated even more conservative thresholds. See David A. Grimes & Kenneth F. Schulz, False alarms and pseudo-epidemics: the limitations of observational epidemiology, 120 OBSTET. & GYNECOL. 920 (2012) (“Most reported associations in observational clinical research are false, and the minority of associations that are true are often exaggerated. This credibility problem has many causes, including the failure of authors, reviewers, and editors to recognize the inherent limitations of these studies. This issue is especially problematic for weak associations, variably defined as relative risks (RRs) or odds ratios (ORs) less than 4.”)