TORTINI

For your delectation and delight, desultory dicta on the law of delicts.

Reconstructing Professor Sanders on Expert Witness Gatekeeping

May 5th, 2013

Last week, I addressed two papers from a symposium organized by the litigation industry to applaud the First Circuit’s decision in Milward v. Acuity Products Group, Inc., 639 F.3d 11 (1st Cir. 2011), cert. denied, 132 S.Ct. 1002 (2012).  Professor Joseph Sanders also contributed to the symposium, in a paper that is a bit more measured, scholarly, and disinterested than the other papers in the group.  Joseph Sanders, “Milward v. Acuity Specialty Products Group: Constructing and Deconstructing Sciences and Law in Judicial Opinion,3 Wake Forest J. L & Policy 141 (2013).  PDF  Still, the industry sponsor, the so-called Center for Progressive Reform, has reasons to be satisfied with the result.

Sanders argues that the Milward opinion is important because it highlights what he characterizes as a “rhetorical conflict that has been ongoing, often below the surface, since the United States Supreme Court’s 1993 opinion in Daubert v. Merrell Dow Pharmaceuticals, Inc.”  Id. at 142.  The argument is overly kind to the judiciary.  There has not been so much as a rhetorical conflict as a reactionary revolt against evidence-based decision making in the federal courts.  Milward simply represents the highwater mark of this revolt against law and science.  See, e.g., David Bernstein on the Daubert Counterrevolution (April 19, 2013).

Sanders invokes the ghost of Derrida and his black brush of deconstruction to suggest that the Daubert process is nothing more than the unraveling of the scientific enterprise, with the goal of showing that it is arbitrary and subjective.  Sanders at 143-44.  According to Sanders, radical deconstruction pushes towards a leveling of “distinctions between fact and faction … more akin to poetry and music than to evidence and argument.” Id. at 145 (citing Stephan Fuchs & Steven Ward, “What Is Deconstruction and Where and When Does It Take Place? Making Facts in Science, Building Cases in Law,” 59 Am. Soc. Rev. 481, 482-83 (1994)).

Lawyers sometimes realize the cost of radical deconstruction is a nihilism that undermines their own credibility and their ability to claim or defend factual assertions.  Sometimes, of course, lawyers ignore these considerations and talk out of both sides of their mouths. In re Welding Fume Prods. Liab. Litig., No. 1:03-CV-17000, MDL 1535, 2006 WL 4507859, *33 (N.D. Ohio 2006) (“According to plaintiffs, the rate of PD [Parkinson’s disease] mortality is so poor a proxy for measuring the rate of overall PD incidence, that the Coggon study proves nothing. In the next breath, however, plaintiffs set forth an unpublished statistical analysis (by Dr. Wells) of PD mortality data collected by the National Center for Health Statistics, arguing it proves that welders, as a group, suffer earlier onset of PD than the general population.77 Of course, the devil is in the details, discussion of which is beyond the scope of this opinion (and perhaps beyond the scope of understanding of the average juror),78 but this example shows how hard it is to tease out whether the limitations of a given study make it unreliable under Daubert.”).

Sanders gives a nod to Sheila Jasanoff, whom he quotes with apparent approval:

“[t]he adversarial structure of litigation is particularly conducive to the deconstruction of scientific facts, since it provides both the incentive (winning the lawsuit) and the formal means (cross-examination) for challenging the contingencies in their opponents’ scientific arguments.”

Sanders at 147 (quoting Shelia Jasanoff, “What Judges Should Know About the Sociology of Science,” 32 Jurimetrics J. 345, 348 (1992)).

With his acknowledgment that adversarial litigation of scientific issues involves  “deconstruction,” or just good, old-fashioned rhetorical excess, Sanders points to the Daubert trilogy as the Supreme Court’s measured response to the problem.

Sanders is, however, not entirely happy about the judiciary’s attempt to rein in the rhetorical excesses of adversarial litigation of scientific issues. Daubert barely scratched the surface of the scientific validity and reliability issues in the Bendectin record, but Sanders asserts that Chief Justice Rehnquist went too far in looking under the hood of the lemon that Joiner’s expert witnesses were trying to sell:

“perhaps Chief Justice Rehnquist erred in the other direction in Joiner when he systematically reviewed the animal studies and four separate epidemiological studies cited by the plaintiff as supporting the position that exposure to PCBs either caused or ‘promoted’66 the plaintiff’s lung cancer.67

Sanders at 154.  Horrors!  A systematic review!! Perish the thought.

Sanders seems to fault the Chief Justice’s approach of picking off “one-by-one” the studies relied upon by Joiner’s expert witnesses as a deconstructive exercise.  Id. at 154 – 55.  As Sanders notes, the Court’s opinion delved into the internal and external validity of the four cited epidemiologic studies to recognize that the:

“studies did not support the plaintiff’s position because the authors of the study said there was no evidence of a relationship, the relationship was not statistically significant, the substance to which the subjects were exposed was not the same as that to which Mr. Joiner was exposed, and the subjects were simultaneously exposed to other carcinogens.70

Id. at 155.  To be fair, not all of these were dispositive considerations, but they represent a summary of a district court’s extensive consideration of the scientific record.

Channeling Susan Haack, Sanders argues that the one-by-one approach (which Professor Green pejoratively called a “Balkanized” approach, and Sanders calls “atomistic”) ignores that a wall is made up of constituent bricks.  Sanders might have done better to study a more accomplished philosopher-scientist:

“[O]n fait la science avec des faits comme une maison avec des pierres; mais une accumulation de faits n’est pas plus une science qu’un tas de pierres n’est une maison.”

Jules Henri Poincaré, La Science et l’Hypothèse (1905) (chapter 9, Les Hypothèses en Physique)(“Science is built up with facts, as a house is with stones. But a collection of facts is no more a science than a heap of stones is a house.”).  Poincaré’s metaphor is more powerful and descriptive than Sander’s because it acknowledges that interlocking pieces of evidence may cohere as a building, or they may be no more than a pile of rubble.  Deeper analysis is required. Poorly constructed walls revert to the pile of bricks from which they came.  Furthermore, the mason must look at the individual bricks to see whether they are cracked, crumbling, or crooked before building a wall that must endure. We want a wall that will endure at least long enough to stand on, or put a roof on. Much more is required than simply invoking “mosaics,” “walls from bricks,” or “crossword puzzles” to transmute a pile of studies into a “warranted” claim to knowledge. Litigants, either plaintiff or defendant, should not be allowed to pick out isolated findings in a variety of studies, and throw them together as if that were science.  This is precisely the rhetorical excess against which Rule 702, with its requirement of “knowledge,” should protect judges, juries, and litigants.

Indeed, as Sanders eventually concedes, the Joiner Court noted the appropriateness of considering the four relied-upon epidemiologic studies, individually or collectively:

“We further hold that, because it was within the District Court’s discretion to conclude that the studies upon which the experts relied were not sufficient, whether individually or in combination, to support their conclusions that Joiner’s exposure to PCB’s contributed to his cancer, the District Court did not abuse its discretion in excluding their testimony.71

General Elec. Co. v. Joiner, 522 U.S. 136, 146-47 (1997).

Professor Sanders might have raised a justiciable argument against the gatekeeping process in Joiner if he had shown, or if he had adverted to other analyses, that the four relied-upon studies collectively meshed to overcome each other’s clear inadequacies.  The silence of Sanders, and other critics, on this score is telling.  Was Rabbi Teitelbaum (one of the Joiners’ expert witnesses) simply insightful or prescient in reading the early returns on PCBs and lung cancer?  What has happened subsequently?  Has the IARC embraced PCBs as a known cause of lung cancer?  Have well-conducted studies and meta-analyses vindicated Teitelbaum’s claims, or have they further confirmed that the gatekeeping in Joiner successfully excluded witnesses who were advancing pathologically weak, specious claims, by pushing and squeezing data until they fit into a pre-determined causal conclusions.

The complaints about judicial “deconstruction” are unfair and empty without looking at these details.  It behooves evidence scholars who want to write in this area to roll up their sleeves and look at the evidence that was in front of the courts, and to learn something about science.  Of course, scientists did not stop exploring the PCB/lung cancer hypothesis after Joiner was decided.  See, e.g, Avima Ruder, Misty Hein, Nancy Nilsen, et al., “Mortality among Workers Exposed to Polychlorinated Biphenyls (PCBs) in an Electrical Capacitor Manufacturing Plant in Indiana: An Update,” 114 Envt’l Health Perspect. 18 (2006) (study by National Institute for Occupational Safety and Health, showing reduced rates of respiratory cancer among PCB-exposed workers, with age-standardized risk ratio of 0.85, and a 95% confidence interval, 0.6 to 1.1)

Stevens’ partial dissent in Joiner of course invested deeply in the mosaic theory, which we now know was the brainchild of plaintiffs’ counsel, Barry Nace. Michael D. Green, “Pessimism about Milward,” 3 Wake Forest J. L & Policy 41, 63 (2013) (reporting that Barry Nace acknowledged having “fed” this rhetorical device to expert witness Alan Done to support arguments for manufacturing certainty in the face of an emerging body of exonerative evidence).  Justice Stevens also cited the EPA as employing “weight of the evidence,” which simply makes the point that WOE is a precautionary approach to scientific evidence, not one for serious causal determinations.  Justice Stevens, and Professor Sanders, might have done better to have looked at what the FDA requires for health claims.  See, e.g., FDA, Guidance for industry evidence-based review system for the scientific evaluation of health claims (2009) (articulating an evidence-based approach). Justice Stevens’ argument fundamentally misconstrues the scientific enterprise of determining causation of health outcomes by reducing it to a precautionary enterprise of regulating possible harms.  Professor Sanders is unclear whether he is restating Justice Stevens’ view, or his own, when he writes:

“Chief Justice Rehnquist was wrong to show the flaws of individual bricks, because it is the wall as a whole that makes up the plaintiff’s case.74

Sanders at 157 (citing Stevens’ opinion in Joiner, 522 U.S. 136, 153n.5  (1997).  If this is Professor Sanders’ view, it is profoundly wrong.  Looking at the individual bricks is necessary to determine whether it can support the plaintiff’s case.  Of course, to the extent it was Justice Stevens’ view, it was a dissent, not a holding, and it was superseded by a statute when Rule 702 was revised in 2000.

 

Milward’s Singular Embrace of Comment C

May 4th, 2013

Professor Michael D. Green is one of the Reporters for the American Law Institute’s Restatement (Third) of Torts: Liability for Physical and Emotional Harm.   Green has been an important interlocutor in the on-going debate and discussion over standards for expert witness opinions.  Although many of his opinions are questionable, his writing is clear, and his positions, transparent.  The seduction of Professor Green and the Wake Forest School of Law by one of the litigation-industry’s organizations, the Center for Progressive Reform, is unfortunate, but the resulting symposium gave Professor Green an opportunity to speak and write about the justly controversial comment c.   Restatement (Third) of Torts: Liability for Physical and Emotional Harm § 28, cmt. c (2010).

Mock Pessimism Over Milward

Professor Green professes to be pessimistic about the Milward decision, but his only real ground for pessimism is that Milward will not be followed.  Michael D. Green, “Pessimism about Milward,” 3 Wake Forest J. L & Policy 41 (2013).  Green describes the First Circuit’s decision in Milward as “fresh,” “virtually unique and sophisticated,” and “satisfying.” Id. at 41, 43, and 50.  Green describes his own reaction to the decision in terms approaching ecstasy:  “delighted,” “favorable,” and “elation.”  Id. at 42, 42, and 43.

Green interprets Milward to embrace four comment c propositions:

  1. “Recognizing that judgment and interpretation are required in assessments of causation.52
  2. Endorsing explicitly and taking seriously weight of the evidence methodology,53 against the great majority of federal courts that had, since Joiner, employed a Balkanized approach to assessing different pieces of evidence bearing on causation.54
  3. Appreciating that because no algorithm exists to constrain the inferential process, scientists may reasonably reach contrary conclusions.55
  4. Not only stating, but taking seriously, the proposition that epidemiology demonstrating the connection between plaintiff’s disease and defendant’s harm is not required for an expert to testify on causation.56 Many courts had stated that idea, but very few had found non-epidemiologic evidence that satisfied them.57

Id. at 50-51.

Green’s points suggest that comment c was designed to reinject a radical subjectivism into scientific judgments allowed to pass for expert witness opinions in American courts.  None of the points is persuasive.  Point (1) is vacuous.  Saying that judgment is necessary does not imply that anything goes or that we will permit the expert witness to be the judge of whether his opinion rises to the level of scientific knowledge.  The required judgment involves an exacting attention to the role of random error, bias, or confounding in producing an apparent association, as well as to the validity of the data, methods, and analyses used to interpret observational or experimental studies.  The required judgment involves an appreciation that not all studies are equally weighty, or equally worthy of consideration for use in reaching causal knowledge.  Some inferences are fatally weak or wrong; some analyses or re-analyses of data are incorrect.  Not all judgments can be blessed by anointing some of them “subjective.”

Point (2) illustrates how far the Restatement process has wondered into the radical terrain of abandoning gatekeeping altogether.  The approach that Green pejoratively calls “Balkanized” is a careful look at what expert witnesses have relied upon to assess whether their conclusions or claims follow from their relied upon sources.  This is the approach used by International Agency for Research on Cancer (IARC) working groups, whose method Green seems to endorse.  Id. at 59.  IARC working groups discuss and debate their inclusionary and exclusionary criteria for studies to be considered, and the validity of each study and its analyses, before they get to an assessment of the entire evidentiary display.  (And several of the IARC working groups have been by no means free of the conscious bias and advocacy that Green sees in party-selected expert witnesses.)  Elsewhere, Green refers to the approach of most federal courts as “corpuscular.”  Id. at 51. Clearly, expert witnesses want to say things in court that do not, so to speak, add up, but Green appears to want to give them all a pass.

Point (3) is, at best, a half truth.  Is Green claiming that reasonable scientists always disagree?  His statement of the point suggests epistemiologic nihilism. Although there are no clear algorithms, the field of science is littered with abandoned and unsuccessful theories from which we can learn when to be skeptical or dismissive of claims and conclusions.  Certainly there are times when reasonable experts will disagree, but there are also times when experts on one side or the other, or both, are overinterpreting or misinterpreting the available evidence.  The judicial system has the option and the obligation to withhold judgments when faced with sparse or inconsistent data.  In many instances, litigation arises because the scientific issues are controversial and unsettled, and the only reasonable position is to look for more evidence, or to look more carefully at the extant evidence.

Point (4) is similarly overblown and misguided.  Green states his point as though epidemiology will never be required.  Here Green’s sympathies betray any sense of fidelity to law or science.  Of course, there may be instances in which epidemiologic evidence will not be necessary, but it is also clear that sometimes only epidemiologic methods can establish the causal claim with any meaningful degree of epistemic warrant.

ANECDOTES TO LIVE BY

Anthony Robbins’ Howler

Professor Green delightfully shares two important anecdotes.  Both are revealing of the process that led up to comment c, and to Milward.

The first anecdote involves the 2002 meeting of the American Law Institute.  Apparently someone thought to invite Dr. Anthony Robbins as a guest. (Green does not tell us who performed this subversive act.)  Robbins is a member of SKAPP, the organization started with plaintiffs’ counsel’s slush fund money diverted from MDL 926, the silicone-gel breast implant litigation.

Robbins rose at the meeting to chastise the ALI for not knowing what it was talking about:

“clear, in my opinion, misstatements of . . . science” or reflected a misunderstanding of scientific principles that “leaves everyone in doubt as to whether you know what you are talking about . . . .”

Id. at 44 (quoting from 79th Annual Meeting, 2002 A.L.I. PROC. at 294).  Pretty harsh, except that Professor Green proceeds to show that it was Robbins who had no idea of what he was talking about.

Robbins asserted that the requirement of a relative risk of greater than two was scientifically incorrect. From Green’s telling of the story, it is difficult to understand whether Robbins was complaining about the use of relative risks (greater than two) for inferring general or specific causation.  If the former, there is some truth to his point, but Robbins would be wrong as to the latter.  Many scientists have opined that relative risks provide information about attributable fractions, which in turn permit inferences about individual cases.  See, e.g., Troyen A. Brennan, “Can Epidemiologists Give Us Some Specific Advice?” 1 Courts, Health Science & the Law 397, 398 (1991) (“This indeterminancy complicates any case in which epidemiological evidence forms the basis for causation, especially when attributable fractions are lower than 50%.  In such cases, it is more probable than not that the individual has her illness as a result of unknown causes, rather than as a result of exposure to hazardous substance.”).  Others have criticized the inference, but usually on the basis that the inference requires that the risk be stochastically distributed in the population under consideration, and we often do not know whether this assumption is true.  Of course, the alternative is that we must stand mute in the face of even very large relative risks and established general causation.  See, e.g., McTear v. Imperial Tobacco Ltd., [2005] CSOH 69, at ¶ 6.180 (Nimmo Smith, L.J.) (“epidemiological evidence cannot be used to make statements about individual causation… . Epidemiology cannot provide information on the likelihood that an exposure produced an individual’s condition.  The population attributable risk is a measure for populations only and does not imply a likelihood of disease occurrence within an individual, contingent upon that individual’s exposure.”).

Robbins second point was truly a howler, one that suggests his animus against gatekeeping may grow out of a concern that he would never pass a basic test of statistical competency.  According to Green, Robbins claimed that “increasing the number of subjects in an epidemiology study can identify small effects with ‘an almost indisputable causal role’.” Id. at 45 (quoting Robbins).  Ironically, lawyer and law professor Green was left to take Robbins to school, to educate him on the differences between sampling error, bias, and confounding.  Green does not get the story completely right because he draws an artificial line between observational epidemiology and experimental clinical trials, and incorrectly implies that bias and confounding are problems only in observational studies.  Id. at 45 n. 24.  Although randomization is undertaken in clinical trials to control for bias and confounding, it is not true that this strategy always works or always works completely.  Still, here we have a lawyer delivering the comeuppance to the scolding scientist.  Sometimes scientists really have no good basis to support their claims, and it is the responsibility of the courts to say so.  Green’s handling of Robbins’ errant views is actually a wonderful demonstration of gatekeeping in action.  What is lovely about it is that the claims and their rebuttal were documented and reported, rather than being swept away in the fog of a jury verdict.

Professor Green’s account of Robbins’ foolery should be troubling because, despite Robbins’ manifest errors, and his more covert biases, we learn that Robbins’ remarks had “a profound impact” on the ALI’s deliberations. Courts that are tempted by the facile answers of comment c should find this impact profoundly disturbing.

Alan Done’s Weight of the Evidence (WOE) or Mosaic Methodology

Professor Green relays an anecdote that bears repeating, many times.  In the Bendectin litigation, plaintiffs’ expert witness, Alan Done testified that Bendectin caused birth defects in children of mothers who ingested the anti-nausea medication during pregnancy.  Done had a relatively easy time spinning his speculative web in the first Bendectin trial because there was only one epidemiologic study, which qualitatively was not very good.  In his second outing, Done was confronted by the defense with an emerging body of exonerative epidemiologic research. In response, he deployed his “mosaic theory” of evidence, of different pieces or lines of evidence that singularly do not show much, but together paint a conclusive picture of the causal pattern. Id. at 61 (describing Done’s use of structure-activity, in vitro animal studies, in vivo animal studies, and his own [idiosyncratic] interpretation of the epidemiologic studies).  Done called his pattern a “mosaic,” which Green correctly sees is none other than “weight of the evidence.”  Id. at 62.

After this second trial was won with the jury, but lost on post-trial motions, plaintiffs’ counsel, Barry Nace, pressed the mosaic theory as a legitimate scientific strategy to demonstrate causation, and the appellate court accepted the strategem:

“Like the pieces of a mosaic, the individual studies showed little or nothing when viewed separately from one another, but they combined to produce a whole that was greater than the sum of its parts: a foundation for Dr. Done’s opinion that Bendectin caused appellant’s birth defects. The evidence also established that Dr. Done’s methodology was generally accepted in the field of teratology, and his qualifications as an expert have not been challenged.103

Id. at 61(citing Oxendine v. Merrell Dow Pharm., Inc., 506 A.2d 1100, 1110 (D.C. 1986).  Green then drops his bombshell:  the philosopher of science who developed the “mosaic theory” (WOE) was the plaintiffs’ lawyer, Barry Nace. According to Green, Nace declared the mosaic idea “Damn brilliant, and I was the one who thought of it and fed it to Alan [Done].”  Id. at 63.

Green attempts to reassure himself that Milward does not mean that Done could use his WOE approach to testify today that Bendectin causes human birth defects.  Id. at 63.  Alas, he provides no meaningful solution to protect against future bogus cases.  Green fails to come to grips with the obvious truth that Done was wrong ab initio.  He was wrong before he was exposed for his perjurious testimony.  See id. at 62 n. 107, and he was wrong before there was a “solid body” of exonerative epidemiology.  His method never had the epistemic warrant he claimed for it, and the only thing that changed over time was a greater recognition of his character for veracity, and the emergence of evidence that collectively supported the null hypothesis of no association.  The defense, however, never had the burden to show that Done’s methodology was unreliable or invalid, and we should look to the more discerning scientists who saw through the smokescreen from the beginning.

I Don’t See Any Method At All

May 2nd, 2013

Kurtz: Did they say why, Willard, why they want to terminate my command?
Willard: I was sent on a classified mission, sir.
Kurtz: It’s no longer classified, is it? Did they tell you?
Willard: They told me that you had gone totally insane, and that your methods were unsound.
Kurtz: Are my methods unsound?
Willard: I don’t see any method at all, sir.
Kurtz: I expected someone like you. What did you expect? Are you an assassin?
Willard: I’m a soldier.
Kurtz: You’re neither. You’re an errand boy, sent by grocery clerks, to collect a bill.

* * * * * * * * * * * * * * * *

The Royal Society, the National Academies of Science, the Nobel Laureates have nothing on the organized plaintiffs’ bar.  Consider the genius and the accomplishments of these men and women.  They have discovered and built a perpetual motion machine — the asbestos litigation.  They have learned how to violate the law of non-contradiction with impunity (e.g., industry is evil, and (litigation) industry is good).  In the realm of the sciences, especially as applied in the courtroom, they have demonstrated the falsity of one of core beliefs: ex nihilo nihil fit.  We have a lot to learn from the plaintiffs’ bar.

WOE to Corporate America

Steve Baughman Jensen is a plaintiffs’ lawyer and he justifiably gloats over his success as lead counsel in Milward v. Acuity Products Group, Inc., 639 F.3d 11 (1st Cir. 2011), cert. denied, 132 S.Ct. 1002 (2012).  In a recent article for the litigation industry’s scholarly journal, Jensen touts Milward as Ariadne’s thread, which will take plaintiffs out of the mazes and traps set for them by the benighted law of expert witnesses.  Steve Baughman Jensen, “Reframing the Daubert Issue in Toxic Tort Cases,” 49 Trial 46 (Feb. 2013).  Jensen alleged that his client worked with solvents that contained varying amounts of  benzene, which caused him to develop Acute Promyelocytic Leukemia (APL), a subtype of Acute Myeloid Leukemia (AML).  The district court excluded plaintiffs’ expert witnesses’ causation opinions; the First Circuit reversed.  Jensen crows about his accomplished feat.

Weight of the Evidence (WOE) — Let Them Drink Ripple

Jensen, with help from philosopher of popular science Carl Cranor and toxicologist Martyn Smith persuaded the appellate court that a “weight of the evidence” (WOE) analysis necessarily involves scientific judgment. (Millward, 639 F.3d at 18), and that this “use of judgment in the weight of the evidence methodology is similar to that in differential diagnosis, which we have repeatedly found to be a reliable method of medical diagnosis.” Id. (internal citations omitted).

Is this what judicial gatekeeping of scientific expert opinion has come to?  Phrenology, homeopathy, aroma therapy, and reflexology involve medical judgment, of sorts, and so they too are now reliable methodologies.  Ripple makes red wine, and so does Chateau Margaux.  Chateau Margaux is based upon judgment in oenology, and so is Ripple.  That only one of these products will stand the test of time is irrelevant; both are the product of oenological judgment.  It’s all a question of the weight you would assign the differing qualities of Ripple and a premier cru bordeaux.

Jensen never defines WOE; the closest he comes to describing WOE is to tell us that it essentially involves a delegation to expert witnesses to validate their own subjective weighing of the evidence. As in the King of Hearts, Jensen rejoices that the inmates are now running the asylum.

Too Much of Nothing

Jensen complains about a “divide and conquer” strategy by which defendants take individual studies, one at a time, pronounce them inadequate to support a judgment of causality, and then claim that the aggregate evidence fails to support causality as well.  Surely sometimes that approach is misguided; yet sometimes the evidentiary display collectively represents “too much of nothing.”  In some litigations, there are hundreds of studies, which despite their numbers, still fail to support causation.  In General Electric v. Joiner, the Supreme Court discerned that the studies relied upon were largely irrelevant or inconclusive, and that taken alone or together, the cited studies failed to support plaintiffs’ claim of causality.  In the silicone-gel breast implant litigation, the plaintiffs’ steering committee submitted banker boxes of studies and argument to the court’s appointed expert witnesses, in an attempt to manufacture causation.  The committee, however, took its time and saw that the evidence taken individually or collectively did not amount to a scientific peppercorn.

Let Ignorance Rule

One of Jensen’s clever attempts to beguile the judiciary involves the transmutation of scientific inference into personal credibility.  “Second-guessing an expert’s application of scientific judgment necessarily requires assessing that expert’s credibility, which is the jury’s role.” Jensen, 49 Trial at 49.  Jensen attempts to reduce the “battle of experts” to a credibility contest and thus outside the purview of judicial gatekeeping.  His argument conflates credibility with methodology and its application. Because the expert witness will predictably opine that he applied the methodology faithfully, Jensen asserts that the court is barred from examining the correctness of the expert witness’s self-validation.

But scientific inference is scientific because it does not depend upon the person drawing it.  The inference may be weak, strong, erroneous, valid, or invalid.  How we characterize the inference will turn on the data and their analysis, not on the witness’s say so.

Jensen cites comment c, to Section 28 of Restatement (Third) of Torts, as supporting his reactionary arguments for abandoning judicial gatekeeping of expert witness opinion testimony.  “Juries, not judges, should determine the validity of two competing expert opinions, both of which typically fall within the realm of reasonable science.” Jensen, 49 Trial at 51 (emphasis added).  The law, however, requires trial courts to assess the validity vel non of would-be testifying expert witnesses:

“[A] trial judge, acting as ‘gatekeeper’, must ‘ensure that any and all scientific testimony or evidence admitted is not only relevant, but reliable’.  This requirement will sometimes ask judges to make subtle and sophisticated determinations about scientific methodology and its relation to the conclusions an expert witness seeks to offer— particularly when a case arises in an area where the science itself is tentative or uncertain, or where testimony about general risk levels in human beings or animals is offered to prove individual causation.”

General Elec. Co. v. Joiner, 522 U.S. 136, 147–49 (1997) (Breyer, J., concurring) (citations omitted).  Not only is Jensen’s argument contrary to the law, the argument is based upon a cynical understanding that juries will usually have little time, experience, or aptitude for assessing  validity issues, and that delegating validity issues to juries ensures that the legal system will not be able to root out pathologically weak evidence and inferences. The resolution of validity issues will be hidden behind the secretive walls of the jury room, rather than in the open sight of reasoned, published opinions, subject to public and scholarly commentary.   See, e.g., In re Welding Fume Prods. Liab. Litig., No. 1:03-CV-17000, MDL 1535, 2006 WL 4507859, *33 n.78 (N.D. Ohio 2006).  (“even the smartest and most attentive juror will be challenged by the parties’ assertions of observation bias, selection bias, information bias, sampling error, confounding, low statistical power, insufficient odds ratio, excessive confidence intervals, miscalculation, design flaws, and other alleged shortcomings of all of the epidemiological studies.”)

Martyn Smith

Jensen extols the achievements of Dr. Martyn Smith, his expert witness who was excluded by the trial court in Milward.  A disinterested reader might mistakenly think that Smith was among the leading benzene researchers in the world, but a little Googling would turn up that Milward was not his first litigation citation.  Smith has been pulled over for outrunning his expert-witness headlights in several other litigations, including:

  • Jacoby v. Rite Aid, Phila. Cty. Ct. Common Pleas (Order of April 27, 2012; Opinion of April 12, 2012) (excluding Smith as an expert witness on the toxicity of Fix-o-Dent)
  • In re Baycol Prods. Litig., 495 F. Supp.2d 977 (D. Minn. 2007)
  • In re Rezulin Prods. Liab. Litig., MDL 1348, 441 F.Supp.2d 567 (S.D.N.Y. 2006)(“silent injury”)

None of these other cases involved benzene, but they all involved speculative opinions.

The Milward Symposium

Jensen took another victory lap at the Milward Symposium Organized By Plaintiffs’ Counsel and Witnesses.  The presentations from this symposium have now appeared in print:  Wake Forest Publishes the Litigation Industry’s Views on MilwardSee Steve Baughman Jensen, “Sometimes Doubt Doesn’t Sell:  A Plaintiffs’ Lawyer’s Perspective on Milward v. Acuity Products,” 3 Wake Forest J. L. & Policy 177 (2013).  Jensen’s contribution was mostly a shrill ad hominem against corporations, as well as their lawyers and scientists who complicitly support an alleged campaign to manufacture doubt.

Perhaps someday the law journal’s faculty advisors and editors will feel some embarrassment over the lack of balance and scholarship in Jensen’s contribution to the symposium.  Corporations are bad; get it?  They manufacture doubt about the litigation industry’s enterprise.  Don’t pay attention to massive litigation fraud, such as faux silicosis, faux asbestosis, faux fen-phen heart disease, faux product identification, etc.  See Larry Husten, “79-Year-Old Cardiologist Sentenced To 6 Years In Prison For Fen-Phen Fraud,” Forbes (Mar. 27, 2013).   Forget that ATLA/AAJ is one of the most powerful rent-seeking lobbies in the United States.  Litigants have a constitutional right to extrapolate as they please.  If a substance causes one disease at a very high dose, then it causes every ailment known to mankind at moderate or low doses.  Specific disease entails general disease, etc.  What you balk?  You must be a doubt mongerer.

Jensen assures us that many scientists support and agree with Martyn Smith, both in his methodology and in his conclusions.  Jensen’s articles are sketchy on details, and of course, the devil is in the details.  See Amended Amicus Curiae Brief of the Council for Education and Research on Toxins et al., In Support of Appellants, in Milward.  This Council seems to fly under the internet radar, but I suspect that its membership and that of the Center for Progressive Reform overlaps somewhat.

Jensen’s article is just one of several published in the Wake Forest Journal of Law & Policy.  Let’s hope the remaining articles have more substance to them.

Further Musings on U.S. v. Harkonen

April 15th, 2013

Epistemic Crimes

In U.S. v. Harkonen, the government prosecuted a physician, company CEO, for issuing a press release that stated a clinical trial “demonstrated” benefit when the government believed that the clinical trial was inconclusive.  No doubt the government was intent upon punishing what it thought was off-label promotion in the same press release, but the jury acquitted on the charge of misbranding, and convicted on the wire fraud count.  The trial court denied post-trial motions, and recently, the United States Court of Appeals, for the Ninth Circuit, affirmed, in an unpublished per curiam opinion.  United States v. Harkonen, No. 11-10209, No. 11-10242, 2013 WL 782354, 2013 U.S. App. LEXIS 4472 (9th Cir. March 4, 2013).

A Gedanken Experiment

An expert witness writes a report that X, a drug therapy, causes Y, a benefit in survival, for a disease, Z.

The expert witness sent his report by email, and regular mail, to counsel, who then served it upon his adversary.  The report set out some of the support for the opinion, as follows.

The expert witness relied upon a randomized clinical trial, conducted with one primary and nine secondary endpoints.  The multiple endpoints were chosen because of uncertainty of how the anticipated benefit would manifest.  Mortality (survival), although obviously a very important endpoint, was not made primary endpoint because the scientists who conducted the trial did not anticipate sufficient deaths over the course of the trials to see a statistically significant benefit.

This clinical trial had surprising results. Although the trial did not show a difference on the primary endpoint, a composite defined in terms of various pulmonary functional changes or death, the trial did “demonstrate,” according to the witness, a survival benefit.  Indeed, the survival benefit was clinically significant.  Patients randomized to therapy experienced a 40% decrease in mortality, compared to those randomized to placebo. (p = 0.084).  The expert witness pointed out, in his report, that the survival benefit was even stronger in a subgroup of the clinical trial, which consisted of the patients who had mild- to moderate-disease at the time of randomization.  For this subgroup, the decrease in mortality was even more dramatic, 70%, p = 0.004.  The witness’s report did not clearly label this subgroup as “post-hoc,” although a discerning reader might well have assessed it as such.

The expert witness was not relying upon only one clinical trial.  His report identified an earlier trial, published in a leading clinical medical journal, which reported benefit from the drug, p < 0.001.  This trial was extended, with continuing strong evidence of differential survival.  In terms of survival at five years, the earlier trial showed survival in the therapy group at 77.8%, compared to 16.7% in the control groups, p = 0.009.

The expert witness’s report did not explicitly reference clinical experience, or the in vitro and in vivo mechanistic evidence that the therapy, X, plays a role in inhibiting processes that are clearly involved in producing the disease, Z.  The expert witness could have written a stronger expert witness report with these references, but did not expect that this level of completeness was required.  The expert witness did note that he would marshal the data in more detail at a later time. The expert witness further relied upon the assessment of the principal investigator of the later clinical, who had written that the benefit against mortality of X was “compelling,” and that the finding was “a major breakthrough.”  The principal investigator of the trial noted that X was “the first treatment ever to show any meaningful clinical impact in this disease in rigorous clinical trials, and these results would indicate that [X] should be used early in the course of this disease in order to realize the most favorable long-term survival benefit.”  The report went on to note, accurately, that there are no FDA-approved therapies for Z.

Adversary counsel, receiving this report, moved pursuant to Federal Rules of Evidence 702 and 703, to exclude the expert witness’s report and his opinions.  The motion to exclude was made in advance of the deposition, and without a preliminary motion for more detail about the supporting data.  In particular, the motion to exclude claimed that the expert witness was unjustified in concluding that a benefit had been “demonstrated,” as opposed to being merely suggested.

What would be the challenger’s chances of success on the Rule 702 motion?  The outcome, Y, was not “statistically significant” at the conventional two-tail 5% (but would have been on a one-tail test).  The subgroup that sported a p-value of 0.004 was not clearly marked as a post-hoc subgroup, although the challenger could discern that it was likely exploratory, and challenged it as uncorrected for multiple testing.  The challenger, however, did not attempt to offer a modified p-value that took into account of multiple testing.  The essence of this challenge was that the expert witness’s statement that a benefit had been “demonstrated” was not supported by sufficient evidence, and that the low p-value of 0.004 was not truly “significant” because the result emerged from an analysis that was not pre-planned.

My hunch, based upon published judicial opinions on both state and federal Rule 702 motions, is that many judges would allow the challenged expert witness to testify.  There would be the usual judicial hand waving about the challenge’s going to the weight not the admissibility of the expert witness’s opinion.  Perhaps an occasional judge might order additional discovery.  I believe that most judges would not find that this expert witness had engaged in pathologically bad science such that the party proponent should be denied its expert witness.

Transmuting Disputed Causal Inferences Into Criminal Fraud

Instead of moving to exclude the expert witness’s opinion, why not turn the report over to the U.S. Attorney’s office to prosecute for wire or mail fraud?  Even if a trial court were to brand the opinion “inadmissible,” that outcome would hardly suggest that the opinion was the kind of speech that could qualify as fraudulent under federal wire or mail fraud statutes. Branding a scientist as a fraudfeasor, however, was exactly the result reached in U.S. v. Harkonen, where the Ninth Circuit upheld a wire fraud conviction of a physician whose written statements would likely have been admissible in most federal courtrooms, under Federal Rule 702.  As much as I would like to see more stringent gatekeeping of expert witness opinions, there is something unseemly about the government’s efforts here to criminalize scientific opinions with which it disagrees.

Dr. Harkonen has petitioned the Ninth Circuit for reconsideration, in a brief filed by attorneys, Mark Haddad and colleagues, of Sidley Austin.  Petition for Rehearing En Banc (filed 29, 2013).  The case raises important First Amendment and due process issues, which were addressed by the party and amici briefs before the Panel.

The case also raises the specter of prosecutions of scientists for speech in various contexts, including grant applications and reports, under the False Claims Act, for witness perjury for testimony in judicial, administrative, or legislative proceedings, or for wire or mail fraud for manuscript submissions to journals. On April 8th, Professor Robert Makuch, of Yale University, Professor Timothy Lash, of Emory University, and I filed an amicus brief, which addresses the government’s controversial branding statements “false as a matter of statistics.” The government has gone from one extreme of painting, broad brush, that statistical significance is not important or necessary (in Matrixx Initiatives Inc. v. Siracusano), to the other extreme that statistical significance is so important that a scientist who his opinion on causality on evidence the government believes is not statistically significant has committed fraud (in Harkonen).  Both extreme positions are untenable.

New York Breathes Life Into Frye Standard – Reeps v. BMW

March 5th, 2013

In Lumpenepidemiology, I detailed how one federal judge, the Hon. Helen Berrigan, was willing to “just say no” to bad epidemiology and bad science, and to shut an expert witness’s attempt to distort and subvert scientific methodology.  Judge Berrigan closely examined the plaintiffs’ claim that a mother’s ingestion of Paxil caused a child’s heart defect, and found the proffered expert witness testimony to fail legal and scientific standards. Frischhertz v. SmithKline Beecham Corp., 2012 U.S. Dist. LEXIS 181507 (E.D.La. 2012).  The plaintiffs’ key expert witness, Dr. Shira Kramer, attempted to provide plaintiffs with a necessary association by “lumping” all birth defects together in her analysis of epidemiologic data of birth defects among children of women who had ingested Paxil (or other SSRIs).  Given the clear evidence that different birth defects arise at different times, based upon interference with different embryological processes, the trial court discerned this “lumping” of end points to be methodologically inappropriate.  Id. at *13 (citing Chamber v. Exxon Corp., 81 F. Supp. 2d 661 (M.D. La. 2000), aff’d, 247 F.3d 240 (5th Cir. 2001).

Frischhertz was decided in December 2012, the same month that another trial judge, right here in New York City, caught Dr. Shira Kramer in the commission of similar lumpenepidemiology, in Reeps v. BMW of North America, LLC, New York S.Ct., Index No. 100725/08 (New York Cty. Dec. 21, 2012) (York, J.).   See William Ruskin, “Frye Decision in BMW Case Results in Exclusion of Plaintiff’s Experts(Jan. 17, 2013). Reeps was also a birth defects case.  Debra Reeps claimed that during the first trimester of her pregnancy, she was exposed to gasoline fumes from a fuel-line leak in her BMS 525i. She also claimed that her son’s adverse birth outcomes (which included severe mental retardation, severe cerebral palsy, and a congenital heart defect) were caused by her inhalation of gasoline fumes.  Heading the plaintiffs’ team of expert witnesses in support of these claims, Epidemiologist Shira Kramer opined that all of the boy’s problems were caused by the mother’s exposure to unleaded gasoline fumes.

Kramer’s opinions read like the Berenstain Bears’ guide to epidemiology.  She asserted that gasoline vapors and  its constituents (toluene, benzene, solvents, etc.), individually or collectively, cause “birth defects” generally, and Sean Reeps’ defects specifically.  Kramer also asserted that she used a “weight-of-evidence assessment,” which included a consideration of Bradford Hill’s criteria for judging causality. BMW moved to exclude plaintiffs’ witnesses, including Kramer, on grounds that the witnesses’ evidence and methods were “novel, unorthodox, unreliable and not generally accepted in the relevant scientific communities.”  Reeps slip op. at 5.  The number of ways that Kramer’s opinions ran afoul of New York law of expert witness opinions is remarkable.

Animal Studies

The animal studies found no relevant adverse birth effects, even at high gasoline fume exposure levels.  Kramer and her posse nonetheless cited animal studies involving cancer, miscarriage, and anemia for the general claim that gasoline fumes causes birth defects, as though such defects could all be lumped together.

Case Reports

Kramer relied upon two published papers of case reports in which women were exposed to leaded gasoline and then gave birth to children with malformations.  Given that the exposures reported were to leaded gasoline, the case reports were dubious in the first instance.  Furthermore, the reported defects were not even the same as those experienced by Sean Reeps.  Although the court seemed willing to engage in a discussion of what these case reports might offer towards a synthesis of all evidence, it ultimately recognized that, pace Raymond Wolfinger, plural of anecdote is not data.  “Courts have recognized that … case reports are not generally accepted in the scientific community on questions of causation.” Slip op. at 17 (quoting from Heckstall v. Pincus, 19 A.D.3d 203, 205, 797 N.Y.S.2d 445 (1st Dept. 2005)).

Exposure Assessment

The chemical components in gasoline, blamed by Kramer, make up no more than two percent of gasoline vapor. Reeps, slip at 6. One of plaintiffs’ expert witnesses asserted, without measurements, that Debra Reeps experienced atmospheric concentrations of gasoline at least 1,000 p.p.m.  Plaintiffs claimed that this level of exposure was tantamount to recreational solvent abuse, in an attempt to rely upon studies of solvent exposure at very high levels. BMW showed that the witnesses’ speculation was unfounded and implausible.  The fuel-line leak would have to leak about a gallon per mile driven to generate 1,000 p.p.m. in the passenger compartment.  Id. at 8.

Teratology Principles

Because certain structures, organs, and tissues in a developing embryo or fetus form at predictable stages of pregnancy, the science of teratology plays close attention to when the exposure to the putative teratogen occurred in the time course of a pregnancy.  Late exposures to known teratogens cannot very well explain harms that can result only from exposure early in pregnancy.  Similarly, early exposures cannot explain harms that arise only out of teratogenic exposures.  Debra Reeps’ claimed gasoline exposure occurred in her first trimester.  Despite Dr. Shira Kramer’s efforts, the neurological deficits and injuries in Sean Reeps thus cannot be explained by his mother’s early term exposures, even if gasoline fumes had the claimed teratogenic properties.

Ipse Dixit

Debra Reeps had a history of herpes simplex infection, which could explain her son’s cerebral palsy.  Slip op. at 7.  Dr. Kramer asserted that there were no alternative causes.

Epidemiology

Apparently no analytical epidemiologic study (either cohort or case-control) found an association between gasoline fume exposure in pregnancy and Sean Reeps’ birth defects. Kramer attempted to claim that the “[f]ailure to detect a statistical association does not establish that there is no association between an exposure and an outcome.”  Slip op. at 15.  Absence of evidence may not show evidence of absence, but it also does not show evidence of harm.

In one episode of Seinfeld, Jerry Seinfeld chides a rental car clerk for not honoring a reservation.  “You know how to take the reservation; you just don’t know how to hold the reservation.  And that’s really the most important part.”  Scientific methodology is similar to making reservations.  Anyone can claim to be following Sir Austin Bradford Hill’s causal criteria, but actually applying the criteria faithfully is really what the “methodology” is all about.  The abridged form favored by Dr. Kramer is indeed unorthodox, novel, unreliable, invalid, and unacceptable, scientifically and legally, as Justice York found in Reeps.  In essence, the plaintiffs argued that all their expert witnesses need show is that they are aware of proper methodologies, not that they actually used the methodologies properly.  Shira Kramer, and the other plaintiffs’ expert witnesses, offered a pastiche of a method, in the hopes that this would be sufficient.  Absent was a systematic review, and a proper analysis of the evidence. The Reeps case rejoined:  the law requires the real thing.

New York’s adherence to a Frye standard creates a potential roadblock to meaningful gatekeeping.  If an expert witness could evade gatekeeping by simply claiming to be following epidemiologic methods, regardless of how badly, that witness could undermine the interests of the justice system in weeding out speculative, unreliable, or invalid opinions.  The New York Court of Appeals demonstrated its unwillingness to tolerate such evasions.  See, e.g., Parker v. Mobil Oil Corp., 7 N.Y.3d 434, 857 N.E.2d 1114, 824 N.Y.S.2d 584 (2006) (excluding testimony of Dr. Bernard Goldstein, and dismissing leukemia (AML) claim based upon claimed low-level benzene exposure from gasoline) , aff’g 16 A.D.3d 648 (App. Div. 2d Dep’t 2005).

In Reeps, Justice York makes clear that it is the “plaintiff’s burden to prove the methodology applied to reach the conclusions will not be rejected by specialists in the field.”  Slip op. at 11.  The trial court recognized that a Frye hearing in New York must determine whether plaintiffs’ expert witnesses are faithfully applying a methodology, such as the Bradford Hill criteria, or whether they are they are “pay[ing] lip service to them while pursuing a completely different enterprise.”  Id.  To be sure, litigants might not welcome this level of scrutiny for their expert witnesses.  Justice York’s recognition that the court must examine a proffered opinion to determine whether it “properly relates existing data, studies or literature to the plaintiff’s situation, or whether, instead, it is connected to existing data only by the ipse dixit of the expert,” carries with it, an acknowledgment that New York law, like federal Rule 702, requires an assessment of the validity and sufficiency of the evidence and inferences that make up an expert witness’s opinions.  Id. (internal quotations omitted).

Leaving Las Vegas

February 24th, 2013

The Journal of the National Cancer Institute recently published a curious article about what appears to be unpublished research that suggests a non-asbestos environmental cause of malignant mesothelioma in Clark County, Nevada.  Leslie Harris O’Hanlon, “Researchers Explore Possible Link Between Mesothelioma and Dust Emissions in Southern Nevada,” J. Nat’l Cancer Instit., doi: 10.1093/jnci/djt033,  published ahead of print (Feb. 12, 2013).

The researcher appears to have been Francine Baumann , an epidemiologist at the University of Hawaii Cancer Center, who has worked with Michele Carbone, on occasion.  Analyzing Nevada’s cancer registry data from 1995 to 2008, Baumann found what she believed to be an increase in earlier age at diagnosis, and a reduced ratio of male-to-female cases for Clark County.   She interpreted these data to show that an environmental exposure was at work, but she professed ignorance of what the exposure might be.

The article also quotes the Nevada state epidemiologist, Ihsan Azzam, M.D., Ph.D., as saying:

“We analyzed the data and used the same data set as the researcher and came to completely different conclusions and findings. Their interpretation of data and their representation of it is wrong.”

The article presents no data or statistical analysis.  Given that Baumann’s work is unpublished, and apparently contradicted, it is curious that the Journal would publish any story about it.  Some of the raw data can be found online at Nevada Central Cancer Registry, including an online database, and Reports From The Office of Public Health Informatics and Epidemiology.

The O’Hanlon article is even more curious considering the nature of the research.  There are 16 counties in Nevada,  so Baumann presumably was canvassing counties without a pre-specified hypothesis as to whether Clark County was different from the others, or from the national rates.  This seems like post-hoc data dredging, but the Journal does not provide sufficient information to assess the validity of Baumann’s work.

The O’Hanlon article bizarrely talks about an unknown environmental cause in Clark County, but does not mention erionite, a zeolite.  The article discusses erionite-associated mesothelioma in Turkey, and an investigation into erionite occurrences in the United States.  Remarkably, O’Hanlon fails to mention that erionite occurs in Clark County, and in many other counties, throughout Nevada.  The NIOSH Science Blog fills in the missing information by showing how widespread erionite deposits are throughout Nevada.  See David Weissman, MD, and Max Kiefer, MS, CIH, “Erionite: An Emerging North American Hazard,” (Nov. 22, 2011).  Of course, the widespread deposits argue against erionite as a causal explanation for the putative environmental trigger in Clark County.  See also Arthur J. Gude & Richard Sheppard, “Wooly Erionite from the Reese River Zeolite Deposit, Lander County, Nevada, and its Relationship to Other Erionites,” 29 Clays and Clay Minerals, 378-384 (1981); Keith Papke, “Erionite and Other Associated Zeolites in Nevada,” Bulletin 79, Nevada Bureau of Mines and Geology (1972).

Erionite occurs in several mineralogical forms, including non-fibrous and various fibrous forms.  The erionite associated with environmental cases in Turkey has been studied and found to be fibrous, but there are many variations in fibers, including length, and length-to-diameter aspect ratio.  Erionite is a zeolite mineral and has the ability to absorb metal ions, including chromate, uranyl, and other ions, which may be an independent source of potential carcinogenicity.

There are many reasons to leave Las Vegas, but Dr. Baumann probably has not found a new one.

Reanalysis of Epidemiologic Studies – Not Intrinsically WOEful

December 27th, 2012

A recent student law review article discusses reanalyses of epidemiologic studies, an important, and overlooked topic in the jurisprudence of scientific evidence.  Alexander J. Bandza, “Epidemiological-Study Reanalyses and Daubert: A Modest Proposal to Level the Playing Field in Toxic Tort Litigation,” 39 Ecology L. Q. 247 (2012).

In the Daubert case itself, the Ninth Circuit, speaking through Judge Kozinksi, avoided the methodological issues raised by Shanna Swan’s reanalysis of Bendectin epidemiologic studies, by assuming arguendo its validity, and holding that the small relative risk yielded by the reanalysis would not support a jury verdict of specific causation. Daubert v. Merrell Dow Pharm., Inc., 43 F.3d 1311, 1317–18 (9th Cir. 1995).

There is much that can, and should, be said about reanalyses in litigation and in the scientific process, but Bandza never really gets down to the business at hand. His 36 page article curiously does not begin to address reanalysis until the bottom of the 20th page. The first half of the article, and then some, reviews some time-worn insights and factoids about scientific evidence. Finally, at page 266, the author introduces and defines reanalysis:

“Reanalysis occurs ‘when a person other than the original investigator obtains an epidemiologic data set and conducts analyses to evaluate the quality, reliability or validity of the dataset, methods, results or conclusions reported by the original investigator’.”

Bandza at 266 (quoting Raymond Neutra et al., “Toward Guidelines for the Ethical Reanalysis and Reinterpretation of Another’s Research,” 17 Epidemiology 335, 335 (2006).

Bandza correctly identifies some of the bases for judicial hostility to re-analyses. For instance, some courts are troubled or confused when expert witnesses disagree with, or reevaluate, the conclusions of a published article. The witnesses’ conclusions may not be published or peer reviewed, and thus the proffered testimony fails one of the Daubert factors.  Bandza correctly notes that peer review is greatly overrated by judges. Bandza at 270. I would add that peer review is an inappropriate proxy for validity, a “test,” which reflects a distrust of the unpublished.  Unfortunately, this judicial factor ignores the poor quality of much of what is published, and the extreme variability in the peer review process. Judges overrate peer review because they are desperate for a proxy for validity of the studies relied upon, which will allow them to pass their gatekeeping responsibility on to the jury. Furthermore, the authors’ own conclusions are hearsay, and their qualifications are often not fully before the court.  What is important is the opinion of the expert witness who can be cross-examined and challenged.  SeeFOLLOW THE DATA, NOT THE DISCUSSION.” What counts is the validity of the expert witness’s reasoning and inferences.

Bandza’s article, which by title advertises itself to be about re-analyses, gives only a few examples of re-analyses without much detail.  He notes concerns that reanalyses may impugn the reputation of published scientists, and burden them with defending their data.  Who would have it any other way? After this short discussion, the article careens into a discussion of “weight of the evidence” (WOE) methodology. Bandza tells us that the rejection of re-analyses in judicial proceedings “implicitly rules out using the weight-of-the-evidence methodology often appropriate for, or even necessary to, scientific analysis of potentially toxic substances.” Bandza at 270.  This argument, however, is one sustained non-sequitur.  WOE is defined in several ways, but none of the definitions require or suggest the incorporation of re-analyses. Re-analyses raise reliability and validity issues regardless whether an expert witness incorporates them into a WOE assessment. Yet Bandza tells us that the rejection of re-analyses “Implicitly Ignores the Weight-of-the-Evidence Methodology Appropriate for the Scientific Analysis of Potentially Toxic Substances.” Bandza at 274. This conclusion simply does not follow from the nature of WOE methodology or reanalyses.

Bandza’s ipse dixit raises the independent issue whether WOE methodology is appropriate for scientific analysis. WOE is described as embraced or used by regulatory agencies, but that description hardly recommends the methodology as the basis for a scientific, as opposed to a regulatory, conclusion.  Furthermore, Bandza ignores the ambiguity and variability of WOE by referring to it as a methodology, when in reality, WOE is used to describe a wide variety of methods of reasoning to a conclusion. Bandza cites Douglas Weed’s article on WOE, but fails to come to grips with the serious objections raised by Weed in his article to the use of WOE methodologies.  Douglas Weed, “Weight of Evidence: A Review of Concept and Methods,” 25 Risk Analysis 1545, 1546–52 (2005) (describing the vagueness and imprecision of WOE methodologies). See also “WOE-fully Inadequate Methodology – An Ipse Dixit By Another Name.”

Bandza concludes his article with a hymn to the First Circuit’s decision in Milward v. Acuity Specialty Products Group, Inc., 639 F.3d 11 (1st Cir. 2011). Plaintiffs’ expert witness, Dr. Martyn Smith claimed to have performed a WOE analysis, which in turn was based upon a re-analysis of several epidemiologic studies. True, true, and immaterial.  The re-analyses were not inherently a part of a WOE approach. Presumably, Smith re-analyzed some of the epidemiologic studies because he felt that the data as presented did not support his desired conclusion.  Given the motivations at work, the district court in Milward was correct to look skeptically and critically at the re-analyses.

Bandza notes that there are procedural and evidentiary safeguards in federal court against unreliable or invalid re-analyses of epidemiologic studies.  Bandza at 277. Yes, there are safeguards but they help only when they are actually used. The First Circuit in Milward reversed the district court for looking too closely at the re-analyses, spouting the chestnut that the objections went to the weight not the admissibility of the evidence.  Bandza embraces the rhetoric of the Circuit, but he offers no description or analysis of the liberties that Martyn Smith took with the data, or the reasonableness of Smith’s reliance upon the re-analyzed data.

There is no necessary connection between WOE methodologies and re-analyses of epidemiologic studies.  Re-analyses can be done properly to support or deconstruct the conclusions of published papers.  As Bandza points out, some re-analyses may go on to be peer reviewed and published themselves.  Validity is the key, and WOE methodologies have little to do with the process of evaluating the original or the re-analyzed study.

 

 

Lumpenepidemiology

December 24th, 2012

Judge Helen Berrigan, who presides over the Paxil birth defects MDL in New Orleans, has issued a nicely reasoned Rule 702 opinion, upholding defense objections to plaintiffs expert witnesses, Paul Goldstein, Ph.D., and Shira Kramer, Ph.D. Frischhertz v SmithKline Beecham EDLa 2012 702 MSJ Op.

The plaintiff, Andrea Frischhertz, took GSK’s Paxil, a selective serotonin reuptake inhibitor (SSRI), for depression while pregnant with her daughter, E.F. The parties agreed that E.F. was born with a deformity of her right hand.  Plaintiffs originally claimed that E.F. had a heart defect, but their expert witnesses appeared to give up this claim at deposition, as lacking evidential support.

Adhering to Daubert’s Epistemiologic Lesson

Like many other lower federal courts, Judge Berrigan focused her analysis on the language of Daubert v. Merrell Dow Pharmaceuticals Inc., 509 U.S. 579 (1993), a case that has been superseded by subsequent cases and a revision to the operative statute, Rule 702.  Fortunately, the trial court did not lose sight of the key epistemological teaching of Daubert, which is based upon Rule 702:

“Regarding reliability, the [Daubert] Court said: ‘the subject of an expert’s testimony must be “scientific . . . knowledge.” The adjective “scientific” implies a grounding in the methods and procedures of science. Similarly, the word “knowledge” connotes more than subjective belief or unsupported speculation’.”

Slip Op. at 3 (quoting Daubert, 509 U.S. at 589-590).

There was not much to the plaintiffs’ expert witnesses’ opinion beyond speculation, but many other courts have been beguiled by speculation dressed up as “scientific … knowledge.”  Dr. Goldstein relied upon whole embryo culture testing of SSRIs, but in the face overwhelming evidence, Dr. Goldstein was forced to concede that this test may generate hypotheses about, but cannot predict, human risk of birth defects.  No doubt this concession made the trial court’s decision easier, but the result would have been required regardless of Dr. Goldstein’s exhibition of truthfulness at deposition.

Statistical Association – A Good Place to Begin

More interestingly, the trial court rejected the plaintiffs’ expert witnesses’ efforts to leapfrog finding a statistically significant association to parsing the so-called Bradford Hill factors:

“The Bradford-Hill criteria can only be applied after a statistically significant association has been identified. Federal Judicial Center, Reference Manual on Scientific Evidence, 599, n.141 (3d. ed. 2011) (“In a number of cases, experts attempted to use these guidelines to support the existence of causation in the absence of any epidemiologic studies finding an association . . . . There may be some logic to that effort, but it does not reflect accepted epidemiologic methodology.”). See, e.g., Dunn v. Sandoz Pharms., 275 F. Supp. 2d 672, 678 (M.D.N.C. 2003). Here, Dr. Goldstein attempted to use the Bradford-Hill criteria to prove causation without first identifying a valid statistically significant association. He first developed a hypothesis and then attempted to use the Bradford-Hill criteria to prove it. Rec. Doc. 187, Exh. 2, depo. Goldstein, p. 103. Because there is no data showing an association between Paxil and limb defects, no association existed for Dr. Goldstein to apply the Bradford-Hill criteria. Hence, Dr. Goldstein’s general causation opinion is not reliable.”

Slip op. at 6.

The trial court’s rejection of Dr. Goldstein’s attempted end run is particularly noteworthy given the Reference Manual’s weak-kneed attempt to suggest that this reasoning has “some logic” to it.  The Manual never articulates what “logic” commends Dr. Goldstein’s approach; nor does it identify any causal relationship ever established with such paltry evidence in the real world of science. The Manual does cite several legal cases that excused or overlooked the need to find a statistically significant association, and even elevated such reasoning into legally acceptable, admissibility method.  See Reference Manual on Scientific Evidence at 599 n. 141 (describing cases in which purported expert witnesses attempted to use Bradford Hill factors in the absence of a statistically significant association; citing Rains v. PPG Indus., Inc., 361 F. Supp. 2d 829, 836–37 (S.D. Ill. 2004); ); Soldo v. Sandoz Pharms. Corp., 244 F. Supp. 2d 434, 460–61 (W.D. Pa. 2003).  The Reference Manual also cited cases, without obvious disapproval, which completely dispatched with any necessity of considering any of the Bradford Hill factors, or the precondition of a statistically significant association.  See Reference Manual at 599 n. 144 (citing Cook v. Rockwell Int’l Corp., 580 F. Supp. 2d 1071, 1098 (D. Colo. 2006) (“Defendants cite no authority, scientific or legal, that compliance with all, or even one, of these factors is required. . . . The scientific consensus is, in fact, to the contrary. It identifies Defendants’ list of factors as some of the nine factors or lenses that guide epidemiologists in making judgments about causation. . . . These factors are not tests for determining the reliability of any study or the causal inferences drawn from it.“).

Shira Kramer Takes Her Lumpings

The plaintiffs’ other key expert witness, Dr. Shira Kramer, was a more sophisticated and experienced obfuscator.  Kramer attempted to provide plaintiffs with a necessary association by “lumping” all birth defects together in her analysis of epidemiologic data of birth defects among children of women who had ingested Paxil (or other SSRIs).  Given the clear evidence that different birth defects arise at different times, based upon interference with different embryological processes, the trial court discerned this “lumping” of end points to be methodologically inappropriate.  Slip op. at 8 (citing Chamber v. Exxon Corp., 81 F. Supp. 2d 661 (M.D. La. 2000), aff’d, 247 F.3d 240 (5th Cir. 2001) (unpublished).

Without her “lumping”, Dr. Kramer was left with only a weak, inconsistent claim of biological plausibility and temporality. Finding that Dr. Kramer’s opinion had outrun her headlights, Judge Berrigan, excluded Dr. Kramer as an expert witness, and granted GSK summary judgment.

Merry Christmas!

 

EPA Post Hoc Statistical Tests – One Tail vs Two

December 2nd, 2012

EPA 1992 Meta-Analysis of ETA & Lung Cancer – Part 2

In 1992, the U.S. Environmental Protection Agency (EPA) published a risk assessment of lung cancer (and other) risks from environmental tobacco smoke (ETS).  See Respiratory Health Effects of Passive Smoking: Lung Cancer and Other Disorders EPA/600/6-90/006F (1992).  The agency concluded that ETS causes about 3,000 lung cancer deaths each year among non-smoking adults.  See also EPA “Fact Sheet: Respiratory Health Effects of Passive Smoking,” Office of Research and Development, and Office of Air and Radiation, EPA Document Number 43-F-93-003 (Jan. 1993).

In my last post, I discussed  how various plaintiffs, including tobacco companies, challenged the EPA’s conclusions as agency action that violated administrative and statutory procedures. “EPA Cherry Picking (WOE) – EPA 1992 Meta-Analysis of ETA & Lung Cancer – Part 1” (Dec. 2. 2012). The plaintiffs further claimed that the EPA had manufactured its methods to achieve the result it desired in advance of the analyses. A federal district court agreed with the methodological challenges to the EPA’s report, but the Court of Appeals reversed on grounds that the agency’s report was not reviewable agency action.  Flue-Cured Tobacco Cooperative Stabilization Corp. v. EPA, 4 F. Supp. 2d 435 (M.D.N.C. 1998), rev’d 313 F.3d 852, 862 (4th Cir. 2002) (Widener, J.) (holding that the issuance of the report was not “final agency action”).

One of the grounds of the plaintiffs’ challenge was that the EPA had changed, without explanation, from a 95% to a 90% confidence interval.  The change in the specification of the coefficient of confidence was equivalent to a shift from a two-tailed to a one-tailed test of confidence, with alpha set at 5%.  This change, along with gerrymandering or “cherry picking” of studies, allowed the EPA to claim a statistically significant association between ETS and lung cancer. 4 F. Supp. 2d at 461.  The plaintiffs pointed to EPA’s own previous risk assessments, as well as statistical analyses by the World Health Organization (International Agency for Research on Cancer), the National Research Council, and the Surgeon General, all of which routinely use 95% intervals, and two-tailed tests of significance.  Id.

In its 1990 Draft ETS Risk Assessment, the EPA had used a 95% confidence interval, but in later drafts, changed to a 90% interval.  One of the epidemiologists on the EPA’s Scientific Advisory Board, Geoffrey Kabat, criticized this post hoc change, noting that the use of 90% intervals are disfavored and that the post hoc change in statistical methodology created the appearance of an intent to influence the outcome of the analysis. Id. (citing Geoffrey Kabat, “Comments on EPA’s Draft Report: Respiratory Health Effects of Passive Smoking: Lung Cancer and Other Disorders,” II.SAB.9.15 at 6 (July 28, 1992) (JA 12,185).

The EPA argued that its adoption of a one-tailed test of significance was justified on the basis of an a priori hypothesis that ETS is associated with lung cancer.  Id. at 451-52, 461 (citing to ETS Risk Assessment at 5–2). The court found this EPA argument hopelessly circular.  The agency postulated its a priori hypothesis, which it then took as license to dilute the statistical test for assessing the evidence.  The agency, therefore, had assumed what it wished to show, in order to achieve the result it sought.  Id. at 456.  The EPA claimed that the one-tailed test had more power, but with dozens of studies aggregated into a summary result, the court recognized that Type I error was a larger threat to the validity of the agency’s conclusions.

The EPA also advanced a muddled defense of its use of 90% confidence intervals by arguing that if it used a 95% interval, the results would have been incongruent with the one-tailed p-values.  The court recognized that this was really no discrepancy at all, but only a corollary of using either one-tailed 5% tests or 90% confidence intervals.  Id. at 461.

If the EPA had adhered to its normal methodology, there would have been no statistically significant association between ETS and lung cancer. With its post hoc methodological choice, and highly selective approach to study inclusions in its meta-analysis, the EPA was able to claim a weak statistically significant association between ETS and lung cancer.  Id. at 463.  The court found this to be a deviation from the legally required use of “best judgment possible based upon the available evidence.”  Id.

Of course, the EPA could have announced its one-tailed test from the inception of the risk assessment, and justified its use on grounds that it was attempting to reach only a precautionary judgment for purposes of regulation.  Instead, the agency tried to showcase its finding as a scientific conclusion, which only further supported the tobacco companies’ challenge to the post hoc change in plan for statistical analysis.

Although the validity issues in the EPA’s 1992 meta-analysis should have been superseded by later studies, and later meta-analyses, the government’s fraud case, before Judge Kessler, resurrected the issue:

“3344. Defendants criticized EPA’s meta-analysis of U.S. epidemiological studies, particularly its use of an ‘unconventional 90 percent confidence interval’. However, Dr. [David] Burns, who participated in the EPA Risk Assessment, testified that the EPA used a one-tailed 95% confidence interval, not a two-tailed 90% confidence interval. He also explained in detail why a one-tailed test was proper: The EPA did not use a 90% confidence interval. They used a traditional 95% confidence interval, but they tested for that interval only in one direction. That is, rather than testing for both the possibility that exposure to ETS increased risk and the possibility that it decreased risk, the EPA only tested for the possibility that it increased the risk. It tested for that possibility using the traditional 5% chance or a P value of 0.05. It did not test for the possibility that ETS protected those exposed from developing lung cancer at the direction of the advisory panel which made that decision based on its prior decision that the evidence established that ETS was a carcinogen. What was being tested was whether the exposure was sufficient to increase lung cancer risk, not whether the agent itself, that is cigarette smoke, had the capacity to cause lung cancer with sufficient exposure. The statement that a 90% confidence interval was used comes from the observation that if you test for a 5% probability in one direction the boundary is the same as testing for a 10% probability in two directions. Burns WD, 67:5-15. In fact, the EPA Risk Assessment stated, ‘Throughout this chapter, one-tailed tests of significance (p = 0.05) are used …’ .”

U.S. v. Philip Morris USA, Inc., 449 F. Supp. 2d 1, 702-03 (D.D.C., 2006) (Kessler, J.) (internal citations omitted).

Judge Kessler was misled by Dr. Burns, a frequent testifier for plaintiffs’ counsel in tobacco cases.  Burns should have known that with respect to the lower bound of the confidence interval, which is what matters for determining whether the meta-analysis excludes a risk ratio of 1.0, there is no difference between a one-tailed 95% confidence interval and a two-tailed 90% interval.  Burns’ sophistry hardly saves the EPA’s error in changing its pre-specified end point and statistical analysis, or the danger of unduly increasing the risk of Type I error in the EPA meta-analysis. SeePin the Tail on the Significance Test” (July 14th, 2012)

Post-script

Judge Widener wrote the opinion for a panel of the United States Court of Appeals, for the Fourth Circuit, which reversed the district court’s judgment, enjoining the EPA’s report.  The Circuit’s decision did not address the scientific issues, but by holding that the agency action was not reviewable, the basis for the district court’ review of the scientific and statistical issues was removed.  For those pundits who see only self-interested behavior in judging, the author of the Circuit’s decision was a life-time smoker, who grew Burley tobacco on his farm, outside Abingdon, Virginia.  Judge Widener died on September 19, 2007, of lung cancer.

EPA Cherry Picking (WOE) – EPA 1992 Meta-Analysis of ETS & Lung Cancer – Part 1

December 2nd, 2012

Somehow, before the Supreme Court breathed life into Federal Rule of Evidence 702, parties sometimes found a way to challenge dubious scientific evidence in court.  One good example is the challenge to the United States Environmental Protection Agency’s risk assessment of passive smoking, also known as environmental tobacco smoke (ETS).  In 1992, the Environmental Protection Agency (EPA) published a risk assessment of lung cancer (and other) risks from ETS.  See Respiratory Health Effects of Passive Smoking: Lung Cancer and Other Disorders EPA/600/6-90/006F (1992).  The agency concluded that ETS causes about 3,000 lung cancer deaths each year among non-smoking adults in the United States.  See also EPA “Fact Sheet: Respiratory Health Effects of Passive Smoking,” Office of Research & Development; EPA Document Number 43-F-93-003 (Jan. 1993).

Various plaintiffs, including tobacco companies, challenged the EPA’s conclusions as agency action that violated administrative and statutory procedures.  The plaintiffs further claimed that the EPA had manufactured its methods to achieve the result it desired in advance of the analyses. In other words, plaintiffs asserted that the EPA’s issuance of the ETS report violated the Administrative Procedures Act’ procedural requirements, as well as the requirements of the specific enabling legislation, the Radon Gas and Indoor Air Quality Research Act, Pub.L. No. 99–499, 100 Stat. 1758–60 (1986) (codified at 42 U.S.C. § 7401 note (1994)).  A federal district court agreed with the methodological challenges to the EPA’s report, but the Court of Appeals reversed on grounds that the agency’s report was not reviewable agency action.  Flue-Cured Tobacco Cooperative Stabilization Corp. v. EPA, 4 F. Supp. 2d 435 (M.D.N.C. 1998), rev’d on other grounds, 313 F.3d 852, 862 (4th Cir. 2002) (Widener, J.) (holding that the issuance of the report was not “final agency action”). The district court’s assessment of the validity issues were not addressed by the appellate court.

Notwithstanding the district court’s findings, the EPA continues to claim that it had reached valid scientific conclusions using a “scientific approach”:

“EPA reached its conclusions concerning the potential for ETS to act as a human carcinogen based on an analysis of all of the available data, including more than 30 epidemiologic (human) studies looking specifically at passive smoking as well as information on active or direct smoking. In addition, EPA considered animal data, biological measurements of human uptake of tobacco smoke components and other available data. The conclusions were based on what is commonly known as the total weight-of-evidence” rather than on any one study or type of study.

The finding that ETS should be classified as a Group A carcinogen is based on the conclusive evidence of the dose-related lung carcinogenicity of mainstream smoke in active smokers and the similarities of mainstream and sidestream smoke given off by the burning end of the cigarette. The finding is bolstered by the statistically significant exposure-related increase in lung cancer in nonsmoking spouses of smokers which is found in an analysis of more than 30 epidemiology studies that examined the association between secondhand smoke and lung cancer.”

EPA “Fact Sheet: Respiratory Health Effects of Passive Smoking,”  Office of Research and Development; EPA Document Number 43-F-93-003, January 1993 (emphasis added).

A prominent feature of the EPA’s analysis was a meta-analysis of epidemiologic studies of ETS and lung cancer.  Interestingly, the tobacco industry plaintiffs did not appear to challenge the legitimacy of the basic meta-analytic enterprise, which was still controversial at the time.  See, e.g., “Samuel Shapiro, Meta-analysis/Smeta-analysis,” 140 Am. J. Epidem. 771 (1994); Alvan Feinstein, “Meta-Analysis: Statistical Alchemy for the 21st Century,” 48 J. Clin. Epidem. 71 (1995).  Their challenge went straight to the validity of the EPA’s meta-analysis, and a documented post hoc change in the agency’s statistical plan for analyzing the meta-analysis results.  Only a few years earlier, the defense in polychlorobiphenyl (PCB) litigation broadly challenged a plaintiffs’ expert witness’s use of meta-analysis of observational epidemiologic studies, only to have the Third Circuit reject the challenge and to direct the district court to review the validity of the meta-analysis as conducted by the witness.  In re Paoli RR Yard PCB Litig., 706 F. Supp. 358, 373 (E.D. Pa. 1988), rev’d, 916 F.2d 829, 856-57 (3d Cir. 1990), cert. denied, 499 U.S. 961 (1991); see also Hines v. Consol. Rail Corp., 926 F.2d 262, 273 (3d Cir. 1991).

The EPA report was not the first attempt to use meta-analysis for the epidemiology of ETS and lung cancer.  In 1986, the National Academy of Sciences reported a meta-analysis on the subject.  See National Research Council, National Academy of Sciences,  Environmental tobacco smoke: measuring exposures and assessing health effects (Wash. DC 1986).  This earlier meta-analysis was also controversial.  Indeed, some of the early concerns over the use of meta-analysis for observational epidemiologic studies arose in the context of studies of ETS.  See, e.g., Joseph L. Fleiss & Alan J. Gross, “Meta-Analysis in Epidemiology, with Special Reference to Studies of the Association between Exposure to Environmental Tobacco Smoke and Lung Cancer:  A Critique,” 44 J. Clin. Epidem. 127 (1991) (criticizing the National Research Council 1986 meta-analysis of ETS and lung cancer studies as unwarranted based upon the low quality of the studies included).  These concerns were heightened by politicized use of meta-analyses in regulatory agencies to overclaim scientific conclusions from weak, inconclusive data.

In the EPA’s meta-analysis, statistical significance was achieved only by changing the criterion of significance, post hoc, from a two-tailed to a one-tailed 5% test.  Perhaps more disturbing was the scientific gerrymandering that took place as to which studies to include and exclude from the meta-analysis.

In its first review of the EPA’s draft report, a committee of the agency’s Scientific Advisory Board, the IAQC [the Indoor Air Quality/Total Human Exposure Committee] found that the EPA’s ETS risk assessment violated one of the necessary criteria for a valid meta-analysis – a “precise definition of criteria used to include (or exclude) studies.”  4 F. Supp. 2d at 459 (citing EPA, An SAB Report: Review of Draft Environmental Tobacco Smoke Health Effects Document, EPA/SAB/IAQC/91/007 at 32–33 (1991) (SAB 1991 Review) (JA 9,497–98)).  The agency had not provided specific criteria for including studies. The IAQC also noted that it was important to evaluate the consequences of having excluded studies in the form of sensitivity studies. In a later review, in 1992, both the EPA and the IAQC dropped this critique of the agency’s meta-analysis, without explanation.  Id. at 459.

By the time the EPA released its ETS report in 1993, there were about 58 published epidemiologic studies available for inclusion in any meta-analysis.  The EPA included only 31.  The agency limited its analysis to nonsmoking women married to smoking spouses.  There were 33 studies of this exposed group; the EPA included 31 of the 33.  There were also available 12 studies of women exposed to ETS in their workplace, and 13 studies of women who were exposed to ETS as children.  Id. at 458. There were three late-breaking studies of women with spousal exposures, but the EPA excluded two, without explanation.  Id. at 459.

In reviewing the plaintiffs’ challenge, the district court noted that the EPA had given a bare, unconvincing explanation for excluding the childhood and workplace studies.  Id.  The EPA argued that there was less data in the childhood and workplace studies, but this assertion struck the court as an evasive rationale when one of the purposes of conducting a meta-analysis was to incorporate the data from smaller, less powerful studies.  Id. 458-59.  The primary author of the disputed chapter of the EPA report, Kenneth Brown, called the disputed studies “inadequate,” without providing a rational basis or explanation.  The IAQC, in its earlier review of a 1991 draft report, recognized that the excluded studies provided less information, but concluded that the agency’s “the report should review and comment on the data that do exist… .” Id. at 459.

The court found the EPA’s selection of studies for inclusion in a meta-analysis to be “disturbing”:

“First, there is evidence in the record supporting the accusation that EPA ‘cherry picked’ its data. Without criteria for pooling studies into a meta-analysis, the court cannot determine whether the exclusion of studies likely to disprove EPA’s a priori hypothesis was coincidence or intentional. Second, EPA’s excluding nearly half of the available studies directly conflicts with EPA’s purported purpose for analyzing the epidemiological studies and conflicts with EPA’s Risk Assessment Guidelines. See ETS Risk Assessment at 4–29 (“These data should also be examined in the interest of weighing all the available evidence, as recommended by EPA’s carcinogen risk assessment guidelines (U.S.EPA, 1986a) ….” (emphasis added)). Third, EPA’s selective use of data conflicts with the Radon Research Act. The Act states EPA’s program shall ‘‘gather data and information on all aspects of indoor air quality….’’ Radon Research Act § 403(a)(1) (emphasis added). In conducting a risk assessment under the Act, EPA deliberately refused to assess information on all aspects of indoor air quality.”

4 F. Supp. 2d at 460.

The court was no doubt impressed by the duplicity of the agency’s claim to have used a “total weight of the evidence” approach to the question of causality, and its censoring of the analysis in a way that appeared to game the result.  Id. at 454  The EPA’s guidelines called for basing conclusions on all available evidence.  EPA’s Guidelines for Carcinogen Risk Assessment, 51 Fed. Reg. 33,996, 33,999-34,000 (1986).

Using evidence selectively, with a post hoc adoption of a one-tailed test of statistical significance, the EPA reported a summary estimate of risk of 1.19, and categorized ETS as a “Group A” carcinogens. In most of its previous Group A classifications, the agency had based its decisions upon much higher relative risks.  Indeed, the agency had rejected Group A classifications when relative risks were found to be less than three.  4 F. Supp. 2d at 461.  The sum total of the agency’s methodological laxity was too much for the district court, which struck the chapters of the EPA report.  Four years later, the Fourth Circuit of the U.S. Court of Appeals reversed, on grounds that the EPA report was not reviewable agency action.

The EPA report became a lightning rod for methodological criticism of meta-analysis for observational studies, and the EPA’s use of meta-analysis.  Critics argued that the EPA had succumbed to political pressure from the anti-tobacco lobby.  See, e.g., Gio B. Gori & John C. Luik, Passive Smoke: The EPA’s Betrayal of Science and Policy (Vancouver, BC: The Fraser Institute 1999); John C. Luik, “Pandora’s Box: The Dangers of Politically Corrupted Science for Democratic Public Policy,” Bostonia 54 (Winter 1999-94).  See also Elizabeth Fisher, “Case law analysis. Passive smoking and active courts: the nature and role of risk regulators in the US and UK.  Flue-cured Tobacco Co-op v US Environmental Protection Agency,” 12 J. Envt’l Law 79 (2000).

The federal government has been trying to defend the EPA’s 1992 report, ever since.  In 1998, upon listing ETS as a known carcinogen, the Department of Health and Human Services noted that “[t]he individual studies were carefully summarized and evaluated”  in the 1992 EPA report.  U.S. Dep’t of Health & Human Services, National Toxicology Program, Final Report on Carcinogens – Background Document for Environmental Tobacco Smoke: Meeting of the NTP Board of Scientific Counselors – Report on Carcinogens Subcommittee at 24 (Research Triangle Park, NC 1998).  Anti-tobacco scientists, including scientists involved in the EPA report, have attacked the motives of the industry, and of the scientists who have challenged the report.  See, e.g., Jonathan M. Samet & Thomas A. Burke, “Turning Science Into Junk: The Tobacco Industry and Passive Smoking,” 91 Am. J. Pub. Health 1742 (2001); Monique E. Muggl, Richard D. Hurt, and James Repace, “The Tobacco Industry’s Political Efforts to Derail the EPA Report on ETS,” 26 Am. J. Prev. Med. 167 (2004); Deborah E. Barnes & Lisa A. Bero, “Why review articles on the health effects of passive smoking reach different conclusions,” 279 J. Am. Med. Ass’n 1566 (1998).

Of course, science did not remain status quo 1992.  Later studies were published, and the controversy continued, such that the 1992 meta-analysis is now largely scientifically irrelevant.  See James Enstrom & Geoffrey Kabat, “Environmental tobacco smoke and tobacco related mortality in a prospective study of Californians, 1960-98,” 326 Br. Med. J. 1057 (2003); G. Davey Smith, “Effect of passive smoking on health: More information is available, but the controversy still persists,” 326 Br. Med. J. 1048–9 (2003).

A troubling implication of those who attack the tobacco industry is that the industry was not allowed to raise methodological challenges to the EPA’s purported use of a scientific method.  The EPA defenders rarely engage with the specifics of the methodological challenge or the district court’s review. Another implication is that the EPA’s meta-analysis remains a clear example of where a regulatory agency could have acted upon a precautionary principle, but chose to dress up its analysis as something it was not:  a scientific conclusion of causality. Given that the agency was not even engaged in reviewable agency action, and that it had plenty of biological plausibility for a precautionary finding that ETS causes lung cancer, the agency could easily have avoided the vitriolic debate it engendered with its 1992 report.