# United States v. Vulcan Soc. Inc.

> District Court, E.D. New York · July 22, 2009 · 637 F. Supp. 2d 77

URL: https://www.frixlaw.com/law-library/cases/2342626

## Case

- **Full name:** UNITED STATES of America, Plaintiff, and the VULCAN SOCIETY, INC., Marcus Haywood, Candido Nuñez, Roger Gregg, Plaintiff-Intervenors, v. the City of New York, Fire Department of the City of New York, New York City Department of Citywide Administrative Services, and Mayor Michael Bloomberg and New York City Fire Commissioner Nicholas Scoppetta, in Their Individual and Official Capacities, Defendants
- **Court:** District Court, E.D. New York
- **Decided:** July 22, 2009
- **Citations:** 637 F. Supp. 2d 77; 106 Fair Empl. Prac. Cas. (BNA) 1561; 2009 U.S. Dist. LEXIS 63153
- **Precedential status:** Published
- **Opinion:** Opinion by Garaufis
- **Judges:** Nicholas G. Garaufis
- **Cited by:** 20 later opinions in the Frix Law Library

## Citator (automated)

- No negative treatment found by the automated citator. That is not the same as a confirmation that the case is good law; read the citing cases.
- Full citator and citing cases: https://www.frixlaw.com/law-library/cases/2342626

## How later opinions describe it (automated extraction)

- finding the exam there unrepresentative of the content of the- job after determining that the job included "a host of requirements other than cognitive abilities,” and that only one of the 21 "job clusters” identified by the test proponents in the job analysis required cogniti…
- finding that defendant’s job analysis for a test given to firefighter candidates was inadequate because no effort had been made to explain the relationship between the knowledge, skills, and *517 abilities being tested on the exam and the tasks involved in being a firefighter
- finding that defendant’s failure to provide evidence linking knowledge, skills, and abilities to job tasks “undermines the court’s confidence that ‘the pertinent abilities have been selected for measurement’ ” (quoting Guardians, 630 F.2d at 96)
- observing that “Guardians contains an unusually complete discussion of the details of test validation,” but that, in creating the exams, the City “ignored the Second Circuit’s guidance” and “appears to be relying on the same practices for which it was criticized by the Second …
- finding that report summarizing the steps of defendant’s job analysis did not satisfy defendant’s burden of proving that it had conducted a suitable analysis

## Opinion text

MEMORANDUM & ORDER
NICHOLAS G. GARAUFIS, District Judge.
From 1999 to 2007, the New York City Fire Department used written examinations with discriminatory effects and little relationship to the job of a firefighter to select more than 5,300 candidates for admission to the New York City Fire Academy. These examinations unfairly excluded hundreds of qualified people of color from the opportunity to serve as New York City firefighters. Today, the court holds that New York City’s reliance on these examinations constitutes employment discrimination in violation of Title VII of the Civil Rights Act of 1964.
I. INTRODUCTION
In recent years, black and Hispanic residents of New York City (the “City”) have come to comprise a substantial portion of the City’s population, but their representation in the New York City Fire Department (“FDNY”) has remained extraordi
*80
narily low.
1
In 2002, the New York City Department of City Planning identified 25% of the City’s residents as black and 27% of its residents as Hispanic.
2
At the same time, however, only 2.6% of its firefighters were black and 3.7% of its firefighters were Hispanic.
3
When this litigation commenced in 2007, the percentages of black and Hispanic firefighters had increased to just 3.4% and 6.7%, respectively.
4
In other words, on a force of 8,998 firefighters, there were just 303 black firefighters and 605 Hispanic firefighters. These numbers stand in stark contrast to some of the nation’s other large cities, such as Los Angeles, Chicago, Philadelphia, and Houston, where minority firefighters have been represented in significantly higher percentages.
5
In this case, Plaintiff United States of America (the “Federal Government”) as well as the Vulcan Society, Inc., Marcus Haywood, Candido Núñez and Roger Gregg (the “Intervenors”), have sued to enforce the right of black and Hispanic candidates to be treated fairly in the application process for positions in the FDNY. Specifically, the Federal Government and the Intervenors (“Plaintiffs”)
6
challenge the City’s reliance on two written examinations that are used to appoint entry-level firefighters to classes at the New York City Fire Academy (“Academy”). These examinations — Written Examination 7029 and Written Examination 2043 — were administered from 1999 to 2007, and the City has appointed more than 5,300 entry-level firefighters based upon their results. Although Plaintiffs identify approximately 3,100 of the examination candidates as black and approximately 4,200 of the examination candidates as Hispanic, the City has appointed just 184 black firefighters and 461 Hispanic firefighters from the challenged examinations.
{See
Section III.A,
infra.)
Plaintiffs assert that the City’s reliance on Exams 7029 and 2043 in selecting entry-level firefighters has had a disparate impact on black and Hispanic candidates in violation of Title VII of the Civil Rights Act of 1964, 42 U.S.C. § 2000e,
et seq.
(“Title VII”). The Intervenors also claim,
*81
under a disparate treatment theory, that the City, two city agencies, the Mayor and the Fire Commissioner “have long been aware of the discriminatory impact on blacks of their examination process,” and that their “continued reliance on and perpetuation of these racially discriminatory hiring processes constitute intentional race discrimination .... ” (Intervenors’ Compl. (Docket Entry # 48) ¶ 51.)
To remedy these claimed violations, Plaintiffs seek various forms of injunctive and monetary relief. The Federal Government seeks to enjoin the City from engaging in discriminatory practices “against blacks on the basis of race and against Hispanics on the basis of national origin,” and seeks a specific injunction against the practices challenged in this case.
(See
Compl. (Docket Entry # 1) ¶ 38.) It also asks the court to order the City to take “appropriate action to correct the present effects of its discriminatory policies and practices” and to enjoin it from failing to “make whole” those harmed by the City’s policies and practices.
(Id.)
The Intervenors seek similar, but broader relief, including an injunction requiring the City to “appoint entry-level firefighters from among qualified black applicants in sufficient numbers to offset the historic pattern and practice of discrimination against blacks in testing and appointment to that position.” (Int. Compl., Prayer For Relief ¶ 3(d).) The Intervenors seek to require the City to “recruit black candidates and implement and improve long-range recruitment programs” and to “provide ... future test scores, appointment criteria, eligibility lists, appointment data, and all other information necessary to conduct an adverse impact and job-relatedness analysis of the examination and selection process.”
(Id.
11113(e), (f).) The Intervenors also seek damages and other fees.
(Id.
¶¶ 4-9.)
This is not the first time the City has been brought to federal court to defend its entry-level firefighter examinations against charges of discrimination. In the early 1970s, Judge Weinfeld in the Southern District of New York found that the City’s written and physical examinations for entry-level firefighters violated the Equal Protection Clause of the Constitution because of their discriminatory impact on black and Hispanic applicants.
See Vulcan Soc’y of New York City Fire Dep’t, Inc. v. Civil Serv. Comrn’n,
360 F.Supp. 1265, 1269 (S.D.N.Y.1973),
affirmed in relevant part by
490 F.2d 387 (2d Cir.1973). Following Judge Weinfeld’s decision in
Vulcan Society,
the City contracted with a private consulting firm to construct valid written and physical examinations; these contracts were cancelled three years later, however, apparently on account of a budget crisis.
See Berkman v. City of New York,
536 F.Supp. 177, 184 (E.D.N.Y. 1982).
7
At the time of
Vulcan Society,
Judge Weinfeld noted the “overwhelming disparity between minority representation in the [FDNY] (5%) and in the general population of New York City within the age group eligible for appointment (32%).”
Vulcan Soc’y,
360 F.Supp. at 1269 . In the three decades that have followed, these minority groups have come to represent an even greater share of the City’s population. Despite these changes, the overwhelmingly monochromatic composition of the FDNY has stubbornly persisted.
8
This court has already issued several decisions in the case. I have bifurcated
*82
the liability and relief phases
(see
Docket Entry # 47), permitted intervention by the Intervenors
(see id.),
denied the Intervenors’ motion to amend the Intervenors’ Complaint
(see
Docket Entry # 182), declined to dismiss the Intervenors’ Complaint on timeliness grounds
(see
Docket Entry # 231), and certified a class consisting of black applicants for the position of entry-level firefighter
(see
Docket Entry # 281).
Now before the court are Motions for Summary Judgment by the Federal Government and the Intervenors. (Docket Entries ##251, 260.) The Federal Government and the Intervenors have moved for summary judgment on the prima facie case of disparate impact, and the Federal Government has joined the Intervenors’ Motion for Summary Judgment on the City’s business necessity defense.
9
(See Docket Entries ## 251, 260, 263.)
Upon consideration of the parties’ submissions and oral argument, the court concludes that Plaintiffs have established a prima facie case that the City’s use of the
*83
two written examinations has resulted in a disparate impact upon black and Hispanic applicants for the position of entry-level firefighter. The court also concludes that the City has failed to present sufficient evidence supporting a business justification for its employment practices. I therefore grant Plaintiffs’ Motions for Summary Judgment in their entirety.
In essence, my ruling is premised upon two basic conclusions. First, Plaintiffs have shown that there is no triable issue of fact as to whether the City’s use of Written Exams 7029 and 2043 has resulted in a statistically and practically significant adverse impact on black and Hispanic firefighter applicants. Black and Hispanic applicants disproportionately failed the written examinations, and those who passed were placed disproportionately lower down than white candidates on the hierarchical hiring lists resulting from their scores. Second, although the City has had the opportunity to justify this adverse impact by showing that it used the written examinations to test for the relevant skills and abilities of entry-level firefighters, the City has failed to raise a triable issue on this defense. Under Second Circuit precedent, the evidence presented by the City is insufficient as a matter of law to justify its reliance on the challenged examinations.
Before proceeding to the legal analysis, I offer a brief word about the Supreme Court’s recent decision in
Ricci v. DeStefano,
- U.S. -, 129 S.Ct. 2658 , 174 L.Ed.2d 490 (June 29, 2009). I reference
Ricci
not because the Supreme Court’s ruling controls the outcome in this case; to the contrary, I mention
Ricci
precisely to point out that it does not. In
Ricci ,
the City of New Haven had set aside the results of a promotional examination, and the Supreme Court confronted the narrow issue of whether New Haven could defend a violation of Title VII’s disparate treatment provision by asserting that its challenged employment action was an attempt to comply with Title VII’s disparate impact provision. The Court held that such a defense is only available when “the employer can demonstrate a strong basis in evidence that, had it not taken the action, it would have been liable under the disparate-impact statute.”
Id.
at 2664 . In contrast, this case presents the entirely separate question of whether Plaintiffs have shown that the City’s use of Exams 7029 and 2043 has
actually had
a disparate impact upon black and Hispanic applicants for positions as entry-level firefighters.
Ricci
did not confront that issue.
The
Ricci
Court concluded that New Haven would not likely have been liable under a disparate impact theory.
See id.
at 2681 . In doing so, the Court relied on the various steps that New Haven took to validate its civil service examination.
Id.
at 2678-79 . It is noteworthy, however, that in this case New York City has taken significantly fewer steps than New Haven took in validating its examination. The relevant teaching of
Ricci ,
in this regard, is that the process of designing employment examinations is complex, requiring consultation with experts and careful consideration of accepted testing standards. As discussed below, these requirements are reflected in federal regulations and existing Second Circuit precedent. This legal authority sets forth a simple principle: municipalities must take adequate measures to ensure that their civil service examinations reliably test the relevant knowledge, skills and abilities that will determine which applicants will best perform their specific public duties.
In rendering this decision, I am aware that the use of multiple-choice examinations is typically intended to apply objective standards to employment decisions.
*84
Similarly, I recognize that it is natural to assume that the best performers on an employment test must be the best people for the job. But, the significance of these principles is undermined when an examination is not fair. As Congress recognized in enacting Title VII, when an employment test is not adequately related to the job for which it tests — and when the test adversely affects minority groups — we may not fall back on the notion that better test takers make better employees. The City asks the court to do just that. Regrettably, though, the City did not take sufficient measures to ensure that better performers on its examinations would actually be better firefighters. Accordingly, the court grants the Motions for Summary Judgment and finds that Plaintiffs have established disparate impact liability.
II. SUMMARY JUDGMENT STANDARD
“Summary judgment is appropriate when the pleadings and admissible evidence proffered to the district court show that there is ‘no genuine issue as to any material fact and that the moving party is entitled to a judgment as a matter of law
Major League Baseball Props., Inc. v. Salvino, Inc.,
542 F.3d 290, 309 (2d Cir.2008)
(quoting
Fed.R.Civ.P. 56(c)). “Material facts are those which ‘might affect the outcome of the suit under the governing law,’ and a dispute is ‘genuine’ if ‘the evidence is such that a reasonable jury could return a verdict for the nonmoving party.’ ”
Coppola v. Bear Stearns & Co.,
499 F.3d 144, 148 (2d Cir.2007)
(quoting Anderson v. Liberty Lobby, Inc.,
477 U.S. 242, 248 , 106 S.Ct. 2505 , 91 L.Ed.2d 202 (1986)). Factual disputes that are irrelevant or immaterial to the disposition of a case cannot preclude a grant of summary judgment.
See Loria v. Gorman,
306 F.3d 1271, 1282-83 (2d Cir.2002).
In considering a motion for summary judgment, the court construes the facts “in the light most favorable to the nonmoving party,” and draws “all reasonable inferences in its favor.”
SCR Joint Venture L.P. v. Warshawsky,
559 F.3d 133, 137 (2d Cir.2009). “[T]he moving party bears the burden of demonstrating the absence of a genuine issue of material fact.”
Baisch v. Gallina,
346 F.3d 366, 371-72 (2d Cir.2003)
(citing Celotex Corp. v. Catrett,
477 U.S. 317, 323 , 106 S.Ct. 2548 , 91 L.Ed.2d 265 (1986)). In response, the nonmoving party “ ‘must do more than simply show that there is some metaphysical doubt as to the material facts ....’”
Jeffreys v. City of New York,
426 F.3d 549, 554 (2d Cir.2005)
(quoting Matsushita Elec. Indus. Co. v. Zenith Radio Corp.,
475 U.S. 574, 586 , 106 S.Ct. 1348 , 89 L.Ed.2d 538 (1986)).
III. THE PRIMA FACIE CASE
Plaintiffs seek summary judgment on their prima facie case of disparate impact. As set forth below, summary judgment on Plaintiffs’ prima facie case is warranted. The facts set forth below are undisputed, unless otherwise noted.
A. The Hiring Process
During the relevant period, the hiring of entry-level firefighters from Exams 7029 and 2043 proceeded in several stages. Candidates interested in a position as an entry-level firefighter began by submitting an application to the Department of Citywide Administrative Services (“DCAS”), paid an application fee (unless it had been waived), and received an admission card for a written examination.
(See
USA 56.1 ¶¶ 13, 14; Int. 56.1 ¶ 9.) Each written examination was an 85-question, paper-and-pencil multiple choice test, and an applicant’s raw score on that examination was simply a percentage of the questions answered correctly.
(See
USA 56.1 ¶¶ 16, 23;
*85
Int. 56.1 ¶ 11.) A passing score was set for each examination, and after the results were in, the City notified each applicant of his or her score, as well as whether he or she had passed.
(See
Int. 56.1 ¶ 17.) Versions of each examination (with the same questions, but sometimes different question-orderings) were administered on repeated occasions- — Exam 7029 was administered from 1999 through 2002, and Exam 2043 was administered from 2002 through 2007.
(See
USA 56.1 ¶¶ 17-22.)
Candidates who passed the written examination were allowed to take the physical performance test (“PPT”), but those who failed the written examination could
not
take the PPT.
(See id.
¶ 27; Int. 56.1 ¶ 14.) The PPT consisted of eight physical tasks, and a candidate had to pass a minimum of six tasks to achieve a passing score overall. (Int. 56.1 ¶ 19.) A passing candidate’s score on the PPT was simply a percentage of the number of tasks successfully completed. For example, passing eight tasks resulted in a score of 100%, passing seven tasks resulted in a score of 87.5%, and passing six tasks resulted in a score of 75%.
(See id.)
Candidates who passed both the written examination and the PPT were placed on a “rank-order” eligibility list. (USA 56.1 ¶¶ 28-30; Seeley Deck app. I.) The ordering of the eligibility list was based upon an elaborate process of,
inter alia,
“standardizing,” “combining,” and “transforming” the raw scores.
(See
USA 56.1 ¶ 31; Int. 56.1 ¶ 20.) Specifically, the raw score from the written examination and the PPT would be “standardized” by subtracting the average score for all candidates from an individual candidate’s score and then dividing that number by the standard deviation for the test.
(See
Seeley Deck apps. J, K.) The resulting scores from both the written examination and the PPT would then be divided in half and added together to create a “Combined Weighted Standard Score.”
(Id.)
The Combined Weighted Standard Score was then converted into a “Transformed Score” by multiplying by either 18.472906403940886699 (for Exam 7029) or 12.7226 (for Exam 2043), and then adding either 83.74384236453 (for Exam 7029) or 88.4606 (for Exam 2043).
(Id.)
Finally, the “Adjusted Final Average,” used to rank candidates on the eligibility list, was created by adding any “Residency,” “Legacy,” or “Veteran” points to the Transformed Score.
(Id.; see
USA 56.1 ¶¶ 30, 31; Int. 56.1 ¶20.) This elaborate process resulted in a list of candidates eligible to be appointed to Academy classes in order of rank.
The DCAS and FDNY would determine how many candidates would be needed to fill an upcoming class and would certify a portion of the eligibility list for appointment, beginning with the highest scores. (USA 56.1 ¶¶ 40-41; Int. 56.1 ¶ 23.) The FDNY’s Candidate Investigation Division (“CID”) took steps to process and investigate candidates in order of their ranking on the eligibility list. (USA 56.1 ¶¶ 36-37, 39; Int. 56.1 ¶24.) This investigation involved,
inter alia,
background checks, intake interviews, and medical and psychological examinations by the FDNY’s Bureau of Health Services. (USA 56.1 ¶¶ 35-37; Int. 56.1 ¶24; Seeley Deck app. A (Request for Admission # 101).) The City would fill slots in an Academy class by proceeding down the list of eligible and qualified applicants until the class was filled; once a class was filled, any eligible and qualified candidate still remaining would not be appointed for that class. (USA 56.1 ¶¶ 42, 43; Int. 56.1 ¶25.)
Written Examination 7029 was first administered on February 26, 1999, and versions of it were administered as late as 2002. (USA 56.1 ¶¶ 17-19; Int. 56.1 ¶ 15.)
*86
Approximately 1,750 black applicants and approximately 2,125 Hispanic applicants sat for Exam 7029. (Siskin Report tbls. 1, 2.) The City hired from the eligibility list resulting from Exam 7029 from February 2001 through at least September 2004, and appointed over 3,200 entry-level firefighters from that examination. (USA 56.1 ¶¶ 11, 44.) Of this number, 104 (3.2%) individuals were black and 274 (8.5%) individuals were Hispanic.
(Id.
¶ 11.)
Written Examination 2043 was first administered on December 14, 2002, and versions of it were administered as late as March 2007.
(Id.
¶¶ 20-22; Int. 56.1 ¶ 16.) Approximately 1,390 black applicants and approximately 2,125 Hispanic applicants sat for Exam 2043. (Siskin Report tbls. 5, 6.) The City hired firefighters from the eligibility list resulting from Exam 2043 from May 2004 through at least January 2008, and, as of November 2007, the City had appointed over 2,100 entry-level firefighters from that examination. (USA 56.1 ¶¶ 12, 46.) Of this number, 80 (3.7%) were black and 187 (8.7%) were Hispanic.
(Id.
¶ 12.)
With these undisputed background facts in mind, the court addresses the prima facie case.
B. The Use of Statistics for a Prima Facie Case
A prima facie showing of disparate impact “requires plaintiffs to establish by a preponderance of the evidence that the employer ‘uses a particular employment practice that causes a disparate impact on the basis of race, color, religion, sex, or national origin.’ ”
Robinson v. Metro-North Commuter R.R. Co.,
267 F.3d 147, 160 (2d Cir.2001)
(quoting
42 U.S.C. § 2000e-2(k)(l)(A)(i)). “To make this showing, a plaintiff must (1) identify a policy or practice, (2) demonstrate that a disparity exists, and (3) establish a causal relationship between the two.”
Id.
at 160.
Statistics alone can make out the prima facie case.
See EEOC v. Joint Apprenticeship Comm. of Joint Indus. Bd. of Elec. Indus.,
186 F.3d 110, 117 (2d Cir.1999);
see also Robinson,
267 F.3d at 160 (“[Statistical proof almost always occupies center stage in a prima facie showing of a disparate impact claim.”). In order to do so, “[t]he statistics must reveal that the disparity is substantial or significant.”
Robinson,
267 F.3d at 160 (internal quotation marks omitted). “Moreover, the statistics must be of a kind and degree sufficient to reveal a causal relationship between the challenged practice and the disparity.”
Id.
“[A] plaintiff may establish a prima facie case of disparate impact discrimination by proffering statistical evidence which reveals a disparity substantial enough to raise an inference of causation. That is, a plaintiffs statistical evidence must reflect a disparity so great that it cannot be accounted for by chance.”
Joint Apprenticeship Comm.,
186 F.3d at 117 .
There are at least two widely recognized statistical measures of disparate impact: (1) the 80% or Four-fifths Rule, and (2) statistical significance or standard deviation analysis.
See Atkins v. Westchester County Dep’t of Soc. Serv.,
31 Fed.Appx. 52, 53 (2d Cir.2002) (summary order) (“[i]n evaluating disparate impact claims under Title VII, this Court has primarily relied upon [these] two methods of measuring disparities between groups”). Federal regulations set out the 80% Rule, and courts have recognized it as a “rule of thumb” for statistical analysis of disparate impact.
See, e.g., Joint Apprenticeship Comm.,
186 F.3d at 118 (“This rule is not binding on courts, and is merely a ‘rule of thumb’ to be considered in appropriate circumstances.”);
see also United States v.
*87
New York City Bd. of Educ.,
487 F.Supp.2d 220, 224 (E.D.N.Y.2007). The 80% Rule appears at 29 C.F.R. § 1607 .4D, which states:
A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by Federal enforcement agencies as evidence of adverse impact, while a greater than four-fifths rate will generally not be regarded by Federal enforcement agencies as evidence of adverse impact. Smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms ....
Id.
Essentially, this means that if the minority group performs less than 80% as well as the highest performing group, disparate impact will generally be inferred.
Courts have also relied upon standard deviation analysis (or statistical significance analysis) in determining whether there has been a disparate impact. “Standard deviation analysis measures the probability that a result is a random deviation from the predicted result — the more standard deviations the lower the probability the result is a random one.”
Waisome v. Port Auth. of New York & New Jersey,
948 F.2d 1370 , 1376 (2d Cir.1991);
see also
Barbara Lindemann
&
Paul Grossman, Employment Discrimination Law (“Lindemann”) 94 (3d ed. 1996) (“Tests of statistical significance are commonly used in the social sciences to rule out chance as the cause of observed disparities.”). “Basically, looking at standard deviations indicates how far an obtained result varies from an expected result.”
Smith v. Xerox Corp.,
196 F.3d 358, 365 (2d Cir.1999),
overruled on unrelated grounds by Meacham v. Knolls Atomic Power Lab.,
461 F.3d 134 , 141 (2d Cir.2006). The Second Circuit has “looked to whether the plaintiff can show a statistically significant disparity of two standard deviations” in making a prima facie showing.
Id.
at 365.
“Although courts have considered both the four-fifths rule and standard deviation calculations in deciding whether a disparity is sufficiently substantial to establish a prima facie case of disparate impact, there is no one test that always answers the question.”
Id.
at 366. According to either measurement, “the substantiality of a disparity is judged on a case-by-case basis.”
Id.
C. The Statistics for Plaintiffs’ Prima Facie Case
Plaintiffs allege that four employment practices related to the challenged examinations have had an unlawful disparate impact on black and Hispanic candidates for the position of entry-level firefighter. Specifically, they challenge the City’s use of:
(1) Written Examination 7029 as a pass/ fail screening device with a cutoff score of 84.705;
(2) Rank-order processing and selection of candidates from the Written Examination 7029 eligibility list based on a combination of their scores on Written Examination 7029 and the PPT;
(3) Written Examination 2043 as a pass/ fail screening device with a cutoff score of 70;
(4) Rank-order processing and selection of candidates from the Written Examination 2043 eligibility list based on a combination of their scores on Written Examination 2043 and the PPT.
(USA Mem. 1;
see
Int. Mem. 7.) As this court previously noted, “[t]he use of an examination as a ‘pass/fail screening de
*88
vice’ means the use [of] the examination to exclude from appointment those applicants that have failed the examination.” ( 258 F.R.D. 47, 52 (E.D.N.Y.2009).) “The use of an examination as a part of ‘rank-order processing’ means the use of the examination as a component of the overall score that determines an applicant’s position on a hierarchical hiring list.”
(Id.)
According to the Federal Government, the statistical significance analysis performed by its expert, Bernard R. Siskin, Ph.D., establishes a prima facie case of disparate impact. (USA Mem. 2.) Similarly, the Intervenors argue that the statistical significance analyses performed by Dr. Siskin and their expert, Joel P. Wiesen, Ph.D., establish a prima facie case. (Int. Mem. 7.)
The court briefly reviews the statistics that Plaintiffs have presented on each of the challenged practices. The City does not dispute the statistical calculations of Plaintiffs’ experts, but rather, disputes Plaintiffs’ reliance on statistical significance testing because of assumptions underlying that methodology. The court will address these assumptions after setting out the undisputed statistical calculations of Plaintiffs’ experts.
1. Pass/Fail Use of Exam 7029
The cutoff passing score for Written Examination 7029 was 84.705%. (USA 56.1 ¶ 24; Int. 56.1 ¶ 13.) Based on that cutoff score, the pass rate of white candidates for Exam 7029 was 89.9%, while the pass rate of black candidates was 60.3%. (USA 56.1 ¶ 83; Int. 56.1 ¶ 28.) In other words, out of 12,915 white test takers, 11,613 received a passing score of at least 84.705, whereas out of 1,749 black test takers, only 1,054 received a passing score.
(See
Int. 56.1 ¶27; Wiesen Report, tbl. 3a.) The pass rate of black candidates was, therefore, 67% of the pass rate of white candidates. (USA 56.1 ¶ 89.) Both Dr. Siskin’s and Dr. Wiesen’s standard deviation analysis found that this disparity is equivalent to 33.9 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 4.5 million-billion.
(See id.
¶ 85
(citing, inter alia,
Siskin Report 3, 21);
see also
Int. 56.1 ¶30
(citing, inter alia,
Wiesen Report 18-19).)
The practical effect of this disparity, according to Dr. Siskin, is that 519 black candidates who failed the examination— 74.7% of the black applicants who failed— were eliminated from consideration.
(See
USA 56.1 ¶¶ 86-87
(citing
Siskin Report).) Dr. Wiesen estimated that 457 black candidates would have passed the examination but for the effect of this disparity.
(See
Int. 56.1 ¶38
(citing
Wiesen Report).) Based on Dr. Siskin’s calculation, 114 additional black firefighters would have been appointed absent the disparity. (USA 56.1 ¶ 88
(citing
Siskin Report).) This last calculation was based on the assumption that the black applicants who failed Exam 7029 would have passed the PPT at the same rate as other similarly situated passers, and would have met the other qualifications and been appointed at the same rate as other passers.
10
(Siskin Report 16-17.)
The pass rate for Hispanic candidates taking Exam 7029 was 76.7%, compared with a pass rate of 89.9% for white candidates. (USA 56.1 ¶ 92.) Accordingly, the
*89
pass rate of Hispanic candidates was 85.3% of the pass rate of white candidates. (Sis-kin Report, tbl. 2.) Dr. Siskin’s standard deviation analysis found that this disparity is equivalent to 17.4 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 4.5 million-billion. (USA 56.1 ¶ 94 (citing,
inter alia,
Siskin Report 3, 23).)
The practical effect of this deviation, according to Dr. Siskin, is that 282 Hispanic candidates who failed the examination— 56.9% of the Hispanic applicants who failed — were eliminated from consideration.
(See id.
¶¶ 95-96
(citing
Siskin Report).) Based on Dr. Siskin’s calculation, 62 additional Hispanic firefighters would have been appointed absent the disparity.
(Id.
¶ 97
(citing
Siskin Report).) This last calculation was based on the assumption that the Hispanic applicants who failed Exam 7029 would have passed the PPT at the same rate as other similarly situated passers, and would have met the other qualifications and been appointed at the same rate as other passers. (Siskin Report 16-17.)
2. Pass/Fail Use of Exam 2043
The cutoff passing score for Written Examination 2043 was 70%. (USA 56.1 ¶ 25; Int. 56.1 ¶ 13.) Based on this cutoff score, the pass rate of white candidates taking Exam 2043 was 97.2%, while the pass rate of black candidates was 85.4%. (USA 56.1 ¶ 100;
see
Wiesen Report 42.) In other words, out of 13,877 white test takers, 13,495 received a passing score of at least 70, whereas, out of 1,393 black test takers, 1,190 received a passing score.
(See
Int. 56.1 ¶ 32; Wiesen Report, tbl. 16a.) The pass rate of black candidates was, therefore, 87.8% of the pass rate of white candidates. (Siskin Report, tbl. 5.) Both Dr. Siskin’s and Dr. Wiesen’s standard deviation analysis found that this disparity is equivalent to 21.8 units of standard deviation, meaning that the likelihood that it occurred by chance is less than 1 in 4.5 million-billion.
(See
USA 56.1 ¶ 102
(citing, inter alia,
Siskin Report 5, 26);
see also
Int. 56.1 ¶ 35
(citing, inter alia,
Wiesen Report 42-43).)
The practical effect of this deviation, according to Dr. Siskin, is that 165 black candidates who failed the examination— 81.3% of the black applicants who failed— were eliminated from consideration.
(See
USA 56.1 ¶¶ 103-04
(citing
Siskin Report).) Dr. Wiesen estimated that 150 black candidates would have passed the examination absent the disparity.
(See
Int. 56.1 ¶ 39
(citing
Wiesen Report).) Based op Dr. Siskin’s calculation, 30 additional black firefighters would have been appointed absent the disparity. (USA 56.1 ¶ 105
(citing
Siskin Report).) This last calculation was based on the assumption that the black applicants who failed Exam 2043 would have passed the PPT at the same rate as other similarly situated passers, and would have met the other qualifications and been appointed at the same rate as other passers. (Siskin Report 16-17.)
The pass rate for Hispanic candidates taking Exam 2043 was 92.8%, compared with a pass rate of 97.2% for white candidates. (USA 56.1 ¶ 108.) The pass rate of Hispanic candidates was, therefore, 95.5% of the pass rate of white candidates. (Sis-kin Report, tbl. 6.) Dr. Siskin’s standard deviation analysis found that this disparity is equivalent to 10.5 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 4.5 million-billion. (USA 56.1 ¶ 110
(citing, inter alia,
Siskin Report 5, 27).)
The practical effect of this deviation, according to Dr. Siskin, is that 94 Hispanic candidates who failed the examination— 61.8% of the Hispanic applicants who failed — were eliminated from consider
*90
ation.
(See id.
¶¶ 111-12
(citing
Siskin Report).) Based on Dr. Siskin’s calculation, 17 additional Hispanic firefighters would have been appointed absent the disparity.
(Id.
¶ 113
(citing
Siskin Report).) This last calculation was based on the assumption that the Hispanic applicants who failed Exam 2043 would have passed the PPT at the same rate as other similarly situated passers, and would have met the other qualifications and been appointed at the same rate as other passers. (Siskin Report 16-17.)
3. Rank-Ordering Use of Exam 7029
On the eligibility list created from Written Exam 7029 and the PPT, black candidates were grouped disproportionately lower down than white candidates. For example, only 10.1% of black candidates were in the top 20% of all candidates, while 53.8% of black candidates were in the bottom 40%, and 29.2% of black candidates were in the bottom 20%.
11
(USA 56.1 ¶ 120.) While 33% of white candidates had eligibility list numbers at or above 2000, only 21% of black candidates did, and while 20% of white candidates had list numbers at or below 5001, 30% of black candidates did.
(Id.
¶¶ 118-19.) According to Dr. Wiesen’s calculations, the average ranking of a black candidate was 630 ranking places lower than that of a white candidate, amounting to a disparity of 6.5 units of standard deviation.
(See
Wiesen Report, tbl. 9b.) According to Dr. Siskin’s calculation, the disparity between the placement of black and white candidates on the eligibility list is equivalent to 6.5 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 11 billion. (Siskin Report 24, 25; USA 56.1 ¶ 117.) Dr. Siskin calculated that, on account of this disparity, 68 out of 104 black candidates were delayed in appointment for an aggregate total of approximately 20 years of delayed wages and seniority. (USA 56.1 ¶ 122; Siskin Report, tbl. 3b.)
Similarly, the eligibility list created from the Exam 7029 results placed Hispanic candidates disproportionately lower down than white candidates. Only 14.3% of Hispanic candidates were in the top 20% of all candidates, while 47.8% of Hispanic candidates were in the bottom 40%, and 27.3% of Hispanic candidates were in the bottom 20%.
12
(USA 56.1 ¶ 130.) While 33% of white candidates had ■ eligibility list numbers at or above 2000, only 28% of Hispanic candidates did, and while 20% of white candidates had list numbers at or below 5001, 29% of Hispanic candidates did.
(Id.
¶¶ 128-29.) According to Dr. Siskin’s calculation, the disparity between the placement of Hispanic and white candidates on the eligibility list is equivalent to 4.6 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 204,000.
(Id.
¶ 127.) Dr. Siskin calculated that, on account of this dispari
*91
ty, 86 out of 274 Hispanic candidates were delayed in appointment for an aggregate total of approximately 23 years of delayed wages and seniority.
(Id.
¶ 132; Siskin Report, tbl. 4b.)
4. Rank-Ordering Use of Exam 2043
On the eligibility list created from Written Exam 2043 and the PPT, black candidates were grouped disproportionately lower down than white candidates. For example, only 11.4% of black candidates were in the top 20% of all candidates, while 56.9% of black candidates were in the bottom 40%, and 46.2% of black candidates were in the bottom 20%.
13
(USA 56.1 ¶ 139.) While 28% of white candidates had eligibility list numbers at or above 2000, only 18% of black candidates did, and while 30% of white candidates had list numbers at or below 5001, 50% of black candidates did.
(Id.
¶¶ 137-38.) According to Dr. Wiesen’s calculations, the average ranking of a black candidate was 974 ranking places lower than that of a white candidate, amounting to a disparity of 9.6 units of standard deviation.
(See
Wiesen Report, tbl. 22b.) According to Dr. Siskin’s calculation, the disparity between the placement of black and white candidates on the eligibility list was equivalent to 9.5 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 4.5 million-billion. (USA 56.1 ¶ 136.) Dr. Siskin calculated that, on account of this disparity, 44 out of 80 black candidates were delayed in appointment for an aggregate total of approximately fourteen years of delayed wages and seniority.
(Id.
¶ 146; Siskin Report, tbl. 12b.)
Similarly, the eligibility list created from the results of Exam 2043 placed Hispanic candidates disproportionately lower down than white candidates. Only 17.2% of Hispanic candidates were in the top 20% of all candidates, while 45.4% of Hispanic candidates were in the bottom 40%, and 24.6% of Hispanic candidates were in the bottom 20%.
14
(USA 56.1 ¶ 152.) While 28% of white candidates had eligibility list numbers at or above 2000, only 25% of Hispanic candidates did, and while 30% of white candidates had list numbers at or below 5001, 39% of Hispanic candidates did.
(Id.
¶¶ 150-51.) According to Dr. Siskin’s calculation, the disparity between the placement of Hispanic and white candidates on the eligibility list is equivalent to 4.6 units of standard deviation, meaning that the likelihood it occurred by chance is less than 1 in 186,225.
(Id.
¶ 149.) Dr. Siskin calculated that, on account of this disparity, 51 out of 187 Hispanic candidates were delayed in appointment for an aggregate total of approximately twelve years of delayed wages and seniority.
(Id.
¶ 158; Sis-kin Report, tbl. 14b.)
Dr. Siskin also conducted tests addressing the fact that the eligibility list from Exam 2043 was not exhausted, and, therefore, candidates very low down on that list were never reached. Dr. Siskin determined that those who were never reached “effectively failed” the examination pro
*92
cess.
(See
USA 56.1 ¶ 160.) According to Dr. Siskin’s calculations, out of 95 black candidates and 63 Hispanic candidates who would have ranked high enough to be considered for hire, 42 of the black candidates and 28 of the Hispanic candidates would have been appointed absent a disparity resulting from Written Exam 2043. (Sis-kin Report 33-35; USA 56.1 ¶¶ 145, 157.) Based on the hiring rates of candidates from the Exam 2043 eligibility list, he calculated that the disparity of hiring rates between white and black candidates amounted to 9.7 units of standard deviation, while the disparity of hiring rates between white and Hispanic candidates amounted to 5 units of standard deviation. (USA 56.1 ¶¶ 142,155.)
Dr. Siskin utilized the data relating those who effectively failed in order to calculate an “effective pass rate” for Written Examination 2043, determined to be 70.3% for white candidates, 41.5% for black candidates, and 58.9% for Hispanic candidates.
(Id.
¶¶ 162-63, 167.) This amounts to a statistical disparity between white and black candidates of 21.9 units of standard deviation, and a statistical disparity between white and Hispanic candidates of 10.5 units of standard deviation.
(Id.
¶¶ 164,168.)
D. The Parties’ Respective Positions
Plaintiffs argue that the presented statistics establish a prima facie case of disparate impact for the four challenged employment practices.
(See
USA Mem. 10-17; Int. Mem. 6-12.) They point out that the calculated disparities between black and minority candidates resulting from the challenged practices are much greater than three units of standard deviation. They also emphasize the practical significance of these disparities — for example, the Federal Government relies on the statistical analyses showing that, but for the disparities resulting from the written examinations, “1,060 additional black and Hispanic candidates would have been considered for appointment as FDNY firefighters,” “an estimated 293 additional black and Hispanic candidates would have been appointed as FDNY firefighters,” and “249 black and Hispanic firefighters who were appointed — about 39% of those appointed from the examinations at issue in this case — would have been appointed earlier.” (USA Mem. 2-3.) Finally, Plaintiffs argue that the City has conceded these statistical conclusions.
(See, e.g., id.
at 3.)
In opposition to the Motions, the City offers several iterations of the same basic argument. In essence, the City asks the court to reject Plaintiffs’ statistical significance analysis because it improperly assumes “perfect parity” among groups of people
(see
Def. PF Mem. 1-3, 5-7), and erroneously produces a finding of disparate impact solely on account of large sample sizes
(see id.
at 1, 5, 6, 7). The City asks the court to rely exclusively upon the 80% Rule in determining whether there has been a disparate impact between white and minority candidates.
(See id.
at 1, 2-3.) Because application of this statistical rule would result in a finding of disparate impact for some, but not all, of the challenged employment practices, the City asks the court to deny summary judgment relating to those practices that do not meet the 80% Rule.
(See id.
at 7-8.) The City does not contest the specific calculations in Plaintiffs’ Rule 56.1 Statements, instead attacking the assumptions on which they rely, and denying the “materiality” of the facts presented.
Plaintiffs respond that large sample sizes do not undermine the validity of statistical significance testing; rather, they argue, it is small sample sizes that render statistical significance tests less reliable.
*93
(See
USA Mem. 18-20; Int. Mem. 10.) Plaintiffs also argue that there is no basis for relying on the 80% Rule to the exclusion of statistical significance testing, and that, in fact, all legal authority is to the contrary. (USA Mem. 20-22; USA Reply 3-7; Int. Mem. 13 n. 10; Int. PF Reply 2-4.)
Before addressing the parties’ respective positions, the court notes that the dispute regarding the proper statistical measurement for disparate impact does not relate to all of the challenged employment practices. It is undisputed that the City’s pass/fail use of Exam 7029 has had a disparate impact upon black candidates under both statistical significance testing
and
under the 80% Rule.
(See
USA 56.1 ¶¶ 59, 83, 89.) Moreover, the City offers the 80% Rule only as a means of comparing pass rates, not rank-ordering.
(See
USA 56.1 ¶¶ 67, 69.) It has not presented an alternative statistical measure for the rank-ordering of candidates.
(See
Def. PF Mem. 7-8.) The City’s preference for the 80% Rule, therefore, solely relates to the pass/fail uses of Exam 2043 with respect to black candidates, and the pass/fail uses of Exam 7029 and 2043 with respect to Hispanic candidates.
E. Plaintiffs Have Demonstrated a Prima Facie Case
Plaintiffs have demonstrated a prima facie case of disparate impact by (1) identifying four specific employment practices (each relating to both black and Hispanic applicants), (2) demonstrating that a disparity exists among groups, and (3) establishing a causal relationship between the employment practices and the disparities.
See Robinson,
267 F.3d at 160 . For each employment practice, Plaintiffs have presented analyses from two experts that thoroughly demonstrate the statistical significance of the disparities between groups of candidates. For each of the pass/fail uses of the examinations, these analyses demonstrate that the disparities between the pass rates of whites and minority candidates were between 10.5 and 33.9 units of standard deviation. For each of the rank-ordering uses of the examinations, the analyses demonstrate that the disparities between the rankings of whites and minority candidates were between 4.6 and 9.7 units of standard deviation. These statistical disparities show that black and Hispanic candidates disproportionately failed Written Exams 7029 and 2043, and were placed disproportionately lower on the eligibility lists created from those examinations.
The Second Circuit has repeatedly recognized that standard deviations of more than 2 or 3 units can give rise to a prima facie case of disparate impact because of the low likelihood that such disparities have resulted from chance.
See Malave v. Potter,
320 F.3d 321, 327 (2d Cir.2003) (“courts ‘generally consider this level of significance [i.e., two standard deviations] sufficient to warrant an inference of discrimination.’ ”)
(quoting Smith,
196 F.3d at 365 );
Waisome,
948 F.2d at 1376 (“[a] finding of two or three standard deviations (one in 384 chance the result is random) is generally highly probative of discriminatory tréatment”);
Ottaviani v. State Univ. of New York,
875 F.2d 365 , 372 (2d Cir.1989) (“It is certainly true that a finding of two to three standard deviations can be highly probative of discriminatory treatment.”);
Guardians Assoc. of New York City Police Dep’t, Inc. v. Civil Serv. Comm.,
630 F.2d 79, 86 (2d Cir.1980)
(“Guardians
”) (“[I]n cases involving large samples, ‘if the difference between the expected value (from a random selection) and the observed number is greater than two or three standard deviations,’ a prima facie case is established.”)
(quoting Castaneda v. Partida,
430 U.S. 482 , 496 n. 17, 97 S.Ct. 1272 , 51
*94
L.Ed.2d 498 (1977)). The calculated standard deviations in this case are all well beyond 2 to 3 units, strongly supporting a conclusion of a causal relationship between the observed disparities and the employment practices at issue.
The significance of Plaintiffs’ statistics is bolstered by evidence that the disparities have been significant as a practical matter.
See
Lindemann 94 (“To guard against the possibility that a finding of adverse impact could result from the statistical significance of a trivial disparity or meaningless difference in results, the Uniform Guidelines on Employee Selection Procedures[, 29 C.F.R. § 1607 .4D,] and the courts have adopted an additional test for adverse impact: that a statistically significant disparity also has practical significance.”). As mentioned above, approximately one thousand additional black and Hispanic candidates would have been considered for appointment as FDNY firefighters had it not been for the disparities resulting from the examinations. Further, absent these disparities, approximately 293 additional black and Hispanic candidates would have been appointed from the eligibility lists used from 2001 through 2008, and approximately 249 black and Hispanic applicants who were actually appointed would have been appointed sooner. Given that, in 2007, the FDNY had 8,998 firefighters, including only 303 black firefighters and 605 Hispanic firefighters
(see
Seeley Decl. app. C), it is clear that these disparities have a substantial practical significance. In fact, the disparities are overwhelming.
The accuracy of Plaintiffs’ statistical calculations is not disputed, and the City’s Responses to Plaintiffs’ 56.1 Statements essentially concede the statistical picture establishing a prima facie case. The City specifically concedes that: (1) the disparity created by each of the challenged practices is more than three units of standard deviation
(see
Def. USA 56.1 ¶¶84, 93, 101, 109, 116, 126, 135, 148; Def. Int. 56.1 ¶¶ 30, 31, 35, 37), (2) the Plaintiffs’ calculations of statistical significance are “undisputed”
(see, e.g.,
Def. USA 56.1 ¶¶ 85, 94, 102, 110, 117, 127, 136, 149), and (3) “[o]ne of the City’s experts conducted analyses to attempt to verify Dr. Siskin’s statistical calculations and confirmed the results reported by Dr. Siskin” (USA 56.1 ¶ 77; Def. USA 56.1 ¶ 77). Moreover, the City has essentially admitted the calculations performed by Plaintiffs’ experts showing the disparities’ practical significance.
15
(See
Def. USA 56.1 ¶¶ 86-88, 95-97, 103-105, 111-13, 122, 132, 145, 157; Def. Int. 56.1 ¶¶ 37-39;
see also
Def. USA 56.1 ¶¶ 82, 90, 98, 106 (accepting as undisputed that “a test of statistical significance ... can result in a finding of disparate impact” for pass/fail uses);
id.
¶¶ 114, 123, 134, 147 (accepting as undisputed finding of disparate impact for rank-ordering uses “assuming use of a test of statistical significance”).) These admissions eliminate the existence of any factual dispute over the prima facie case.
F. The City’s Arguments
1.
Large Sample Sizes
Rather than attacking the accuracy of Plaintiffs’ statistics, the City objects to Plaintiffs’ reliance on statistical signifi
*95
canee testing as a general matter. The City raises a number of supposed theoretical problems with such testing. The City’s principal argument is that the size of the populations being tested in this case
(i.e.,
the many thousands of applicants who took each examination) renders a statistical significance test unreliable. This is because, the City contends, the “larger the group we examine the more likely we are to find differences[ ]” among candidates that will cause particular individuals to fail. (Def. PF Mem. 5 (noting that “the larger the group we are examining, the more candidates who sit for the exam, the greater our likelihood that some of them will not do as well as others”).)
The City has it backwards. Rather than undermining confidence in statistical significance testing, large sample sizes make such testing more rehable. Larger sample sizes create a greater likelihood that random differences between individuals will even out among all groups, and a lower likelihood that significant differences between the performance of racial or ethnic groups will have resulted from chance. Existing precedent confirms this principle. Courts have sometimes declined to rely on statistical significance analysis when a sample size was too small.
See, e.g.,
Lindemann 1734 (“Courts have recognized that statistical evidence often is unreliable when the sample size is small.”);
Pietras v. Bd. of Fire Comm’rs,
180 F.3d 468, 475 (2d Cir.1999) (recognizing “authority holding that a disparate impact finding based solely on a sample size as small as the one presented here
[ie.,
7 people] cannot stand”). Yet, the City has pointed to no cases rejecting such testing because a sample size was too large. As the Second Circuit stated in
Guardians, “in cases involving large samples,
‘if the difference between the expected value (from a random selection) and the observed number is greater than two or three standard deviations,’ a prima facie case is established.” 630 F.2d at 86 (emphasis added)
(quoting Castaneda,
430 U.S. at 496 n. 17, 97 S.Ct. 1272 ).
The City’s own admissions support this understanding of large sample sizes. The Federal Government has provided a helpful illustration that the City explicitly accepts as undisputed:
Flipping a coin is a common example that illustrates why sample size
should
affect the number of standard deviations that is equivalent to a given disparity. Flipping a fair coin 10 times will not always result in exactly five heads and five tails; a result of six heads and four tails on ten flips would not indicate with a reasonable degree of certainty that the coin was not fair
(ie.,
that the disparity was not likely due to chance variation). However, if one flipped a fair coin 1,000 times, one would expect that the number of heads and tails would be close to equal, and a result of 600 heads and 400 tails would allow one to conclude with a high degree of certainty that the coin was not fair
(ie.,
that disparity between the rate at which heads came up and the rate at which tails came up was not likely do to chance variation).
Put simply, with a disparity in pass rates of a given size, the bigger the sample
(e.g.,
the more times one flips the coin, or the more applicants who take the test), the more confident one can be that the difference in pass rates in the sample is not due to chance.
(USA 56.1 ¶¶ 56-57 (internal citations omitted);
see
Def. USA 56.1 ¶¶ 56-57.)
Another undisputed statement, which relies on the City’s own expert, further supports the reliability of statistical significance testing when large sample sizes are involved:
*96
With a large sample size, a test of statistical significance using 1% as the standard
(i.e.,
concluding that there is a statistically significant disparity if there is no more than a
1%
likelihood of observing a disparity so large due to chance)
is better than the 80% Rule
at controlling for false positives (situations in which the test used will indicate a disparity when there is no disparity) and false negatives (situations in which the test will indicate there is no disparity when there is a disparity). In other words,
with a large sample size, a test of statistical significance is more likely to produce the “right” answer to the question of whether there is a non-chance disparity between the pass rates of two groups.
(USA 56.1 ¶ 64
(citing
deposition of City’s expert, Dr. Bobko) (emphases added);
see
Def. USA 56.1 ¶ 64.) These undisputed statements plainly support a conclusion that large sample sizes enhance, rather than undermine, the reliability of statistical testing. Accordingly, while the City purports to challenge the use of statistical significance testing based on sample size, the City’s own admissions contradict its position.
16
When an employment examination is used to make hiring decisions for thousands of applicants, seemingly small differences in pass rates can have a substantial effect on large groups of people. Contrary to the City’s position, therefore, it is important to rely upon statistical testing to determine whether such differences have resulted from chance or, rather, from a particular employment practice. In this case, statistical significance testing has been used to show that the disparities between groups of candidates have resulted from the challenged examinations. The City’s arguments to the contrary are unavailing.
2.
Perfect Parity Among Groups
The City also attacks Dr. Siskin’s “shortfall” analysis, which estimates the number of minority candidates who would have passed or been appointed had the written examinations not had a discriminatory impact. The City criticizes the fact that such calculations hypothesize a world of “perfect parity” among racial or ethnic groups. (Def. PF Mem. 5.) In other words, the City argues that Plaintiffs’ analyses inappropriately compare the racial disparity in test results to a hypothetical world in which racial and ethnic groups perform equally well.
(See id.
(“[Statistical significance testing will assume that all people perform at equal levels. However, we know that all individuals do not perform at the same level.”).) This argument misstates the law.
First of all, the court rejects the premise that comparison to a standard of equality among groups provides an improper foundation for statistical testing under Title VII. In order to determine whether a particular employment practice has had a disparate impact on a minority group, statistical tests “ask what the results
would be
for the salient variable ... if there [had been] no discrimination.” Adams
v. Ameritech
*97
Servs., Inc.,
231 F.3d 414, 424 (7th Cir. 2000) (emphasis added). To determine what results “would be,” statistical tests properly assume that racial or ethnic groups will perform equally well absent discrimination.
See Smith,
196 F.3d at 366 (recognizing “null hypothesis” of no difference between compared groups).
17
Statistical significance testing relies on this assumption of equality in assessing whether disparities among groups are based upon chance, or rather, upon some other factor, such as race or national origin.
See Ottaviani, 875
F.2d at 371 (“Statistical significance is a measure of the probability that a disparity is simply due to chance, rather than any other identifiable factor.”);
Joint Apprenticeship Comm.,
186 F.3d at 117 (“a plaintiffs statistical evidence must reflect a disparity so great that it cannot be accounted for by chance”); Lindemann 94 (“Tests of statistical significance are commonly used in the social sciences to rule out chance as the cause of observed disparities.”). In accordance with these principles, Plaintiffs’ statistical evidence shows that the disparities in this case have not been the result of chance; instead, the disparate impact upon black and Hispanic candidates has resulted from the challenged employment practices.
The court similarly rejects the suggestion in the City’s Rule 56.1 Responses that a prima facie case has not been established because disparities between white, black, and Hispanic candidates can be explained by differences in their “capability and preparedness.”
{See, e.g.,
Def. USA 56.1 ¶¶ 86-88, 95-97, 103-05, 111-13.) The City’s Rule 56.1 Response suggests that the City believes black and Hispanic candidates received lower scores on its written examinations because of their lower capability and preparedness for the job of firefighter. But, if the City contends that differences in aptitude relating to the job of firefighter have led to an adverse impact on minority groups, Title VII’s burden-shifting framework allows the City to justify the disparate impact as a matter of business necessity. At the prima facie stage, however, the question is only whether there are disparities attributable to the challenged practices, not whether the City can provide a justification for them. During this stage, the City cannot rebut the existence of disparities by claiming that they are explained by the overall “capability and preparedness” of particular groups.
Therefore, the court rejects the City’s arguments about the assumptions in Plaintiffs’ statistics.
3.
The 80% Rule
Finally, in opposing summary judgment, the City argues that the court should rely on the 80% Rule to the exclusion of statistical significance testing. There is no support for this position. Controlling precedent holds that the 80% Rule is not an exclusive means of proof, and that alternative statistical tests should be considered.
See Joint Apprenticeship Comm.,
186 F.3d at 118 (“[80%] rule is not binding on courts, and is merely a ‘rule of thumb’ to be considered in appropriate circumstances.”);
see also Watson v. Fort Worth
*98
Bank & Trust,
487 U.S. 977 , 995 n. 3, 108 S.Ct. 2777 , 101 L.Ed.2d 827 (1988) (citing criticism of the 80% Rule, recognizing the usefulness of statistical methods in Title VII cases, and endorsing a case-by-case approach);
Bew v. City of Chicago,
252 F.3d 891, 893 (7th Cir.2001) (“The district court properly noted that the 80% guideline may be ignored when other statistical evidence indicates a disparate impact.”). Moreover, the regulation containing the 80% Rule plainly endorses the use of alternative tests. It specifically states: “Smaller differences in selection rate may nevertheless constitute adverse impact,
where they are significant in both statistical and practical
terms.... ” 29 C.F.R. § 1607 .4D (emphasis added). Rather than supporting the use of the 80% Rule to the exclusion of statistical significance testing, this sentence expressly contemplates alternative statistical tests as a means of showing disparate impact. This is precisely the showing Plaintiffs have made in this case.
In support of its preference for the 80% Rule, the City points to the Second Circuit’s decision in
Waisome.
In
Waisome,
the plaintiffs challenged examinations used by the Port Authority in the process of promoting police officers! 948 F.2d at 1372. The standard deviation between the pass rates of white and black candidates in that case was 2.68, and the district court concluded that this was insufficient to find disparate impact because the pass rate of blacks was 87.2% the pass rate of whites.
Id.
at 1375. In approving this part of the district court decision, however, the Second Circuit did not find that the 80% Rule controlled the outcome of the case. Instead, the court also relied upon the fact that the calculated statistical significance, 2.68 standard deviations, was a borderline figure, and that, as a practical matter, “if two additional black candidates passed the written examination the disparity would no longer be of statistical importance.”
Id.
at 1376.
The statistics in this case, however, are uniformly beyond 3 units of standard deviation, and, for many of the analyses performed, drastically beyond that. There are also large sample sizes, which make a finding of statistical significance more reliable. Hundreds of black and Hispanic applicants were affected by the City’s written examinations. Given the practical significance of the disparities here on the actual hiring rates for black and Hispanic applicants, this case is clearly distinguishable from
Waisome.
Finally, it is worth noting that the Second Circuit remanded the case in
Waisome
because it found that there had been a showing of disparate impact.
Id.
at 1372.
In sum, the City has conceded the accuracy of the calculations of Plaintiffs’ experts, which provide ample support for the statistical and practical significance of the disparities at issue. The City’s only defense is to resort to abstract arguments relating to the nature of statistical testing in general.
(See, e.g.,
Mar. 19, 2009 Tr. 29-30 (arguing against statistical significance testing based upon Plato’s “allegory of the cave”).) At the same time, the City has made admissions in its Responses to Plaintiffs’ Rule 56.1 Statements that directly contradict its only purported challenges to Plaintiffs’ proof. Considering its crucial admissions and the lack of legal authority for its position, the City has failed to wage a serious attack on Plaintiffs’ prima facie case.
Under these circumstances, the court finds no material factual dispute relating to the prima facie case. To the extent the City purports to dispute the factual evidence presented by Plaintiffs, it has raised nothing more than metaphysical doubts about the nature of that evidence. Such doubts cannot preclude summary judgment.
See Matsushita,
475 U.S. at 586 ,
*99
106 S.Ct. 1348 . There is no dispute that Plaintiffs have satisfied a statistical standard for a prima facie case of disparate impact that has been repeatedly accepted by the Second Circuit, nor is there any dispute that they have shown that the disparity has had a substantial, practical significance for the composition of the eligibility lists and hiring of entry-level firefighters. Accordingly, the court grants summary judgment for Plaintiffs on their prima facie case of disparate impact.
IV. BUSINESS NECESSITY
While Plaintiffs have shown that the City’s uses of Written Examinations 7029 and 2043 resulted in a disparate impact upon black and Hispanic candidates, the City may defend against Title VII liability by showing that those uses were justified by legitimate business and job-related considerations.
18
The City bears the burden of making this showing.
19
See Gulino v. New York State Educ. Dep’t,
460 F.3d 361 , 385 (2d Cir.2006). In
Gulino,
the Second Circuit explained the business necessity defense as follows:
[T]he basic rule has always been that “discriminatory tests are impermissible unless shown, by professionally acceptable methods, to be predictive of or significantly correlated with important elements of work behavior which comprise or are relevant to the job or jobs for which candidates are being evaluated.” This rule operates as both a limitation and a license for employers: employers have been given explicit permission to use job related tests that have a disparate impact, but those tests must be “demonstrably a reasonable measure of job performance.”
460 F.3d at 383
(quoting Albemarle Paper Co. v. Moody,
422 U.S. 405, 431, 426 , 95 S.Ct. 2362 , 45 L.Ed.2d 280 (1975)).
*100
In setting forth its analysis, this court first reviews the process by which the City created the two challenged examinations. The court then addresses a motion to strike post-discovery submissions made by the City relating to its business necessity defense. Finally, the court sets out and applies the relevant Second Circuit standard for assessing the challenged examinations. The court concludes that the City has not met that standard.
A. Creation of Challenged Examinations
The City used the same development process for Written Exams 7029 and 2043. For Written Exam 7029, a “Test Development Report” was prepared by Matthew Morrongiello (“Morrongiello”), a Tests and Measurement Specialist in the City’s DCAS.
(See
Levy Decl. Ex. EE (“Test Development Report”).) Alberto Johnston (“Johnston”) of DCAS was “primarily responsible” for developing Exam 2043, and testified that he “was told that we probably can use the old job analysis ... from [Exam] 7029 .... ” (Fraenkel Decl. Ex. 6. (“Johnston Dep.”), at 17-18, 19.) Accordingly, the same Test Development Report was relied upon in the City’s development of Written Exam 2043.
(See
Int. 56.1 ¶ 12;
see also id.
¶¶ 85-86; Levy Decl. Ex. CC, at 5 (Admission # 11); Fraenkel Decl. Ex. 10 (“Patitucci Dep.”), at 208-09.)
20
The Test Development Report is a twelve-page document with nine appendices that sets out the process by which the City arrived at the abilities it intended to evaluate on the written examinations, as well as the number of questions it would devote to each ability. The salient features of the process, as set out in that Report, are not in dispute.
21
The Test Development Report states that a series of meetings were held, pursuant to which DCAS concluded that it would conduct a new, “comprehensive” job analysis, determine how many people would be needed for a “job analysis survey,” and convene a panel of 8 to 12 incumbent firefighters to review the task and ability lists and the job analysis questionnaire.
22
(See
Test Development Report 4.) The basic plan for the
*101
process was as follows: first, the tasks and abilities most relevant to the job of firefighter would be determined; second, the relative importance of these tasks and abilities would be assessed; third, clusters of tasks would be matched up with the abilities needed to perform them; and finally, a test would be created to evaluate the identified abilities in the proper proportions. The court briefly reviews this process, as set forth in the Test Development Report.
1.
Deriving a List of Tasks and Abilities
In order to familiarize himself with the job of firefighter, Morrongiello conducted interviews with six incumbent firefighters.
(Id.
at 5.) He used the results of these interviews, coupled with the task list used to develop a prior examination, Exam 0084, to create an “updated task list” reflecting the tasks performed by firefighters.
(Id.)
Morrongiello then compiled a list of 21 cognitive abilities, derived from “Fleishman’s ability list.” (See Test Development Report 5; Levy Deck Ex. X, at 129.) Morrongiello convened a focus group of ten firefighters who reviewed the task and ability lists, and, based on the focus group, “suggested changes to the proposed task list so it more accurately reflected the current job of Firefighter.” (Test Development Report 5.)
The focus group also reviewed a Job Analysis Questionnaire (“JAQ”) that was distributed to 195 incumbent firefighters.
(See id.
at 5, 6, 10.) The JAQ asked the firefighters to assess whether 196 listed tasks (grouped into 21 specific task clusters, plus a miscellaneous cluster) were “4. Critical,” “3. Important,” “2. Somewhat important,” or “1. Not [] important” to the job of firefighter, and, similarly, whether 21 listed abilities were “4. Critical,” “3. Important,” “2. Somewhat important,” or “1. Not relevant” to the firefighter job.
(Id.
at 10
&
app. E (JAQ test results).) The responses to the JAQ by 192 of the 195 surveyed firefighters, excluding 3 defective responses, were used to hone the list of abilities down to 18, and the list of tasks down to 111.
(Id.
at 10.) This was done by removing those tasks and abilities which did not receive at least an average score of 2.5, corresponding to a rating between “important” and “somewhat important.”
(Id.)
2.
Linking Tasks With Abilities
With these results in hand, Morrongiello assembled twelve firefighters into a “Linking Panel,” whose purpose was to “rate— or link — the task clusters to the abilities.”
(Id.)
The specific “clusters” of tasks important to the job of entry-level firefighter were:
•
Initial Response to Incidents/Driving
(tasks that “occur between receiving an alarm and initial fire fighting or emergency activities, including driving apparatus to and from various points”);
•
Size Up
(tasks that involve “evaluating the fire or incident scene to determine actions which should initially be taken and obtaining information needed for evaluation”);
• Ladder Operation
(tasks that involve “stabilizing ladder trucks and elevating and operating aerial ladders and platforms in order to rescue victims, provide access for ventilation, operate master stream devices, etc.”);
•
Climbing and Portable Ladder Activities
(tasks that involve “climbing ladders, stairs and fire escapes, and raising and setting up portable ladders”);
• Building Entry
(tasks that involve “prying open or breaking through doors or otherwise entering buildings in order to search for and rescue victims and provide access to the fire for offensive fire fighting, using axes, hal
*102
ligan tool, hooks, rabbit tools, sledge hammers, power saws, and other tools”);
• Search
(tasks that involve “searching fire or assigned area in order to locate victims and to obtain further information about fire, following standard search procedures”);
• Rescue
(tasks that involve “assisting, carrying or dragging victims from emergency area by means of interior access (stairs, hallways, etc.) or, if necessary, by ladders, fire escapes, platforms, or other means of escape”);
• Ventilation
(tasks that involve “opening or breaking open windows, chopping or cutting holes in roofs, breaking through walls or doors, and hanging fans in windows or doors to remove heat, smoke and gas from burning buildings”);
• Supplies Water for Hose Operation
(tasks that involve “connecting or hooking up engine to fire hydrant and operating pumps to supply water in appropriate pressure and volume for fire fighting, using hydrant wrenches, couplings, hoses, spanner wrenches, and other tools”);
• Hose Operations During Extinguishment
(tasks that involve “stretching line to fire scenes and delivering water to scene of fire”);
•
Overhaul
(tasks that involve “opening up walls and ceilings, cutting or pulling up floors and moving or turning over debris, in order to check for hidden fires which could rekindle or spread, using hooks, axes, saws and pitchforks”);
• Salvage
(tasks that involve “moving and covering furniture, appliances, merchandise and other property, and covering holes in buildings and redirecting or cleaning up water in order to minimize damage, using plastic and canvas covers, ropes, staple guns, mops, squeegees, and other tools”);
•
Clean Up/Pick Up
(tasks that involve “picking up and returning equipment to vehicle and rolling up or folding up hose, so that the company can go back in service”);
• Equipment Maintenance
(tasks that involve “inspecting, cleaning, and maintaining apparatus, equipment carried on the apparatus, and personal gear and equipment”);
•
Inspection of Buildings/Hydrants
(tasks that involve “inspecting buildings for code violations or hazards on a periodic basis or during the course of activities, and inspecting hydrants for operational use”);
• Extrication
(tasks that involve “extricating victims from vehicles, cave-ins, collapsed buildings or other entrapments in order to save lives, using shovels, torches, drills, pry bars, saws, jacks, hurst tools, air bags, and other equipment”);
• Providing Medical Assistance
(tasks that involve “providing first aid and direct medical assistance to persons requiring emergency attention”);
•
Elevator Related Tasks
(tasks that involve “controlling elevators and rescuing persons from stalled elevator cars”);
•
Training
(tasks that involve “participating in drills which simulate important fire or rescue activities, and attending lectures or formal training”);
•
Watch Duties
(tasks that involve “standing watch to receive incoming alarms and information, answering phones, and monitoring access to the station house”);
•
Station Duties and Chores
(tasks that involve “performing routine house
*103
keeping chores or ‘committee work’ ”); and
• Miscellaneous
(tasks that involve “miscellaneous tasks”).
{See id.
app. E, at USA000371-USA000380;
see also
Levy Deck Ex. FF (“Linking Panel Worksheet”).) The 18 abilities to which members of the Linking Panel were supposed to “link” these tasks were:
•
Oral Comprehension
(the ability to “understand spoken English words and sentences”);
•
Written Comprehension
(the ability to “understand written sentences and paragraphs”);
•
Oral Expression
(the ability to “use English words or sentences in speaking so that others will understand”);
•
Written Expression
(the ability to “use English words or sentences in writing so that others will understand”);
•
Fluency of Ideas
(the ability to “produce a number of ideas about a given topic”);
•
Originality
(the ability to “produce unusual or clever ideas about a given topic or situation,” and to “invent creative solutions to problems or to develop new procedures for situations in which standard operating procedures do not apply”);
•
Memorization
(the ability to “remember information, such as words, numbers, pictures and procedures. Pieces of information can be remembered by themselves or with other pieces of information”);
•
Problem Sensitivity
(the ability to “tell when something is wrong or is likely to go wrong. It includes being able to identify the whole problem as well as elements of the problem”);
• Deductive Reasoning
(the ability to “apply general rules to specific problems to come up with logical answers. It involves deciding if an answer makes sense”);
• Inductive Reasoning
(the ability to “combine separate pieces of information, or specific answers to problems, to form general rules or conclusions. It involves the ability to think of possible reasons for why things go together”);
• Information Ordering
(the ability to “follow correctly a rule or set of rules or actions in a certain order. The rule or set of rules used must be given. The things or actions to be put in order can include numbers, letters, words, pictures, procedures, sentences, and mathematical or logical operations”);
•
Speed of Closure
(“involves the degree to which different pieces of information can be combined and organized into one meaningful pattern quickly. It is not known beforehand what the pattern will be. The material may be visual or auditory”);
• Flexibility of Closure
(the ability to “identify or detect a known pattern (like a figure, word, or object) that is hidden in other material. The task is to pick out the disguised pattern from the background material”);
•
Spatial Orientation
(the ability to “tell where you are in relation to the location of some object or to tell where the object is in relation to you”);
•
Visualization
(the ability to “imagine how something would look when it is moved around or when its parts are moved or rearranged. It requires the forming of mental images of how patterns or objects would look after certain changes, such as unfolding or rotation. One has to predict how an
*104
object, set of objects, or pattern will appear after the changes have been carried out”);
• Perceptual Speed
(“involves the degree to which one can compare letter, numbers, objects, pictures, or patterns, quickly and accurately. The things to be compared may be presented at the same time or one after the other. This ability also includes comparing a presented object with a remembered object”);
•
Selective Attention
(the ability to “concentrate on a task one is doing. This ability involves concentrating while performing a boring task and not being distracted”); and
•
Time Sharing
(the ability to “shift back and forth between two or more sources of information”).
{See
Test Development Report app. F, at USA000397-USA000398;
id.
app. E, at USA000381-USA000382;
id.
app. F, at USA000384-USA000394.)
23
The goal of the Linking Panel was to match up these 18 abilities with the 21 identified task clusters.
{See
Test Development Report app. F (“You will ... be asked to rate the importance of each of the eighteen abilities for the performance of each of the twenty-one task clusters.”);
see also
Linking Panel Worksheet.) Members of the Linking Panel each had to come up with a rating to reflect how important each ability was to each cluster.
{See id.)
This rating was either “Critical to the performance of the task cluster,” “Important to the performance of the task cluster,” “Somewhat important to the performance of the task cluster,” or “Not relevant to the performance of the task cluster.”
{See id.)
Although this appears to have been the only step in the process intended to capture the relationship between the tasks of a firefighter and the abilities tested on the written examination, the Test Development Report does not explain how or why particular tasks were matched with particular abilities. In fact, before performing this task, Linking Panel members were not given any explanation about the meaning of the ratings they were supposed to provide. (Int. 56.1 ¶ 95.) No statistical analyses were conducted to confirm the reliability of the ratings or the agreement in ratings among panel members.
{See
Morrongiello Dep. 299-300.)
Although the Linking Panel matched the 21 clusters to 18 abilities, only nine of the 18 abilities were deemed “testable” in a written multiple-choice format: Written Comprehension, Written Expression, Memorization, Problem Sensitivity, Deductive Reasoning, Inductive Reasoning, Information Ordering, Spatial Orientation, and Visualization. (Test Development Report 11.) The two abilities which incumbent firefighters rated highest in importance — Oral Comprehension and Oral Expression— were not among those tested, because “structured interviews” with thousands of candidates (which would help evaluate oral abilities) would not have been feasible.
{See, e.g.,
Fraenkel Deck Ex. 11 (“Patitucci II Dep.”), at 131-32, 274-75;
see also
Int. 56.1 ¶ 106; Def. Int. 56.1 ¶ 106.) Regarding the other seven omitted abilities, Mor
*105
rongiello stated that he “didn’t do anything specific to determine” whether or not they were testable, and that he was “going by ... standard operating procedure at that time in our unit ... that these abilities” would not be tested. (Morrongiello Dep. 443.)
To determine how many examination questions would be devoted to each ability, the “average importance ratings for each of the nine testable abilities within each cluster were determined from the individual ratings given by the linking panel.”. (Test Development Report 11;
see also id.
app. G (setting out column with average importance rating of ability “to task cluster”).) The panel multiplied the average of each importance rating by the rating that the JAQ questionnaires had given to the ability.
(Id.
at 11.) An average rating was then calculated for each ability — that rating was “pro-rated” and rounded based on an 85-question, multiple-choice test.
(Id.)
The result of this process was a test intended to evaluate nine abilities as follows: Written Comprehension (9 questions), Written Expression (6 questions), Memorization (11 questions), Problem Sensitivity (12 questions), Deductive Reasoning (9 questions), Inductive Reasoning (9 questions), Information Ordering (11 questions), Spatial Orientation (10 questions), and Visualization (8 questions).
(See
Test Development Report, at USA000404; Levy Deck Ex. M, at 4 (Admission # 30).) The parties agree that all of these are “cognitive” abilities.
(See
Levy Deck Ex. M, at 4, 6 (Admission ## 29, 36).)
3.
Test Construction
The next step in the process was to construct a written examination based upon the job analysis. For Exam 7029, one Lieutenant and four firefighters were “given training in how to write exams,” were placed on a panel, and then wrote the examination. (Test Development Report 11.) A “Review Panel” was assembled to review the questions on the examination (Int. 56.1 ¶ 124; Test Development Report 12), although it is unclear what this panel reviewed the questions for. One thing that the reviewers did
not
consider was whether “each [question] measured the ability it was originally designed to measure.” (Int. 56.1 ¶ 124.) No analysis of the reading level of the examination was conducted. (Int. 56.1 ¶ 128.)
This test-writing process was essentially the same for Exam 2043.
(See
Johnston Dep. 23-28.)
24
B. Motion to Strike
The determination of whether an employment test is job-related relies heavily upon expert testimony. Although discovery in this case closed in October 2008, the City submitted two new declarations containing expert assertions with its summary judgment papers in February 2009: a declaration from the City’s expert, Dr. Schemmer (Fraenkel Deck Ex. 2 (“Schemmer Deck”)) and a declaration from Dr. Catherine Cline, who participated in the development of Exam 6019, administered after Exams 7029 and 2043 (Fraenkel Deck Ex. 3 (“Cline Deck”)). Plaintiffs have moved to strike these declarations.
(See
Docket Entries ##273, 274.) For the reasons
*106
that follow, the court grants the motion in part and denies it in part.
The Intervenors argue that the court should strike these declarations because the deadline for submitting expert reports was January 21, 2008, and all expert and fact discovery closed on October 31, 2008.
(See
Declaration of Richard Levy dated March 4, 2009 (Docket Entry #274).) The Intervenors point out that Federal Rule of Civil Procedure 26(a) requires a written report with a complete statement of an expert witness’ opinions, including the reasons for them, and that Rule 26(e) requires supplementation of that report “in a timely manner.”
(See
Memorandum of Law In Support of Motion to Strike (Docket Entry #274) (“Strike Mem.”) 2-3.) They further argue that Rule 37(c)(1) prevents a party from relying on information it did not disclose in accordance with Rules 26(a) and (e).
(See id.
at 3.) Because the declarations offer new expert opinions in violation of the discovery rules, the Intervenors ask the court to strike them.
Under Rule 26(a)(2)(B) of the Federal Rules of Civil Procedure, expert testimony must be accompanied by a written report which shall contain,
inter alia,
“a complete statement of all opinions the witness will express and the basis and reasons for them,” “the data or other information considered by the witness in forming them,” and “any exhibits that will be used to summarize or support them.” A party must make these disclosures “at the times and in the sequence that the court orders.” Fed.R.Civ.P. 26(a)(2)(C). Rule 37(c)(1) states that if a party fails to abide by these requirements, “the party is not allowed to use that information ... to supply evidence on a motion, at a hearing, or at a trial, unless the failure was substantially justified or is harmless.” The Second Circuit has construed the language in Rule 37(c)(1) to provide discretion to preclude evidence if “the trial court finds that there is no substantial justification and the failure to disclose is not harmless.”
Design Strategy, Inc. v. Davis,
469 F.3d 284, 294 (2d Cir.2006).
The parties have had ample time to conduct expert discovery. As Magistrate Judge Roanne Mann’s Scheduling Orders make clear, the City was required to make expert disclosures on business necessity by January 7, 2008.
(See
Scheduling Order (Docket Entry #30) ¶4.) That deadline was extended to January 21, 2008.
(See
Revised Schedule (Docket Entry # 66) ¶ 4.) The schedule for the City’s expert depositions was amended several times, and the deadline was last scheduled for the end of March 2008.
(See
Scheduling Order (Docket Entry # 30) ¶ 5; Revised Schedule (Docket Entry # 66) ¶ 5; Modified Scheduling Order (Docket Entry # 86) ¶ 5.) A schedule for Plaintiffs’ rebuttal on business necessity was also ordered by Judge Mann, with all expert and fact discovery to conclude on October 31, 2008.
(See
Modified Scheduling Order (Docket Entry # 181).) The October deadline was set at the direction of this court to ensure that all discovery would be completed by then.
(See
April 10, 2008 Tr. 15-16.)
Clearly, submitting new expert opinions after the close of discovery violates the discovery rules. The City does not dispute this principle, instead arguing that it is simply not making new expert disclosures. Regarding the Cline Declaration, the City states that Dr. Cline is not being offered as an expert witness, but is, rather, only being offered as a fact witness to correct certain remarks about Exam 6019 that were made in Intervenors’ summary judgment papers.
(See
Memorandum of Law in Opposition to Motion to Strike (Docket Entry # 276) (“Strike Opp.”) 4, 6 (contending that Dr. Cline is “not being offered as an expert,” but simply as “a fact witness
*107
concerning her work” developing Exam 6019).) Regarding the Sehemmer Declaration, the City states that its submission merely provides further “detail” on issues already opined on by Dr. Sehemmer.
(Id.
at 2-3.)
The City’s position on both counts is disingenuous. First, the Cline Declaration consists primarily of expert opinions about the validity of Exams 7029 and 2043. Except in a few places, these assertions are based upon specialized knowledge of the art of test construction and validation, rather than the personal knowledge of a lay witness. Although it appears that Dr. Cline might have qualified to serve as an expert — had the City offered her, following the proper procedures — the City has chosen not to do so. If Dr. Cline is not offered as an expert, she may not opine about the validity of Exams 2043 and 7029, about which she has no personal knowledge.
See
Fed.R.Evid. 701;
United States v. Rigas,
490 F.3d 208, 224 (2d Cir.2007) (“Rule 701(c), which prohibits testimony from a lay witness that is ‘based on scientific, technical, or other specialized knowledge,’ is intended ‘to eliminate the risk that the reliability requirements set forth in Rule 702 will be evaded through the simple expedient of proffering an expert in lay witness clothing.’ ”).
Based on the court’s review of the Cline Declaration, it is clear that it is nothing more than an expert declaration submitted following the close of expert discovery. It would be prejudicial to Plaintiffs to have to address these new' expert opinions from Dr. Cline, who was never offered as an expert. The Cline Declaration shall therefore be stricken. Nevertheless, those portions of the Declaration that simply clarify, as a factual matter within Dr. Cline’s personal experience, the preparation of Exam 6019, will not be stricken.
(See id.
¶¶ 11 (first six sentences based on personal knowledge), 12 (first sentence based on personal knowledge), 16 (first sentence based on personal knowledge).) Exam 6019 is not at issue in this litigation, and it would be harmless to supplement the record regarding that examination. The court need not strike those portions of the Cline Declaration.
Second, the court rejects the City’s argument that the Sehemmer Declaration offers no new expert opinions. That Declaration contains numerous paragraphs directly addressing issues for which the City has offered no reference to a timely report or disclosure.
(See, e.g.,
Sehemmer Decl. ¶¶ 3-14.) The numerous conclusory assertions in the Sehemmer Declaration suggest that they were constructed to fill holes in the evidence that the City failed to gather during discovery, and to rebut analyses presented over a year ago in Plaintiffs’ expert reports.
See Point Prods. A.G. v. Sony Music Entertainment, Inc.,
No. 93-cv-4001(NRB), 2004 WL 345551 , at *9 (S.D.N.Y. Feb. 23, 2004) (“To accept the contention that the new affidavits merely support an initial position when they in fact expound a wholly new and complex approach designed to fill a significant and logical gap in the first report would eviscerate the purpose of the expert disclosure rules.”). Dr. Schemmer’s largely conclusory and unsupported statements strongly suggest an attempt by the City to “sandbag” its opponents with new opinions designed to defeat summary judgment.
See Disability Advocates, Inc. v. Paterson,
No. 03-CV-3209 (NGGXMDG), 2008 WL 5378365 , at *11 (E.D.N.Y. Dec. 22, 2008) (“The purpose of [the disclosure rules] is to prevent the practice of ‘sandbagging’ an opposing party with new evidence.”). It would be prejudicial to Plaintiffs to have to address these assertions at this point in the litigation.
*108
The court will not consider new, conclusory opinions by Dr. Schemmer. The City was aware of its burden to demonstrate business necessity during the discovery process, and it is bound by the analysis and opinions offered by Dr. Schemmer during that time.
See Wechsler v. Hunt Health Sys., Ltd.,
381 F.Supp.2d 135, 156 (S.D.N.Y.2003). Indeed, the City does not even attempt to argue that new evidence should be considered, and simply pretends that the Schemmer Declaration offers no new opinions. To allow such new evidence to be presented would undermine the purpose of the discovery rules, circumvent the discovery schedule that was ordered by the court, and prejudice Plaintiffs. Accordingly, the court strikes the Schemmer Declaration in its entirety.
The court now turns to the merits of the City’s business necessity defense.
C.
Guardians
and the Validity of Employment Tests
To be considered job-related, an employment examination must be properly “validated,” and the Second Circuit has identified two sources that help determine their validity: (1) “the testimony of experts in the field of test validation” and (2) “the Equal Employment Opportunity Commission’s ‘Uniform Guidelines on Employee Selection Procedures’ (“EEOC Guidelines”).”
Gulino,
460 F.3d at 382
(citing
29 C.F.R. §§ 1607.1-1607.18 ). Each source is important to a court’s decision. As the Second Circuit stated in
Gulino,
while courts “must take into account the expertise of test validation professionals,” they “must also remain aware that reliance upon the findings of experts in the field of testing should be tempered by the scrutiny of reason and the guidance of Congressional intent.”
Id.
(internal citation and quotation marks omitted).
25
And, although courts must “approach the [EEOC] Guidelines with the appropriate mixture of deference and wariness, thirty-five years of using these Guidelines makes them the primary yardstick by which we measure defendants’ attempt to validate” employment tests.
Id.
at 384 (internal citation and quotation marks omitted).
The governing case in this Circuit for assessing the validity of employment tests is
Guardians Association of the New York City Police Department, Inc. v. Civil Service Commission,
630 F.2d 79, 82 (2d Cir. 1980).
See Gulino,
460 F.3d at 385
(“Guardians
is still the law in this Circuit.”).
Guardians
involved a test administered by the City in 1979 to over 36,000 applicants for positions in the New York City Police Department. “The exam was developed by a fairly elaborate two-stage process,” with stage one involving a job analysis with input from, among others, numerous panels of police officers and questionnaires to thousands of police officers, and stage two involving additional panels and test-question revision from police experts and the New York City Department of Personnel. 630 F.2d at 83-84 . The examination had a disparate impact, and the principal issue on appeal was “whether the defendants have rebutted the plaintiffs’ prima facie case by showing that its test was job-related.”
Id.
at 88 .
*109
In assessing whether the employment test was job-related,
Guardians
recognized that the EEOC Guidelines set forth a “sharp distinction” between
“tests that measure ‘content’
— i.e., the ‘knowledges, skills or abilities’ required by a job — and
tests that purport to measure ‘constructs’
— i.e., the ‘inferences about mental processes or traits, such as ‘intelligence, aptitude, personality, commonsense, judgment, leadership and spatial ability.’ ”
Gulino,
460 F.3d at 384
(quoting Guardians,
630 F.2d at 91-92 ) (emphases added). “To demonstrate ‘content validity,’ the employer must introduce data ‘showing that the content of the selection procedure is representative of important aspects of performance on the job for which the candidates are to be evaluated.’ ”
Id.
at 384 n. 23
(quoting
29 C.F.R. § 1607 .5B and
citing
29 C.F.R. § 1607 .14C). “To demonstrate ‘construct validity’ on the other hand, the employer must introduce data ‘showing that the procedure measures the degree to which candidates have identifiable characteristics which have been determined to be important in successful performance in the job for which the candidates are to be evaluated.’ ”
Id. (quoting
29 C.F.R. § 1607 .5B and
citing
29 C.F.R. § 1607 .14D).
Guardians
criticized the sharp distinction between “content validity” and “construct validity.” The court observed that, under the EEOC Guidelines, “content validation is generally much easier to achieve than construct validation,” even though the test types differ more in degree than in kind.
Id.
at 384 . This is because “content and construct represent a continuum that ‘starts with precise capacities and extends to increasingly abstract ones.’ ”
Id. (quoting Guardians,
630 F.2d at 93 ). Because of the difficulty in showing “construct validity,” the court observed that “a conclusion that construct validation is required would often decide a case against a test-maker, once a disparate racial impact has been demonstrated.”
Guardians,
630 F.2d at 92 .
In response to these concerns, the court tempered the standards set forth in the EEOC Guidelines with a functional approach to test validation.
See Gulino,
460 F.3d at 386 n. 25.
Guardians
established a five-part test to determine the content validity of an employment test, which is flexible enough to encompass concepts of construct validity:
(1) the test-makers must have conducted a suitable job analysis;
(2) they must have used reasonable competence in constructing the test itself;
(3) the content of the test must be related to the content of the job;
(4) the content of the test must be representative of the content of the job; and
(5) there must be a scoring system that usefully selects from among the applicants those who can better perform the job.
Id.
at 384-85
(quoting Guardians,
630 F.2d at 95 ).
The first two requirements relate to the “quality of the test’s development,” while the final three “are more in the nature of standards that the test, as produced and used, must be shown to have met.”
Guardians,
630 F.2d at 95 .
Gulino
instructs that the
Guardians
approach “has the advantage of tracking the [EEOC] Guidelines standards while still allowing the courts to take a more functional approach to the analysis,” and frees “the courts from having to draw sharp distinctions between ‘content’ and ‘construct’ or ‘knowledge’ and ‘ability.’ ” 460 F.3d at 385
(quoting Guardians,
630 F.2d at 93-94 ).
*110
The parties do not dispute that
Guardians
provides the appropriate standard by which to evaluate Written Exams 7029 and 2043.
D. Application of
Guardians
The court must determine whether the City has offered sufficient evidence to create a disputed issue of material fact that Written Examinations 7029 and 2043 are job-related under
Guardians.
The basic question before the court is whether the examinations selected candidates who would be better firefighters.
See Guardians,
630 F.2d at 88 . The more specific question is whether the City has met the detailed requirements of test validation set out by
Guardians. See
Lindemann 151
(“Guardians
contains an unusually complete discussion of the details of test validation ... [and] the validation criteria set forth in
Guardians
are ones that employers should attempt to satisfy in comparable situations.”).
The court rejects the City’s assertion that there are material factual disputes sufficient to preclude summary judgment on job-relatedness. Even considered in the light most favorable to the City, the undisputed evidence paints an extremely troubling picture of the test construction process and the content that the City sought to test. Even under the summary judgment standard, the City has failed to meet its burden to show that its reliance on the challenged examinations was warranted by a valid business justification. For each of the
Guardians
requirements, the City’s arguments are riddled with serious defects, and the facts it presents patently fail to satisfy the demands of test validation. Insufficient evidence is presented for a reasonable fact finder to conclude that the challenged examinations were related to the job of a firefighter and relied upon as a matter of business necessity. The court finds itself compelled to grant summary judgment for Plaintiffs.
As the court sets forth below, the City’s evidence fails to show valid test construction under
Guardians’
first and second requirements, and fails to establish appropriate test content under
Guardians’
fourth requirement. The City’s showing on
Guardians’
third requirement demonstrates only a minimal relationship between the content of its examinations and the content of the job of firefighter. These serious failings culminated in the City’s decision to use the problem-riddled examinations to impermissibly fail and arbitrarily rank firefighter candidates. The imposition of these scoring devices, based upon the results of poorly constructed examinations, means that the City has failed to meet the fifth
Guardians
requirement. The examinations were simply unable to “select from among the applicants those who can better perform the job.”
Guardians,
630 F.2d at 95 . The recurrence of severe deficiencies at every step of the court’s review destroys any pretense that the challenged examinations had “a manifest relationship to the employment in question.”
Albemarle Paper Co.,
422 U.S. at 425 , 95 S.Ct. 2362 (citation omitted).
1.
Job Analysis
“According to the [EEOC] Guidelines, a job analysis involves an assessment ‘of the important work behavior(s) required for successful performance and their relative importance.’ ”
Guardians,
630 F.2d at 95
(quoting
29 C.F.R. § 1607 .14C(2)). In
Guardians,
the court concluded that the City’s extensive job analysis adequately identified 42 important work tasks or behaviors. Specifically, the “work behaviors involved in being a police officer were identified by extensive interviewing, and subjected to serious review ....”
Id.
at 95 ;
see also id.
at 83 (describing process by which New York City Personnel Depart
*111
ment identified tasks performed by police officers and a panel of police officers honed the tasks identified). The relative importance of the 42 tasks were assessed by “means of an extensively distributed questionnaire” used to rank them.
Id.; see also id.
at 83 (noting that over 2,600 police officers answered questionnaires to help rank the tasks’ importance).
Nevertheless,
Guardians
deemed the overall job analysis to have been of “questionable sufficiency.”
Id.
at 96 . In transforming the 42 tasks into the corresponding “knowledge, skills or abilities necessary to the effective performance” of those tasks,' “no effort was made to explain the relationship between any of the ... abilities and the 42 job tasks from which they were ostensibly derived.”
Id.
Because of this shortcoming, the Second Circuit cautioned that, “[o]nly if the relationship of abilities to tasks is clearly set forth can there be confidence that the pertinent abilities have been selected for measurement.”
Id.
at 98 .
This deficiency in the City’s job analysis is also present here. As in
Guardians,
the City conducted a job analysis aimed at identifying the tasks of an entry-level firefighter. The City took measures to develop an extensive task list based on panels and job questionnaires with incumbent firefighters.
(See
Test Development Report 4-10.) It then used those results to pare down the list to 21 specific task clusters and 18 necessary abilities.-
(See id.
app. E, at USA000371-USA000393;
id.
app. E, at USA000381-USA000382;
id.
app. F, at USA000397-USA000398;
see also
Linking Panel Worksheet.)
However, as in
Guardians,
the City has offered no evidence of “the relationship of abilities to tasks.” 630 F.2d at 96 ! The absence of such evidence undermines the court’s confidence “that the pertinent abilities have been selected for measurement.”
Id.
Looking at the 21 task clusters — including such categories as Ladder Operation, Climbing and Portable Ladder Activities, Building Entry, Search, Rescue, Ventilation, Hose Operations During Ex-tinguishment, and Extrication — it is not apparent how they relate to the nine spe^cific abilities identified by the City for testing. Indeed, the City has not offered any explanation or documentation indicating how the task clusters relate to the nine abilities.
Cf. M.O.C.H.A. Soc’y, Inc. v. City of Buffalo,
No. 98-CV-99C(JTC), 2009 WL 604898 , at *14 (W.D.N.Y. Mar. 9, 2009) (noting that procedures, including linking tasks and abilities, were “painstakingly documented”).
Not only is there an absence of evidence supporting the relationship between tasks and abilities, but there is also strong evidence that no such relationship exists. Deposition testimony of Linking Panel members shows a considerable degree of confusion about the process and about the definitions of abilities the members were supposed to evaluate. For example, one member testified that he “probably didn’t know” what Inductive Reasoning meant when providing a rating, “so I gave it a two since I didn’t know what it was, quite frankly.” (Fraenkel Decl. Ex. 20, at 66-67.) When presented with a definition, he was able to explain its importance, but he also stated that he had not been given definitions of the abilities at the time he made his linking determinations.
(Id.)
This panel member also stated that “I don’t believe I really knew what deductive reasoning was at the time of the examination.”
(Id.
at 72.) Other members testified to being unsure of the meanings of various abilities.
(See
Fraenkel Decl. Ex. 21, at 57 (Problem Sensitivity)),
id.
at 57-58 (Deductive Reasoning),
id.
at 58 (Inductive Reasoning),
id.
at 58-59 (Information Ordering),
id.
at 61-62 (Visualization),
id.
*112
at 62 (Time Sharing); Levy Decl. Ex. II, at 54-55 (Visualization); Levy Decl. Ex. JJ, at 47-48 (Inductive Reasoning, Deductive Reasoning). These difficulties seemed to stem from the fact that panel members were not informed by the test-maker of the meaning of the abilities.
(See
Int. 56.1 ¶ 95.)
The parties’ expert submissions highlight these deficiencies in the City’s job analysis. First, although the City bears the burden to show a proper job analysis, the City’s Bobko-Schemmer Report does little to satisfy this burden. The Bobko-Schemmer Report simply summarizes the test construction process, as described in the Test Development Report, and provides a few parenthetical comments about the steps taken. It states that the Test Development Report:
• updated important task statements from [Dr. Landy’s] prior task list;
• collected firefighter importance ratings of tasks and cognitive abilities, as well as links between these two domains (a process that is widely used in industrial-organizational psychology to provide a basis for demonstrating content validity);
• collected the above information using targeted interviews/observations, a focus group, and a job analysis survey completed by a sample of 192 firefighters which included ethnic/racial minority groups and females;
• used an ability taxonomy that is largely based on work by Fleishman (the Fleishman taxonomy provides a consistent framework for researchers to examine incumbent and expert perceptions regarding the ability demand of jobs);
• used nine of these abilities to write [questions] for the written exam (nine abilities also formed the basis for [questions] in the earlier [Dr.] Landy ... written exam); and
• used and trained panels of incumbent firefighters as [question] writers and [question] reviewers, with attention to diversity of these panels. (Using incumbents as preliminary [question] writers has several potential advantages. It should help ensure that the [question] content is consistent with the firefighter job. Also, the language and reading level of the [question] text will tend to be consistent with that in the incumbent population.)
(Bobko-Schemmer Report 28.)
26
According to Dr. Bobko and Dr. Schemmer, “the cognitively-based written exams were developed following standard job analytic and test development procedures — thus speaking to their job relatedness.”
(Id.)
Notably, however, the Bobko-Schemmer Report does not address the deficiencies in the linking process.
By contrast, Plaintiffs’ expert reports provide specific reasons to doubt the validity of the City’s job analysis process. According to the expert opinion of each of Plaintiffs’ experts, “the flaws in the [City’s] job analysis were fatal to the validity of the exams.” (Int. 56.1 ¶ 101.) This is not simply a difference in opinion among experts — Plaintiffs’ experts set out specific problems with the City’s job analysis that the City’s experts never address.
For example, the Jones-Hough Report includes four pages of criticism about the
*113
reliability of judgments made by the Linking Panel.
(See
Jones-Hough Report 26-30.) It states that “[t]he linking panel judgments on which [the City’s] examination development plan was based appear to have been done without sufficient understanding on the part of the linking panel members.” (Jones-Hough Report 35-36.) According to Dr. Jones and Dr. Hough, it was “[t]roubling in this stage of the project [that] information regarding the degree to which the 12 firefighters involved in producing the final written examination specification did not understand and perform their assignment.”
(Id.
at 26.) The Jones-Hough Report summarizes this critique as follows:
[T]hough linking panel members appear to have experienced problems in performing the task they were presented, their judgments were used to determine the number of questions that would be used to assess each of the cognitive abilities measured by the new written examination. In our professional opinion, this represents a fatal flaw in the information used to determine the test development plan for Written Exam 7029.
(Id.
at 29;
see also
Goldstein Report 15 (criticizing the work of the linking panel).) The City never addresses these deficiencies in the Linking Panel judgments, and based on the undisputed evidence, a fact finder could not conclude that abilities and tasks were properly matched.
27
Another deficiency identified Plaintiffs experts relates to the specific tasks and abilities selected for measurement. In his expert report submitted for Plaintiffs, Dr. Goldstein opines that the City inappropriately retained tasks and abilities in its job analysis that did not meet a “Day One” standard — in other words, the City tested for tasks and abilities that could be learned on the job. (Goldstein Report 12.) Citing the EEOC Guidelines, Dr. Goldstein explained that a “content valid test should measure work behaviors, activities, and/or worker [knowledge, skills, abilities or characteristics] that are important for the performance of the job and are needed at entry, rather than learned on the job.”
(Id.
at 12)
(citing
29 C.F.R. § 1607 .14C(1) (“Content validity is also not an appropriate strategy when the selection procedure involves knowledge, skills, or abilities which an employee will be expected to learn on the job.”).) As Dr. Goldstein opined, “content validity models are concerned with establishing that the content of a test reflects the content of a job — in other words, that what a candidate must do to perform well on the test corresponds to what a worker must do to perform well on the job. That critical content validity link is broken if what the worker must do to perform well on the job is learned after entering the job (and, thus, after taking the test).”
(Id.)
This Day One standard is also reflected in 29 C.F.R. § 1607 .5F, which states that employers should “avoid making employment decisions on the basis of measures. of knowledge, skills, or abilities which are normally learned in a brief orientation period, and which have an adverse impact.” The City does not address its failure to comply with the EEOC regulations setting out a Day One standard.
*114
Instead uf attempting to address these deficiencies, the City resorts almost entirely to citing the work performed by Dr. Landy on Exam 0084, a predecessor to the examinations at issue in this case. The City repeatedly argues that it was reasonable for it to rely upon Dr. Landy’s validation study and test plan in devising Exams 7029 and 2043. (Def. BN Mem. 5 (“Exam 7029 was developed in 1999 and was based on the work done by outside consultant Dr. Frank Landy for Exam 0084.”).) Although the City cites extensively to Dr. Landy’s involvement in the development of Exam 0084, the cited evidence shows that his work had only limited effect on Exams 7029 and 2043.
The Test Development Report states that the task list developed by Dr. Landy was used as part of an “updated task list” that reflected new input from Morrongiello for Exam 7029.
(See
Test Development Report 5 (“The information derived from the interviews/observations and the task list from the previous job analysis were incorporated into an updated task list.”).) At his deposition, Morrongiello confirmed that he used Dr. Landy’s work in this way, testifying that he “started out with the task list that was used on 0084 just as a starting point .... ” (Morrongiello Dep. 478-79;
see also
Test Development Report 4 (referring to use of Fleishman’s ability list and use of those abilities on prior exam).) Indeed, Morrongiello stated at his deposition that he had merely “referred” to Dr. Landy’s report, but that he had not “read [it] in detail.” (Morrongiello Dep. 440;
see also id.
at 478 (stating that he relied on Dr. Landy “to a degree”).) There is simply
no
evidence presented that Dr. Landy’s work played any other role in the process of developing the examinations at issue. The undisputed evidence shows that the tasks and abilities lists for Exams 7029 and 2043 used Dr. Landy’s work on Exam 0084 as a starting point, nothing more.
In spite of his limited role, the City nonetheless refers repeatedly to the work of Dr. Landy in arguing for the validity of Exams 7029 and 2043. The City cites several times to Dr. Schemmer’s statement, which has been stricken, that “given Dr. Landy’s stature in the field it would be hard to imagine that one of Dr. Landy’s studies would possess substantial defects.” (Def. BN Mem. 5.) As the EEOC Guidelines explicitly provide, however, reliance on the stature of a test-maker cannot stand in for a proper showing of validity.
See
29 C.F.R. § 1607 .9A (“Under no circumstances will the general reputation of a test or other selection procedures, its author or its publisher, or casual reports of [its] validity be accepted in lieu of evidence of validity.”).
28
The mere presence of Dr. Landy in the process of identifying tasks
*115
and abilities for Exam 0084 does not allow the City to ignore the problems in its job analysis for Exams 7029 and 2043.
Moreover, the City fails to confront the fact that the Landy Report was very different from the analysis and process used to construct Exams 7029 and 2043.
(See
Int. BN Reply 8-10.) The Federal Government points out that “Dr. Landy’s own report regarding the work he did for the City (which is labeled on its front cover a ‘Draft’) explicitly states that he was not able to complete a job analysis because the firefighters’ union refused to cooperate, and only 217 of the 5,500 job analysis questionnaires Dr. Landy sent to FDNY firefighters were completed and returned.” (USA BN Mem. 8-9
(citing
Landy Report 1-2, 14).) Like the other deficiencies identified by Plaintiffs, these remain unaddressed by the City.
In sum, with respect to the first
Guardians
requirement, a reasonable fact finder could conclude that the City’s job analysis adequately began by identifying tasks and abilities important to the job of entry-level firefighter. Yet, the undisputed evidence shows that the City nonetheless failed to establish the
relationship
between the tasks it identified and the abilities it sought to test, and that it failed to rely on a Day One standard in assessing what abilities should be tested. Accordingly the City’s job analysis in this case is, like the showing in Gmardians, of “questionable sufficiency.” 630 F.2d at 96 .
2.
Test Construction Process
The second
Guardians
requirement is a proper test construction process. In analyzing this requirement,
Guardians
set forth two relevant points of guidance. First,
Guardians
explained that civil service examinations should be constructed by testing professionals. Although observing that “the law should not be designed to subsidize specialists,” the Second Circuit cautioned that “employment testing is a task of sufficient difficulty to suggest that an employer dispenses with expert assistance at his peril.”
Id.
Therefore,
Guardians
criticized the City for allowing police officers themselves to write test questions.
Id.
(“The questions were initially framed by police officers, who may have had expertise in identifying tasks involved in their job but were amateurs in the art of test construction.”). Second, in
Guardians
the City never tested its examination questions for reliability, nor did it “perform[ ] the minimal sample testing to ensure that the questions were comprehensible and unambiguous.” 630 F.2d at 96 . The Second Circuit therefore cautioned that examination questions should be tested to ensure their reliability.
Id.
In constructing Exams 7029 and 2043, the City has ignored the Second Circuit’s guidance. The Test Development Report makes clear that no outside expertise was utilized to construct test questions. Instead, the City relied upon panels of firefighters to write the questions for the challenged examinations. While input from incumbent firefighters was crucial in determining what tasks firefighters do, input from testing professionals was needed to devise
questions
that could assess which candidates would better perform those tasks.
See Guardians,
630 F.2d at 97 . However thoroughly a test-maker determines the important tasks of firefighters, the resulting examination will be deficient if its questions fail to connect to those tasks or fail to identify which candidates are best equipped to perform them. The City ignored
Guardians’
warning the municipalities should rely on expert assistance in constructing civil service examinations.
See Guardians,
630 F.2d at 96 ;
cf. Ricci,
129 S.Ct. at 2665-66 (test constructed by exam specialist);
Fickling v. N.Y.S Dep’t of Civil Serv.,
909 F.Supp. 185 , 190
*116
(S.D.N.Y.1995) (tests for New York State welfare eligibility examiner not constructed by testing specialists);
Cuesta v. N.Y.S. Office of Court Admin.,
657 F.Supp. 1084, 1097 (S.D.N.Y.1987) (noting that the Office of Court Administration “duly heeded” the
Guardians
warning to have expert assistance in test construction).
Second, the City has presented no evidence that it performed any sample testing to ensure its examinations adequately and reliably tested the nine identified abilities.
Guardians,
630 F.2d at 96 ;
Cuesta,
657 F.Supp. at 1097-98 (civil service examination questions were “pilot-tested on selected sample populations to measure their difficulty, impact, and validity”). Because the City has not presented evidence of sample testing, the court is left without any confidence that Written Exams 7029 and 2043 reliably tested the abilities identified by the City’s job analysis.
In sum, the City has not offered any evidence of a competent test construction process under the second
Guardians
requirement. Instead, the City argues the same points as it did in support of its job analysis. (Def. BN Mem. 7.) Even were these arguments sufficient to show an adequate job analysis — which they are not— they are insufficient to satisfy the
separate
requirement of competent test construction. Viewed together, these inadequacies in the overall test development process mirror those in
Guardians
— indeed, the City appears to be relying on the same practices for which it was criticized by the Second Circuit thirty years ago.
3.
Direct Relationship
Guardians’
third requirement is that the content of the test be directly related to the content of the job. The third requirement reflects “[t]he central requirement of Title VII” that a test be job-related. 630 F.2d at 97-98 . In
Guardians,
the court was satisfied that the “abilities that were actually tested for ... adequately related to most of the identified tasks.”
Id.
at 98 . The abilities tested were “filling out forms,” “remembering facts,” and applying “general principles to specific fact situations.”
Id.
The Second Circuit was satisfied that these were the abilities needed to be a police officer.
Id.
Here, the City has offered evidence from which a fact finder could conclude that the abilities it attempted to test had some relationship to the job of entry-level firefighter. The nine abilities that the City intend to test on Exams 7029 and 2043 were: Written Comprehension, Written Expression, Memorization, Problem Sensitivity, Deductive Reasoning, Inductive Reasoning, Information Ordering, Spatial Orientation, and Visualization.
(See
Levy Decl. Ex. M, at 4, 6 (Admission ## 28, 35).) Although these are all cognitive abilities, the City has presented sufficient evidence that the nine abilities reflect, to some degree, the job of entry-level firefighter. This uncontroversial point does not appear to be disputed.
A major flaw in the City’s showing, however, is identified in the expert opinion of Dr. Siskin that Exams 7029 and 2043
did not actually test
those nine abilities. To reach this conclusion, Dr. Siskin conducted two statistical analyses. One analysis measured the “correlation” among test questions: this analysis presupposes that, if test questions are actually measuring the ability they are intended to measure, then questions measuring the
same
ability will be more highly “correlated” with each other than with questions measuring
other
abilities.
(See
Siskin II Report 5;
see also
Fraenkel Decl. Ex. 14 (“Cline Dep.”), at 322.) Dr. Siskin measured the correlation of each of the nine abilities with itself and with each other ability for Exams 7029 and 2043.
(See
Siskin II Report 5-6 & tbls. 1
&
2.) This analysis revealed a pattern
*117
showing that the “[questions] intended to measure an individual cognitive ability actually tend[ed] to correlate as or more highly with [questions] intended to measure different cognitive abilities .... ”
(Id.
at 6.) “Four of the nine abilities [had questions] that correlate[d] on average more highly with [questions] intended to measure
different
abilities than with [questions] intended to measure the
same
ability.”
(Id.
(emphases added).) For Exam 7029, “everything except the [questions] intended to measure Spatial Orientation correlate[d] most highly with Written Expression.”
29
(Id.)
These correlation patterns led Dr. Siskin to conclude that “the [questions] on Written Exams 7029 and 2043 do not measure nine distinct abilities, as they were designed to do,” and that “the written examinations fail to measure and weight the nine ability constructs consistent with what the test developer’s job analysis deemed to be relevant to performance.”
(Id.
at 7.)
To further support his conclusion, Dr. Siskin applied a method called “factor a

[Text truncated at 120,000 characters. The full text is on the page linked above.]

---

Source: Frix Law Library, https://www.frixlaw.com/law-library/cases/2342626. Public record. Not legal advice.
