# Bazile v. City of Houston

> District Court, S.D. Texas · February 6, 2012 · 858 F. Supp. 2d 718

URL: https://www.frixlaw.com/law-library/cases/8698171

## Case

- **Full name:** Dwight BAZILE v. CITY OF HOUSTON
- **Court:** District Court, S.D. Texas
- **Decided:** February 6, 2012
- **Citations:** 858 F. Supp. 2d 718; 114 Fair Empl. Prac. Cas. (BNA) 596; 2012 U.S. Dist. LEXIS 14712; 2012 WL 573633
- **Precedential status:** Published
- **Opinion:** Opinion of the court by Rosenthal
- **Judges:** Rosenthal
- **Cited by:** 3 later opinions in the Frix Law Library

## Citator (automated)

- No negative treatment found by the automated citator. That is not the same as a confirmation that the case is good law; read the citing cases.
- Full citator and citing cases: https://www.frixlaw.com/law-library/cases/8698171

## Opinion text

MEMORANDUM AND OPINION
LEE H. ROSENTHAL, District Judge.
This Title VII disparate-impact suit challenges the City of Houston’s system for promoting firefighters to the positions of captain and senior captain. Historically, the City has promoted firefighters based on their years of service with the Houston Fire Department (“HFD”) and their scores on a multiple-choice exam. The format and general content of that exam are set out in the Texas Local Government Code (“TLGC”) and in the collective bargaining agreement (“CBA”) between the City and the firefighters’ union, the Houston Professional Fire Fighters Association (“HPFFA”). Seven black firefighters sued the City, alleging that the promotional exams for the captain and senior-captain positions were racially discriminatory, in violation of the Fourteenth Amendment, 42 U.S.C. § 1981 , and Title VII, 42 U.S.C. § 2000e-2. After mediation in February and March 2010, the City and the seven firefighters reached a settlement that included a proposed consent decree. The decree would require the City to implement changes to the captain and senior-captain promotion exams in two phases. The first phase required minor changes to the November 2010 captain exam. The second phase required more significant changes to the May 2011 senior-captain exam that would apply to future captain and senior-captain exams. The HPFFA intervened in the lawsuit and objected because the proposed consent decree changed the exams in ways inconsistent with the TLGC and the CBA. The HPFFA contended that the City and the plaintiffs had not shown discrimination that would permit this court to approve the consent decree.
This court bifurcated the proceedings to resolve the HPFFA’s objections. The first stage addressed the narrow set of changes proposed for the November 2010 captain exam. The second stage addressed the broader changes proposed for subsequent captain and senior-captain exams. In the first stage, this court, with the HPFFA’s agreement, approved changes to the November exam. This opinion addresses the HPFFA’s objections to the proposed changes to the future captain and senior-captain exams.
The HPFFA vigorously objects that the proposed changes involve “a far-reaching and wholesale restructuring of the entire promotional process that goes beyond anything plaintiffs have even alleged in this lawsuit” and “bypass both the long-established protections of state law and the union’s protected role in being the sole, collective voice for the city’s firefighters.” (Docket Entry No. 89, at 8). The City and the seven individual plaintiffs acknowledge *721 that the changes are far-reaching but argue that they are needed to comply with federal antidiscrimination law. They argue that the current exam system is not “job related for the position[s] in question and consistent with business necessity” as required by 42 U.S.C. § 2000e-2(k)(1)(A)(i).
This court held an evidentiary hearing to consider the proposed changes to the May 2011 senior-captain exam and to future exams. The summary of the evidence shows the welter of expert opinions the parties presented on whether the existing format and content of the City’s promotion exams for the captain and senior-captain positions have a disparate impact on African-American candidates; whether the existing exams are reliable and valid measures of the knowledge and qualities relevant to the promotion decisions; whether the existing exams are reliable and valid ways to compare candidates; and whether the proposed changes to the exams will provide reliable and valid exams and address disparate impact. The experts’ testimony and submissions left the court with a sense of disquiet about the opinions expressed. The science of testing to measure and compare promotion-worthiness is admittedly imperfect. The expert witnesses, particularly for the City, acknowledged some errors and some incomplete aspects of their work in designing and administering the promotion exams. At best, all the witnesses’ opinions amount to uncertain efforts to gauge how well different exam approaches measure, compare, and predict job performance. The analytical steps required by the applicable legal standards must be approached with a recognition of the limits of the expert testimony.
At the same time, courts clearly lack expertise in the area of testing validity. ‘“The study of employment testing, although it has necessarily been adopted by the law as a result of Title VII and related statutes, is not primarily a legal subject.’ Because of the substantive difficulty of test validation, courts must take into account the expertise of test validation professionals.” Gulino v. N.Y. State Educ. Dep’t, 460 F.3d 361 , 383 (2d Cir.2006) (quoting Guardians Ass’n of N.Y.C. Police Dep’t, Inc. v. Civil Serv. Comm’n of City of N.Y., 630 F.2d 79, 89 (2d Cir.1980)). The combination of the lack of judicial expertise in this area and the limits of the expertise of those who do have training and experience support a cautious and careful judicial approach.
Based on the parties’ filings, the evidence, and the applicable law, this court finds that the City and the seven individual plaintiffs have shown that the captain and senior-captain exams violate Title VII. But this court also finds that some of the changes in the proposed consent decree violate the CBA and TLGC and that the City and the plaintiffs have not shown that all these changes are necessary to comply with Title VII. Based on these findings and conclusions, the proposed consent decree is accepted in part and denied in part. The use of situational-judgment questions and an assessment center are justified by the record evidence and are job-related and consistent with business necessity. But other parts of the proposed modified consent decree violate the TLGC and CBA, and the City and the plaintiffs have not shown that they are tailored to respond to the disparate impact alleged. Using the parties’ descriptions of the proposed changes, the provisions that violate the TLGC and CBA without the necessary justification in the record, and which this court does not accept, are as follows:
2. Job-Knowledge Written Test
— Pass/fail test (test designer to determine cut-off score)
— No rank-order list
*722 — Test designer may elect not to use any written job-knowledge cognitive test
3. Scenario-Based Computer-Objective Test
— Rank-order list from which all intended promotions to assessment center will be made 1
4. Assessment Center
— Rank-order list
•— Sliding bands based on test accuracy as determined by consultant
— Fire Chief will document reasons for the selection of each candidate within bands
■— No Rule of Three
(Docket Entry No. 69-2, at 29).
The reasons for finding these aspects of the proposed changes to the promotion examinations invalid, and the remaining aspects supported by the record and the applicable law, are explained below. This opinion first describes the promotion system in place before any changes; reviews the expert and other evidence relevant to assessing disparate impact; and analyzes whether, under the applicable law, the proposed settlement is tailored to remedying the disparate impact that is shown. A hearing is set for February 21, 2012, at 1:30 p.m. to address the issues that remain to be resolved and a timetable for doing so.
Finally, because the terminology used by the HFD and the industrial psychologists who served as experts in this case produced a number of acronyms and abbreviations, a list of the most commonly used is attached to this Memorandum and Opinion.
I. Background
A. The Houston Fire Department
The HFD has approximately 4,000 employees involved in firefighting. Ninety percent are in the Emergency Operations Division (“EOD”). Half of the EOD employees are at the “firefighter” level and perform “task-level jobs” such as retrieving and using fire hoses. The next rank above firefighter is “engineer operator” (“EO”). In addition to performing firefighters’ tasks, EOs drive fire trucks and HFD ambulances and operate ladders and pumps. Firefighters outnumber EOs two-to-one. (Evidentiary Hr’g Tr. 115, Docket Entry No. 130).
Captains are ranked immediately above EOs. HFD captains are the “first line of supervisor position[s] in the fire department.” A captain supervises the operation of fire engines, which are smaller fire trucks that carry hoses and pump water. Each HFD fire station has at least one fire engine and one captain. A captain supervises an EO and two firefighters assigned to an engine. When a captain misses a day of work, an EO may “ride up” and perform the absent captain’s job duties.
Senior captains are ranked immediately above captains. A senior captain supervises the operations of “ladder trucks,” which are large fire trucks with aerial ladders. Only half of the City’s fire stations have a ladder truck with a senior captain in addition to a fire engine and captain. A senior captain may supervise up to eight firefighters, including EOs. When a senior captain misses a day of work, a captain may “ride up.”
*723 During a fire emergency, a district chief — ranked above senior captain — is responsible for developing the firefighting strategy. The district chief may decide, for example, whether firefighters will enter a burning building and address a fire directly or instead contain it by protecting adjacent buildings. Senior captains may participate in the strategy development, but the district chief bears ultimate responsibility. Once a strategy is set, the senior captain and captain are responsible for implementing it. Usually a senior captain and the ladder-truck crew are responsible for forcible entries into a building to ventilate it, for attempting rescues, and for creating ways for other firefighters to enter. The captain and the fire-engine crew are usually responsible for locating, confining, and extinguishing fires. (Id. at 116—20).
To summarize the promotional system that is discussed in detail below, promotion from EO to captain and from captain to senior captain depends largely on a candidate’s score on a multiple-choice test. Any person meeting the experience requirement can take the test. An EO can apply for captain after four years in the fire department. A captain can apply for senior captain after two additional years of service as a captain. Tex. Loc. Gov’t Code § 143.028(a). A candidate’s length of service with the HFD will add some points to the test score, but the test score largely determines promotion.
The City makes promotion decisions based on a rank-order list of the candidates’ test points added to their length-of-service points. For each captain or senior-captain position available during the three years after the exam, the top three candidates’ names and scores are submitted to the HFD fire chief. The presumption is that the fire chief will select the candidate with the highest test score. If the fire chief selects the second or third highest scoring candidate, the chief must explain his reasons in writing. If a candidate is not selected for promotion within the three-year period, the candidate must retake the exam. These promotional procedures for the captain and senior-captain positions are based on the TLGC and the CBA.
1. The Texas Local Government Code
The City of Houston adopted the Fire Fighter and Police Civil Service Act (“CSA”), codified as Chapter 143 of the TLGC, on January 31, 1948. 2 The CSA’s “fundamental principle” is ensuring that public-service appointments and promotions are made “according to merit and fitness, ascertained by competitive examinations.” Klinger v. City of San Angelo, 902 S.W.2d 669, 671 (Tex.App.-Austin 1995, writ denied). The Texas legislature passed the CSA “to secure efficient fire and police departments composed of capable personnel who are free from political influence.” Tex. Loc. Gov’t Code § 143.001(a). The TLGC requires a test-based promotional system for firefighters. Section 143.021(c) states that positions within fire departments must be filled “from an eligibility list that results from an examination held in accordance with [the CSA].” The TLGC contains detailed rules describing the eligibility list, exam, and procedure for selecting firefighters for promotion.
The promotional process begins when a city posts notice of an upcoming examination. Municipalities like the City of Houston, with populations greater than 1.5 mil *724 lion, must post notice in plain view on a bulletin board located in’City Hall’s main lobby and in the Firefighters’ and Police Officers’ Civil Service Commission office by the 90th day before the date a promotional exam is scheduled. This 90-day notice must show the positions to be filled and the date, time, and place of the exam. Tex. Loc. Gov’t Code § 143.107(a). The 90-day notice must also list the sources from which the exam questions are taken. Id. § 143.029(a). By the 30th day before the date a promotional exam is scheduled, another notice must be posted in the same locations. Id. § 143.107(b). The 30-day notice must state the number of newly created positions and may “include the name of each source used for the examination, the number of questions taken from each source, and the chapter used in each source.” Id. § 143.029(c).
The TLGC requires that the test be in writing and forbids tests that “in any part consist of an oral interview.” Id. § 143.032(c). The questions must “test the knowledge of the eligible promotional candidates about information and facts.” Id. § 143.032(d). The information-and-fact questions “must” be based on:
(1) the duties of the position for which the examination is held;
(2) material that is of reasonably current publication and that has been made reasonably available to each member of the fire or police department involved in the examination; and
(3) any study course given by the departmental schools of instruction.
Id. The questions must also be taken from the sources identified in the posted notices. Id. § 143.032(e). Finally, the “examination questions must be prepared and composed so that the grading of the examination can be promptly completed immediately after the examination is over.” Id. § 143.032(f).
The exam grade determines whether the candidate will be placed on a promotion-eligibility list. Grading begins as soon as an individual candidate completes the exam. The candidate may remain present during the grading. Id. § 143.033(a). The multiple-choice exam score is based on a maximum grade of 100 points and is determined by the correctness of the answers to the questions. Id. § 143.033(c). Each candidate also receives one point for each year of seniority, with a maximum of 10 points. Id. § 143.033(b). In municipalities like Houston, a candidate must score at least 70 points on the exam to be eligible for promotion. Id. § 143.108(a).
All scores must be posted within 24 hours of the exam. Id. § 143.033(d), Each candidate may see the answers, grading, and source materials after the exam and can appeal a score within 5 days. Id. § 143.034(a). The City has 60 days to decide the appeal. Id. § 143.1015(a). A candidate who appeals is entitled to a hearing. Id. § 143.1015(b).
Once the scores are finalized, all candidates who pass are listed in rank order on a promotion-eligibility list. See id. § 143.021(c); id. § 143.108(f). When vacancies occur, the names of the three persons with the highest scores for the position are certified and provided to the head of the department with the vacancy. Id. § 143.036(b). This is known as the “Rule of Three.” The TLGC provides that “[u]n-less the department head has a valid reason” for not doing so, “the department head shall appoint the eligible promotional candidate having the highest grade on the eligibility list.” Id. § 143.036(f). If the candidate with the highest grade is not selected, the department head must personally discuss the reason with that candidate and file a written explanation. Id.
*725 2. The Collective Bargaining Agreement
Texas law establishes firefighters’ right to collective bargaining. Tex. Loc. Gov’t Code § 174.002(b) (“The policy of this state is that fire fighters and police officers, like employees in the private sector, should have the right to organize for collective bargaining, as collective bargaining is a fair and practical method for determining compensation and other conditions of employment. Denying fire fighters and police officers the right to organize and bargain collectively would lead to strife and unrest, consequently injuring the health, safety, and welfare of the public.”); id. § 143.204(a) (stating that a firefighter association submitting a petition signed by the majority of the paid firefighters in the municipality “may be recognized ... as the sole and exclusive bargaining agent for all of the covered fire fighters”). The HPFFA is the sole and exclusive bargaining agent for the City’s firefighters.
The TLGC allows the City and the HPFFA to enter into a written agreement binding when ratified by both. Id. § 143.206(a). Such an agreement can supersede the TLGC’s provisions “concerning wages, salaries, rates of pay, hours of work, and other terms and conditions of employment to the extent of any conflict with the [written agreement].” Id. § 143.207(a). The agreement “preempts all contrary local ordinances, executive orders, legislation, or rules adopted by the state.” Id. § 143.207(b).
The 2009-2010 CBA between the City of Houston and the HPFFA made few departures from the TLGC’s exam provisions. Like the TLGC, the CBA required a grade of at least 70% for promotion eligibility. The CBA specified that the test must consist of “not less than 100 and not more than 150 questions.” (Docket Entry No. 69-6, at 20). Unlike the TLGC, the CBA allowed only a .5-point increase in the score for each year of service, with a maximum of 10 points. The CBA also allowed a .5-point increase for each year of service for certain ranks. For example, an engineer or operator applying to be a captain is awarded .5 points for each year of service as an engineer. (Id.). Aside from these changes, the 2009-2010 CBA provided that the TLGC “remain[s] in full force in the same manner as on the date [the CBA] became effective.” (Id. at 13).
B. Title VII
“Congress enacted Title VII of the Civil Rights Act of 1964, 42 U.S.C. § 2000e et seq., to assure equality of employment opportunities by eliminating those practices and devices that discriminate on the basis of race, color, religion, sex, or national origin.” Alexander v. Gardner-Denver Co., 415 U.S. 36, 44 , 94 S.Ct. 1011 , 39 L.Ed.2d 147 (1974). Title VIPs prohibitions include using “a particular employment practice that causes a disparate impact on the basis of race, color, religion, sex, or national origin” unless the employment practice “is job related for the position in question and consistent with business necessity.” 42 U.S.C. § 2000e-2(k)(1)(A)(i). The plaintiffs alleged that the City’s promotional procedures for captain and senior captain violated Title VTI’s disparate-impact provision.
“Congress intended voluntary compliance to be the preferred means of achieving the objectives of Title VII.” Local No. 93, Int’l Ass’n of Firefighters, AFL-CIO v. City of Cleveland, 478 U.S. 501, 515 , 106 S.Ct. 3063 , 92 L.Ed.2d 405 (1986). To help employers comply with Title VII, Congress authorized the Equal Employment Opportunity Commission (“EEOC”) to issue compliance guidelines (the “Guidelines”). The Guidelines “are not administrative regulations promulgated pursuant to formal procedures established by the Con *726 gress. But ... they do constitute ‘(t)he administrative interpretation of the Act by the enforcing agency,’ and consequently they are ‘entitled to great deference.’ ” Albemarle Paper Co. v. Moody, 422 U.S. 405, 431 , 95 S.Ct. 2362 , 45 L.Ed.2d 280 (1975) (citing Griggs v. Duke Power Co., 401 U.S. 424, 433-34 , 91 S.Ct. 849 , 28 L.Ed.2d 158 (1971)).
The Guidelines require employers who make promotional decisions based on test scores to maintain records of tests and test results. 29 C.F.R. § 1607.4 (A). The Guidelines’ rule of thumb for determining disparate impact is the “4/5 Rule.” Under this Rule:
A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact, while a greater than four-fifths rate will generally not be regarded by Federal enforcement agencies as evidence of disparate impact.
Id. § 1607.4(D). There are exceptions to the 4/5 Rule. The Guidelines state:
Smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms or where a' user’s actions have discouraged applicants disproportionately on grounds of race, sex, or ethnic group. Greater differences in selection rate may not constitute adverse impact where the differences are based on small numbers and are not statistically significant, or where special recruiting or other programs cause the pool of minority or female candidates to be atypical of the normal pool of applicants from that group. Where the user’s evidence concerning the impact of a selection procedure indicates adverse impact but is based upon numbers which are. too small to be reliable, evidence concerning the impact of the procedure over a longer period of time and/or evidence concerning the impact which the selection procedure had when used in the same manner in similar circumstances elsewhere may be considered in determining adverse impact. Where the user has not maintained data on adverse impact as required by the documentation section of applicable guidelines, the Federal enforcement agencies may draw an inference of adverse impact of the selection process from the failure of the user to maintain such data, if the user has an underutilization of a group in the job category, as compared to the group’s representation in the relevant labor market or, in the case of jobs filled from within, the applicable work force.
Id.
If analyzing an employer’s test results under the 4/5 Rule shows “that the total selection process for a job has an adverse impact, the individual components of the selection process should be evaluated for adverse impact.” Id. § 1607.4(C). The method for evaluating individual components is a “validity study.” Id. § 1607.3(A). The Guidelines describe three types of validity studies: criterion-related-validity studies; content-validity studies; and construct-validity studies. Id. § 1607.5(A). A criterion-related-validity study analyzes whether test results correlate to “criteria that [are] predictive of job performance.” Mark R. Bandsuch, Ten Troubles with Title VII and Trait Discrimination Plus One Simple Solution (A Totality of the Circumstances Framework), 37 Cap. U.L. Rev. 965, 1089 (2009). A content-validity study analyzes whether test results correlate to “the knowledge, skills, and abilities related to that job.” Id. A construct-validity study examines whether test results correlate to “general characteristics important to job performance.” Id.
*727 The Guidelines also describe the evidence each type of validity study requires. A criterion-related-validity study requires “empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance.” 29 C.F.R. § 1607.5 (B). A content-validity study requires “data showing that the content of the selection procedure is representative of important aspects of performance on the job for which the candidates are to be evaluated.” Id. A construct-validity study requires “data showing that the procedure measures the degree to which candidates have identifiable characteristics which have been determined to be important in successful performance in the job for which the candidates are to be evaluated.” Id.
One court has summarized the content-validity and criterion-validity methods for evaluating a promotion or other employment test, as follows:
[Ejmployers can establish job-relatedness by one of three methods, including “content validity,” which entails showing that the test measures the job or adequately reflects the skills or knowledge required by the job. A typing test for secretaries exemplifies this kind of approach. This method does not require empirical evidence, but instead “should consist of data showing that the content of the selection procedure is representative of important aspects of performance on the job.” 29 C.F.R. § 1607.5 (B). In contrast, the “criterion related” approach evaluates whether a test is adequately correlated with future job performance and is constructed to measure traits thought to be relevant to future job performance. An IQ test is a typical criterion-related method. Unlike content validity, this ' method requires “empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance.” 29 C.F.R. § 1607.5 (B).
Banos v. City of Chicago, 398 F.3d 889, 893 (7th Cir.2005) (citations and internal quotations marks omitted). 3
Before conducting a validity study, an employer should conduct a “job analysis.” 29 C.F.R. § 1607.14 (A). Each type of validity study requires a different type of job analysis. Criterion-related-validity studies require “reviewing job information to determine measures of work behavior(s) or performance that are relevant to the job or group of jobs in question.” Id. § 1607.14(B)(2). “These measures or criteria are relevant to the extent that they represent critical or important job duties, work behaviors or work outcomes as developed from the review of job information”; “[b]ias should be considered.” 4 Id. Con *728 tent-validity studies should include “an analysis of the important work behavior(s) required for successful performance and their relative importance and, if the behavior results in work product(s), an analysis of the work product(s). Any job analysis should focus on the work behavior(s) and the tasks associated with them.” Id. § 1607.14(C)(2). Construct-validity studies “should show the work behavior(s) required for successful performance of the job, or the groups of jobs being studied, the critical or important work behavior(s) in the job or group of jobs being studied, and an identification of the construct(s) believed to underlie successful performance of these critical or important work behaviors in the job or jobs in question.” Id. § 1607.14(D)(2). “Each construct should be named and defined, so as to distinguish it from other constructs.” Id.
If one or more validity studies produces evidence “sufficient to warrant use of the procedure for the intended purpose under the standard of these guidelines,” the promotional procedure is “properly validated.” Id. § 1607.16(X). But if no validity study produces sufficient evidence, an employer “should initiate affirmative steps to remedy the situation.” Id. § 1607.17(3). These steps, “which in design and execution may be race, color, sex, or ethnic ‘conscious,’ include, but are not limited to,” the following:
(a) The establishment of a long-term goal, and short-range, interim goals and timetables for the specific job classifications, all of which should take into account the availability of basically qualified persons in the relevant job market;
(b) A recruitment program designed to attract qualified members of the group in question;
(c) A systematic effort to organize work and redesign jobs in ways that provide opportunities for persons lacking “journeyman” level knowledge or skills to enter and, with appropriate training, to progress in a career field;
(d) Revamping selection instruments or procedures which have not yet been validated in order to reduce or eliminate exclusionary effects on particular groups in particular job classifications;
(e) The initiation of measures designed to assure that members of the affected group who are qualified to perform the job are included within the pool of persons from which the selecting official makes the selection;
(f) A systematic effort to provide career advancement training, both classroom and on-the-job, to employees locked into dead end jobs; and
(g) The establishment of a system for regularly monitoring the effectiveness of the particular affirmative action program, and procedures for making timely adjustments in this program where effectiveness is not demonstrated.
Id.
C. The Procedural History of this Case
On August 4, 2008, seven firefighters sued the City of Houston, alleging that the 2006 captain and senior-captain exams had a discriminatory effect on their promotion opportunities, in violation of § 1981 and 42 U.S.C. § 2000e-2. The seven individual plaintiffs contended that the 2006 exams had a disparate impact on the promotion of black firefighters to captain and senior-captain positions compared to white firefighters. Four plaintiffs — Dwight Bazile, Johnny Garrett, Trevin Hines, and Mundo Olford — were lieutenants denied promotion to captain. Three plaintiffs — George Runnels, Dwight Allen, and Thomas Ward— *729 were captains denied promotion to senior captain. (Docket Entry No. 1).
The City and the plaintiffs settled. (Docket Entry No. 64). The HPFFA was not a party to the negotiations or settlement. The City agreed to promote Bazile, Olford, and Hines to captain; to promote Allen to senior captain; to allow Garrett to retire as a captain; and to allow Runnels and Ward to retire as senior captains. The City also agreed to pay each plaintiff backpay in amounts ranging from $376.80 to $23,075.46. (Docket Entry No. 69-2, at 2-8, 17-22, 26-27).
The settlement also contained a proposed consent decree to be submitted to the court for approval. (Id. at 9-10, 28-30). The decree required the City to implement changes to the captain and senior-captain exams in two phases. In the first phase, the City agreed to implement “modest” changes to the November 2010 captain exam. In the second phase, the City agreed to implement broader changes, beginning with the May 2011 senior-captain exam and applying to all future captain and senior-captain exams. The settlement agreement required the parties to give notice to the HPFFA of “this conceptual agreement” and to “meet in person or conference call [with the HPFFA] to explore potential adjustments of union suggestions prior to final settlement meeting.” (Id. at 10). The settlement agreement also required approval by the Houston City Council and by this court.
The parties notified this court of the settlement and their intent to file the proposed consent decree. Before filing the decree, the parties moved to join the HPFFA to the suit because the proposed changes to the promotion exams conflicted with the TLGC and the CBA. (Docket Entry No. 69). The HPFFA moved to intervene and asked this court to bifurcate review of the proposed consent decree. The first step would be to consider the HPFFA’s objections to the proposed changes to the November 2010 captain exam. The second stage would be to consider the HPFFA’s objections to the proposed changes to the subsequent senior-captain and later captain and senior-captain exams. This court granted the motion and entered a scheduling order. (Docket Entry Nos. 70 & 71).
1. The November 2010 Captain Exam
The HPFFA objected to certain proposed changes to the November 2010 captain exam. (Docket Entry No. 75). This court heard arguments and evidence on the objections on September 16, 2010. On the same date, and with the HPFFA’s agreement, this court found that the existing captain exam disparately impacted black firefighters and entered an order allowing the City to implement the consent decree provisions changing the November 2010 captain exam. (Docket Entry Nos. 82 & 85). The proposed consent decree described those changes to the 2010 captain exam, as follows:
Hybrid written examination
Content-validated job-knowledge portion of exam
— weighted to Houston Departmental material, plus
— carefully selected directly relevant test material,
— supported by job analysis and incumbent/supervisor feedback in direct interviews
— job analysis
Content-validated multiple-choice situational-judgment exam
— designed by an industrial/organization psychologist
— HFD departmental scenario based
— zero, partial, and full-credit options
— additional points to stay the same as under the Collective Bargaining *730 Agreement in an effort to minimize adverse impact potential
— rank order for selection process by Chief
— Rule of Three for promotion
(Docket Entry No. 69-2, at 28-29).
The most significant change to the November 2010 captain exam was the inclusion of multiple-choice “situational-judgment” questions in addition to the “job-knowledge” questions used on previous exams. Situational-judgment questions present hypothetical situations encountered on the job and ask candidates how they would respond. 5 By contrast, the job-knowledge questions that made up the previous captain exams, mandated by the TLGC, see Tex. Loc. Gov’t Code § 143.032(d), ask about facts related to the job, such as the content of applicable regulations or specific HFD policies and procedures. 6
This court’s order approved the inclusion of situational-judgment multiple-choice questions for the November 2010 captain exam based on a finding that “[t]he continued exclusive use of questions based on ‘fact’ and ‘information’ as stated in Local Government Code § 143.032(d) is likely to continue to result in adverse impact.” (Docket Entry No. 85, at 2). The situational-judgment questions included in the November 2010 captain exam were developed by industrial-psychology consultants selected by the City, the plaintiffs, and the HPFFA (the “consultants”). 7 These consultants developed the questions by creating a job analysis for the captain position. See 29 C.F.R. § 1607.14 (A) (describing a job analysis). To create the job analysis, the consultants interviewed “subject-matter experts” (“SMEs”); analyzed HFD materials related to the captain position, such as internal job descriptions and policies and procedures; and analyzed ex *731 ternal source materials such as published industry standards and firefighting textbooks. The consultants interviewed both internal SMEs — incumbent HFD captains and their supervisors — and external SMEs — individuals with similar experience who did not work for HFD. Through the job analysis, the consultants identified the “knowledge, skills, abilities and other characteristics” (“KSAOs”) required for successful performance in the captain position and designed the situational-judgment questions to measure the identified KSAOs.
This court also approved an additional consent-decree provision inconsistent with the TLGC and CBA. The TLGC and CBA allow a promotional candidate to be present while the candidate’s exam is scored and require that scores be posted within 24 hours of the exam. Tex. Loc. Gov’t Code § 143.033(a), (d). The consent decree required an “item analysis” of the score before it was finalized. Item analysis requires the consultants to aggregate data related to each question — or “item”— to eliminate questions that did not reliably measure an individual candidate’s exam performance. 8 Because item analysis requires collecting data from all promotional candidates’ exams and time to evaluate this data, the parties agreed to post the raw scores from the exam within 24 hours to meet the TLGC and the CBA requirements, but these raw scores would not be the final scores until the item analysis was completed. The promotional candidates would not remain throughout the item analysis.
Many of the TLGC and CBA requirements remained in place under the consent decree for the November 2010 captain exam. The consent decree still required job-knowledge questions. The exam that *732 was administered contained 75 job-knowledge questions. The job analysis was included as a “source material” in the posted notices and made available before the exam. See Tex. Loc. Gov’t Code § 143.029(a) (requiring posting of source material). The consent decree also required a rank order of the candidates based on their exam scores and required that promotions be made according to the Rule of Three set out in the TLGC. See id. § 143.036(b), (f) (describing the Rule of Three).
The City administered the captain exam on November 17, 2010. The consultants conducted an item analysis after the exam. A panel consisting of representatives for the City, the individual plaintiffs, and the HPFFA met to review scoring. Initially, based on the item analysis, the consultants recommended giving candidates full credit for seven job-knowledge questions and for fourteen situational-judgment questions, effectively eliminating those questions as a way to differentiate among the candidates. In addition, the City’s internal SMEs recommended giving full credit for one job-knowledge question and for six situational-judgment questions. The panel agreed with the SMEs’ recommendation. (Docket Entry No. 94, at 2; Docket Entry No. 94-1 at 2-3). The panel also agreed that a candidate’s score on the job-knowledge portion and the situational-judgment portion would be weighted equally in calculating the final score. (Docket Entry No. 94, at 2). Based on these decisions, a rank-order results list for the exam was created. 9 Some discrepancies emerged in the statistical calculations and the consultants recommended a credit adjustment for additional questions. On January 10, 2011, the City submitted a revised rank-order list of candidates who passed the exam. (Docket Entry No. 104, at 2).
On January 12, 2011, the City advised this court that there were more than 200 appeals by the promotional candidates. On January 14, the City moved for additional time to finalize the scores and rank-order list. (Docket Entry Nos. 110 & 112). The City sought more time- than the 60 days the TLGC allowed to decide whether to sustain the appeals. The HPFFA did not object to the request, and this court granted the motion. (Docket Entry No. 117). This court has not been updated on the status of the appeals or on promotions to captain under the November 2010 exam.
2. The May 2011 Senior-Captain Exam and Future Exams
The HPFFA filed its objections to the proposed changes to the May 2011 senior-captain exam and to subsequent captain and senior-captain exams. (Docket Entry No. 89). The proposed changes are described as follows:
*733 1. Officer Development Program
— 2 years in grade
— Educational Courses
• Officer Development I, II
• Available online at stations for all to participate
2. Job-Knowledge Written Test
— New job analysis
— Designed by an industrial/organizational psychologist
— HFD departmental based
— Pass/Fail Exam (test designer to determine cut off score)
— No rank-order list
— Test designer may elect not to use a written job-knowledge cognitive test
3. Scenario-Based Computer-Objective Test
— Situational-judgment exam -with HFD departmental scenarios
— Computer simulations, such as in-basket exercises or incident-scenario judgment
— Zero, partial, and full credit answers may be used
— Same responses receive same points
— Scored test
— Rank-order list from which all of intended promotions will proceed to assessment center
4. Assessment Center
— Rank-order list
— Sliding bands based on test accuracy as determined by consultant
— Fire Chief will document reasons for selection of each candidate within bands
— No Rule of Three
— Two-year eligibility list
— Will be applied to the May 2011 Senior-Captain exam
4a. [Blank]
— Additional points in effort to minimize disparate impact potential
■ All points stay the same as the Collective Bargaining Agreement until May 1, 2011
■ The City will propose and will bargain for a point system which does not exceed the following points:
• 10 points for seniority
• 5 points for time in rank
• Education/Certification
O 1 point-intermediate certification
O 2 points-Advanced certification
O 3 points-Masters certification or Associates degree
O 4 points-Bachelor’s degree
O 5 points-Masters degree
(Docket Entry No. 69-2, at 29-30).
The parties agree that many of these proposed changes violate the TLGC and the CBA. Under the consent decree, the test designer “may” elect to use a “written job-knowledge test” depending on the job analyses for the captain and senior-captain positions. Whether such questions are included depends on the importance of the “knowledge” component compared to the skills, abilities, and other characteristics identified for the positions. If the test designer elects to use written job-knowledge questions, they are scored on a pass/ fail basis. Only candidates who get a passing score remain promotion-eligible. A candidate’s specific score on the job-knowledge test is otherwise irrelevant. The score is not used to produce a rank-order list and the Rule of Three is abandoned as to this part of the promotional process.
The remaining two parts of the captain and senior-captain exam are not questions based exclusively on facts and information. 10 One part is a scenario-based eom *734 puter objective test. The second part uses an assessment center.
The scenario-based computer objective test is based on situational-judgment concepts, using a computer to present hypothetical situations that captains and senior captains would likely encounter on the job. One type of situational-judgment question identified in the consent decree is an “in-basket exercise.” In such an exercise, a candidate is given documents or other information creating a hypothetical fact pattern and is asked to analyze or describe a response. An in-basket exercise testing training abilities might ask the candidate to review a firefighter’s performance evaluations and identify what training that firefighter needs to improve. (Dr. Brink Report 48). The consent decree allows for scoring these situational-judgment questions on a full-credit, partial-credit, and zero-credit basis, provided that the “same responses” receive the same credit. The consent decree requires ranking the candidates according to their scores on this part. In the initial settlement agreement, the candidates’ scores on this situational-judgment component determined whether the candidate could proceed to the final phase of the exam, but the modified settlement agreement provides that all candidates advance. (Compare Docket Entry No. 69-2, at 29, with Docket Entry No. 86-1, at 3).
The final exam component is an assessment center. “An assessment center consists of multiple exercises simulating job activities that are designed to allow trained observers, or assessors, to make judgments about candidates’ behaviors as related to job performance.” (Dr. Brink Report 47). One type of simulation used in assessment centers is a “role play.” A role play “is a simulation of a face-to-face meeting between the candidate (playing the role of a job incumbent) and a trained role player acting as a person incumbents frequently encounter on the job (such as a subordinate or citizen).” (Id.). Assessment-center activities such as role play violate the CBA and TLGC. See Tex. Loc. Gov’t Code § 143.032(c) (forbidding tests that “in any part consist of an oral interview”). “Assessors” score promotional candidates’ performance on the assessment-center activities. Although there are preset criteria distinguishing better from worse performance, the scoring system is subjective and violates the TLGC and the CBA.
Another inconsistency between the TLGC and the CBA on the one hand and the consent decree provisions on the other is the requirement in the consent decree to “band” the promotional candidates’ assessment center scores. “Banding” scores means adjusting the individual test scores based on statistical analyses showing the likelihood that: (1) a candidate could score higher or lower on the same exam; and (2) the individual assessor could have given the candidate a higher or lower score for the same performance. Banding tends to convert individualized score differences into homogenized “bands” of more uniform scores. For example, three candidates’ scores of 85, 86, and 87 might be “banded” as one score of 86, depending on the results of the statistical analysis. Banding is like converting individual scores of 95%, 97%, and 100% on a 100-question multiple choice test into three “As” that are viewed as identical. The conversion is based on statistical analysis showing that an individual scoring 95% on the exam has the same chance of scoring 100% on the exam as the person scoring 100% on the exam has of scoring 95%. 11 Banding is based on the *735 assumption that small differences in scores do not reliably demonstrate superiority in the KSAOs the exam is supposed to measure. One of the City’s expert witnesses, Dr. Morris, summarized banding as follows:
“[A] band is ... saying if I made 87 and someone else made 85, is it possible that the next day I could have made 85 and they could have made 87? So a band— the band we’re trying to calculate the standard error of measurement is simply saying that certain number of times that band is going to fall within a standard error of measurement that we calculate. So, it’s a reasonable thing. And most people in our field accept using bands as a way to minimize the error that could be assumed in the minds of the decision-makers.
(Evidentiary Hr’g Tr. 100, Docket Entry No. 130).
Banding is inconsistent with the Rule of Three. Under the proposed consent decree, the final promotion decision is based on the banded assessment-center scores. Names are submitted by score “bands,” not subject to the Rule of Three that would have applied under the TLGC and the CBA. Under the Rule of Three, if the three highest scores were 85, 86, and 87, the names of those applicants would be submitted. The person who scored the 87 would be selected unless the decision-maker provided a written reason for selecting the person who scored the 86 or 85. Under the banding system, the three individuals would be treated by the decision-maker as having the same score. The “band” might also be larger than three persons; its size would be determined by statistical analyses rather than a preset number. The consent decree requires the decision-maker to select one within the band and to provide a written explanation for the selection.
The consent decree does not state the role of a candidate’s race or how the decision-maker may consider race in choosing who to promote within a band. There was testimony that using race as a factor to select a candidate within a band could reduce the exam’s disparate impact on African-American applicants. Within a band, all applicants are viewed as equal. (Evidentiary Hr’g Tr. 107-08, Docket Entry No. 130). But the consent decree does not explicitly authorize race-based promotional decisions.
II. The Evidence in the Record
At an evidentiary hearing, the parties presented evidence as to (1) whether the senior-captain exam disparately impacted black firefighters, and (2) whether the proposed changes to the captain and senior-captain exams were justified by business necessity.
A. The Evidence as to Disparate Impact of Past Exams
On February 8, 2006, the City of Houston administered the senior-captain exam to 221 promotional candidates. Of the 221 candidates taking the exam, 172 were white, 15 were black, 33 were Hispanic, and 1 was “other.” The 212 candidates who passed by scoring above 70 consisted of 166 Caucasians, 13 African-Americans, 32 Hispanics, and 1 “other.” The City promoted 70 candidates based on the rank-order list of those who passed the exam. Of those promoted, 59 were Caucasian, 2 were African-American, 8 were Hispanic, and 1 was in the “other” category. (Evidentiary Hr’g Ex. 7, Dr. McPhail Report, at 5).
The following experts submitted reports or testified as to whether the 2006 senior-captain exam disparately impacted black firefighters:
*736 • Dr. S. Morton McPhail, an industrial-psychology consultant, on behalf of the City. Dr. McPhail is licensed by the Texas State Board of Examiners of Psychologists and is a Fellow of the Society for Industrial and Organizational Psychology (“SIOP”). He has served as an adjunct faculty member in the psychology departments of Rice University and the University of Houston. Dr. McPhail has “authored scholarly articles and [has] given many symposia, presentations, and continuing education workshops for peers on issues relating to employment, and in several instances, on the topics of job analysis.” (Docket Entry No. 37-1, at 3).
• Dr. Kyle Brink, an industrial-psychology consultant and a tenure-track assistant professor in the management department of the Bittner School of Business at St. John Fisher College, testified on the plaintiffs’ behalf. Dr. Brink’s doctorate is in industrial psychology. He has experience developing and validating promotion procedures at both private companies and governmental organizations. Recently, he worked with the Personnel Board of Jefferson County, Alabama, helping end a federally imposed consent decree. (Dr. Brink Report 5).
• Dr. Kathleen Lundquist, an industrial-psychology consultant, testified for the City. Dr. Lundquist is the president and CEO of APT, Inc. She has “extensively researched, designed and conducted statistical analyses and provided consultation in the areas of job analysis, test validation, performance appraisal and research design” for “major corporations in the banking, financial services, retail, electronics, aerospace, pharmaceutical,, telecommunications, and electric utility industries, as well as for federal, state, and local agencies.” (Dr. Lundquist Aff. 1, Docket Entry No. 93-2). She has a Ph.D. in psychometrics from Fordham University.
• Dr. Winfred Arthur, a full professor of psychology and management at Texas A & M University, testified for the HPFFA. Dr. Arthur has a Ph.D. in industrial/organizational psychology from the University of Akron and “over 20 years of practical experience in the areas of test development, selection, public safety testing, and training.” He is a SIOP fellow. (Dr. Arthur Aff. 1, Docket Entry No. 89-1).
• Dr. David M. Morris, an industrial-psychology consultant, testified for the City. Dr. Morris is the president of Morris & McDaniel, Inc., an industrial-psychology consulting firm he started in 1976. He received his Ph.D. in psychology, with a specialization in industrial/organizational psychology, from the University of Southern Mississippi. Dr. Morris has authored numerous scholarly articles and is a member of the industrial/organizational division of the American Psychological Association and also a member of SIOP. (Evidentiary Hr’g Ex. 14).
All the experts agreed that the “total selection process” for promoting HFD captains to senior captain showed disparate impact under the 4/5 Rule. The Guidelines require that “[a]dverse impact is determined first for the overall selection process for each job.” Adoption of Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures, 44 Fed.Reg. 11966, 11998 (1979) [hereinafter Guidelines Questions & Answers], “The ‘total selection process’ refers to the combined effect of all selection procedures leading to the final employment decision such as hiring or promoting.” Id. The experts agreed that the rate *737 of blacks promoted to senior captain— 13.3% — is less than 4/5ths the selection rate for whites — 34.3%. (See, e.g., Dr. Brink Report 9-10). 12 But only Dr. Brink and Dr. Lundquist found disparate impact for the senior-captain exam.
Both parties’ experts testified that a 4/5 Rule violation is an unreliable basis to find disparate impact when the population size of one group is small. Only 17 black firefighters were eligible for promotion to senior captain. The experts agreed that this is too small a number to make a 4/5 Rule violation sufficient to find disparate impact. (Dr. Brink Report 10; Dr. Lundquist Aff. 6; Dr. Arthur Aff. 2-3; Dr. McPhail Report 6-7; Evidentiary Hr’g Ex. 12, Dr. Morris Report, at 0010447). There was general agreement among the experts that when group populations are small, statistical analyses should be used to determine whether the 4/5 Rule violation is the product of “chance.” This requires determining the statistical significance of the 4/5 Rule violation. Dr. Morris’s report noted that the 4/5 Rule risks “sampling error,” which statistical-significance analysis mitigates. Dr. Morris’s report stated:
[T]he 4/5ths Rule has two major limitations, precision and sampling error. The 4/5ths Rule provides a descriptive ratio; it is not a statistical test. As such, the 4/5ths Rule cannot determine if an observed disparity is the result of mere chance or an indication of underlying bias. Use of the 4/5ths Rule is limited further by sample error. Unlike statistical tests, the 4/5ths Rule does not make adjustments for sampling error and, in cases where sample sizes are small, may fail to detect disparities. More problematic, the 4/5ths Rule has proven to falsely show adverse impact when no adverse impact exists.
When sample sizes are small, results from the 4/5ths Rule will vary, often dramatically, because the composition of candidates vary each time a personnel decision is made. Statistically speaking, this variation is referred to as sampling error. The 4/5ths Rule is insensitive to sampling error. Boardman (1979) and Greenberg (1979) demonstrated that the 4/5ths Rule is susceptible to either falsely identifying adverse impact when none exists or failing to identify adverse impact when it does exist. At their very heart, statistical tests directly address sampling error.
(Dr. Morris Report 0010446-47). Dr. Brink’s report discussed how the 4/5 Rule risks “Type I” error by leading to the conclusion “that adverse impact exists, when in reality the difference in selection rates is a result of sampling error (or chance).” (Dr. Brink Report 12). Dr. Brink’s report explained how statistical tests can “control the potential amount of Type I error”:
Type I error in this context is defined as concluding that adverse impact exists, when in reality the difference in selec *738 tion rates is a result of sampling error (or chance). A statistically significant result is one in which the probability of incorrectly concluding that adverse impact exists (i.e., a Type I error) is less than a specified level; this specified level is referred to as an alpha level.... Statistical tests produce a probability value ... that determines or estimates the probability of obtaining the sample result assuming there were no differences in the population. If the [probability value] resulting from the statistical test is less than the specified alpha level, we say the result is statistically significant and would decide, based on the test, that there is adverse impact. For example, if an alpha level of .05 is chosen and the [probability value] resulting from the statistical test is less than .05, then there is less than a 5% probability that the difference is due to chance (i.e., there is less than a 5% probability of making a Type I error) and we say the result is statistically significant. Conversely, you can conclude that there is a 95% probability that the difference is not due to chance.
(Id.).
The experts also identified peer-reviewed journal articles critical of the 4/5 Rule. In 1979, Anthony Boardman and Irwin Greenberg authored analyses showing that the 4/5 Rule could lead to both Type I (falsely identifying disparate impact when none exists) and Type II (failing to identify disparate impact when it does exist) statistical errors. See Irwin Greenberg, An Analysis of the EEOCC ‘Four-Fifths’ Rule, 25 Mgmt. Sci. 762 (1979); Anthony E. Boardman, Another Analysis of the EEOCC ‘Four-Fifths’ Rule, 25 Mgmt. Sci. 770 (1979). A recent article similarly concluded that “there is a fairly high false-positive rate for the 4/5ths rule used by itself.” Phillip L. Roth et al., Modeling the Behavior of the I/5ths Rule for Determining Adverse Impact: Reasons for Caution, 91 J. Applied Psychol. 507, 519 (2006). The article’s authors cautioned that “other factors (e.g., sample size) were quite important” and recommended using “a test such as Fisher’s exact test or a chi-square test to mitigate false-positives.” Id. Also noting the 4/5 Rule’s shortcomings, Scott Morris and Russell Lobsenz recently proposed a “more complex” statistical technique, the Zir test, for evaluating disparate impact. See Scott B. Morris & Russell Lobsenz, Significance Tests and Confidence Intervals for the Adverse Impact Ratio, 53 Personnel Psychol. 89 (2000). 13
The experts applied three tests to measure statistical significance: the Fisher exact; the Pearson chi-square; and the Zd. The Fisher exact test provides “the exact probability of obtaining the observed frequency table (or one more extreme) under the null hypothesis [that no disparate impact exists]” and is particularly suited to analyses involving small sample sizes. Michael W. Collins & Scott B. Morris, Testing for Adverse Impact When Sample Size Is Small, 93 J. Applied Psychol. 463, 464 (2008). Dr. Brink reported the Fisher exact test probability to be .15 — that is, a 15% chance that the observed results were due to pure chance. (Dr. Brink Report 15). Social scientists generally require the p-value — the probability that the observed results are due to pure chance — to be less than .05 for results to be considered statistically significant.
The Pearson chi-square test estimates the probability of obtaining the observed frequency table under the null hypothesis *739 that no disparate impact exists. (Id. at 13). Dr. Brink reported the Pearson chi-square probability to be .10 — that is, a 10% chance that the observed results were due to pure chance. (Id. at 15). Like the Fisher exact test’s p-value, the Pearson chi-square p-value is greater than .05 and is viewed as statistically insignificant.
Unlike the Pearson chi-square and Fisher exact tests, the Zd test does not provide a p-score, the probability that the observed results were due to pure chance. Instead, the Zd test yields a Z statistic. The difference between the selection rates of the two groups being compared is statistically significant if the absolute value of the Z statistic is greater than 1.96. See Collins & Morris, supra, at 464. Using the Zd test, Dr. Brink reported a statistically insignificant Z statistic of — 1.66. (Dr. Brink Report 15).
Neither the Fisher exact, the Pearson chi-square, nor the Zd test demonstrated that the 4/5 Rule violation for the total selection process for senior captain was statistically significant. (Dr. Brink Report 12; Dr. Arthur Aff. 3; Dr. McPhail Report 8; Dr. Morris Report 0010448). Only one test of statistical significance, the Zir test, showed that the 4/5 Rule violation was statistically significant. Unlike the Zd test, which evaluates the observed difference in selection rates, the Zir test evaluates the difference in selection-rate ratios. Dr. Brink argued for the use of the Zir test in these circumstances because the test uses the same comparison as the 4/5 Rule and is “slightly more powerful than the Zd or chi-square tests, especially as the proportion of minorities is smaller.” (Dr. Brink Report 13-14). But the use of the Zir test is not as well supported in the literature as the other tests, especially in the context of very small sample sizes. 14 One peer-reviewed journal article explained that while the Zir test was “interesting and deserve[s] greater thought,” further research was needed on the test’s ability to evaluate disparate impact. Roth et al., supra, at 520 . The test’s creators acknowledged that the Fisher exact test will “provide a more accurate evaluation of statistical significance” than the Zir test when the smallest expected value in the analysis is less than five. Morris & Lobsenz, supra, at 97. Using a formula provided by Morris and Lobsenz, Dr. McPhail calculated the smallest expected value in the analysis to be 4.75 and concluded that use of the Zir test was inappropriate. 15 (Dr. McPhail Report 9).
Dr. Brink offered other grounds to support his conclusion that under the 4/5 Rule, there was valid evidence of disparate impact. One was the “N of 1” or “flip-flop rule.” Dr. Brink stated that the “N of 1 rule calculates an adjusted impact ratio assuming one more person from the minority group ... and one less person from the majority group were hired (and, consequently, one less minority and one more majority were not hired). If the resulting selection rates are such that the minority selection rate is now larger than the majority selection rate, selection rate differences may be attributed to small sample sizes.” (Dr. Brink Report 10).
The Guidelines illustrate the application of the flip-flop rule. The Guidelines present a hypothetical in which 80 Caucasians *740 apply for a job and 20 African-Americans apply for the job. In the hypothetical, 16 Caucasians are hired and 3 African-Americans are hired. The Guidelines use the flip-flop rule to answer the following question: “[i]s evidence of adverse impact sufficient to warrant a validity study or an enforcement action where the numbers involved are so small that it is more likely than not that the difference could have occurred by chance?”
No. If the numbers of persons and the difference in selection rates are so small that it is likely that the difference could have occurred by chance, the Federal agencies will not assume the existence of disparate impact, in the absence of other evidence. In this example, the difference in selection rates is too small, given the small number of black applicants, to constitute adverse impact in the absence of other information (see Section 4D). If only one more black had been hired instead of a white the selection rate for blacks (20%) would be higher than that for whites (18.7%). Generally, it is inappropriate to require validity evidence or to take enforcement action where the number of persons and the difference in selection rates are so small that the selection of one different person for one job would shift the result from adverse impact against one group to a situation in which that group has a higher selection rate than the other group.
Guidelines Questions & Answers, 44 Fed. Reg. at 11999.
Dr. Brink’s report showed that applying the flip-flop rule discussed in the Guidelines to the February 2006 senior-captain exam, the ratio of Caucasians compared to AMcan-Americans selected for promotion did not change. Dr. Brink concluded that this was further evidence that the sample size was not too small to invalidate the disparate-impact evidence provided by the 4/5 Rule.
Dr. Brink also cited the “one-person rule” as additional support for this conclusion. The “one-person rule is computed by taking the difference between actual minority hires ... and the expected minority hires .... If the difference is less than 1, then violations of the 4/5ths rule are likely due to small sample sizes.” (Dr. Brink Report 10). When the one-person rule is applied to the results of the 2006 senior-captain exam, the difference between actual minority hires and expected minority hires is not fewer than one. Dr. Brink concluded that “[i]n all cases, the one-person rule indicates that the violations are not due to small samples.” (Id.).
Dr. Brink’s report deemphasized the value of statistical analysis to determine disparate impact for the total selection process for senior-captain promotions. (Id. at 15-16). Dr. Brink cautioned against “dogmatic adherence” to the social scientists’ use of .05 as the statistical-significance level, stating that the .05 level should not be used in all contexts and noting that lower statistical results can be meaningful. .(Id.). Statistical analyses estimate only the likelihood that if a different pool of applicants applied for promotion to senior captain, adverse impact would result. But in the context of disparate-impact analysis, the applicant pool is fixed; there is no other relevant pool. Dr. Brink quoted from an article by Collins and Morris stating that:
when evaluating a promotion decision, the pool of candidates is relatively fixed. If the decision were repeated at a different point in time, the set of candidates under consideration would be mostly the same. In such cases, probabilities based on randomly sampling from a population ... would not apply. Similarly, probabilities based on random reassignment of participants ... would not be appropriate. Without some theoretical process for producing different patterns of *741 data ... statistical significance cannot be defined.
(Id. at 12). Dr. Brink also testified that research from Collins and Morris suggested that the Fisher exact test was “overly conservative” and recommended abandoning the test as a measure of disparate impact. (Evidentiary Hr’g Tr. 157, Docket Entry No. 130).
Dr. Lundquist analyzed data relating to the senior-captain promotion process beyond the 2006 senior-captain exam. 16 The data is summarized in the table below. The proportion of white candidates promoted to black candidates is the Adverse Impact Ration (“AIR”) set out in the last column. The 4/5 Rule requires an AIR less than .8 to show disparate impact.
Black Candidates White Candidates Black Candidates White Candidates Promoted Promoted AIR
1993 17 147 0 36 0.00
1996 13 84 1 21 0.31
1999 12 122 0 31 0.00
2002 8 104 2 64 0.41
2006 15 172 2 59 0.39
2009 17 14 136 2 31 063~
Overall 79 765 7 242
Dr. Lundquist testified that a historical analysis using data from multiple test administrations provides “a much more accurate picture of adverse impact of a promotional process.” (Dr. Lundquist Aff. 6, Docket Entry No. 93-2). Dr. Arthur, however, was dismissive of any historical aggregation approach because the data would necessarily extend beyond the 2006 exam that was at issue. Dr. Arthur did state that if one were to look to historical data, “the correct analysis would be a Mantel-Haenszel chi-square test.” (Dr. Arthur Aff. 3, Docket Entry No. 89-1). Dr. Lundquist agreed with Dr. Arthur that the Mantel-Haenszel test was the most methodologically appropriate analysis. The Mantel-Haenszel test allows statisticians to investigate the consistency of data trends over time while avoiding errors due to aggregation. Applied to the senior-captain exams, Dr. Lundquist found that the Mantel-Haenszel test showed a statistically significant pattern of adverse impact against African-Americans.
B. The Validity Evidence
1. The City’s Job Descriptions
The City has not conducted a full validity study of its promotional exams. (Dr. Brink Report 20). The City has, however, produced job descriptions for both the captain and senior-captain positions. (Docket Entry No. 96-1, at 54). The job descriptions contain a detailed listing of the responsibilities for each. The responsibilities for the captain position include: (1) supervising “the emergency response of their assigned apparatus to ensure a safe and timely response to alarms”; (2) super *742 vising “the proper development of an adequate water supply when ordered to do so or fire conditions dictate”; (3) providing “leadership in performing search and rescue operations”; (4) presenting “programs to the community on safety and fire prevention topics”; (5) assisting “in training new employees”; (6) maintaining “proper staffing for all apparatus assigned to their fire station”; and (7) maintaining “the personnel records of members assigned to their station and shift.”
Under the “specifications” section are subheadings for “basic knowledge,” “specific knowledge,” “advanced skills,” and “ability to.” Bullet points beneath the subheadings describe the KSAOs for captains. The “basic” and “specific” knowledge identified for captains include: (1) knowledge directly related to firefighting — such as knowledge about municipal and private fire protection, building construction, and water supplies; (2) knowledge about federal, state, and local laws; (3) knowledge about HFD operational procedures; and (4) other types of knowledge, such as about Robert’s Rule of Order. (Docket Entry No. 96-1, at 56). The advanced skills include management and supervision, organization, problem-solving, firefighting strategy and tactics, teaching, and public relations. (Id.). The abilities include: communication abilities, such as establishing and maintaining working relationships with subordinates as well as “impromptu public speaking”; tactical abilities, such as implementing firefighting strategies; leadership abilities, such as recognizing and responding to individual and group needs; and administrative abilities. (Id. at 57). The senior-captain KSAOs are similarly organized and described.
2. Dr. James C. Sharf s Report
Dr. James C. Sharf prepared a report for the HPFFA concluding that the promotional exams are valid based on “validity generalization,” a validation method not described in the Guidelines. Dr. Sharf, an employment consultant specializing in risk management, has published “a dozen professional publications including two peer-reviewed chapters: 1) in The Society for Industrial and Organizational Psychology Practice Series (2000), and 2) in American Psychological Association Books (2008).” Dr. Sharf served as Special Assistant to the Chairman of the EEOC from 1990 to 1993 and as the EEOC’s chief psychologist from 1974 to 1978. Dr. Sharf has “over three decades’ experience developing, implementing and defending selection and appraisal systems in both the public and private sector.” He is a Fellow of the Society for Industrial and Organizational Psychology, a Fellow of the Association for Psychological Science, and a Fellow of the American Psychological Association. (Dr. Sharf Report 2, 5).
Acknowledging that the Guidelines identify content-related, criterion-related, and construct-validity studies as proper validation methods, Dr. Sharf argued that these studies establish only a starting point for validity analysis. Dr. Sharf pointed out that the EEOC did not intend the Guidelines to preclude “other professionally acceptable techniques with respect to validation of selection procedures” because test-validity science has evolved since the Guidelines’ publication. 29 C.F.R. § 1607.14 ; see also Guidelines Questions & Answers, 44 Fed.Reg. at 12002 (“The validation provisions of the Guidelines are designed to be consistent with the generally accepted standards of the psychological profession.”). Dr. Sharf also pointed out that since the Guidelines were published, the American Psychological Association has revised the Standards for Educational and Psychological Tests (“APA Standards”) and the Society for Industrial and Organizational Psychology has revised the Principles for the Vali *743 dation and Use of Personnel Selection Procedures (“SIOP Standards”). (Dr. Sharf Report 8). In light of the revisions to the APA and SIOP Standards, and recent trends in industrial-psychology scholarly publications, Dr. Sharfs report concludes that validity generalization shows that the captain and senior-captain exams are valid, and that validity generalization better analyzes the validity of a test than the methods identified in the Guidelines.
Dr. Sharf argues that one basis for deemphasizing the Guidelines’ validation methods is the criticism by industrial psychologists of the Guidelines’ recommended approach to job analysis. Dr. Sharf characterized the Guidelines’ job analysis as requiring a detailed list of job tasks based on “observable behaviors.” Dr. Sharfs report identified several scholarly articles published in psychology journals suggesting that job analyses based on detailed task descriptions are unreliable or unhelpful. For example, a 1981 article in the Journal of Applied Psychology by Schmidt, Hunter, and Pearlman, found that detailed job analyses based on observable behaviors created the appearance of large differences between jobs “that are not of practical significance in selection.” (Id. at 11). Similarly, the SIOP Standards require only a “general” description of KSAOs and allow a “less detailed analysis ... when there is already information descriptive of the work.” (Id. at 15).
A second basis for Dr. Sharfs criticisms of the Guidelines was their emphasis— reflected, for example, in the job-analysis provisions — on “observable behaviors.” Dr. Sharf contrasted “observable behaviors” with “unobservable cognitive skills,” which the Guidelines do not emphasize. Dr. Sharf argued that the Guidelines’ focus on validity studies measuring observable behaviors produces less reliable results than validity studies measuring cognitive skills. Dr. Sharfs report cited a number of scholarly articles finding that a promotional candidate’s cognitive skills better predict performance after promotion than observable behaviors; (Id. at 30-31). He summarized the articles’ findings as follows: “The conclusion from these studies is that pencil and paper tests of cognitive ability such as verbal, quantitative and technical/problem solving abilities not only predict job performance but that they predict job performance better than any alternative — the general case of validity generalization research empirically built upon measures of cognitive ability.” (Id. at 32). Dr. Sharf acknowledged that some job learning occurs for any candidate, but he argued that recent research shows that individuals with greater cognitive ability will acquire the skills necessary to perform a job successfully. (Id at 36-37). In light of these studies, Dr. Sharf concluded that the Guidelines’ “emphasis on ‘observable behavior’ is both illogical and out of touch with contemporary industrial psychology because there is no knowledge, skill or ability which does not depend on unobservable mental processes involving cognitive abilities.” (Id. at 11).
Dr. Sharf urged that the more reliable method for analyzing validity is validity generalization. “Validity generalization is industrial psychology’s science of the general case demonstrating empirically that the cognitive abilities most studied in industrial psychology — verbal, quantitative and technical abilities — are also the best predictors of job performance.” (Id. at 14). Generally, validity generalization analyzes whether a test reliably measures the verbal, quantitative, and technical skills a job requires rather than its KSAOs. The SIOP Standards recognize validity generalization as one method for validating a cognitive-based test. A former EEOC senior attorney has also argued that validity generalization is a valid measure of a test’s validity. (Id. at 13).
*744 Dr. Sharf used validity generalization to analyze the captain and senior-captain exams. Dr. Sharf first analyzed whether the HFD’s job descriptions for the captain and senior-captain positions were valid as a job analysis for validity generalization. His focus was on whether the job descriptions reflected the “verbal, quantitative and technical/problem solving cognitive abilities” required for the positions. (Id. at 16). Describing the HFD job descriptions as “among the most comprehensive, thorough, succinctly described job descriptions that I have ever come across,” (id.), Dr. Sharf categorized the responsibilities and the KSAOs identified in the job descriptions into verbal, quantitative or technical/problem-solving skills. For example, Dr. Sharf classified a senior captain’s responsibility to assist “in department supervisory and administrative activities as assigned” as a “verbal ability,” (id. at 17); a senior captain’s responsibility to maintain “inventories of all station, apparatus, and equipment” as a “quantitative ability,” (id. at 19); and a senior captain’s responsibility to assume “command at Emergency Operations incidents” as a “technical/problem solving ability,” (id. at 20).
Dr. Sharf noted that the additional responsibilities identified in the HFD job description for the senior-captain position related to assuming control, maintaining control, coordinating, supervising, and evaluating the “most effective use.” .(Id. at 17). A captain is responsible for “siz[ing]-up the scene at emergency medical calls, as a first responder, in order to begin providing needed emergency medical intervention to mitigate the problems encountered within the HFD guidelines.” (Id. at 19). A senior captain is also responsible for ássuming “control of medical emergencies when arriving first,” maintaining “control of patient care until the arrival of a higher medical authority,” and supervising or performing “medical intervention in accordance with one’s level of training.” (Id. at 19). Based on the emphasis on control and supervision in,the HFD job description and a 1986 article published by Hunter & Hunter in Psychological Bulletin suggesting that “the more complex the job, the better cognitive ability predicted job performance,” Dr. Sharf concluded that the senior-captain position requires greater cognitive ability than the captain position.
Dr. Sharf also compared the City’s job descriptions to the job analysis for municipal firefighters created by the United States Department of Labor (the “DOL Analysis”). 18 Dr. Sharf found that the HFD’s captain and senior-captain job descriptions were consistent with the DOL Analysis. Based on these similarities and “generally accepted principles and practices of industrial psychology,” Dr. Sharf concluded that the HFD had conducted a valid job'analysis for developing an “objective, job-related Captain’s and Senior Captain’s exam.” (Id. at 22).
Using the recategorized HFD job descriptions and the DOL' Analysis, Dr. Sharf prepared a “combined job analysis.” (Id. at 22-30). Dr. Sharf then discussed whether the HFD captain and senior-captain exams validly measured the cognitive skills identified in the combined job analysis and found that they did. Dr. Sharfs report does not provide the analysis that led to his conclusion or refer to statistical studies to provide support. Dr. Sharf relies on scholarly articles arguing that individual differences in cognitive performance *745 correlate to job performance to support his conclusion that the captain and senior-captain exams are valid under validity generalization.
Dr. Sharf s expert report also criticized noncognitive measures of test validity. He argued that video- or situational-based assessments are poor simulations of actual situations a captain or senior captain encounters. He colorfully explained that “[a] video depiction is hardly the stress of an adrenalin rush from the danger of a whiff of noxious chemicals or a lung full of searing smoke.” (Id. at 37). Dr. Sharf also emphasized that the knowledge a candidate brings to the position of captain — • “what you think” — -will impact how he responds to emergency situations. He argued that having the technical knowledge required to respond to certain emergency situations is a prerequisite to making the appropriate response.
3. Dr. McPhail’s Report
Dr. McPhail conducted a criterion-related validation of the 2006 captain exam. He did not conduct a similar validation study for the 2006 senior-captain exam. Dr. McPhail’s criterion-related validation study compared candidates who were promoted to captain based on the 2006 exam with those who were not but who “rode up” as captain after the exam. The validation study analyzed whether there was a relation between success on the exam and performance as captain by comparing the promotional candidates’ 2006 exam scores with performance evaluations created by Dr. McPhail and filled out by supervisors. Dr. McPhail concluded that the validation study showed only “equivocal” results as to the captain exam’s validity.
Initially, Dr. McPhail identified “a set of performance dimensions appropriate and important to effective performance as a Captain,” using the HFD’s job description for the position, the exam, and information from internal SMEs. (Docket Entry No. 37-1, at 6). Dr. McPhail’s report identified nine specific performance dimensions: “emergency operations”; “station management”; “technical knowledge”; “management resources”; “supervision”; “problem solving”; “interpersonal effectiveness”; “professional orientation & commitment”; and “overall job performance.” 19 (Id. at 37) . After more discussions with the SMEs, Dr. McPhail developed “performance evaluation behavioral anchors” related to each dimension. Dr. McPhail then asked the SMEs to evaluate each behavioral anchor’s relative importance. Based on these evaluations, the behavioral anchors were classified according to a five-point rating scale. For example, a five-point “exemplary rating” for “emergency operations” requires the following behavioral anchors: applying appropriate triage to prioritize transport of injured persons, quickly sizing up “the scene and us[ing] resources to prioritize evacuation of threatened occupants of multi-story building while simultaneously maintaining communication with IC and initiating interior fire attack,” and identifying “immediate rescue situation in burning building and prioritizing] deployment of truck line to protect tactically correct extrication of victim prior to laying supply line.” (Id. at 38) . A three-point “meets job requirements” rating requires only one behavioral anchor: laying “proper supply line before entering building to search for unknown *746 possible victims and ensure safety of fire fighters.” (Id.).
Using the performance dimensions, the behavioral anchors, and the SME evaluations, Dr. McPhail created a Performance Dimension Rating Form (“PDRF”). (Id. at 40). Each PDRF asks a supervisor to analyze one performance dimension of the individual to be scored. The performance dimension is identified at the top of the PDRF. The five rating categories are placed on a left-hand column and the behavioral anchors are listed beneath each. Within each rating category are twelve possible scores. The possible scores start at one, which corresponds to the “unacceptable” rating category — the lowest rating possible — and end at sixty, which corresponds to the “exemplary” rating. (Id.). An individual whose performance in emergency operations is “unacceptable” can receive a score from one to twelve and an individual whose performance is exemplary can receive a score from forty-nine to sixty. (Id.).
The PDRFs were circulated to HFD district chiefs, who supervise captains. The district chiefs were asked to score the performance dimensions for captains promoted after the January 2006 captain exam and EOs who were not selected for captain after the exam but had ridden up as captains after that. In January 2006, 438 firefighters took the exam and 157 were promoted to captain. Of those who took the exam but were not promoted, 281 EOs rode up as captains. Of the 84 district chiefs asked to evaluate performances, 77 submitted evaluations. The results of the supervisors’ assessments of the captains and EOs was then compared to the scores on the January 2006 captain exam. (Id. at 60). The raw data for the PDRFs showed a mean score of 79.35, with a 10.13 standard deviation, for the 438 firefighters who took the January exam; a mean score of 90.10 with a 3.83 standard deviation, for the 157 firefighters promoted to captain based on the January exam; and a mean score of 73.35, with a 7.13 standard deviation, for EOs who took the January exam but who were not promoted and later rode up as captains. (Id. at 45).
Dr. McPhail constructed a “validation sample” of 199 firefighters to test the validity of the 2006 exam. Using the validation sample, Dr. McPhail conducted statistical analyses “to evaluate evidence for the validity of the promotional examination.” (Id. at 52). The analyses included “zero-order (bivariate) correlations for three different samples: the entire validation sample, only those promoted to captain, and only those not promoted to captain.” (Id.). Dr. McPhail found that the bivariate correlations “appeared to provide supporting evidence for the validity of the examination.” He noted that the correlations “were all significant and ranged from r = .37 to r = .51.” (Id.). But when Dr. McPhail placed the results of the bivariate analysis on a scatter plot, he noted that the relationships between the “examination scores and criteria indicated barbell shaped bivariate distributions, in which most of the performance ratings for captains were located in the upper end of the distribution and most of the performance ratings for engineers/operators were located in the lower end of the distribution.” (Id.). He explained that this supported at least two inferences: (1) “the captain promotional examination effectively taps the intended construct domain which results in the observation that those scoring higher on the exam tend to have higher performance”; and (2) “because promotional examination scores were used as a basis for promotion ... scores should be correlated with captain performance because those at the formal captain rank have a greater opportunity to acquire knowledge and skills integral to effective functioning.” (Id.).
*747 Dr. MePhail also separated the validation sample into those who were promoted to captain and the EOs who were not. He then conducted bivariate correlations between the exam scores and the performance criteria for each group. For the captain sample, there were no significant correlations between the promotional exams and the performance criteria. For the EO sample, the results were “inconsistent.” Dr. MePhail found that the promotional exam correlated to station management, management of resources, and problem solving, “but were not significantly correlated with other criteria.” (Id. at 54). 20
Dr. MePhail also used “a number of multiple regression models” to measure the January 2006 captain exam’s validity. He explained these analyses, as follows:
For each criterion, a variable identifying individuals as either captains or [EOs] and the captain promotional examination variable were entered as predictors in the regression equation. If the regression weight associated with the promotional examination variable was significant, the results suggested that the promotional examination predicts the respective criterion variable above and beyond the prediction provided by the variable identifying rank. That is, the effect of being a captain versus an [EO] on the criterion variable is accounted for, and any additional prediction provided by the promotional examination variable can be interpreted as the unique impact of promotional examination scores.
(Id. at 55). The regression analyses showed that the January 2006 captain exam significantly predicted station management, resource management, and problem solving. (Id.).
Dr. MePhail concluded that the statistical analyses showed “equivocal evidence of the predictive capability of the 2006 examination.” (Id. at 60). He noted that the analyses of the entire sample showed “substantial and statistically significant correlations ... between the test and rated performance,” but that the bivariate scatter plot moderated the correlations. (Id.). He also noted that both the bivariate analysis and the multiple-regression analysis showed significant correlations within the EO subgroup with station management, management of resources, and problem solving. Dr. MePhail concluded that “among a xnuch less restricted sample, test scores provided incremental prediction of performance ... even after accounting for the relationship of promotion status with the performance ratings.” (Id. at 61).
4. Dr. Brink’s Report and Testimony
Dr. Brink’s report, prepared for the City, used the Guidelines, SIOP Standards, and scholarly articles to criticize the 2006 captain and senior-captain exams. Dr. Brink’s report criticized the job descriptions offered by the HFD as job analyses, the “linkage” between the HFD’s job descriptions and the exam, the reliability of the exam, and the processes the City used to establish the promotional system. Based on these criticisms, Dr. Brink concluded that the captain and senior-captain exams were not content-valid and that the City’s promotional process violated the Guidelines. Dr. Brink’s expert report also identified alternative evaluation measures.
Dr. Brink’s report identified numerous shortcomings in the HFD job analyses. The report explained that both the Guidelines and SIOP Standards emphasize the importance of a job analysis to determine content validity. The Guidelines contain *748 detailed requirements for the job analysis prepared for a content-validity study. See generally 29 C.F.R. § 1607.15 (C). Similarly, the SIOP Standards state that “[evidence for validity based on content rests on demonstrating that the selection procedure adequately samples and is linked to the important work behaviors, activities, and/or worker ... knowledge, skills, abilities, and other characteristics ... defined by the analysis of work.” (Dr. Brink Report 20). The report proceeded to detail the weaknesses of the City’s job descriptions as job analyses in light of the Guidelines and the SIOP Standards.
The report first discussed the City’s failure to maintain records of “required information” for documenting validity. Dr. Brink described the records the City produced to document the promotional tests’ validity as “almost 8,000 mishmash pages.” (Id. at 20). Under the Guidelines, the City should record the users, locations, and dates of a job analysis and any purposes related to the analysis. 29 C.F.R. § 1607.15 (C)(1), (2). 21 The City did produce questionnaires used to develop the job descriptions, the job descriptions themselves, and “incumbent frequency ratings, supervisor criticality ratings, and computer overall criticalities” of the KSAOs identified in the job descriptions. (Dr. Brink Report 23). Dr. Brink, however, found that these documents did not adequately validate the City’s promotional processes for the captain and senior-captain positions.
Dr. Brink’s report concluded that the job descriptions based on questionnaire responses were insufficient as job analyses. A job analysis should incorporate information from a number of sources, including background research, observing SMEs performing the job, interviewing SMEs, SME focus groups, and job-analysis questionnaires. The more sources are incorporated, the stronger the job analysis will be. (Id. at 21). Dr. Brink found that the City relied exclusively on questionnaires. He noted that the job descriptions and questionnaires the City supplied stated that “any one position may not include all the duties listed, nor do the examples listed necessarily include all duties performed.” (Id.). Dr. Brink also noted that it was not clear whether internal SMEs had participated in creating the job descriptions. Dr. Brink argued that the lack of SME input would raise concerns about accuracy. Dr. Brink noted as an example that a specific knowledge identified in the captain job description was “Robert’s Rule of Order,” but that “several” captains did not know what this meant. (Id. at 22). Finally, Dr. Brink’s report emphasized the vagueness of both the questionnaires and the job descriptions. For example, the job descriptions for both captain and senior captain list “training” as a “knowledge.” (Id.). Dr. Brink argued that these descriptions failed to meet the Guidelines’ requirement that “an operational” definition should be provided for each KSAO. 29 C.F.R. § 1607.15 (C)(3).
Dr. Brink’s report also found that the City’s “incumbent frequency ratings, supervisor criticality ratings, and computer overall criticalities” of the KSAOs identified in the job descriptions showing their relative importance did not correlate with the questions asked on the captain and senior-captain exams. His criticisms for the captain job-description assessments included the following:
• ‘Responsibilities’ AA and BB (the third and fourth most critical responsibilities) were not assessed by any items[;]
• ‘Responsibilities’ P and Q (the 2 least critical of the 35 responsibilities) were *749 each ... assessed by 12 items; more items than any of the five most critical responsibilities^ and]
• Sixteen of the 18 ‘specific knowledges’ were not assessed by any items; four of them ... had some of the highest criticalities.
(Dr. Brink Report 23). He had similar criticisms for the senior-captain job-description assessments:
• ‘Responsibility’ N (the fifth lowest criticality) was ... assessed by 25 items while AA and BB (the third and fourth most critical responsibilities) were assessed by 13 items each and X (the sixth most critical) was assessed by only 1 item[;]
• The three most critical ‘specific knowledge[s]’ ... are ... assessed by 0, 16, and 2 items ... whereas other less critical specific knowledge is ... assessed by as many as 36 items[; and]
• Six of the nine ‘advanced skills’ ... were ... assessed by 39 or 40 items regardless of criticality.
(Id.).
After consulting with SMEs, Dr. Brink found that 63% of the captain exam content and 86% of the senior-captain exam content did not reflect knowledge or skills necessary for the first day of work. (Id. at 24). Dr. Brink found that this was evidence that the test violated the SIOP Standards requirement that a “selection procedure should be based on an analysis of work that defines the balance between the work behaviors, activities, and/or [KSAOs] the applicant is expected to have before placement on the job.” (Id. at 24). The Guidelines similarly require that “[f]or any selection procedure measuring a knowledge, skill, or ability the user should show that (a) the selection procedure measures and is a representative sample of that knowledge, skill, or ability; and (b) that knowledge, skill, or ability is used in and is a necessary prerequisite to performance of critical or important work behavior(s).” 29 C.F.R. § 1607.14 (C)(4). Referring to this as the “necessary-upon-promotion” requirement, Dr. Brink concluded that the test poorly assessed whether a promotional candidate has the KSAOs required to begin work as a captain or senior captain.
According to Dr. Brink, “[p]erhaps the most condemning fact regarding the job analysis is that it was completely irrelevant.” (Dr. Brink Report 26). The report stated:
The exams were developed based on the information provided in the text books that were on the ‘source materials list.’ The source materials lists for Captain and Senior Captain were established on August 17 and September 20, 2005, respectively. The Captain and Senior Captain jobs were announced ... on October 4 and October 26, 2005, respectively. It is not documented when the Captain job analysis questionnaire was commenced or completed; however, it was still ongoing well into October .... Therefore, the job analysis ... was not even completed for Captain or even started for Senior Captain until after the source materials lists were determined and announced to candidates.... This backwards approach to test validation is clearly inappropriate and invalid ....
(Id.).
Dr. Brink’s report also criticized the exam itself, concluding that the “linkage” of test questions to the captain and senior-captain positions was “too abstract.” (Id. at 31). This was inconsistent with the Guidelines, which state as follows:
There must be a defined, well recognized body of information, and knowledge of the information must be prerequisite to performance of the required work behaviors. The work behavior(s) to which each knowledge is related should be identified on an item by item *750 basis. The test should fairly sample the information that is actually used by the employee on the job, so that the level of difficulty of the test items should correspond to the level of difficulty of the knowledge as used in the work behavior.
Guidelines Questions & Answers, 44 Fed. Reg. at 12007. Dr. Brink stated that the identified responsibilities should have been linked to KSAOs and that the responsibilities and KSAOs should in turn have been linked to specific exam questions. He noted that there was no documentation linking the source materials to individual questions. Although the City provided “matrices linking responsibilities and [KSAOs] to source material lists,” Dr. Brink found that many of the matrices did not correlate with the cited portions of the source materials. (Dr. Brink Report 31). The City also provided documents linking responsibilities and KSAOs to the exam questions, but Dr. Brink found that the questions rarely correlated with the responsibilities and KSAOs. Dr. Brink gave the following examples:
For Senior Captain, responsibilities E (search and rescue) and M (assumes command) and basic knowledge K (safety accident prevention) and L (incident management systems) each were supposedly assessed by over half of the exam ... and all were supposedly assessed by items 1-45 and 71-80. Many of these questions have nothing to do with any of these responsibilities/KSA[0]s (e.g., question 25 asks about the primary cause of cardiac arrest in infants and children), much less assess all of these as well as the many other abilities that are supposedly assessed by them.
(Id. at 32).
Dr. Brink also faulted the exam for failing to assess the identified KSAOs in “the context in which they are used on the job.” (Id.). The basis for this conclusion was Dr. Brink’s experience that objective multiple-choice exams poorly evaluate supervisory, leadership, and communication skills and that such exams fail to simulate situations using job-related abilities. During the evidentiary hearing, Dr. Brink distinguished between “high fidelity” tests, which closely resemble actual job behaviors, and “low fidelity” tests, which do not resemble job behaviors. Dr. Brink testified that multiple-choice tests are “low fidelity” and that only high-fidelity tests are likely to have content validity. (Evidentiary Hr’g Tr. 149, Docket Entry No 130).
Dr. Brink’s report also stated that low-fidelity tests, are inconsistent with the Guidelines. The Guidelines state that:
The closer the content and the context of the selection procedure are to work samples or work behaviors, the stronger is the basis for showing content validity. As the content of the selection procedure less resembles a work behavior, or the setting and manner of the administration of the selection procedure less resemble the- work situation, or the result less resembles a work product, the less likely the selection procedure is to be content valid, and the greater the need for other evidence of validity.
29 C.F.R. § 1607.14 (C)(4). Similarly, the Q & As in the Guidelines note that:
Paper-and-pencil tests which are intended to replicate a work behavior are most likely to be appropriate where work behaviors are performed in paper and pencil form (e.g., editing and bookkeeping). Paper-and-pencil tests of effectiveness in interpersonal relations (e.g., sales or supervision), or of physical activities (e.g., automobile repair) or ability to function properly under danger (e.g., firefighters) generally are not close enough approximations of work behaviors to show content validity.
*751 Guidelines Questions & Answers, 44 Fed.Reg. at 12007.
Dr. Brink identified a number of questions on both the captain and senior-captain exams to illustrate his arguments against the exclusive use of multiple-choice questions. (Dr. Brink Report 33). A candidate’s ability to delegate is examined by asking for the definition of “delegate”; a candidate’s ability to manage a station is examined by asking about the difference between “managers” and “leaders”; and a candidate’s ability to ensure the safety of personnel is measured by asking about the definition of “human factors theory of accident causes.” (Id.). Dr. Brink summarized this aspect of his findings about the exams:
There is no evidence that memorization of trivial management and diversity facts and definitions is related to success as Captain or Senior Captain.... The exams assess the ability to learn a body of information; at best, this is only one of many determining factors with respect to success as a Captain or Senior Captain. The exams also assess characteristics such as reading comprehension, test-taking ability, motivation to study, ability to memorize, leisure time, and disposable income (Captain source material cost $217.29 and Senior Captain source materials cost at least $177.23 ...). These characteristics are all important for success in school; however, none of these characteristics are included in either job analysis.
(Id. at 34).
Dr. Brink also examined the City’s use of time limits and a cutoff score, finding no validity support for either. As to time limits, the Guidelines state that “[establishment of time limits, if any, and how these limits are related to the speed with which duties must be performed on the job, should be explained.” 29 C.F.R. § 1607.15 (C)(5). As to cutoffs, the Guidelines state the following:
The methods considered for use of the selection procedure (e.g., as a screening device with a cutoff score, for grouping or ranking, or combined with other procedures in a battery) and available evidence of their impact should be described (essential). This description should include the rationale for choosing the method for operational use, and the evidence of the validity and utility of the procedure as it is to be used (essential). The purpose for which the procedure is to be used (e.g., hiring, transfer, promotion) should be described (essential). If the selection procedure is used with a cutoff score, the user should describe the way in which normal expectations of proficiency within the work force were determined and the way in which the cutoff score was determined (essential). In addition, if the selection procedure is to be used for ranking, the user should specify the evidence showing that a higher score on the selection procedure is likely to result in better job performance.
Id. § 1607.15(C)(7). The Guidelines also state that:
Where cutoff scores are used, they should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force. Where applicants are ranked on the basis of properly validated selection procedures and those applicants scoring below a higher cutoff score than appropriate in light of such expectations have little or no chance of being selected for employment, the higher cutoff score may be appropriate, but the degree of adverse impact should be considered.
Id. § 1607.5(H).
Dr. Brink also found that the City violated the Guidelines and SIOP Standards *752 requirement that an employer consider alternative measures for promotional systems. His report stated that both the Guidelines and the SIOP Standards require an employer, after conducting a job analysis, to determine how to assess the KSAOs identified. Because there was no evidence that the City considered any format other than an objective, multiple-choice test, Dr. Brink concluded that the City’s approach was inconsistent with the SIOP Standards and Guidelines. (Dr. Brink Report 29).
Using the City’s data on test scores, Dr. Brink also reviewed the “item analysis” the City conducted. An “item analysis” measures a test’s reliability by looking at its measurement error. “[F]or an employment test to accurately predict job performance, it must be reliable; but having a reliable test does not guarantee accurate prediction of job performance.” (Id. at 39). Dr. Brink measured reliability using a formula proposed in an article by Ghiselli, Campbell, and Zedeck. The formula produces a “validity coefficient” that ranges from 0 to 1. The higher the coefficient, the more valid the test. (Id. at 38). Relying on an article by Nunnally and Bernstein, Dr. Brink acknowledged that “the level of reliability that is considered satisfactory depends on how a test is being used,” but that in all cases, the validity coefficient should be at least .70. (Id. at 39). Dr. Brink stated that for “high stakes testing” like the captain and senior-captain exams, which provide the most important aspect of the promotional decision, the validity coefficient should be .90 at a “bare minimum,” and .95 is “desirable.” (Id.).
The City has conducted “some” item analysis for the captain exam and has conducted an item analysis for 7 items on the senior-captain exam. The City’s item analysis showed validity coefficients ranging from .458 to .646 for criteria measured on those exams. Dr. Brink criticized the City’s item analysis because it analyzed broad criteria. For example, on the senior-captain exam, the City measured validity against the following criteria: “strategic & tactical considerations on the fire-ground” (items 1-20); “fire service first responder” (items 21-45); “supervisor” (items 46-70); “terrorism response” (items 71-80); and “Houston Fire Department Guidelines” (items 81-100). (Id.). Dr. Brink’s report stated that the City’s item analysis failed to include any analysis based on the specific KSAOs within each large criteria group. Dr. Brink also stated that the relevant literature shows that measuring large numbers of items at the same time increases the validity coefficient. The City measured the validity of all 100 test items instead of measuring the test validity within each subcriteria, which would inflate the validity coefficient while failing to capture important criteria.
Dr. Brink also measured the “item difficulty” of the exams. “Item difficulty is a statistic used in test analysis that indicates the percentage of applicants who answered an item correctly.” (Id. at 40). Item difficulty measures reliability, not validity. Dr. Brink stated that “[t]he purpose of a valid promotional exam is to differentiate candidates based on job-related criteria; if all or most candidates get an exam question correct or incorrect, the item is useless for this purpose.” (Id.). Items with difficulties above .9 (90% applicants answered correctly) should be eliminated unless the exam is designed to separate the bottom 10% of applicants from the top 90%. Dr. Brink found that 34 items on the captain exam had item difficulties over .9 and that 60 items for the senior-captain exam had item difficulties above .9.
Another way to measure whether an item differentiates between candidates who will perform well after promotion and candidates who will not is through “item *753 discrimination.” As the label indicates, item discrimination “provides an indication of whether an item appropriately discriminates or differentiates between those examinees who perform well on the test and those who perform poorly.... If performance on an item is unrelated to performance on the overall exam, then that item is not differentiating among candidates as intended.” (Id. at 41).
The two common methods for measuring item discrimination are the index of discrimination and the item-total correlation. The index of discrimination is computed by first dividing the examinees into upper and lower groups based on overall test scores, then subtracting the proportion of the lower group who answered the item correctly from the proportion of the upper group who answered the item correctly. This produces a value, “D.” Crocker and Algina argue that questions with a “D” value lower than .2 should be eliminated. Dr. Brink found that the captain exam had 53 items with a “D” value below .2 and the senior-captain exam had 66 items with a “D” value below .2. (Id.).
The item-total correlation “represents the correlation between an item and the rest of the test (i.e., the correlation between the item and the total score on the exam calculated excluding that item).” (Id.). A low item-total correlation means that an item has little relationship with the overall test and does not discriminate between those who perform well and those who perform poorly. Dr. Brink’s report stated that items with “low item-total correlations should be dropped ... because they are not operating in the intended manner and do not improve reliability.” (Id.). Nunnally and Bernstein wrote an article concluding that items with an item-total correlation below .05 is “a very poorly discriminating item” and that items with an item-total correlation less than .2 “are at least moderately discriminating.” (Id. at 42). Dr. Brink found that 45 items on the captain exam and 46 items on the senior-captain exam had item-total correlations below .2. He also found that 6 items on the captain exam had negative item-total correlation, indicating that high-scoring candidates were more likely to get the item incorrect than were low-scoring candidates.
Finally, Dr. Brink faulted the exams for producing statistically significant performance differences between black and white examinees. Dr. Brink found that 32 items on the captain exam and 12 items on the senior-captain exam showed statistically significant differences based on a chi-square analysis using p-values of less than .05. (Id.). Dr. Brink argued that “[although changes to tests should not be made based solely on significant group differences, these items should be [the] focus of further evaluation to ensure that they are functioning appropriately.” (Id.). The City had conducted no such evaluation.
Dr. Brink also calculated item bias through differential-item functioning (“DIF”). The SIOP Standards state that test developers should attempt to detect and eliminate aspects of test design, content, and format that may bias test scores for particular groups. DIF is intended to measure such sources of bias. Dr. Brink used the Mantel-Haenszel method to examine DIF. He stated that “[r]ace groups may differ with respect to performance on a particular item due to true differences (i.e., for some reason, there are real job-related differences between Blacks and Whites with respect to performance on the item within the sample) or race bias (i.e., there are not real job-related differences between Blacks and Whites with respect to performance on one item; for some reason, performance differences are occurring on the item because the item is biased *754 against one of the race groups).” (Id. at 43). Dr. Brink’s report stated that the Mantel-Haenszel analysis determines the extent to which black-white test-result differences are due to bias by examining the extent to which such differences exist after taking into account overall test performance. Relying on an article by Biddle, Dr. Brink stated that items with DIFs with p- values below .1 should be considered “meaningful” and those below .05 should be considered “even more substantial,” but that items at both levels should be considered for removal from the test. (Id.). Dr. Brink found that 16 questions on the captain exam and 10 questions on the senior-captain exam showed bias. He believed that the City should have conducted DIF analysis and considered dropping these questions.
During the evidentiary hearing, Dr. Brink was asked about his findings on the multiple-choice examination format. Dr. Brink admitted that “to some degree,” multiple-choice questions can measure more than job knowledge. (Evidentiary Hr’g Tr. 211, Docket Entry No. 130). But he also testified that skills such as communication and “interpersonal type of abilities” are poorly measured through such job-knowledge tests. (Id. at 212). Dr. Brink also testified that while written tests can measure leadership, command presence, and decision-making ability to a degree, there are better ways to measure these skills and abilities. (Id. at 221). He also testified that situational-judgment questions better measure these skills and abilities than written job-knowledge questions, though he preferred “high-fidelity” exercises such as those performed at an assessment center over any type of written questions. (Id. at 222).
Dr. Brink also cited empirical research by Dr. Arthur. Dr. Brink testified that this research showed that “written tests with open-ended responses [were] actually more reliable than a written test with the closed-ended multiple-choice type responses.” (Id. at 223). Dr. Arthur’s study used a criterion-related validity study to compare the reliability of a multiple-choice exam to a “constructed response exam” that required the individuals to generate— rather than select — responses to exam questions. The construct-response questions were short-answer questions with a structured response format scored according to preestablished criteria. Winfred Arthur, Jr. et al., Multiple-Choice and Constructed Response Tests of Ability: Race-Based Subgroup Performance Differences on Alternative Paper-and-Pencil Test Formats, 55 Personnel Psychology 985, 996 (2002). The study found that the construct-response questions had higher reliability measures than multiple-choice questions. Id. at 998, 1000. The study also found that there was less subgroup difference from the construct-response questions, though the authors acknowledged that the sample size was small. Id. at 1001-02, 1004.
5. Dr. Lundquist’s Affidavit and Testimony
Dr. Lundquist, a witness produced by the City, provided an affidavit and testimony on the validity of multiple-choice exams. She also testified about the validity of assessment centers. Dr. Lundquist concluded that the City’s exclusive use of multiple-choice job-knowledge questions should be abandoned in favor of a promotional exam system incorporating assessment centers.
Pointing to the Guidelines, Dr. Lundquist stated that “[t]he emphasis for any promotional process should be on assessing the critical knowledge, skills, abilities, and other personal characteristics (KSAOs) identified through a job analysis as being required to perform the essential duties of the job.” (Dr. Lundquist Aff. 3, Docket Entry No. 93-2). She acknowl *755 edged that the City’s multiple-choice test could validly assess “the technical knowledge” required for the captain and senior-captain positions, but argued that such a test “inadequately captures the range of KSAOs required for successful performance in a position such as Senior Captain.” (Id. at 4). Specifically, she argued that a multiple-choice test fails to test “supervisory and leadership skills and abilities.” (Id.).
Dr. Lundquist’s affidavit stated that situational-judgment questions and assessment centers better measure many of the abilities senior captains need. Dr. Lundquist argued that the literature shows that situational-judgment questions assess leadership and supervisory skills and abilities and should supplement, not replace, the multiple-choice job knowledge questions. Dr. Lundquist pointed to journal articles supporting the fairness and validity of assessment centers, as well as their ability to minimize disparate impact. During the evidentiary hearing, Dr. Lundquist discussed an article by Dr. Arthur concluding that assessment centers can validly measure “organization and planning and problem solving, ... [and] influencing others.” (Evidentiary Hr’g Tr. 233, Docket Entry No. 130). Dr. Arthur’s study used “meta-analysis to empirically assess the criterion-related validity of separate dimensions tapped by assessment centers.” Winfred Arthur, Jr. et di, A Meta-Analysis of the Criterion-Related Validity of Assessment Center Dimensions, 56 Personnel Psychology 125, 128 (2003). The study found “true validities” for the following performance dimensions: problem-solving, influencing others, and organizing and planning. Id. at 140.
Dr. Lundquist was asked about the objection that assessment centers produce “subjective” scores. She testified that through scoring standards and effective assessor training, assessment centers can produce scores approximating the objectivity of multiple-choice tests. (Evidentiary Hr’g Tr. 260-61, Docket Entry No 130). Dr. Lundquist admitted that scoring an assessment-center exercise is more subjective than scoring a multiple-choice test. But she argued that “reliability and consistency can be produced by certain controls in the design of the ... assessment center exercise itself.” (Id.). She also explained that objectivity and subjectivity are best understood as existing on a continuum, and that subjective scoring becomes more “objective” by structuring the scoring to minimize an assessor’s subjective evaluation of a promotional candidate’s performance. (Id. at 262-63). Dr. Lundquist testified that providing assessors with “very detailed examples of what is high performance, what is average performance, [and] what is low performance” for a particular performance dimension minimizes the assessor’s subjective evaluation. (Id. at 264-65). She also testified that using multiple assessors reduces subjectivity. (Id. at 265).
On cross-examination, Dr. Lundquist admitted that “supervisory skill and planning and coordination skills” are difficult to measure and noted that “[w]hether or not it’s a good assessment depends entirely on how well-written the test is and how similar it is to the requirements of the job.” (Id. at 268). She testified that other methods besides a written test might be better measures of such skills. (Id.).
In response to questions about the importance of measuring cognitive skills, Dr. Lundquist acknowledged that industrial-psychology literature shows that “cognitive ability ... underlies a lot of the performance, a lot of the learning that goes on in terms of any kind of job.” (Id. at 235). But she emphasized that there are different types of cognitive skills and not all are important for the captain and senior-cap *756 tain positions. She explained that “you can think of cognitive ability as akin to an IQ test.... But if I think about a particular job, I don’t want to know your IQ score. I want to know how well you do planning and decision making.” (Id,.). She testified that an assessment center would test the specific cognitive skills and abilities related to the captain and senior-captain jobs. (Id. at 236). . For example, to test supervisory skills, “there will be some level of cognitive ability involved.” (Id.). Dr. Lundquist, citing empirical studies, argued that “cognitively loaded” tests tend to produce subgroup differences. (Id. at 271). Dr. Lundquist testified that situational-judgment questions by themselves do not sufficiently measure such noncognitive skills and abilities as command presence, interpersonal communication, and leadership, because the questions are “cognitively loaded.” Dr. Lundquist explained that situational-judgment questions require “a lot of reading.” (Id. at 256). While asking such questions is “intended to measure application of knowledge,” it “just does it in a way that is fairly complex in terms of getting to the question that people are having to answer.” (Id.). She testified that “if you’re trying to measure something that is not essentially cognitive, like interpersonal skills, ... and you do that by requiring somebody to do a lot of reading, you’ll get a cognitive component to the way the person performs on that test that’s unrelated to what you’re really trying to measure in the first place, which is interpersonal skills.” (Id.). Dr. Lundquist testified that the situational-judgment test used for the November 2010 captain exam was “not that different from the job knowledge test” because “[ijt’s another written multiple-choice test and even though you’re calling it something different and trying to measure something more applied, it’s not different enough to be producing a result that covers more of the space of the skills required.” (Id. at 258).
6. Dr. David Morris’s Testimony
Dr. Morris testified about the jo

[Text truncated at 120,000 characters. The full text is on the page linked above.]

---

Source: Frix Law Library, https://www.frixlaw.com/law-library/cases/8698171. Public record. Not legal advice.
