# Appendix — Personnel Board of Jefferson County v. United States

> Briefs, arguments, decisions, and more.

URL: https://www.frixlaw.com/law-library/documents/brief%3Amicro_IA40385007_0425%3A2

## Record

- **Collection:** Supreme Court brief
- **Document type:** Appendix
- **Published:** January 1, 1980
- **Citation:** 449 U.S. 1061

## Text

a 0 -4 0 6 | mess ‘> : S.

No. SEP 13 1980

Supreme Court of the United States

OCTOBER TERM, 1980

THE PERSONNEL BOARD OF JEFFERSON COUNTY, ALABAMA,
Petitioner,
versus
THE UNITED STATES OF AMERICA,
Respondent.

THE PERSONNEL BOARD OF JEFFERSON COUNTY, ALABAMA,
Petitioner,
versus
ENSLEY BRANCH OF THE NATIONAL ASSOCIATION
For THE ADVANCEMENT OF COLORED PEOPLE, ef al.,
Respondents.

THE PERSONNEL BOARD OF JEFFERSON COUNTY, ALABAMA,
Petitioner,
versus
JOHN W. MarrTIN, ef al.,
Respondent.

THE PERSONNEL BOARD OF JEFFERSON COUNTY, ALABAMA,
Petitioner,
versus
Lucy WALKER, ef al.,
Respondents.

APPENDIX TO PETITION

FOR A WRIT OF CERTIORARI
to the United States Court of Appeals
for the Fifth Circuit

HuBertT A. GrissSoM, JR.
Davip P. WHITESIDE, JR.

Of Counsel: Attorneys for Petitioners

JOHNSTON, BARTON, PROCTOR,
SWEDLAW & NAFF
Twelfth Floor, Bank for Savings Building
Birmingham, Alabama 35203

St. Louis Law Printing Co., Inc., 411 No. Tenth Street 63101 314-231-4477

~-.

TABLE OF CONTENTS

Page
Appendix A - Order of U.S. Court of Appeals Denying
Petition For Rehearing ............... A-l
Appendix B - Opinion of U.S. Court of Appeals ..... A-3
Appendix C - Opinion of U.S. District Court ........ A-28
Appendix D - Order of U.S. District Court .......... A-64
Appendix E - Civil Rights Act of 1964 (42 U.S.C.
SS oS ca cane aci enon ey Suhaeees A-66

Appendix F - Uniform Guidelines on Employee Selec-
tion Procedures (1978); 28 C.F.R.
§50.14 (1978), 29 C.F.R. §1607 (1979) .. A-76

Appendix G - Department of Justice Guidelines on
Employee Selection Procedures (28
on es RE | ere A-138

Appendix H - E.E.O.C. Guidelines on Employment
Selection Procedures (29 C.F.R. §1607
ERR ga AP OL Seder nae 2 Sea A-171

—_ = ee

APPENDIX A

IN THE UNITED STATES COURT OF APPEALS
FOR THE FIFTH CIRCUIT

No. 77-1819

Ensley Branch of the N.A.A.C.P., et al.,
Plaintiffs-Appellees,

versus
George Seibels, et al.,
Defendants,
Personnel Board of Jefferson County,
Defendant-Appellant.

John W. Martin, et al.,

Plaintiffs-Appellees
Cross Appellants,

versus
City of Birmingham, et al.,
Defendants,

Personnel Board of Jefferson County,

Defendant-Appellant
Cross Appellee.

United States of America,

Plaintiff-Appellee
Cross-Appellant,

versus
Jefferson County, et al.,
Defendants,
Personnel Board of Jefferson County,

Defendant-Appellant,
Cross Appellee.

a me

Lucy Walker, et al.,
Plaintiffs-Appellees,

versus
Jefferson County Home, et al.,
Defendants,
Personnel Board of Jefferson County,
Defendant-Appellant.

Appeals from the United States District Court for the
Northern District of Alabama

ON PETITION FOR REHEARING

Before GODBOLD, RONEY and ANDERSON, Circuit
Judges.

PER CURIAM:

IT IS ORDERED that the petition for rehearing filed in the
above entitled and numbered cause be and the same is hereby
denied.

Dated June 16, 1980.

es eel ome

APPENDIX B

Ensley Branch of the N.A.A.C.P., et al.,
Plaintiffs-Appellees,
v.
George Seibels, et al.,
Defendants,
Personnel Board of Jefferson County,
Defendant-Appellant.
John W. Martin, et al.,
Plaintiffs-Appellees
Cross Appellants,
v.
City of Birmingham, et al.,
Defendants,
Personnel Board of Jefferson County,
Defendant-Appellant
Cross Appellee.

United States of America,
Plaintiff-Appellee
Cross-Appellant,

Vv.

Jefferson County, et al.,
Defendants,

Personnel Board of Jefferson County,

Defendant-Appellant,
Cross Appellee.

a Se

Lucy Walker, et al.,
Plaintiffs-Appellees,

Vy

Jefferson County Home, et al.,
Defendants,

Personnel Board of Jefferson County,
Defendant-Appellant.

No. 77-1819

UNITED STATES COURT OF APPEALS
FIFTH CIRCUIT

May 8, 1980

Appeals from the United States District Court for the Nor-
thern District of Alabama.

Before GODBOLD, RONEY and ANDERSON, Circuit
Judges.

R. LANIER ANDERSON, III, Circuit Judge:

This case is a consolidation of four separate actions brought
by the United States and private parties against the Personnel
Board of Jefferson County, Alabama, and other local govern-
ment agencies. The district court held that the Persoanel
Board’s use of two examinations in screening and certifying
candidates for jobs as police officers and firefighters violates Ti-
tle VII of the Civil Rights Act of 1964, as amended, 42 U.S.C.
§2000e et seq. The court ordered certain remedies, the remedies
being tailored to its conclusion that use of the police officer test
did not violate Title VII until April 25, 1975, and that use of the
firefighter test did not violate Title VII until July 8, 1976. We
agree with the district court that the Personnel Board’s use of
both tests abridges Title VII; however, we believe that the court

ay ee

failed to make essential findings in determining the time at
which liability commenced. Accordingly, we affirm in part and
reverse and remand in part with instructions.

On January 4, 1974, the Ensley Branch of the National
Association for the Advancement of Colored People, together
with certain named individuals, for themselves and on behalf of
others similarly situated, filed a complaint in the United States
District Court for the Northern District of Alabama, against
George Seibels (then Mayor of Birmingham, Alabama), the City
of Birmingham, the members of the Personnel Board of Jeffer-
son County, and the Personnel Director of that Board, alleging
that the defendants engage in discriminatory hiring practices
against blacks in violation of the Fourteenth Amendment, 42
U.S.C. §§ 1981, 1983, and 2000e et seq. (Title VII). A suit rais-
ing the same constitutional and statutory allegations was filed
on January 7, 1974, by John W. Martin and other named plain-
tiffs against the City of Birmingham, Jefferson County, and the
Personnel Board of Jefferson County.' On May 27, 1975, the
United States brought suit against the Jefferson County Person-
nel Board and the municipal and other governmental jurisdic-
tions within Jefferson County’ alleging a pattern or practice of
discriminatory employment practices against blacks and women
in violation of Title VII, the Omnibus Crime Control and Safe
Streets Act of 1968, as amended, 42 U.S.C. § 3766(c), the State
and Local Fiscal Assistance Act of 1972, as amended, 31 U.S.C.
§ 1242, the Fourteenth Amendment and 42 U.S.C. § 1981. On

' Also named as defendants in the Martin suit are the Director of
the Jefferson County Personnel Board and three Jefferson County
Commissioners.

? Defendants in the government’s suit other than the Personnel
Board are Jefferson County, the County Department of Public
Health, and the cities of Bessemer, Birmingham, Fairfield, Fulton-
dale, Gardendale, Homewood, Huey Town, Midfield, Mountain-
brook, Pleasant Grove, Tarrant, and Vestavia Hills.

re

February 20, 1976, Lucy Walker filed suit challenging the
employment practices of the Jefferson County nursing home
under Title VII and 42 U.S.C. § 1981. All four cases were con-
solidated for trial.

On December 20-22, 1976, trial was held on the merits of the
limited issue of whether the two tests used by the Personnel
Board to screen and rank applicants for positions as police of-
ficers and firefighters are discriminatory and violative of the
constitutional or statutory rights of blacks.’ All other issues
under the complaints were reserved until a later date.

On January 10, 1977, pursuant to Fed.R.Civ.P. 54(b), the
district court entered a final order on this limited issue.‘ The
court held that the police and firefighter tests do not violate the
Constitution,’ but do violate Title VII. The court determined
that the Title VII violation commenced on April 25, 1975, with
respect to the police test and on July 8, 1976, with respect to the
firefighter test. The Court ordered certain remedies to correct

> At trial, plaintiffs dropped their attack on the Office Worker test
administered by the Personnel Board and also dropped their claim
that the police and fire tests discriminate on the basis of sex.

‘ The district court’s opinion and order is reported at 13
Empl.Prac.Dec. (E.P.D.) 411,504.

* The district court found that the Personnel Board’s use of the
tests in issue is not motivated by intentional discrimination. The par-
ties do not dispute this finding. This appeal, therefore, is restricted to
Title VII issues.

a Sp

the discrimination caused by use of the tests since those dates.°
From this judgment, the Personnel Board filed a notice of ap-
peal, and the United States and the plaintiffs in the Martin ac-
tion filed a joint notice of cross-appeal. The Personnel Board
contends that the tests do not violate Title VII; the United States
and the Martin plaintiffs contest the court’s determination as to
when Tiule VII liability commenced.

This case has been ably argued on appeal by counsel for both
sides,’ and we also have the benefit of a well-reasoned and com-
prehensive opinion by Judge Pointer of the court below.

FACTS

The Personnel Board of Jefferson County is required under
Alabama law to administer examinations to applicants for posi-
tions with local government agencies. Two such examinations
are at issue here: the 10—C test administered to applicants for
positions with the fire department, and the 20—B test admin-
istered to applicants for positions with the police department.

* Specifically, the Court ordered that blacks be referred for open-
ings on the police and firefighter forces at the rate at which they took
the tests when most recently administered. To accomplish this, the
Court ordered that the names of a sufficient number of blacks be
added to the current police and firefighter eligibility lists so that the
lists shall be representative of the racial composition of the test-takers,
i.e., 28 and 14 percent black for police and firefighter lists, respective-
ly; that, one-third of future certifications, i.e., referrals from the lists
for actual employment, are to be black until, considering all certifica-
tions since the relevant 1975 and 1976 dates, the numbers of certifica-
tions become representative of the racial composition of the test-
takers. Thereafter, blacks are to be certified in accordance with their
representation of the lists, i.e., 28 and 14 percent of certifications for
policemen and firefighters, respectively, will be black. Similarly,
referrals from future lists will be a function of the rate at which blacks
take the examinations on which the lists are based, until or unless
defendants develop valid tests. 13 E.P.D. at 6807-08.

’ The Educational Testing Service, Inc. filed a brief as amicus
curiae urging reversal of the district court’s holding that the tests
violate Title VII.

—” Eo

Both tests were developed by the International Personnel
Management Association. Each test is a paper and pencil in-
strument consisting of 120 multiple choice questions.’

Applicants who pass'® the 10—C or 20—B are placed on an
eligibility list for positions with the police or fire department,
and are ranked on the list in the order of their scores. If a single
vacancy occurs, the three persons at the top of the list are cer-
tified to the appropriate department for final selection. If multi-
ple vacancies occur, the number of persons certified is two more
than the number of vacancies to be filled. Thus, merely passing
the tests, and thereby being entered on the relevant eligibility
list, is not nearly so important as obtaining a score sufficiently
high to be placed high enough on the list to be actually certified
to the police or fire department.

ADVERSE IMPACT

A prima facie Title VII case against an employment test
may be built with statistics showing that use of the test has an
adverse racial impact. This showing shifts to the employer the
burden of proving that the test is job-related. Albemarle Paper
Co. v. Moody, 422 U.S. 405, 425, 95 S.Ct. 2362, 2375, 45
L.Ed.2d 280 (1975); Griggs v. Duke Power Co., 401 U.S. 424,
432, 91 S.Ct. 849, 854,28 L.Ed.2d 158 (1971); Scott v. City of
Anniston, 597 F.2d 897, 901 (Sth Cir. 1979).

The district court found that use of the 1O—C and 20—B has
an adverse impact upon blacks. The evidence shows that since

* The International Personnel Management Association was
formerly the Public Personnel Association.

* Since April 10, 1974, only 80 of the 120 questions have been used
in grading the 10—C test. We have examined the tests and note in
passing that a very substantial number of the questions seem to be
aimed primarily at verbal or mathematical ability.

'© The raw score which constitutes a passing grade fluctuates,
depending on the number of vacancies anticipated and other factors.

mae lem

March 24, 1972, the date Title VII became applicable to public
employers such as the Personnel Board, only 51 blacks (6.6%)
have been hired out of a total of 768 blacks who took the police
test. The comparable figure for whites is 455 whites hired
(23.3%) out of the total of 1,953 whites who took the police
test. Similarly, only 9 black firefighters have been hired, which
represents 3.2% of the black applicants, as compared to 215
white firefighters hired, which represents 14.1% of the white ap-
plicants. With respect to the number of blacks and whites who
passed the exam, the court found that the pass rates for blacks
(48.6% for the police test and 24.2% for the firefighter test) are
substantially less than the pass rates for whites (90.2% for the
police test and 82.5% for the firefighter test). The Personnel
Board does not in this appeal contest the district court’s finding
that the two tests have an adverse impact; the Board argues only
that it has carried its burden of showing that the tests are job-
related.

JOB-RELATEDNESS

[2] To prove the job-relatedness of the police and firefighter
tests, the Personnel Board employed criterion-related validity
studies.'' The district court held that the studies failed to

'' **Validation’’ is the process of determining whether a selection
device is sufficiently job-related to comply with the requirements of
Title VII. See Uniform Guidelines on Employee Selection Procedure
(hereinafter referred to as Uniform Guidelines], 43 Fed.Reg. 38290,
38291 (August 25, 1978). There are three basic methods of validation:
** ‘criterion’ validity (demonstrated by identifying criteria that indicate
successful job performance and then correlating test scores and the
criteria so identified); ‘construct’ validity (demonstrated by examina-
tions structured to measure the degree to which job applicants have
identifiable characteristics that have been determined to be important
in successful job performance); and ‘content’ validity (demonstrated
by tests whose content closely approximate tasks to be performed on
the job by the applicant).’’ Washington v. Davis, 426 U.S. 229, 247,
96 S.Ct. 2040, 2051, 48 L.Ed.2d 597 n.13 (1976); See Uniform
Guidelines, § 5, 43 Fed.Reg. at 38298.

— A-10 —

validate either test.'?

The studies used three criterion measures: academy grades,
i.e., relative class rank among students who successfully com-
plete academy training; efficiency ratings of incumbent
employees (police and firefighters) by their supervisors; and ex-
perimental ratings of incumbent employees by their supervisors,
based on 12 categories, the categories relating to personality
characteristics, job knowledge and job-related abilities.

The Uniform Guidelines cited in this footnote were adopted on
August 25, 1978, by the Equal Employment Opportunity Commission
(EEOC), the Civil Service Commission (CSC), the Department of
Labor (DOL), and the Department of Justice (DOJ), and are codified
in the 1979 editions of 29 C.F.R. § 1607 (EEOC); 5 C.F.R. § 300.103
(c) (CSC); 41 C.F.R. § 60-3 (DOL); and 28 C.F.R. § 50.14 (DOJ). The
Guidelines, effective September 25, 1978, supersede the guidelines on
employee selection procedure previously issued by the EEOC and the
DOJ (the DOJ Guidelines were issued jointly with the DOL and CSC).
With the exception of its provision on interim use of unvalidated tests,
discussed below, the Uniform Guidelines are consistent, in all matters
pertinent to this case, with the EEOC and DOJ Guidelines. This opin-
ion shall refer to the EEOC and DOJ Guidelines as well as to the
Uniform Guidelines; citations to the former will be to the 1978
C.F.R., citations to the latter will be to the Federal Register.

'? The Personnel Board, in 1970, initiated a preliminary, in-house,
validation study of the 10—C police officer test. This study,
developed by Eugene Williams, Chief Examiner of the Board was not
designed to meet the validation requirements of Title VII and was not
relied upon by the district court in assessing the validity of the 1O—C.

The studies upon which the district court based its holding of no
job-relatedness, and which are the subjects of our review, were con-
ducted by Drs. William McLaurian and William Farrar, psychology
professors at the University of Alabama in Birmingham. The
McLaurian-Farrar studies examined both the 10—C police officer
tests and the 20—B firefighter test. The final results of these studies,
(the studies began in 1972) were reported to the Personnel Board on
August 25, 1975, with respect to the 1O—C and on July 8, 1976, with
respect to the 20—B.

— A-ll —

Acade 1y Grades. The district court found that the scores on
the 10O—t and 20—B tests bear statistically significant'’ correla-
tion to grades in the training academies and it further found that
the training academies furnished skills or knowledge needed for
performance on the job. However, the district court also found
that academy grades are not valid predictors of job perfor-

'S Explanation of a few statistical concepts is in order. We quote at
length from the government’s brief:

Statistically, the degree of correlation between two variables
(e.g., entrance exam scores and subsequent school grades) is ex-
pressed as a ‘correlation coefficient’ on a scale running from
+1.0 to —1.0. A perfect positive correlation (e.g., entrance
exam scores exactly predict subsequent school grades, with the
higher exam scores predicting the best grades) would be ex-
pressed as +1.0, and a perfect negative correlation (e.g., en-
trance exam scores exactly predict subsequent school grades, ex-
cept in reverse, with the lower exam scores predicting the best
grades) would be expressed as —1.0. Where the two variables
had absolutely no relationship to each other, the correlation
coefficient would be .0. The closer a correlation coefficient is to
either +1.0 or —1.0, the ‘higher the magnitude’ of the correla-
tion; and the closer it is to .0, the ‘lower the magnitude.’
Mueller, Schuessler & Castner, Statistical Reasoning in
Sociology, 2d Ed., at p. 315.

Because a purely random drawing of a sample is liable to produce a
correlation coefficient which is somewhat off an absolute .0, the con-
cept of statistical significance becomes relevant. The concept is tied to
the statistical theory of probability and is dependent upon the number
of people in the sample. Generally, if a correlation coefficient is so
low that, on the basis of the sample size involved, more than | in 20
random drawings could be expected to produce a correlation coeffi-
cient is considered not to be statistically significant, or simply to be the
same as a correlation coefficient of .0. On the other hand, if the ob-
tained coefficient could be expected to reoccur no more than once in
20 random drawings, it is considered statistically significant, the
statistical indication for whichis p.< .05. Acorrelation coefficient of
the obtained magnitude which could not be expected to occur by
chance more than once in 100 random drawings is expressed as
p.< .01. Mueller, et al., pp. 394, ef seq.

— A-12 —

mance.'* The court concluded from these findings that while
completion of the academy is a valid criterion measure,
academy grades are not. Finally, the court found that even
though completion of the training academies is a valid criterion
measure, neither test predicts successful completion of the
academies; even extremely low test scores do not predict failure
at the academies.

Efficiency Ratings. Drs. Farrar and McLaurian, the experts
who conducted the validation studies, testified that the effi-
ciency ratings are not trustworthy assessments of the employees’
actual performance. Largely based on this testimony, the court
held that the efficiency ratings are not a valid criterion measure.

Experimental Ratings. Although determining that the ex-
perimental ratings system is an appropriate criterion measure
for determining job-relatedness, the court discovered fatal defi-
ciencies in the relationship between the test scores and the
ratings. With respect to the 20—B test, the court found that
while there is a statistically significant positive correlation be-
tween test scores and the experimental ratings for firefighters
having less than three years’ experience, there is a significant
negative correlation for firefighters having more than three
years’ experience, thus ‘‘suggesting that over time the lower
scoring applicants may be better employees.’’ Accordingly, the
court held that the 20—B is not a valid predictor of the ex-
perimental ratings. With respect to the 10—C test, the court
found that there is a statistically significant correlation between
test scores and experimental ratings, but that the correlation is

'* The court pointed out that ‘‘for the most part the correlation be-
tween academy grades and measures of job performance are not
significant and, in the few instances where significant correlations are
found, the findings are mixed, some being positive and others being
negative. A negative correlation, of course, indicates that the higher
the academy grades, the lower the performance ratings tend to be.’” 13
E.P.D. at 6802.

— A-13 —

of very low magnitude and lacks practical significance.'’ Accor-
dingly, the court held that the correlation between the 10—C
and the experimental ratings does not validate the 10—C for
operational use in screening or ranking applicants.

The Personnel Board raises numerous objections to the
above findings by the district court respecting the validity
studies. We have reviewed these finding under the clearly er-
roneous standard. See Wade v. Mississippi Cooperative Exten-
sion Service, 528 F.2d 508, 516 (Sth Cir. 1976); United States v.
City of Chicago, 549 F.2d 415, 429 (7th Cir.), cert. denied, 434
U.S. 875, 98 §.Ct. 225, 54 L.Ed.2d 155 (1977); Bridgeport
Guardians, Inc. v. Members of Bridgeport Civil Service Com-
mission, 482 F.2d 1333, 1337 (2d Cir. 1973), cert. denied, 421
U.S. 991, 95 S.Ct. 1997, 44 L.Ed.2d 481 (1975); Washington v.
Davis, 426 U.S. 229, 256, 96 S.Ct. 2040, 2055, 48 L.Ed.2d 597
(1976) (Stevens, J., concurring). Because we are not left with a
‘‘definite and firm conviction’’ that the court committed a
mistake in holding that the studies do not demonstrate job-
relatedness, Wade v. Mississippi Cooperative Extension Service,
supra, 528 F.2d at 516, we reject the Board’s objections.'*

'S Practical or operational significance is required by the guidelines.
Uniform Guidelines, § 14B(6), 43 Fed.Reg. at 38301; EEOC Guide-
lines, 29 C.F.R. § 1607.5(c)(2); DOJ Guidelines 28 C.F.R. Part 50.14
§ 12B(5). Judge Pointer’s finding that the 10—C lacks practical
significance was reached through painstaking application of several
different statistical approaches, see 13 E.P.D. at 6803-06; additional-
ly, as permitted by the guidelines, the judge’s inquiry into practical
significance considered, inter alia, the degree of adverse impact of the
10—C. Id. at 6806.

'6 The Board’s principal criticism of the findings below is the al-
leged failure of Judge Pointer to apply a ‘‘corrected coefficient,’’
which would make the statistics relating to the police test more
favorable to the Board. The Board has not convinced us that the
district court failed to give proper consideration to the ‘‘corrected
coefficient.’’ A careful reading of Judge Pointer’s opinion makes it
apparent that the judge recognized that the ‘‘corrected coefficient’’ is
based on several assumptions, and the judge, accordingly and proper-

— so

APPLICATION OF WASHINGTON
V. DAVIS

The Personnel Board’s principal argument on appeal is
based on the Supreme Court’s opinion in Washington v. Davis,
426 U.S. 229, 96 S.Ct. 2040, 48 L.Ed.2d 597 (1976). The Board
focuses on the district court’s finding that both the police test
and the firefighter test bear a statistically significant correlation
to training academy grades. With that focus the Board argues
that the Davis case compels a decision in its favor. The Board
contends that the Davis case stands for the general proposition
that a test can be validated by showing that it predicts grades in
a job-relevant training program, without regard to the test’s
ability to predict job performance.'’ As discussed below, the

ly, used the ‘‘corrected coefficient’? with due regard for its
dependence on these assumptions. The crucial finding of the district
court, which was reached after careful and comprehensive considera-
tion of all the facts and circumstances, is that the validation study
showed that the police test is predictive of better job performance, but
that the magnitude of the positive prediction is so low that the test is
worthless for all practical purposes. We cannot conclude that this
finding is clearly erroneous.

The Board’s only other significant attack on the district court’s
findings is its suggestion that several of the statistical approaches used
by Judge Pointer in analyzing the operational utility of the police test
was inappropriate. The Board has not convinced us in this regard.
However, even if we should eliminate from our consideration all of
the statistical approaches used by the district judge except the two ap-
proved by the Board, namely, the Taylor-Russell approach and the
Anastasi approach, we still could not label as clearly erroneous the
district court’s finding that the police test is not approvriate for opera-
tional use.

'' The Board contends that this interpretation of Washington v.
Davis was adopted by the Supreme Court in its summary affirmance
of the three-judge District Court decision in United States v. South
Carolina, 445 F.Supp. 1094 (D.S.C.) aff'd. 434 U.S. 1026, 98 S.Ct.
756, 54 L.Ed.2d 775 (1978). The Board’s reliance on United States v.
South Carolina is misplaced for three reasons. (1) ‘‘[T]he precedential
effect of a summary affirmance can extend no farther than ‘the
precise issues presented [in the jurisdictional statement required by
Supreme Court Rule 15] and necessarily decided by those actions.’”’

aw Aal§ a

Davis case held that the test there was properly validated
because it was shown to predict whether those tested had the
minimum reading and verbal skills necessary to complete a job-

Illinois State Board of Elections v. Socialist Workers’ Party, 440 U.S.
173, 182, 99 S.Ct. 983, 990, 59 L.Ed.2d 230 (1979); Washington v.
Confederated Bands and Tribes of the Yakima Indian Nation, 439
U.S. 463, 477 n.20, 99 S.Ct. 740, 750 n.20, 58 L.Ed.2d 740 (1979);
Mandel v. Bradley, 432 U.S. 173, 176, 97 S.Ct. 2238, 2240, 53
L.Ed.2d 199 (1977). We have examined the jurisdictional statement
filed with the Supreme Court in United States v. South Carolina and
determined that neither of the two issues presented specifically ad-
dresses the Washington v. Davis question of validation against train-
ing. (2) A summary affirmance affirms only the judgment and not the
reasoning of the court below. Jilinois State Board of Elections v.
Socialist Workers’ Party, supra, 440 U.S. at 182-83, 99 S.Ct. at
989-990; Washington v. Confederated Bands and Tribes of the
Yakima Indian Nation, supra, 439 U.S. at 477 n.20, 99 S.Ct. at 750
n.20; Mandel v. Bradley, supra, 432 U.S. at 176, 97 S.Ct. at 2240.
Therefore, the three-judge District Court’s interpretation of
Washington v. Davis, to the extent that interpretation differs from
our own, is not binding on this Court. (3) United States v. South
Carolina is factually distinguishable from the instant case. The State
of South Carolina used minimum score requirements on the National
Teachers Examination to certify and determine the pay levels of
teachers within the State. A content validity study demonstrated that
the content of the exam matched the content of teacher training pro-
grams in South Carolina, and that the minimum score requirement
correlated to the minimum amount of knowledge necessary to effec-
tive teaching. 445 F.Supp. at 1112-14; See also Id. at 1107. After a cer-
tain date, examinees who did not achieve the minimum score, and thus
lacked the minimum amount of knowledge to teach effectively, were
not certified. Before that date, most examinees who did not achieve
the minimum score were certified, but their compensation was less
that those who scored the minimum. /d. at 1105-06 & n.12. Thus, the
test in South Carolina was used solely as a measure of minimal com-
petence to teach, much as the test in Washington v. Davis was used to
determine a minimum level of required competence. In contrast, the
instant tests were used not to determine minimum competence to per-
form as policemen or firefighters, but for ranking purposes unrelated
to minimum competence.

— A-16 —

relevant training program. We do not believe the Davis'® ra-
tionale can be extended, as the Board urges, to the general prop-
osition that any test can be validated by showing a relationship
to training. More specifically, we reject the Board’s suggested
extension of the Davis holding to this case, where the tests were
not used to ascertain the minimum skills necessary to complete
job-relevant training'’, but rather were used to rank job ap-
plicants according to their test scores and to select only the
highest test scorers for job placement.

In Washington v. Davis, two blacks (respondents in the
Supreme Court but referred to here as plaintiffs), whose ap-
plications to become police officers in the District of Columbia
Police Department had been rejected, claimed, inter alia, that a
written test (‘‘Test 21’’) used. by the Department for recruiting
purposes violated their rights under the due process clause of
the Fifth Amendment, 42 U.S.C. § 1981, and D.C.Code §
1-320. Test 21, developed by the Civil Service Commission and
administered throughout the federal service, was designed to
test the verbal ability, vocabulary, reading and comprehension
of potential recruits for the police department. In order to gain
entry into the training program, a grade of at least 40 out of 80
was required on Test 21. Plaintiffs moved for summary judg-
ment seeking a declaration that Test 21 unlawfully

'* Some authorities suggest that Washington v. Davis has no ap-
plication to Title VII because the action in Davis was brought under
the due process clause of the Fifth Amendment, 42 U.S.C. § 1981, and
D.C. Code § 1-320, and not Title VII. See, e.g., Guardians Associa-
tion of the New York City Police Department, Inc. v. Civil Service
Commission of the City of New York, 431 F.Supp. 526 (S.D.N.Y.),
vacated and remanded on other grounds, 562 F.2d 38 (2nd Cir. 1977);
See also Davis, 426 U.S. at 255, 96 S.Ct. at 2054 (Stevens, J., con-
curring). Since we conclude that Washington v. Davis is factually
distinguishable, we do not reach the question of its application to Title
VII.

'? In this case, Judge Pointer found that neither test predicts suc-
cessful completion of the training programs. See n.23, infra.

an Att —

discriminated in violation of the Fifth Amendment. The defen-
dants filed a counter-motion for summary judgment, asserting
that plaintiffs were entitled to relief on neither constitutional
nor statutory grounds.

The district court granted defendants’ and denied plaintiffs’
motions. Davis v. Washington, 348 F.Supp. 15 (D.D.C. 1972).
The district court found, among other things, that a higher
percentage of blacks failed Test 21 than whites, and that the test
had not been validated to establish its reliability for measuring
subsequent job performance. These findings were held suffi-
cient to shift the burden of proof to defendants. Jd. at 16. The
district court found that defendants had met their burden by
proving that ‘‘the Test is directly related to a determination of
whether the applicant possesses sufficient skills requisite to the
demands of the curriculum a recruit must master at the police
academy.”’ /d. at 17 (emphasis added).*° Given its relationship
to the training program, the court held that the lack of job per-
formance validation did not defeat the test.

The Court of Appeals reversed on constitutional grounds.
512 F.2d 956 (D.C. Cir. 1975). Applying the Title VII principles
developed in Griggs v. Duke Power Co., supra, 401 U.S. at 424,
91 S.Ct. at 849, to the Fifth Amendment claim, it held, inter
alia, that without proof by defendants that Test 21 had been
validated in regard to job performance, defendants had not
rebutted plaintiffs’ prima facie case of discriminatory impact.
Id. at 961-65.

20 Defendants’ argument to the district court was that,
‘‘Undeniably, a police recruit must have the minimum reading skills in
order to complete the curriculum at the Training Academy . . . Test
21 merely makes certain that a potential recruit has the minimum skills
in reading and verbal ability.’” Memorandum in Support of Defen-
dants Hampton, Spain and Andolsek’s Motion for Summary Judg-
ment, U.S. Supreme Court Records, Briefs at 97, FO—0097—75
(1975).

— A-18 —

The Court of Appeals was in turn reversed by the Supreme
Court. After holding that the D. C. Circuit had erred in apply-
ing Title VII standards to the Fifth Amendment claim of
discrimination, the Supreme Court considered the statutory
issues raised by the summary judgment motion. The Court ap-
proved the use of Test 21 to determine the minimum skills
necessary for satisfactory progress in the training program. For
the majority, Justice White wrote:

The advisability of the police recruit training course infor-
ming the recruit about his upcoming job, acquainting him
with its demands, and attempting to impart a modicum of
required skills seems conceded. It is also apparent to us, as
it was to the District Judge, that some minimum verbal and
communicative skill would be very useful, if not essential,
to satisfactory progress in the training regimen. Based on
the evidence before him, the District Judge concluded that
Test 21 was directly related to the requirements of the
police training program and that a positive relationship
between the test and training-course performance was suf-
ficient to validate the former, wholly aside from its possi-
ble relationship to actual performance as a police officer.
This conclusion of the District Judge that training-
program validation may itself be sufficient is supported by
regulations of the Civil Service Commission, by the opi-
nion evidence placed before the District Judge, and by the
current views of the Civil Service Commissioners who [are]
parties to [this] case. Nor is the conclusion foreclosed by
either Griggs or Albemarle Paper Company v. Moody, 422
U.S. 405, 95 S.Ct. 2362, 45 L.Ed.2d 280 (1975); and it
seems to us the much more sensible consturction of the
job-relatedness requirement.

426 U.S. at 250-251, 96 S.Ct. at 2052-2053 (emphasis added).
Earlier in the opinion Justice White discussed the various ways a
test could be validated:

— A-19 —

It is necessary, in addition, that they be ‘validated’ in terms
of job performance in any one of several ways, perhaps by
ascertaining the minimum skill, ability, or potential
necessary for the position at issue.

426 U.S. at 247, 96 S.Ct. at 2051 (emphasis added). Similarly,
Justice Stevens, concurring, stated:

The test serves the neutral and legitimate purpose of re-
quiring all applicants to meet a uniform minimum stan-
dard of literacy. Reading ability is manifestly relevant to
the police function, there is no evidence that the required
passing grade was set at an arbitrarily high level.

426 U.S. at 254, 96 S.Ct. at 2054 (emphasis added). And, later
in his concurring opinion:

As a matter of law, it is permissible for the police depart-
ment to use a test for the purpose of predicting ability to
master a training program even if the test does not other-
wise predict ability to perform on the job. I regard this as a
reasonable proposition and not inconsistent with the
Court’s prior holdings.

426 U.S. at 256, 96 S.Ct. at 2055. Further explaining his own
opinion for the Davis majority, Justice White, in his dissent
from the summary affirmance of United States v. South
Carolina, 434 U.S. 1026, 98 S.Ct. 756, 54 L.Ed.2d 775 (1977),
stated:

Washington v. Davis . . . was thought by the District Court
to have warranted validating the test in terms of the appli-
cant’s training rather than against job requirements; but
Washington v. Davis, in this respect, held only that the test
there involved, which sought to ascertain whether the ap-
plicant had the minimum communication skills necessary
to understand the offerings in a police training course,
could be used to measure eligibility to enter that program.
The case did not hold that a training course, the comple-

— sp

tion of which is required for employment, need not itself
be validated in terms of job relatedness. Nor did it hold
that a test that a job applicant must pass and that is de-
signed to indicate his mastery of the materials or skills
taight in the training course, can be validated without
reference to the job. Tests supposedly measuring an appli-
cant’s quaiifications for employement, if they have dif-
ferential racial impact, must bear ‘some manifest relation-
ship to the employment in question,’ Griggs v. Duke
Power Co., 401 U.S. 424, 432-[, 91 S.Ct. 849, 854, 28
L.Ed.2d 158] (1971)...

434 U.S. at 1027-28, 98 S.Ct. at 757 (emphasis added).

Thus, Washington v. Davis holds that a selection device
may be validated if it is shown to predict whether an applicant
has the minimum amount of reading and verbal skills necessary
to complete a job-relevant training program. We decline the
Personnel Board’s invitation to extend the Davis rationale by
holding that any test can be validated against training, without
respect to the test’s ability to predict job performance.”' Such
an extension would violate the requirement of job performance
validation enunciated in Griggs and Albemarle, as well as the
agency guidelines”? elaborating upon that requirement.

7! Other courts similarly refuse to read Washington v. Davis as
establishing that selection devices may be validated for Title VII pur-
poses without regard to job performance. See, e.g., Blake v. City of
Los Angeles, 595 F.2d 1367, 1382 n.17 (9th Cir. 1979), petition for
cert. filed, 48 U.S.L.W. 3125 (U.S. July 12, 1979) (No. 79-54); Guar-
dians Association of the New York City Police Department, Inc. v.
Civil Service Commission of the City of New York, 431 F.Supp. 526,
548-49 (S.D.N.Y.), vacated and remanded on other grounds, 562 F.2d
38 (2d Cir. 1977). But see, United States v. Virginia, 454 F.Supp.
1077, 1100-01 (E.D. Va. 1978).

*2 The agency guidelines require that a criteria-related validity
study, such as the one at bar, show a positive relationship between the
challenged test and job performance. Uniform Guidelines, § 5B, 43

— A-21 —

Unlike the test upheld in Washington v. Davis, the tests
used by the Personnel Board are not used to predict whether an
applicant has the minimum amount of knowledge necessary to
complete training”’; rather, the tests are used to rank applicants
according to their scores. Only those at the top of the eligibility
list, those with the highest test scores, are certified for job place-
ment. Even those who score a passing grade, and are deemed by
the Personnel Board to possess the capacity to complete train-
ing, are not hired unless they are among the highest scorers. Use
of a test for such ranking purposes,*‘ rather than as a Davis-like
device to screen out candidates without minimum skills, is
justified only if there is evidence showing that those with a
higher test score do better on the job than those with a lower test

Fed.Reg. at 38298. These guidelines, while not binding on the courts,
are entitled to great deference. Albemarle Paper Company v. Moody,
422 U.S. 405, 431, 95 S.Ct. 2362, 2378, 45 L.Ed.2d 280 (1975);
Thomas v. E.I. duPont de Nemours & Company, 574 F.2d 1324, 1331
n.9 (Sth Cir. 1978); Pettway v. American Cast Iron Pipe Company,
494 F.2d 211, 221 (Sth Cir. 1973), cert. denied, 439 U.S. 1115, 99
S.Ct. 1020, 59 L.Ed.2d 74 (1979); United States v. Georgia Power
Company, 474 F.2d 906, 913 (Sth Cir. 1973).

2 Indeed, the district court found the tests inappropriate even for
the limited purpose of determining whether an applicant has the
minimum capacity to complete training. The court found that neither
test is a valid predictor of ability to pass the training curriculum. 13
E.P.D. at 6802-03. It should also be noted that the Personnel Board
has not contended on this appeal for such limited, Davis-sanctioned,
use of the tests.

** The guidelines also distinguish tests used for screening from tests
used for ranking. To validate the latter under a criterion-related study,
‘the user need show mathematical support for the proposition that
persons who receive higher scores on the [selection] procedure are like-
ly to perform better on the job.’’ Questions and Answers to Clarify
and Provide a Common Interpretation of the Uniform Guidelines on
Employee Selection Procedures, 44 Fed.Reg. 11996, 12005 (Question
& Answer 62).

= fig =

score.” Such evidence is utterly lacking here. The Board’s
validation studies show that higher test scores do not predict
better job performance.

Having accepted the district court’s finding that the Person-
nel Board did not carry its burden of showing that the two tests
are job-related, and having rejected the Board’s suggestion that
the Davis case can be extended to justify validation of these tests
against training grades without regard to job performance, we
affirm the district court’s holding that the Personnel Board’s
use of the two tests violates Title VII. We turn next to the
remedy issue and the subsidiary question of when the Board’s
Title VII liability commenced.

COMMENCEMENT OF LIABILITY

The Personnel Board became subject to the requirements of
Title VII on March 24, 1972. Equal Employment Opportunity
Act of 1972, Pub.L. No. 92-261, 86 Stat. 103. On that date, and
for several years earlier**, the Board was using the 10—C and
20—B to screen and rank applicants. However, the district
court held that use of the police test did not begin to violate Ti-
tle VII until April 25, 1975, and that use of the firefighter test
did not constitute a violation until July 8, 1976. Those were the

25 If academy grades are to be the criterion to which the tests are to
be compared, the criterion itself must be shown to be a good measure
of job performance. See Blake, supra, 595 F.2d at 1382; Vulcan
Society of the New York City Fire Department, Inc. v. Civil Service
Commission, 490 F.2d 387, 396 n.11 (2d Cir. 1973). ‘* ‘The entire ra-
tionale of a criterion-related study requires that the criterion with
which the test results are compared be a good measure of job perfor-
mance.’’’ James v. Stockham Valves and Fittings Co., 559 F.2d 310,
340 (Sth Cir. 1977), cert. denied, 434 U.S. 1034, 98 S.Ct. 767, 54
L.Ed.2d 781 (1978), quoting United States v. City of Chicago, supra,
549 F.2d at 433.

** The Board started using the police and fire tests on August 18,
1967, and October 23, 1968, respectively.

— A-23 —

dates on which the final results of the validation studies for the
two tests (the studies were started in late 1972) were reported to
the Board. The court explained:

The preliminary reports from the consultants, made while
more trustworthy measures of job performance were being
developed [i.e., the experimental ratings], contained signs
of potential validity and recommended continued usage of
the test pending the additional studies. Not until April 25,
1975, with respect to the 10O—C, and July 8, 1976, with
respect to the 20—B were the studies using these new
criterion measures completed and reported to the Board. It
was on these respective dates that, in the court’s opinion, it
should have been concluded that provisional use of the
tests was no longer permissible. Prior thereto, the Board
was, in the court’s opinion, justified in continuing to use
the tests (and the eligibility lists generated therefrom) in
anticipation of favorable results from those studies.

13 E.P.D. at 6807.

It appears from the court’s reference to ‘‘provisional use’’
that Judge Pointer was relying on § 1607.9 of the EEOC Guide-
lines in holding the Board free from liability pending its
receipt of the final validation studies report.”’ Section 1607.9

?? It appears also that in fixing the date of Title VII liability, the
district court considered the good faith and lack of discriminatory in-
tent of the Board in adopting and using the 10O—C and 20—B. The
court observed that the Board was compelled by state law to ad-
minister employment tests to screen and rank applicants; that the
Board selected the 10O—C and 20—B in the late 1960’s ‘‘as the best
tests then available, with the hope that black applicants would fare
better than under previous tests’’; that the Board conducted a
preliminary, in-house, validity study of the tests even before it became
subject to Title VII; and that, ‘‘at least since 1965 the Board has not
intentionally discriminated against black applicants.’’ 13 E.P.D. at
6806-6807. We do not question the court’s finding that the Board
acted in good faith and without discriminatory intent. However, to
the extent that this finding influenced the court’s determination as to

—

‘*authorize[s] provisional use of tests, pending new validation
efforts, in certain very limited circumstances.’’ A/bemarle
Paper Co. v. Moody, supra, 422 U.S. at 436, 95 S.Ct. at 2380.
Section 1607.9 reads in full:

§ 1607.9 Continued use of tests.

Under certain conditions, a person may be permitted to
continue the use of a test which is not at the moment fully
supported by the required evidence of validity. If, for ex-
ample, determination of criterion-related validity in a
specific setting is practicable and required but not yet ob-
tained, the use of the test may continue: Provided: (a) The
person can cite substantial evidence of validity as described
in § 1607.7(a) and (b); and (b) he has in progress validation
procedures which are designed to produce, within a
reasonable time, the additional data required. It is ex-
pected also that the person may have to alter or suspend
test cutoff scores so that score ranges broad enough to per-
mit the identification of criterion-related validity will be
obtained.

Section 1607.7 of the EEOC Guidelines provides that where it is
not feasible to conduct a proper validation study,

evidence from validity studies conducted in other organiza-
tions, such as that reported in test manuals and profes-
sional literature, may be considered acceptable when: (a)
The studies pertain to jobs which are comparable (i.e.,
have basically the same task elements), and (b) there are no
major differences in contextual variables or sample com-
position which are likely to significantly affect validity.

when liability commenced, error was committed. Absence of discrim-
inatory intent is no defense in a Title VII disparate impact case. Griggs
v. Duke Power Co., supra, 401 U.S. at 432, 91 S.Ct. at 854;
Albemarle Paper Co. v. Moody, supra, 422 U.S. at 422, 95 S.Ct. at
2373.

— A-25 —

Any person citing evidence from other validity studies as
evidence of test validity for his own jobs must substantiate
in detail job comparability and must demonstrate the
absence of contextual or sample differences cited in
paragraphs (a) and (b) of this section.

To justify use of the 1O—C and 20—B under § 1607.9, the Per-
sonnel Board should have shown (1) that validation studies
done elsewhere, as described in § 1607.7, provided substantial
evidence of validity”* and, (2) that at ali relevant times the Board
had in progress validation procedures designed to produce,
within a reasonable time, the additional data required. The
district court did not determine whether either of these condi-
tions was satisfied. Because we believe that this determination
‘tis a matter best decided in the first instance by the District
Court,’’ Albermarle Paper Co. v. Moody, supra, 422 U.S. at
436, 95 S.Ct. at 2380, we remand to the district court to ascer-
tain whether the conditions were met.

If the district court, upon remand, finds that the use of the
tests was permissible under § 1607.9 until the Board received the

** The first condition for interim use under § 1607.9 is that ‘‘(a) The
person can cite substantial evidence of validity as described in §
1607.7(a) and (b).’’ (emphasis supplied). Thus, unless an employer
cites to other validity studies which meet the requirements of § 1607.7,
he has not satisfied § 1607.9(a). This condition lends objectivity to §
1607.9(a)’s substantial evidence requirement, and does not place an
onerous burden on employers. Compare, Friend v. Leidinger, 446
F.Supp. 361, 370 (E.D.Va. 1977) (alternative holding sustained the
City of Richmond’s use of firefighter tests under § 1607.9 where tests
had been proven job-related in a California validity study), aff'd. 588
F.2d 61 (4th Cir. 1978); Buckner v. Goodyear Tire and Rubber Co.,
339 F.Supp. 1108, 1115 (N.D.Ala. 1972), aff’d. (and district court
opinion adopted), 476 F.2d 1287 (Sth Cir. 1973) (provisional use of
tests permitted where validation studies at another Goodyear plant
were available). § 1607.9’s incorporation of the requirements of §
1607.7 is especially reasonable in light of the fact that § 1607.9 permits
use of a test, possibly under the cloak of immunity as discussed below,
which has a discriminatory impact.

—< oo

final results of the validation studies, and if the court also finds
that the Board in good faith relied upon § 1607.9, then the
district court would be correct in starting Title VII liability on
April 25, 1975 (police test), and July 8, 1976 (firefighter test).
The natural reading of EEOC Guideline § 1607.9 (as it was in ef-
fect at ail times relevant to this proceeding’’), especially when
read in light of Title VII, § 713(b), is that an employer will be
immune from liability during the period of permissible provi-
sional use of an unvalidated test. Section 713(b) of Title VII
provides a defense to an employer who complies with, and relies
in good faith upon, EEOC Guidelines, such as § 1607.9.*° Sec-
tion 713(b) provides:

In any action or proceeding based on any alleged unlawful
employement practice, no person shall be subject to any

** After the dates fixed by the district court for the commencement
of Title VII liability, two new sets of guidelines on employee selection
procedure were issued: the Department of Justice Guidelines (effective
November 17, 1976) and the Uniform Guidelines (effective September
25, 1978). See supra, n.11. Both the DOJ Guidelines and the Uniform
Guidelines contain provisions on interim use which expressly deny an
employer immunity in the event his pending validation study does not
ultimately establish validity. See DOJ Guidelines § 5(h), 28 C.F.R.
Part 50.14 (‘‘If the additional studies do not produce the data required
to demonstrate validity, the user is not relieved of or protected against
any obligations arising under federal law.’’); Uniform Guidelines,
supra, n.ll, § 5J, 43 Fed.Reg. at 38298 (‘‘If the study does not
demonstrate validity, this provision of these guidelines for interim use
shall not constitute a defense in any action, nor shall it relieve the user
of any obligations arising under Federal law’’). Had the DOJ or
Uniform Guidelines been in effect during the pendency of the Person-
nel Board’s validation studies, a different issue would be presented.
However, at all times relevant to the instant proceedings, only the
EEOC Guidelines, as they were in effect before the adoption of the
Uniform Guidelines, were in force. The EEOC Guidelines do not con-
tain a disclaimer of immunity similar to that contained in the DOJ and
Uniform Guidelines.

*° Friend v. Leidinger, 446 F.Supp. 361, 370(E.D. Va. 1977), aff'd.
588 F.2d 61 (4th Cir. 1978) (Title VII, § 713(b) extends immunity dur-
ing § 1607.9 interim use).

ET

liability or punishment for or on account of (1) the com-
mission by such person of an unlawful employement prac-
tice if he pleads and proves that the act or omission com-
plained of was in good faith, in conformity with, and in
reliance on any written interpretation or opinion of the
[EEOC].?’

If, however, the district court, upon remand, finds that the
conditions for § 1607.9 provisional use were not satisfied or that
the Board did not in good faith rely upon that section while
awaiting the final results of the validation studies, then the court
must mark March 24, 1972, as the date of violation of Title VII,
with respect to both tests. A new remedy must then be fashioned
by the court to correct the discrimination caused by use of the
tests from that date forward.

AFFIRMED IN PART, REVERSED IN PART AND
REMANDED.

*! By holding that the Board would not be subject to liability during
the period it satisfied and relied upon § 1607.9, we are not, as plain-
tiffs contend, assuming that a grace period should be implied for
public employers similar to the one-year grace period expressly ac-
corded private employers (Title VII, § 716(a)), when Title VII was first
enacted. See Blake v. City of Los Angeles, supra, 595 F.2d at 1376-77
(refusing to imply grace period for public employers). Rather, our
holding is based on the immunity which Title VII, § 713(b) extends to
an employer who complies with, and relies in good faith upon, the
EEOC Guidelines.

— A-28 —

APPENDIX C

Ensley Branch of the N.A.A.C.P., Plaintiff v. George
Seibels, et al., Defendants. Civil Action No. 74-Z-12-S.

John W. Martin et al., Plaintiffs v. City of Birmingham et
al., Defendants. Civil Action No. CA 74-Z-17-S.

United States of America, Plaintiff v. Jefferson County et
al., Defendants. Civil Action No. CA 75-P-0666-S.

Lucy Walker et al., Plaintiffs v. Jefferson County Home et
al., Defendants. Civil Action No. CA 76-M-2047-S.

United States District Court, Northern District of Alabama,
Southern Division. January 10, 1977.

Memorandum of Opinion

POINTER, D.J.: Since 1945 the Personnel Board of Jeffer-
son County has been charged under state law with the duty of
periodically administering examinations to ‘‘fairly test the
relative capacity and fitness’’ of applicants for positions with
local governmental agencies.' 1940 Ala. Code Appx. §§645, ef
seg. (Recomp. 1958). Those who pass are ranked on an eligibil-
ity list in the order of their exam scores.” As vacancies occur, the
three persons then at the top of the list are certified to the
employing agency for final selection, the appointments being
probationary in nature for the first twelve months.’ An appli-
cant’s name may be removed from the eligibility list after having
three times been certified and refused employment.

This litigation challenges the employment practices of the
governmental agencies as discriminatory on the basis of race,
color, and sex, and includes an attack upon the examinations
administered by the Personnel Board. Presently at issue, follow-
ing a trial held December 20-22, 1976, are the tests currently
used to screen applicants for positions as police officers, deputy
sheriffs,‘ and firefighters.°

An attack upon the police and firefighters exams is certainly
understandable when one considers that, although the relevant

— A-29 —

labor pool is over 25% black, yet on June 30, 1976, only 56 (or
6.5%) of the 860 police officers were black and only 9 (or 1.4%)
of the 630 firefighters were black. These statistics may,
however, be misleading for*purposes of this lawsuit because
they include the historical results of hiring practices employed
long before passage of the Equal Employment Opportunity Act
of 1972 or, indeed, before utilization of the tests under scrutiny
at this time.

The principal focus should rather be upon the events of more
recent years, with particular attention upon practices subse-
quent to March 24, 1972, when Title VII of the Civil Rights Act
of 1964 was made applicable to the Personnel Board and the
governmental agencies which it serves. Likewise, information as
to the general labor pool in the area is of only marginal import-
ance when one has, as we do, extensive data as to actual ap-
plicants for positions and there is no evidence that minority ap-
plications have been depressed by prior employment practices.

Adoption of the Current Tests

In late 1965, following an independent study as to why no
blacks were then employed as police officers in the City of Birm-
ingham, the Personnel Board decided to replace its police and
firefighter exams with tests developed by the Public Personnel
Association, now known as the International Personnel Man-
agement Association. IPMA tests were being widely used in
other parts of the country and were considered by the Board as
superior to other tests then available. The change was part of a
multi-faceted program intended to increase black participation
in governmental positions. (See Appendix C to X-342). Police-
man Test 1O—C and Firefighter Test 20—B have been in use
since August 18, 1967, and October 23, 1968, respectively, as
the screening examinations for these positions under the state-
mandated selection procedure,* although at times other tests
have been administered for experimental purposes or for valida-
tion studies. Since April 10, 1974, a modified scoring key (based

— A-30 —

upon only 80 of the 120 test items) has been employed in grading
the 10—C test for purposes of the eligibility list. This modifica-
tion was made at the recommendation of qualified independent
consultants who, after study, concluded that the scoring change
would increase validity of the test for black applicants.

Intent

It is clear that the Personnel Board, in performing its func-
tions as an employment agency for the various local govern-
ments, has not intentionally discriminated against blacks.
Indeed, at least since 1965, the Board has not only sought to
provide non-discriminatory opportunities for black applicants,
but also attempted, within the limits of its statutory duties, to
rectify the racial imbalances in local government employment.
Some mention of these aims and efforts is appropriate.

It was the Board’s hope that adoption of the tests now in issue
would benefit black applicants, while nevertheless providing a
fair ‘‘test of the relative capacity and fitness’’ of all applicants,
as required by state law. Immediately, a study was undertaken
to ascertain whether the IPMA policemen test, although a paper
and pencil test, would correlate positively and significantly with
a widely used non-verbal performance test of general in-
telligence, the Revised Beta Examination—and it did. As suc-
cessful applicants were employed by the city of Birmingham,
were trained at the police academy, and entered performance of
their duties, information was incorporated into the Board’s on-
going validation studies—which, while lacking sufficient blacks
in the sample (only 6 in the 10—C sample) to permit full
analysis, were considered by the Board as justifying further
usage of the 10—C.’ These studies are presented by the Board
not as satisfying the requirements of the EEOC or Department
of Justice guidelines on tests, but rather as indicating its efforts
to see that its examinations were fair predictors of job perfor-
mance even at a time when it was not subject to the provisions
of Title VII. As already noted, when, in 1974, it was advised by ©

v

— A-31 --

independent consultants that a modification of the scoring of
the 10—C exam would improve the validity for black ap-
plicants, it immediately put that change into effect.

Since 1965 the Board, with the cooperation of local civic
groups and some of the employing agencies, has been actively
engaged in recruitment efforts to attract black applicants. In
1966 it began assuming the $10.00 medical examination costs
for newly hired persons; and in 1967 it was successful in spon-
soring legislation to eliminate the $1.50 examination fee
previously required and to eliminate the priority previously
given applicants who resided within an employing agency’s
jurisdiction.’ It has experimented with a lowering of the raw
score used to measure a ‘‘passing’’ grade on the exams where it
could justify that approach on the basis of ‘‘supply’’ and
**demand’’.

In short, in its selection, administration and use of the 10—C
and 20—B tests, there has been no design or intent on the part
of the Board to discriminate on the basis of race or color. How-
ever, the standard under Title VII of the Civil Rights Act of
1964 is not so limited’—rather, if the operational effect of test
usage is to discriminate against blacks, then it is proscribed
unless it is shown to be a ‘‘job related’’ requirement,'® with a
‘‘manifest relation to the employment in question.’’'' See
U.S.C.A. § 2000e-2(h). In this inquiry the court is to follow the
guidelines adopted by the EEOC and, more recently (November
17, 1976), by the Department of Justice (DOJ), absent some
‘cogent reason.’’'? See Watkins v. Scott Paper Co., 530 F.2d
1159 (CAS 1976). Also instructive are the 1974 A.P.A. Stan-
dards for Educational & Psychological Tests and the 1975 Prin-
ciples for the Validation and Use of Personnel Selection Pro-
cedures of the A.P.A.’s Division 14.

Adverse Impact

Where the total selection process has an adverse impact upon
a substantial racial group in the labor market, the individual

— A-32 —

components of that process—such as a screening test—are also
to be evaluated for adverse impact. DOJ Guidelines §4b. For
purpose of this two-step analysis, data can be extracted from the
evidence pertaining to administrations of the 10—C and 20—B
tests which have been used for employment decisions after
March 24, 1972."°

10—C 20—B
Black White Black White
Failing test 395 191 216 267
Passing test 373 1,762 69 1,263
Hired 51 455 9 215

According to the DOJ Guidelines, §4b, ‘‘A selection rate for
any racial * * * group which is less than four-fifths (4/5) (or
eighty percent) of the rate for the group with the highest rate
will generally be regarded as evidence of adverse impact * * *
Greater differences in selection rate would not necessarily be
regarded as constituting adverse impact where the differences
are based on small numbers and are not statistically significant,
or where special recruiting or other programs cause the pool of
minority * * * candidates to be atypical of the normal pool of
applicants from that group.’”’

So far as the total selection process is concerned, one finds
from the above data that the hiring rates for blacks (6.6% of the
black applicants on 10—C and 3.2% of the black applicants on
20—B) are substantially less than eighty percent of the hiring
rates for whites (23.3% and 14.1%, respectively). These greater
differences in selection rates cannot be explained on the basis of
inadequate numbers and, according to the court’s calculations,
are statistically significant: the o coefficient for the 10—C test is
.193 and for the 20—B is .121, both of which are significant at
p< .001.

Looking at the data pertinent to the test component of the
selection process, one finds again that the pass rates for blacks

— A-33 —

(48.6% for 10—C and 24.2% for 20—B) are substantially less
than eighty percent of the pass rates for whites (90.2% and
82.5%, respectively). And again, according to the court’s
calculations, these greater differences in pass rates, which are
based upon samples of adequate size, are statistically signifi-
cant: the o coefficient for the 1O—C test is .46 and for the 20—B
is .48, both being significant at p< .001. Also of importance is
the fact that, of the blacks who did pass the tests, 85.4% placed
in the lower half of the initial eligibility lists for police officers
and 89.9% placed in the lower half of the firefighter lists.'*

Some concern can justifiably be expressed that the special
recruiting efforts undertaken by the Personnel Board and other
groups to attract black applicants—while commendable as an
affirmative action to overcome racial imbalance in the police
and firefighter forces—may at the same time have resulted in an
atypical pool of blacks taking the test, producing distortion in
the test performance of the black applicants. For example, the
black applicants may have included many who were not serious-
ly interested or motivated with respect to the jobs in question,
thereby affecting their test performances. Absent, however, any
hard data to support such an hypothesis or to indicate its
magnitude, the court, impressed with the substantial differences
in hire rates and pass rates for the two racial groups, must con-
clude that the overall selection procedures in effect since March
24, 1972, and as a component part thereof the tests used for
those purposes, have had an adverse impact on blacks.

Validation Studies

According to EEOC Guidelines § 1607.3, ‘‘the use of any test
which adversely affects hiring * * * of classes protected by Title
VII constitutes discrimination unless (a) the test has been
validated and evidences a high degree of utility as hereinafter
described * * *.’’ For the purpose of making such validation
studies of its many tests, the Board in 1972 contracted with Drs.
William E. Farrar and William A. McLaurin, Professors in the

—

Psychology Department of the University of Alabama at Birm-
ingham. Both had experience with personnel selection pro-
cedures in public employment systems. Priority, but not ex-
clusive attention was to be given to the police and firefighter
tests, and their work on these tests began in late 1972. Their
studies respecting the two tests continued even to the time of
trial, with various reports being made in each of the years 1973,
1974, 1975 and 1976. That their work was not complete before
trial does not suggest inattention; rather, it is indicative that
their studies were intended to be thorough and were directed to
numerous tests.'*

Psychometric Analyses

The 10—C and 20—B tests are paper-and-pencil instruments,
each consisting of 120 multiple choice items.'* The initial con-
cern of Drs. Farrar and McLaurin was directed to the reli-
ability,'’ item difficulty,'* and item discrimination’® of the tests.
The following findings were made:

reliability items items
satisfactorily satisfactorily
‘sp ‘kr difficult _ discriminating

10—C total (N=479) .95 .95 88 116
10—C black (N=176) .90 .9 77 106
10—C white (N=303) .95 .93 70 115
20—B total (N=507) .92 .91 64 115
20—B black (N = 108) .85 .85 61 66
20—B white (N=399) .86 .86 53 108

Inquiry into reliability is a proper first step, because, while no
test is perfectly reliable, a test which is not reliable is not valid
for any purpose. The consultants found the reliability coeffi-
cients for both tests to be of sufficient magnitude to indicate
satisfactory reliability. They did acknowledge that the methods
selected for this purpose were essentially measures of internal
consistency (and with the KR-20 formula, of content homo-

— A-35 —

geneity), but apparently believed it either not feasible or not
necessary to investigate error variance due to time sampling.
The court agrees as to reliability and notes that possible lack of
stability over time is, in a sense, mitigated by the fact that ap-
plicants may take an exam on more than one administration.

Analyses of item difficulty and discrimination have no direct
bearing upon the validation studies before the court. However,
they do reflect an investigation into possible modification or
supplementation of the tests to improve their utility and reduce
the extent of adverse impact, which is a recommended pro-
cedure.”® See DOJ Guidelines, § 3c.

Documentation and Methodology

The EEOC Guidelines, at §§1607.5(b)(2,3,5) and 1607.6, re-
quire that various items of information (e.g. copies of tests,
manuals, rating forms and instructions and representations of
Statistical data) be included in the report of the study or other-
wise available for inspection. Following the 1974 A.P.A. Stan-
dards, a more extensive list of documentation requirements is
specified in the DOJ Guidelines at §§4a and 13b, involving some
twenty-four ‘‘essential’’ items and several other desirable items.
The Farrar-McLaurin studies satisfy the EEOC requirements,
which were the only ones in effect when their studies were con-
ducted and (so far as then feasible) completed and, indeed,
when supplemented by evidence presented immediately before
and during trial, they also substantially satisfy the DOJ re-
quirements, which became effective on November 23, 1976.?'

The studies include presentations of the following statistics:
For the 10—C test:

@ intercorrelation coefficients, r and r,, for 109 Birming-
ham police officers (without separation by race) respecting their
10—C scores, police academy scores (school average and course
grades) and latest efficiency ratings.

— A-36 —

@ means, standard deviations, and ¢ tests for difference in
means for 10—C scores of 38 black and 101 white Birmingham
police officers.

@ means, standard deviations, and r coefficients for 59
Birmingham police officers (without separation by race) respec-
ting their 10O—C scores, academy scores (average and courses)
and latest efficiency ratings.

@ means, standard deviations, r coefficients, and ¢ tests for
the following:

@ @ 20 black and 76 white Birmingham police officers
respecting their 10O—C scores, academy averages, and latest effi-
ciency ratings.

@ @ 8 black and 140 white Birmingham police officers
respecting their 10—C scores, academy scores (average and
courses), latest efficiency ratings (overall and by sub-parts), and
experimental ratings weighted average and by components.

@ @ 49 black and 140 white police officers respecting their
10—C scores and academy scores (average and courses). (Also
included are data for analyzing significance of differences in
correlation coefficients through z transformations.)

@ @ 83 Jefferson County deputy sheriffs (without separa-
tion by race) respecting their 10—C scores and academy
averages.

@ @ 77 police officers (without separation by race) from
other cities served by the Personnel Board respecting their
10—C scores and academy averages.

For the 20—B test: means, standard deviations, and r coeffi-
cients for the following:

@ 162 Birmingham firefighters (without separation by race)
respecting the 20—B scores, training academy average, and
latest efficiency ratings (overall and by sub-parts).

on eT on

@ 196 Birmingham firefighters (without separation by race)
respecting their 20—B scores, academy averages, latest efficien-
cy ratings, and experimental ratings (overall and by com-
ponents). Statistics are reported separately for short-tenure and
long-tenure firefighters, using three years of experience as the
point of division.

Statistics found to be significant at p < .05 and p < .01 are so
identified in the report.

Some commment should be made about selection and com-
position of the different samples. Each sample contained all the
persons for whom, so far as was known at the time by the con-
sultants, the data needed for that study was available. The dif-
ferent studies were, however, conducted over a period of several
years as either the need was recognized or the particular inquiry
became technically feasible; and during the time intervals the
work force had changed. Some of the studies involved concern
with additional factors (e.g., performance on the Raven and
PAS tests, which have been under consideration for use as sup-
plemental or alternative screening instruments), for whom the
data existed only for a limited number of applicants or em-
ployees. The result is that a particular sample may contain
some, but not necessarily all, of the persons in another sample
and may also contain some persons who were not in the other
sample. This lack of autonomy or consistency complicates
somewhat the process of analysis, but, under the circumstances,
is acceptable. There is no hint of contrivance in selection of the
samples or of lack of representatives of the sample subjects.”?

Criteria

The several criterion-measures (academy grades, efficiency
ratings, and experimental ratings) have certain common factors:
(1) None appears to be ‘‘contaminated’”’ (i.e., affected by know-
ledge by the rater or scorer of prior score on the 10—C or
20—B). (2) None has been subjected to special statistical

— A-38 —

scrutiny to detect or control possible bias among raters or
graders.”? (3) None has been subjected to special statistical
scrutiny for reliability.** (4) Each has been analyzed by the con-
sultants for relevancy (i.e., the extent to which it may be con-
sidered as a measure of critical or important work behaviors).
Each of the measures has, of course, its own special character-
istics and limitations, which will be described separately. It must
be emphasized that, in a criterion-related validation study, one
is attempting to estimate the extent to which a score on a
‘*predictor’’ (e.g., 10—C test) can predict job performance
(i.e., as a police officer) through evaluating its ability to predict
scores or ratings on a ‘‘criterion’”’ (e.g., academy average)—and
hence the study is subject to any limitations which those same
criteria have in either predicting or assessing job performance.

(1) Academy Grades.—Where, as here, new employees are
required to complete special training before performing their
duties, successful completion of that training may properly be
used as a criterion-measure, if, that is, the training is intended
to, and does, provide skills or knowledge needed for perfor-
mance of the job. Based upon the evidence presented, including
testimony of the directors of the Birmingham police and fire
academies, the court finds that the two schools do serve that
purpose and function.

Relative standing or ranking among students who successfully
complete such training is not, however, as such, an appropriate
criterion.**’ Rather, to be relevant as a criterion, such measures
must be shown, empirically or otherwise, to be themselves ap-
propriate predictors of job performance. This, in essence,
means a two-step correlation study; and, in a situation where
one has data on test scores, academy grades, and measures of
job performance for the same group of persons, the more direct
inquiry (correlation between test scores and measures of job
performance) would be preferred to the two-step approach.
Grades on particular courses in the academy must also be

— A-39 —

analyzed for compatibility with findings respecting grades on
other courses.

So far as the evidence indicates, academy grades—provided
they are passing scores—have no impact on job opportunities,
benefits, etc. If this be the case, then, while helpful in prevent-
ing ‘‘contamination’’ during validity studies, academy grades
are likely to be influenced by motivational considerations not
present in actual job performance. The emphasis in the
academies on paper-and-pencil multiple choice items, while pro-
viding objectivity, may also reflect a relationship to the paper-
and-pencil screening exam not found in job performance. These
concerns should cause one to be cautious in making non-
empirical judgments about the usefulness of relative academy
grades as a criterion-measure.

(2) Efficiency Ratings.—The efficiency ratings given
periodically on all employees by their supervisors are direct and,
ostensibly, appropriate measures of job performance. Drs. Far-
rar and McLaurin have, however, acknowledged that these
ratings 2re not trustworthy assessments of the employees’ actual
performance. In addition to other problems, the ratings must be
discussed between the rating supervisor and the employee and
can have important consequences for the employee. These
ratings were, it seems, used in the early studies because of their
availability, in the anticipation that other measures could be
developed and administered in due course.

(3) Experimental Ratings.—By review of existing job descrip-
tions, by interviews to determine ‘‘critical incidents’’ of the
jobs, and by technical assistance and consultation with advisory
committees consisting of representative incumbents and super-
visory personnel, new ‘‘experimental’’ rating forms were
developed for use in the Farrar-McLaurin studies. The forms
consist of twelve rating categories for each of the two jobs, the
categories relating to personality characteristics, job knowledge,
and abilities found through the process to be relevant to job per-

ye

formance. Each is rated on a seven-point scale (poor = 1 to
outstanding = 7), with 3 being fixed as adequate. For the
police form, weights were developed by the advisory committee
to indicate relative importance of the categories to overall job
performance.

Raters—the employees’ supervisors—were given, in person
and in writing, standardized instructions for use of the forms.
To prevent the ‘‘halo’’ effect, supervisors rated all their subor-
dinates on one category before proceeding to rate them on the
next, etc. The raters were told that their evaluations were con-
fidential and would not be used for any purpose other than the
evaluation of the tests.

The court is impressed that the experimental rating method so
developed represents an appropriate criterion measure for the
jobs in question. These jobs are not ones which lend themselves
to some objective measure, such as the sales produced by a sales
representative. The principal limitations with the ratings so ob-
tained are the lack of evidence as to reliability and the lack of
special steps to detect or control possible bias.

— Atl =

Study Findings

Key findings from the Farrar-McLaurin studies are tabulated
below. Correlations are shown only where presented in the body
or exhibits of their reports. Statistically significant differences
in means between two sub-groups (blacks and whites; short-
terure and long-tenure firefighters) are indicated by so
de. .gnating the lower of the two means.

Police

Applicants
T =479
W =33
B=176

Officers
T=109

Officers
T=139

W=101
B=38

Officers
T=59

Officers
T=%
W=76
B=20

Officers
T=148
W=140
B=8
(B= 49)
Deputies
T =83
Officers
T=77

Firefighters
Applicants
T=507

Employees
T = 162

Employees
T=1%
ST = 103
LT =93

test academy ave efficiency rating experimental rating
mean mean r mean r acad fr mean r acad f
65.70
75.21
49.33°*
82.20 .46** ole
69.35
75.03
54.24°*
81.00 88.66 779° 81.73 319°
80.46 87.94 -72°* 81.10 .20 9
84.54 89.00 .64°* 81.44 .0S .08
64.95** 83.90°* 46° 79.81° 17 ~ 05
82.35 87.38 .45°* 84.06 .08 433.79 ale
83.11 87.58 .42°* 84.11 434.75
69.12%* 83.90°° 83.18 417.00
66.78°* 83.82°* .47°*
82.58 87.36 .729°
79.71 87.50 Or"
70.99
83.90 90.60 .43°* 78.27 12 —.23°°
81.99 91.27 .45°* 79.45 ° .20°* -.06 61.24 .08 - .09
81.01 91.98 379° 77.149° 249° Al $8.65°° .259° 21%
83.09 90.47* .62°* 82.02 .05 -.01 64.11 -.20° - .20
*pe .0S

ee

—s on

Fairness and Differential Validity

If members of one racial group generally obtain lower test
scores than members of another group and those differences are
not reflected in differences in measures of job performance,
there is a need, where technically feasible, to investigate for
possible unfairness of the test to the first group. See DOJ
Guidelines, § 12b(7) (noting that this need increases the greater
the severity of the adverse impact on the lower-scoring group).
As the tabulation indicates, the Farrar-McLaurin studies do
show that blacks as a group have scored lower on the 10—C
than have whites. Indeed, in each of their studies where those
scores are reported separately for the two racial groups, the dif-
ferences are significant as p< .O1.

One possibility is that the predictive validity of the test for
one racial group is significantly different than for the other
group. The inquiry here, as emphasized in A.P.A. Standard E9,
is not whether there are differences in the correlation coeffi-
cients or whether one coefficient is statistically significant while
the other is not. Rather, the proper statistical procedure is to
test for significant differences in the coefficients.

In the one study in which a sufficient?* number of both blacks
and whites are involved, Drs. Farrar and McLaurin have per-
formed such an analysis. Using the report of test scores and
academy scores for 140 whites and 49 black police officers, they
tested the correlation coefficients (where the coefficient for
either subgroup was significant at p< .05) after z transforma-
tions, for significance of difference. None was significant at
p< .05, and with respect to only one course (accident investiga-
tion) was the coefficient significant even at p< .10.

Failure, however, to reject the hypothesis that the correlation
coefficients are the same for both groups is not by itself suffi-
cient to demonstrate fairness. Where, as in the present case, test
scores by two groups are used in the same manner for members
of both groups, it is on the assumption that in general an in-

= Akg =

dividual’s test score will appropriately predict his standing on
the criterion whether he is a member of one group or the other.
The predictive relationship between the test score and the
criterion can be represented by a regression line?’ formula for
converting a test score into a predicted criterion score. While
regression lines can be calculated separately for the two groups
and will almost always be somewhat different, it is important to
know whether the slopes or intercepts (or both) of the lines are
sufficiently different to call for abandonment of a common line
for two groups. Otherwise, the common regression line (which
is the effect of using test scores in the same way for both groups)
may systematically underpredict for members of one group
(while overpredicting for the other) their criterion score from a
particular test score. The method for this inquiry, called
analysis of variance, involves use of the F distribution tables for
statistical significance. If desired, one can determine for what
test scores the common regression line should, and should not,
be abandoned.

Significantly, different regression lines may have the same or
similar correlation coefficients. In such a situation comparison
of the coefficients will not reveal the inappropriateness of using
a common regression line. Thus, with the sample containing 76
white and 20” black police officers, comparison of the coeffi-
cients respecting 10—C scores (and modified 10—C scores) and
either academy averages or efficiency ratings does not, as to any
comparison, lead to rejection of the hypothesis of the coeffi-
cients being the same. Yet, when the same data are reviewed by
analysis of covariance, as the court has done”’ for the hypo-
thesis that a common regression line fits both whites and blacks,
the obtained F ratios are 18.38 (1O—C and academy average),
3.97 (10O—C and efficiency rating), 26.91 (modified 10—C and
academy average), and 4.18 (modified 10—C and efficiency
rating), each of which (with n, = 1 and n=93) is significant at
p< .05. It is interesting that in the development of the
modified 80-item scoring key for the 10O—C, adopted to increase

— = le

the validity coefficient for blacks, the evidence becomes
stronger that a common regression line should not be used for
both groups.

In view of the fact that covariance analysis suggests rejection
of acommon regression line for both racial groups, one is temp-
ted—since the differences between white and black means ap-
pear to be far greater on test scores than on the criterion-
measures—to conclude that the performance of blacks is being
underpredicted by the 10—C. However, if regression lines are
computed separately for blacks and whites using the results of
any of the studies where their test scores and criterion scores are
reported separately, it will be found that the lines cross and that
for test scores below that crossing point the criterion scores
thereby predicted for blacks are less than for whites for the same
test scores. Above that point there would be underprediction for
blacks, but the intersections occur at such high test scores (the
lowest point from any of the data is at a raw test score of 87)
that few blacks would actually be affected. If one looks to see
where any overprediction or underprediction is statistically
significant, it is found that the only significant range of scores is
for lower scores, where blacks are being overpredicted by the
10—C.”*°

The net result is that use of 10—C test scores in the same
menner for both blacks and whites does not appear to be under-
predicting the performance of blacks at the academy or on effi-
ciency ratings. This analysis does not, of course, deal with the
possibility of bias affecting the scores blacks obtain at the
academy or on efficiency ratings; it only serves as a foundation
for concluding that the 10—C is not to be criticized on the basis
of differential validity inquiries. The 20—B cannot be subjected
to these inquiries at the present time for lack of sufficient blacks
in any study group.

— Ads =

Operational Utility

It is not, however, sufficient that, as here, a test is shown to
have a statistically significant relationship to one or more
criteria and not to be differentially unfair to the adversely af-
fected racial group. In addition, in words of the EEOC Guide-
lines, § 1607.5(c), the relationship between the test and the
criterion must have ‘‘practical significance’’ or, in the words of
the DOJ Guidelines, § 12b(5), the usage of the test must be
evaluated ‘‘to assure that it is appropriate for operational use.’’
With different words, the two Guidelines are raising the same
concern.

In concluding that the 1O—C and 20—B tests are valid screen-
ing instruments, Drs. Farrar and McLaurin have emphasized
their significant relationship to grades in the training academies,
and this relationship cannot be doubted. However, as already
indicated, it is the court’s couclusion that relative standing in
the academies, as distinguished from successful completion of
academy training, is not an appropriate criterion unless it also »
be demonstrated that those academy grades are themselves valid
predictors of job performance. The studies reflect, however,
that for the most part the correlation between academy grades
and measures of job performance are not significant and, in the
few instances where significant correlations are found, the find-
ings are mixed—some being positive and other being negative.
A negative correlation, of course, indicates that the higher the
academy grades, the lower the performance ratings tend to be.
Although an employer is permitted to select the best person for
the job despite resulting impact on a racial group, it is not per-
mitted to engage in such selection procedures merely to employ
the best person for training.

Nor has it here been demonstrated that either test is a valid
predictor of successful completion of the required training
courses. According to the director of the policy academy, only
11 of the 733 cadets attending the academy since 1962 have

on Dt

failed to complete the training because of inadequate grades—
and no data has been presented as to their 10O—C test scores. Ac-
cording to the director of the firefighters academy, no student
has failed because of poor grades, of course, since historical use
of screening tests (whether the present ones or their predeces-
sors) has imposed a restriction of range, the conclusion does not
necessarily follow that every applicant could complete the train-
ing no matter how low his 10—C or 20—B score. However, it
should be noted that on ocassion the raw test scores used to
determine hiring eligibility have been substantially reduced,
without apparent impact on their successful completion of the
academies. And--while recognizing that, as noted by Dr.
McLaurin, this is not the regression lines developed from any of
the studies would predict passing academy averages even for
persons scoring zero on the 10O—C and 20—B tests.

As earlier discussed, the regular efficiency ratings are not
trustworthy criterion-measures of actual job performance. Even
if they were, the studies provide inconclusive findings with
respect to the 20—B (a significant correlation with a group of
196 firefighters, though only of a magnitude of .20, and an in-
significant correlation of .12 with a group of 162 firefighters),
and even more dubious results with respect to the 10—C.*?

The experimental ratings are, as previously indicated, con-
sidered by the court as an appropriate criterion measure. The
correlation, however, between the 20—B and these ratings is
found to be .08, which is, of course, not significant; and, while
a significant positive correlation is found with respect to the 103
firefighters having less than 3 years service, a significant
negative correlation is found for the 93 having at least 3 years of
tenure. Presumably, higher scores on the 20—B (which carried
over to higher scores during academy training) resulted in better
job performance for the first few years. Had this advantage
merely been erased after more time on the job, this would be
one matter—but, as stated, the findings actually showed a
significant negative correlation for the longer tenured fire-

—_ Pe

fighters, suggesting that over time the lower scoring applicants
made the better employees. Absent any indication that during
the first few years the lower scoring applicants had been inade-
quate on the job, one is hard pressed to conclude that the higher
scoring 20—B applicants are in fact the better persons to hire.
Further study might, of course, lead to other interpretations,
such as a determination that recent improvements in the training
given at the firefighters academy will result in better employees
not only initially but also over time—but no such conclusions
can be supported on the present evidence.

The correlation between 10—C and the experimental ratings
is, with 148 in the sample, significant at p< .05, but has a
magnitude of only .21. What do these figures mean? To begin
with, it should be understood that for a correlation coefficient
to be found significant at p< .05 is equivalent to saying that, if
in fact no relationship between the two variables exists for the
‘*population’’, the obtained results could be expected to occur
only one¢ in twenty such samples—and that therefore one can
be 95% confident that for the population (of applicants) there is
some correlation (or relationship) between the two variables. It
does not mean that one can be 95% confident that the popula-
tion coefficient is .21. Indeed, to state the population coeffi-
cient with only a 5% chance of error (i.e., # < .05) requires use
of a confidence interval: here, with a sample of 148, that the
true coefficient lies somewhere between .0504 and .3592. The
coefficient obtained from the sample is but an estimate of that
true population coefficient.

A second consideration is to look at the test scores in the par-
ticular sample in comparison with the scores of all persons in the
population. Where, as here, there is a restriction in the range of
test scores of those in the sample because of prior use of the test,
a statistical technique, called correction for restriction of range,
may be appropriate for determining the magnitude of the cor-
relation. This ‘‘corrected’’ coefficient, as reported by Drs. Far-
rar and McLaurin, is .36. It may be noted that utilization of the

—_

correlation involves the assumption that the two variables (test
scores and experimental ratings) are for the total population
‘*normally distributed;’’ and, insofar as rating scores are con-
cerned, it is just that—an assumption. Hence, it involves the
same type of risk as does the use of a regression formula for
values beyond the sample on which based—a technique which,
during the trial, provoked Dr. McLaurin’s criticism.

In general, other factors remaining the same, the greater the
magnitude of the coefficient the more likely it is that the test will
be appropriate for use. See DOJ Guidelines, §12b(5). The im-
portance of the size of the correlation coefficient can perhaps
best be understood by reference to certain basic statistical con-
cepts. The square of the correlation coefficient, called the
‘coefficient of determination,’’ gives the proportion of the
variance of the criterion scores which is accountable by
reference to variance of scores on the predictor test. Thus, with
a cori elation coefficient of .21, the study indicates that 4.4% of
the variance among experimental ratings is explainable by
reference to the variance in test scores, while 95.6% is not. Us-
ing the ‘‘corrected’’ coefficient of .36, still only 13% of the
variance among experimental ratings could be accounted for by
test score variance. By another formula, the correlation coeffi-
cient can be converted into a ‘‘coefficient of alienation’’, which
gives the size of the error in a‘ empting to predict experimental
rating scores from test scores relative to the error that would
result from a mere guess, i.e., by not using the test. This calcula-
tion, based on a correlation coefficient of .21, reflects that use
of the test predicts experimental rating scores with a margin of
error that is only 2% smaller than it would be without the test,
and, if based on the ‘‘corrected’’ coefficient of .36, indicates
that the margin of error is only 6.7% less than what would occur
by mere guess.

Anastasi comments, and quite properly so, that evaluation of
a test in terms of the error of estimate will for many testing pur-
poses be unrealistically stringent. Anastasi, PsyCHOLOGICAL

— A-49 —

TESTING, p. 166 (4th Ed. 1976).*? She notes that even tests with
an unusually high validity of .80 would appear to be inefficient
if used to predict individuals’ relative standing on some
criterion, but that most tests are merely used to determine which
individuals will exceed a given minimum standard of perfor-
mance or cutoff point in the criterion. The 10O—C, of course, is
utilized here both to screen applicants (cutoff scores) and to
rank those passing applicants.

While the magnitude of the correlation coefficient is obvious-
ly of great importance, there is no minimum coefficient ap-
plicable to all employment situations. See DOJ Guidelines §
12h(5). ‘‘Under certain circumstances, even validities as low as
.20 or .30 may justify inclusion of the test in a selection pro-
gram.’’ Anastasi, op. cit., p. 166.

Another approach towards evaluation of the relationship
found is to investigate the meaning of differences in test scores
in relation to the differences in criterion scores thereby
predicted. This involves use of the regression formula, which
cam be calculated from the correlation coefficient and the means
and standard deviations of the two variables. Thus, Dr. Roland
Ramsey, the plaintiffs’ expert witness, questioned the practical
value of the 1O—C by noting, from the Farrar-McLaurin study
involving 140 whites and 49 blacks, that an increase in raw test
scores of 40 points produced, under the regression lines given,
less than 5 points increase in predicted academy averages.

The common regression line computed for the 148 officers
with both test scores and experimental ratings is y =
300 + .162x, where x represents a given test score and y is the
rating predicted thereby. At first glance, this regression line
does not appear to be subject to the criticism made by Dr.
Ramsey respecting the other study, for it will be seen that, for
example, a test score difference of 10 will predict an experimen-
tal rating difference at 16. However, it should be understood
that the linear regression formula (as well as the criterion

— A-50 —

mean and criterion standard deviation, although not the cor-
relation coefficient) varies in direct proportion to any factor by
which criterion scores in the sample have been multiplied. The
Farrar-McLaurin study reports the overall experimental rating
as the summation of weighted components which comprise the
rating. For example, a sample subject rated as 5 (very good) on
each of the 12 rating components would, because of the
method, be reported as having a rating score of 499, while
another subject identically rated except for scores of 4 (good) on
the appearance and dependability components would receive an
overall rating of 483, with 698 representing a ‘‘perfect’’ score of
7’s on all components.

To prevent potential misinterpretation, it is well to consider
the regression line not only in the form in which.expressed in the
Farrar-McLaurin studies, but also in a form which does not
contain the inflation caused by the weighting procedure. This
can be done, while still retaining the concept of the components
having differing weights, by expressing the weights in a manner
in which the average weight is 1. That is, instead of the ‘‘com-
munication’’ component having a weight of 8.78 (as reported in
the study), of ‘‘problem solving’’ a weight of 9.44, of
‘‘learning’’ a weight of 8.22, etc., they can be shown as having
weights of 1.056, 1.136, and .989, etc., respectively. Then, a
summation of weighted component scores for a ‘‘perfect’’ score
of 7 on all 12 components would result in 84, identical with a
‘*perfect’’ score if not weighted. Not only does this method re-
tain the concept of weighing the different components, but the
transformation (whether of means, standard deviations, or
regression line) can simply be made by dividing the reported
results by a constant, here 8.31333. The obtained regression line
is y = 36.087 + 195x, where x represents the test score and ¥ is
the predicted rating (using the new method of expressing
weights). It will now be seen that, as with the other studies
criticized by Dr. Ramsey, a large difference in test scores pro-
duces only a small difference in predicted (un-inflated) ex-

— A-51 —

perimental rating scores, e.g., a 40 point raw score difference on
the test gives less than 8 poinis difference on the rating score.

Another method for evaluation, which is not complicated by
the weighting procedure, is to consider the ‘‘standard error of
estimate’’, which, for the data analyzed, is computed to be
92.63. Use of this statistic is demonstrated as follows: while a
raw test score of 70 through the regression formula predicts an
experimental rating of 413, one can through use of the standard
error of estimate determine, at p< .05, the experimental rating
actually to lie within the range of 231 to 595. Similarly, the
predicted experimental rating, at p.< .05, from a raw test score
of 40 is found to be in the range of 183 to 547. Obviously, there
is a potential overlap, where persons with raw scores of 40 and
70 on the test may nevertheless obtain the same experimental
rating score. It is, furthermore, possible to determine how much
of a difference in test scores is required for one to be able to
predict, atp< .05, that the higher-scoring applicant will receive
an experimental rating which also is higher’*—and this calcula-
tion results in a finding that a difference in test scores of over 86
raw points is necessary for such a conclusion to be reached. It
should be noted that the total range of raw test scores to this
date used to rank successful applicants (/.e., from a low raw
score of 48 to the perfect score of 120, which has not been ob-
tained by any) is yet too limited to enable one to say, at the .05
level, that the highest-scoring applicant would be predicted to
obtain a higher experimental rating than the lowest-scoring
applicant.

Since the 10—C is utilized not only in an attempt to rank the
successful candidates, but also to screen the unsuccessful, it is
appropriate to analyze the study results with respect to
minimum experimental ratings and to predictions for persons
scoring at, and below, the test cut-off scores. A test score of 48,
the lowest used as a cut-off, yields a predicted experimental
rating of 378, or an average unweighted rating on each com-
ponent of the experimental rating of 3.78 (3 = adequate,

aw ASE =

4 = good).’*. Common regression lines can, of course, also be
computed separately for each of the twelve components of the
rating. When this is done, one finds that a test score of 48 will as
to each component predict an unweighted rating of 3 or above,
i.e., at least ‘‘adequate.’’ While recognizing the risk in ex-
trapolation beyond the range of sample scores,’* but having lit-
tle else available for comparable analysis, one can look to
estimates of experimental ratings predicted by the regression
lines for test scores below 48. If this is done, it appears that even
with a test score of 0 the predicted rating is ‘‘adequate’’ or
above for the rating as a whole and for seven of the twelve com-
ponents. Even as to the five components for which the
estimated unweighted experimental rating from a 0 test score
would be less than 3, the predicted rating cannot be said to be
less than ‘‘adequate’’ at p< .0S5.

A technique for evaluating tests which employ cut-off scores
for screening purposes is to consider ‘‘false positives’ (persons
scoring below the cutoff but nevertheless scoring above the ac-
ceptable level of performance on the criterion) in relation to
‘‘false acceptances’ (persons scoring above the cutoff but
below the acceptable level of performance), thereby leading to a
comparison between the relative percentage of successful em-
ployees above and below the cut-off scores. However, neither
this method nor the Taylor-Russell tables (used to estimate net
gain in selection accuracy through test usage) can be directly us-
ed in the present case because all employees for whom
‘*success’’ data are available have been screened by the test. It is
possible to project the ‘‘base rate’’ (through use of the regres-
sion line, standard error of estimate, and normal distribution
curves) and then to conduct such inquiries, and, if this be done,
one finds any incremental validity to be negligible.

Still another approach is to estimate the effect of the test not
on the percentage of persons exceeding minimum performance,
but on overall performance of the selected persons. A table

— A-53 —

given by Anastasi, op. cit. at p. 173, gives the expected rise in
criterion scores through test usage in relation to its validity coef-
ficient and the selection ratio. In the present case, with a .21
coefficient and a selection rate of .19 (506 officers hired from
2721 applicants), and with a standard deviation of known
criterion scores of 94.74, one finds from the table that use of the
test (had the applicants actually been hired in the order of the
test scores) would probably have produced an average gain of 27
points in the total weighted experimental rating (over the rating
expected had the test not been used). This gain is equivalent to
being rated one point higher on three of the 12 rating com-
ponents. The average unweighted rating on each of the com-
ponents would, without the test, have been 4.07 (4 = good),
which compares to 4.35, using the test.

The assessment of utility of a test which, like the 1O—C, has a
statistically significant validity, albeit of very low magnitude,
must include certain value judgments. One of these involves
consideration of the nature of the job in question and the conse-
quences of a faulty hiring decision. There can be little dispute
that police officers perform a vital, and sensitive, function in
our society. The desirability of ‘‘upgrading’’ of law enforce-
ment has been emphasized in two reports received in evidence,
the 1967 Task Force Report on the Police, issued by the Presi-
dent’s Commission on Law Enforcement and the Administra-
tion of Justice, and the 1973 Report on Police, issued by the Na-
tional Advisory Commission on Criminal Justice Standards and
Goals. Economic costs are also involved, particularly in view of
the cost of academy training of officers and the restraints placed
upon discharge of marginal officers under the civil service laws.

Without demeaning the importance of law enforcement of-
ficials, however, it can hardly be said that the possibility of oc-
casional selection of an inept officer presents the same type of
daily economic and human risk factors as is involved, for exam-
ple, in the employment of airline pilots or bus-drivers. Cf.
Spurlock v. United Airlines, Inc., [5 EPD 4 7996] 475 F.2d 216,

— AS =

219 (CA7 1972); Usery v. Tamiami Trail Tours, Inc., [11 EPD
910,916] 531 F.2d 224 (CAS 1976) (Age Discrimination in
Employment Act case, with somewhat related question). A
twelve months’ probationary period is provided, during which
time the occasional incompetent may be detected and dismissed
and during part of which time the employee is undergoing train-
ing rather than being ‘‘on the street’’. The principal public con-
cern, it would appear, is not so much that the most able officers
be employed (though that. certainly, would be desirable) as that
the emotionally unfit not be employed. In this context, it is
perhaps noteworthy that the 1967 Presidential Commission’s
report contained a recommendation for use of psychological
tests to detect applicants with personality defects (but no such
recommendation respecting aptitude tests); and the 1973 Na-
tional Advisory Commission’s report, while acknowledging the
desirability of valid aptitude tests, was skeptical as to the results
of research to that date. So far as the court has been informed,
the 10O—C was not designed, and has not been validated, for use
in detecting emotional disorders or defects.

The DOJ Guidelines § 12b(b), provide that, in determining
operational appropriateness, one should consider ‘‘the degree
of adverse impact of the procedure, the availability of other
selection procedures of greater or substantially equal validity,
and the need of an employer, required by law or regulation to
follow merit principles, to have an objective system of
selection.’’ Obviously this latter factor (requirement under civil
service law to give some ubjective test) cannot by itself suffice as
justification for a test which has, as here, substantial adverse
impact on a racial group. While it can be said that no other
available’’ selection with greater validity than the 10—C has
been found, yet it must also be said—considering the minimal
benefits resulting from the 10—C in the context of this employ-
ment situation—that no-test-at-all has ‘‘substantially’’ the same
practical validity as the 10O—C.

In summary, the 20—B Firefighter test has not been shown to
be a valid predictor of a job-relevant criterion measure and the

— A-55 —

10—B Policeman test, while having a statistically significant
relationship of a very low magnitude with a job-relevant
criterion measure, has not been shown to be appropriate for
operational use in screening or ranking applicants.**

Violation and Remedy

Having concluded that use of the 10O—C and 20—B has hadan
adverse impact upon black applicants and that the studies
presented fail to demonstrate job-relatedness, the court must
nevertheless determine when the requirements of law were
violated and what relief is appropriate therefor. This inquiry
should involve no less care then consideration of the tests
themselves.

The requirements of Title VII first became applicable to the
Personnel Board in March 1972. At that time, and for many
years earlier, the Board was required by state law to administer
appropriate tests to screen and rank applicants—a requirement
which continues to the present time, subject to any over-riding
proscriptions of Title VII. It had several years earlier selection
the 10O—C and 20—B tests as the best tests then available, with
the hope that black applicants would fare better than under
previous tests. By March 1972 a preliminary, in-house validity
study had been conducted, which reflected some improvement
in hiring of blacks and the indication of appropriate validity
based upon relationship with existing criterion measures. An in-
depth independent validation study was immediately under-
taken, including investigation of alternative or supplemental
selection procedures to improve the predictive validity or
decrease adverse impact upon blacks. At least since 1965 the
Board has not intentionally discriminated against black ap-
plicants but, to the contrary, has attempted to increase black
employment within the options available under state law, in-
cluding modification of the scoring key for the 10—C when
recommended by the consultants as a method for increasing
validity of the test for black applicants.

— A-56 —

The preliminary reports from the consultants, made while
more trustworthy measures of job performance were being
developed, contained signs of potential validity and recom-
mended continued usage of the test pending the additional
studies. Not until April 25, 1975, with respect to the 10—C, and
July 8, 1976, with respect to the 20—B, were the studies using
these new criterion measures completed and reported to the
Board. It was on these respective dates that, in the court’s opin-
ion, it should have been concluded that provisional use of the
tests was no longer permissible. Prior thereto, the Board was, in
the court’s opinion, justified in continuing to use the tests (and
the eligibility lists generated therefrom) in anticipation of
favorable results from those studies. Use of the tests (or of the
eligibility lists therefrom) was thereafter, however, contrary to
the requirements of Title VII, which override state law inconsis-
tent therewith.

The remedy should be appropriate to the violation found. In
this case, from X-11, it is found that, for the two administra-
tions of the 10O—C from which eligibility lists used after April
25, 1975, were formed, 658 (or 88%) of the 747 white applicants
were placed on the eligibility lists. Had a like percentage of
black applicants been so placed, a total of 252 would have been
on the lists—128 more than actually placed on the list. Accord-
ingly, to the extent they are still interested, an additional 128
blacks from the prior administrations of the test should be
added to the present eligibility lists. This remedy only relates to
prohibited use of the 10O—C as a screening instrument. An addi-
tional measure is needed to correct for the improper use of the
test as ranking procedure. Had the eligibility lists been represen-
tative of the applicant group and had certifications from the list
likewise been representative of the racial composition of the list,
approximately 28% of the persons certified would have been
black. It is clear that there has been ‘‘under-certification’’ of
blacks by this standard, although the precise degree cannot be
determined from evidence before the court, which gives such in-
formation only by calendar years. The Board is directed to

oe

ascertain the extent of such under-certification and in future
certifications to include at least 1 black candidate for every 3
certified until such time that, considering the certifications after
April 25, 1975, total number of blacks certified becomes 28% of
the total number of persons certified. Thereafter (and until
some new selection procedures are adopted which are sufficient-
ly job-related or which have no adverse impact upon blacks) at
least 2 of every 7 persons certified by the Board from the revised
present list shall be black, provided there be a sufficient number
of black applicants interested.

A similar investigation of X-11 with respect to the 20—B,
where only one eligibility list has been in effect since July 8,
1976 (the date of the report involving the experimental ratings),
results in a conclusion that 91 black applicants should be added
to the present eligibility list for firefighters, that at least 1 of
every 3 persons hereafter certified shall be black until such time
that (considering certifications after July 8, 1976) the total
number of blacks certified becomes 14%** of the total number
of persons certified, and that thereafter (pending adoption of
some other valid or nondiscriminatory selection instrument) at
least 1 of every 7 certified by the Board from the revised current
list shall be black.

This order does not preclude use of the 10O—C or 20—B as a
device for ranking one white as against another white, or one
black as against another black. Such a use may be invade by the
Board, if it so desires, without any discriminatory impact on a:
racial group. The order does not prevent the Board from new
administrations of the 10—C or 20—B (or other tests) or from
forming new eligibility lists from time to time; provided,
however, that, unless and uatil a selection instrument is found
which either has no adverse impact racially or is sufficiently
valid, the test results shall be used in a manner consistent with
this opinion, i.e., the eligibility list and certifications to be
representative racially of the applicant group regardless of test
scores.

— A-58 —

APPENDIX C

Footnotes

' Fourteen separate county and municipal employers are covered by
the law. Cities with a population of under 5,000 are excluded.

? With multiple vacancies, the number of persons certified is two
more than the number of vacancies to be filled.

> Not presently at issue are requirements (such as age or education)
which may be imposed as conditions to taking an examination, nor are
specifications (such as residence within Jefferson County) which may
give preference to certain applicants.

* Unless otherwise noted, reference to police officers in the balance
of this opinion will also refer to deputy sheriffs.

* Under F.R.Civ.P. Rule 42, the four actions were consolidated
with respect to challenges to Personnel Board tests and a separate trial
was scheduled respecting the attacks on the Policeman 10—C, Fire-
fighter 20—B, and Office Worker 30—B tests. At the trial the plain-
tiffs indicated that the attack on the Office Worker 30—B test was
dropped for lack of evidence of adverse impact and that any attack on
the 1O—C and 20—B tests based on sex was likewise dropped for lack
of evidence.

* The 10—C and 20—B tests were adopted by the Board after ini-
tially experimenting, commencing in January 1966, with alternate
forms of the IPMA tests.

’ The Board’s studies resulted in selection of the 10—C form
because of its significant and positive correlation with a greater
number of the selected criteria measures than did the alternate IPMA
form. The study indicated a significant and positive correlation be-
tween 10—C scores and training academy average (a. well as several
course grades in the academy) and between the training academy
average and the officers’ latest efficiency ratings.

* Not until 1968 was the residency requirement of the City of Birm-
ingham removed by the city ordinance. A Jefferson County
preference remains in effect, but this can hardly disadvantage blacks,
who constitute a larger portion of the Jefferson County population
than of neighboring counties.

— A-59 —

* An intent to discriminate would presumably be required for there
to be a violation of 42 U.S.C. § 1983, if not of 42 U.S.C. § 1981. See
Washington v. Davis,.—U.S.—(June 7, 1976).

‘© Albemarle Paper Co. v. Moody, [9 EPD 4 10,230) 422 U.S. 405,
425 (1975).

'' Griggs v. Duke Power Co., (3 EPD 48137] 401 U.S. 424, 432
(1971).

'2 There are some conflicts between the EEOC and the DOJ Guide-
lines. However, it is not necessary as to the issues presently before the
court that a choice be made between the two.

' Information for this table has been taken from X-11 and from
data respecting hires supplied by the parties at the court’s request
following formal close of the evidence. Certain caveats should be
noted: The results of the 1O—C exam administered on April 29-30,
1971, have been eliminated because it was not used for employment
decisions after March 24, 1972. The results of the 1O—C exam ad-
ministered on September 30 and October I, 1971 and of the 20—B
exam administered on May 26-27, 1971, have been included in the
tabulation, even though in part the eligibility lists taken therefrom
would have been used prior to March 24, 1972. The number of hires
_ includes those hired in 1972 prior to March 24, 1972. As the Board
points out, the number of hires is affected by voluntary choices of the
candidates (such as declining job offers or waiving consideration),
but, lacking reliable data on such matters for both whites and blacks,
the court has looked to actual hires as the measure of the overall selec-
tion ratios. Finally, it should be noted that, since persons are permit-
ted to take exams more than once, the applicant figures do not com-
pletely accurately reflect the number of different individuals involved.
These limitations do not, in the court’s opinion, prevent meaningful
usage of the data for the purposes indicated.

'* These figures are derived from X-12 and are subject to the ap-
propriate caveats indicated in fn. 13, supra. Moreover, X-12 does not
have any information for one eligibility list and does not contain
percentile information as to several lists. In a very real sense, ‘‘pass-
ing’’ an exam is measured not by obtaining a derived score of at least
70 (and thereby being entered on the eligibility list), but by obtaining a
score sufficiently high to be placed on the eligibility list at a position
where, during use of that list, the candidate will actually be certified to
an employing agency.

ye

'S Until a couple of months prior to trial, the litigation was being
prepared with the anticipation that all tests under challenge were to be
considered at a single hearing. When the decision was made by the
court that the first trial would only concern the 10—C, 20—B, and
30—B tests, counsel and witnesses were freed to shift their attention to
‘““loose-ends’’ on these three tests.

'6 Some items on each test elicit knowledge which apparently would
be needed for performance of job functions; others do not. Some ef-
fort has been made by the test developer to give ‘‘face validity’’, as by
expressing an item involving numerical problem solving in the context
of information with which job occupants would be dealing. Face
validity does not affect validity for usage as a selection procedure so
much as it may overcome motivational resistance by those taking the
test.

'? **Reliability refers to the consistency of scores obtained by the
same persons when reexamined on the same test on different occa-
sions, or with different sets of equivalent items, or under other
variable examining conditions.’’ Anastasi, PsycHOLOGICAL TESTING, Pp.
103 (4th Ed. 1976). The methods used in this study to estimate
reliability were the split-half technique (corrected by the Spearman-
Brown prophecy formula) and the Kuder-Richardson formula 20.

'* A test item correctly answered by too high a proportion of the
applicants is considered not difficult enough; correctly answered by
too low a proportion, it is considered too difficult. In this study items
correctly answered by 30 to 70% of the applicants were considered
satisfactorily difficult.

'? Item ‘‘discrimination’’ is in essence a comparison between scores
made on an individual test item and scores made on the total test,
thereby ascertaining whether particular items ‘‘discriminate’’
significantly in predicting success on the test as a whole. In the study,
significance was established at p< _ .05.

2 As previously indicated, the 10O—

[Text truncated at 120,000 characters. The full text is on the page linked above.]

---

Source: Frix Law Library, https://www.frixlaw.com/law-library/documents/brief%3Amicro_IA40385007_0425%3A2. Public record. Not legal advice.
