# Thomas v. Allen

> District Court, N.D. Alabama · April 21, 2009 · 614 F. Supp. 2d 1257

URL: https://www.frixlaw.com/law-library/cases/2141960

## Case

- **Full name:** Kenneth Glenn THOMAS, Petitioner, v. Richard ALLEN, Commissioner of the Alabama Department of Corrections, Respondent
- **Court:** District Court, N.D. Alabama
- **Decided:** April 21, 2009
- **Citations:** 614 F. Supp. 2d 1257; 2009 U.S. Dist. LEXIS 50825; 2009 WL 1353722
- **Precedential status:** Published
- **Opinion:** Opinion by Smith
- **Judges:** Smith
- **Cited by:** 21 later opinions in the Frix Law Library

## Citator (automated)

- No negative treatment found by the automated citator. That is not the same as a confirmation that the case is good law; read the citing cases.
- Full citator and citing cases: https://www.frixlaw.com/law-library/cases/2141960

## How later opinions describe it (automated extraction)

- holding that "[a] court must consider the Flynn Effect and the standard error of measurement in determining whether a petitioner’s IQ score falls within a range containing scores that are less than” seventy
- noting that, even though the Alabama Supreme Court recognized the legal cut-off score for a finding of significantly subaverage intellectual functioning as an IQ of 70 or below, a court should not look at a raw IQ score as a precise measurement and must consider the “Flynn eff…
- noting that even though recognized legal cutoff score for finding of significantly subaverage intellectual functioning is IQ of 70 or below, a court should not look at raw IQ score as precise measurement and must consider “Flynn effect” and standard error of measurement in det…
- noting that the AAMR manual states at least eight times that the clinical standard for significantly subaverage intelligence functioning is two standard deviations below the mean considering the standard error of measurement for the specific assessment instruments used and the…
- describing “reliability” as the ability of a testing instrument to yield consistent results, and “validity” as its ability to accurately measure the target characteristic

## Opinion text

MEMORANDUM OPINION
C. LYNWOOD SMITH, JR., District Judge.
Petitioner, Kenneth Glenn Thomas, is an inmate in the custody of the Alabama Department of Corrections. He was sentenced to death by the Circuit Court of Limestone County, Alabama, for the intentional murder of Mrs. Flossie McLemore during the course of a burglary.
See
Ala. Code § 13A-5-40(a)(4) (1975). Following exhaustion of his direct appeal rights and post-conviction remedies in the state court system,
1
Thomas filed a petition in this court, seeking relief in the nature of
habeas corpus. See
28 U.S.C. § 2254 . All but one of his claims for relief were dismissed in an order and memorandum opinion entered on March 6, 2007;
2
the sole claim that survived the motion for summary judgment filed by the respondent Commissioner of the Alabama Department of Corrections was Thomas’s contention that he is mentally retarded.
3
If Thomas is retarded, then the Supreme Court’s decision in
Atkins v. Virginia,
536 U.S. 304 , 122 S.Ct. 2242 , 153 L.Ed.2d 335 (2002) — holding under the Eighth Amendment that “death is not a suitable punishment for a mentally retarded criminal,”
id.
at 321 , 122 S.Ct. 2242—will require his death sentence to be vacated.
4
This court originally ordered that Thomas’s
Atkins
contention be remanded to the state court system, with instructions to reevaluate his claim in accordance with standards established by binding authorities.
5
Upon joint motion of the parties,
*1260
however, this court was persuaded to withdraw that remedy in favor of an order amending the judgment to reflect that the claim would be litigated on the merits in this court.
6
An evidentiary hearing commenced on May 19, 2008, and concluded on May 20, 2008.
7
Thereafter, the parties filed post-hearing briefs.
8
Respondent’s brief concedes that the evidence presented during the May hearing establishes that “Thomas’s IQ now falls in the mild mental retardation range,”
9
but goes on to argue that
Atkins
provides him no relief because he failed to prove that “he suffered from significantly sub-average general intellectual functioning or substantial deficits in adaptive behavior in the developmental period.”
10
Respondent thus reduced his argument against a finding of mental retardation to this question:
Has petitioner proven, by a preponderance of the evidence,
11
that he met the criteria for mental retardation prior to the age of eighteen years
(the so-called “developmental period”)? The remainder of this opinion addresses that question, as well as other issues.
12
*1261
Table of Contents
I. The Legal Criteria Defining Mental Retardation............................1262
II. Diagnostic Criteria Defining Mental Retardation ...........................1263
A. Assessment of Intellectual Functioning................................1263
1.
Standardized assessment
instruments...............................1264
a.
The Wechsler Adult Intelligence
Scales...........................1265
b.
The Stanford-Binet Intelligence
Scales...........................1267
2.
Measurement errors and cut
scores..................................1268
a.
The effect of “standard errors of measurement” on “true” IQ
scores......................................................1269
b.
The stipulated SEM and its effect upon determination of petitioner’s
IQ..............................................1271
3.
The “Flynn Effect, ” IQ gains over time, and cut
scores.................1275
4.
Conclusions
......................................................1279
B. Assessment of Adaptive Behavior......................................1281
1.
Standardized assessment
instruments...............................1282
2.
The importance of clinical
judgment.................................1283
3.
The assessment instrument used in this case
.........................1284
III. Expert Witnesses and “Best Practices”.....................................1286
A. Petitioner’s Experts..................................................1286
1.
Dr. Karen
Salekin.................................................1286
2.
Dr. Daniel
Marson................................................1287
B. Respondent’s Expert.................................................1288
C. “Best Practices”.....................................................1289
IV. Assessments of Petitioner’s Intellectual Functioning........................1293
A. Intelligence Assessments Performed Prior to Age Eighteen..............1293
1.
October I, 1968
— age
nine years and seven
months....................1293
2.
October 12,1972
— age
thirteen years and seven
months................1294
3.
June 28,1978
— age
fourteen years and two
months....................1294
4.
March 18,1975
— age sixteen........................................1296
5. Findings.........................................................1296
B. The Intelligence Assessment Performed on April 11, 1977................1297
C. Intelligence Assessments Performed Near the Date of the Offense........1300
1.
March 21, 1985
— age
twenty-six years
...............................1300
2.
January 2j, 1986
— age
twenty-six years and ten
months...............1301
3. Findings.........................................................1301
D. Intelligence Assessment Performed in Preparation for Hearing..........1301
V. Assessments of Petitioner’s Adaptive Behavior..............................1303
A. Petitioner’s Deficiencies Prior to Age Eighteen.........................1303
1.
The home environment
— birth
to age
12..............................1304
2.
Adolescence and foster care
— ages
12 through
18......................1305
3.
Schooling
........................................................1307
4.
Assessments by third-party
informants..............................1309
5. Findings.........................................................1311
B. Petitioner’s Deficiencies at the Time of the Offense.....................1312
C. Petitioner’s Present Adaptive Functioning Abilities.....................1321
VI.Conclusions 1322
*1262
I. THE LEGAL CRITERIA DEFINING MENTAL RETARDATION
The Supreme Court’s decision in
Atkins v. Virginia
did not dictate a national standard for determining whether a criminal defendant is mentally retarded and, for that reason, not subject to the ultimate sanction of the law. Instead, the Court left to the states “the task of developing appropriate ways to enforce the constitutional restriction” upon the execution of mentally retarded convicts.
Atkins,
536 U.S. at 317 , 122 S.Ct. 2242 (citation, internal quotation marks, and footnote omitted).
The Court’s reticence to propound hard and fast rules undoubtedly was grounded in the fact that the statutory definitions of mental retardation adopted by Congress and those states that then prohibited the execution of mentally retarded persons were not identical. Even so, the Court observed that all of the existing statutes generally conformed to diagnostic criteria promulgated by the American Association on Mental Retardation and the American Psychiatric Association.
The American Association on Mental Retardation (AAMR) defines mental retardation as follows:
“Mental retardation
refers to substantial limitations in present functioning. It is characterized by significantly subaverage intellectual functioning, existing concurrently with related limitations in two or more of the following applicable adaptive skill areas: communication, self-care, home living, social skills, community use, self-direction, health and safety, functional academics, leisure, and work. Mental retardation manifests before age 18.” Mental Retardation: Definition, Classification, and Systems of Supports 5 (9th ed.1992).
The American Psychiatric Association’s definition is similar: “The essential feature of Mental Retardation is significantly subaverage general intellectual functioning (Criterion A) that is accompanied by significant limitations in adaptive functioning in at least two of the following skill areas: communication,' self-care, home living, social/interpersonal skills, use of community resources, self-direction, functional academic skills, work, leisure, health, and safety (Criterion B). The onset must occur before age 18 years (Criterion C). Mental Retardation has many different etiologies and may be seen as a final common pathway of various pathological processes that affect the functioning of the central nervous system.” Diagnostic and Statistical Manual of Mental Disorders 41 (4th ed.2000). “Mild” mental retardation is typically used to describe people with an IQ level of 50-55 to approximately 70.
Id.,
at 42-43.
Atkins,
536 U.S. at 309 n. 3, 122 S.Ct. 2242 (emphasis in original). The
Atkins
opinion thus pointed the states in the direction of clinical definitions that have three constituent parts: that is, in order to be diagnosed as “mentally retarded,” the person under evaluation must exhibit (i) before the age of eighteen years
(ii)
significantly sub-average intellectual functioning, accompanied by (iii) significant limitations in adaptive functioning.
Similarly, the Alabama Supreme Court subsequently held that a defendant seeking the benefit of
Atkins
“must have significantly subaverage intellectual functioning (an IQ of 70 or below), and significant or substantial deficits in adaptive behavior. Additionally, these problems must have manifested themselves during the develop
*1263
mental period (i.e., before the defendant reached age 18).”
Ex parte Perkins,
851 So.2d 453, 456 (Ala.2002).
13
The Alabama Supreme Court layered a gloss on the
Perkins
definition in
Smith v. State,
No. 1060427, 2007 WL 1519869 (Ala. May 25, 2007), holding that a defendant must exhibit significantly subaverage intellectual functioning abilities and significant deficits in adaptive behavior during
three periods
of his life: before the age of eighteen; on the date of the capital offense; and currently.
All three factors must be met in order for a person to be classified as mentally retarded for purposes of an
Atkins
claim.
Implicit in the definition is that the subaverage intellectual functioning and the deficits in adaptive behavior must be present at the time the crime was committed as well as having manifested themselves before age 18.
This conclusion finds support in examining the facts we found relevant in
Ex parte Perkins
and
Ex parte Smith
and finds further support in the
Atkins
decision itself, in which the United States Supreme Court noted: “The American Association on Mental Retardation (AAMR) defines mental retardation as follows:
‘Mental retardation
refers to substantial limitations in
present
functioning.’ ” 536 U.S. at 308 n. 3, 122 S.Ct. 2242 , 153 L.Ed.2d 335 (second emphasis added). Therefore, in order for an offender to be considered mentally retarded in the
Atkins
context, the offender must
currently exhibit
subaverage intellectual functioning,
currently exhibit
deficits in adaptive behavior,
and
these problems must have manifested themselves before the age of 18.
Smith,
2007 WL 1519869, at *8 (emphasis supplied).
See also Holladay v. Allen,
555 F.3d 1346, 1353 (11th Cir.2009) (same).
II. DIAGNOSTIC CRITERIA DEFINING MENTAL RETARDATION
A. Assessment of Intellectual Functioning
“The assessment of intellectual functioning is essential to making a diagnosis of mental retardation, as virtually all definitions of mental retardation make reference to significantly subaverage intellectual functioning as one of the diagnostic criteria.” Ruth Luckasson
et al., Mental Retardation: Definition, Classification, and Systems of Supports
51 (Washington, D.C.:
*1264
American Association on Mental Retardation 10th ed.2002) (hereafter, “AAMR,
Mental Retardation
”).
14
[intelligence is not merely book learning, a narrow academic skill, or test-taking smarts. Rather, it reflects a broader and deeper capacity for comprehending our surroundings — catching on, making sense of things, or figuring out what to do. Thus the concept of intelligence represents an attempt to clarify, organize, and explain the fact that individuals differ in their ability to understand complex ideas, to adapt effectively to their environments, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by thinking and communicating.
Id.
at 40 (citation omitted).
1. Standardized assessment instruments
“Although far from perfect, intellectual functioning is still best represented by IQ scores when obtained from appropriate assessment instruments.” AAMR,
Mental Retardation
at 14. The “Wechsler Adult Intelligence Scales — Third Edition” (WAIS-III) and the “Stanford-Binet Intelligence Scales — Fifth Edition” (SB5) are the two most widely used IQ tests,
id.
at 59, and they were utilized by the parties’ expert witnesses to assess petitioner’s current intellectual functioning. Both are “standardized” assessment instruments, meaning that: (a) during the design phase, each was administered to a large, representative sample of the population for which the test was intended to provide reliable, normative data;
15
(6) the reliability and validity of each test has been established over time by cumulative empirical applications and analysis;
16
and (c) each test must be administered, scored, and interpreted by trained examiners in strict accordance with instructions issued
*1265
by the test developers.
17
a.
The Weehsler Adult Intelligence Scales
The first IQ assessment instrument to be named the “Weehsler Adult Intelligence Scales” (WAIS) was published in 1955 as a revision of the “Wechsler-Bellevue Intelligence Scales” developed in 1939 by Dr. David Weehsler, a clinical psychologist, during his association with Bellevue Psychiatric Hospital in New York City.
18
The theoretical basis for the device was Dr. Wechsler’s belief that intelligence is a multifaceted construct that enables an individual to comprehend and deal effectively with the environment in which he or she lives and works.
19
After dividing intelligence into two major types of skill sets, verbal and performance, Weehsler used the statistical technique of factor analysis to determine specific skills within those two major domains.
20
The most recent iteration of this assessment instrument, the so-called “Weehsler Adult Intelligence Scales — Third Edition” (WAIS-III), is a 1997 revision of the “Weehsler Adult Intelligence Scales — Revised Edition” (WAIS-R) published in 1981.
21
It is an individually administered test designed to assess the intelligence of individuals ranging in age from 16 years to 89 years.
*1266
The WAIS-III was standardized on 2,450 adults from the United States. Thirteen separate .standardization groups were created by age classification. Within each group, the number of males and females was roughly equal (except for the 65 to 89 age group, which contained more females), and there was Census-based stratification for race or ethnicity (White, African American, Hispanic), education, and geographic region based on Census reports.
AAMR,
Mental Retardation
at 61. As shown in the following table, fourteen sub-tests are equally divided ■ between seven Verbal Scale subtests and seven Performance Seale subtests. The number preceding each subtest indicates the standardized order of administration:
i.e.,
the “Picture Completion” subtest (on the Performance Scale) is administered first; the “Vocabulary” subtest (Verbal Scale) is given next, and so on in alternating order to assist an examiner in maintaining the' test-subject’s interest.
WAIS-III SUBTESTS (Grouped According to Verbal and Performance Scales)
Verbal Scale
Performance Scale
2. Vocabulary A series of orally and visually presented words that the examinee orally defines.
22
1. Picture Completion A set of color pictures of common objects and settings, each of which is missing an important part that the examinee must identify.
23
4. Similarities A series of orally presented pairs of words for which the examinee explains the similarity of the common objects or concepts they represent._
3. Digit Symbol — Coding A series of numbers, each of which is paired with its own corresponding hieroglyphic-like symbol. Using a key, the examinee writes the symbol corresponding to its number.
6. Arithmetic A series of arithmetic problems that the examinee solves mentally and responds to orally._
5. Block Design A set of modeled or printed two-dimensional geometric patterns that the examinee replicates using two-color cubes._
8. Digit Span A series of orally presented number sequences that the examinee repeats verbatim for Digits Forward and in reverse for Digits Backward._
7. Matrix Reasoning A series of incomplete gridded patterns that the examinee completes by pointing to or saying the number of the correct response from five possible choices._
9. Information
*1267
A series or orally presented questions that tap the examinee’s knowledge of common events, objects, places, and people
*1266
10. Picture Arrangement
*1267
A set of pictures presented in a mixed-up order that the examinee rearranges into a story sequence._
11. Comprehension A series of orally presented questions that require the examinee to understand and articulate social rules and concepts or solutions to everyday problems.
12. Symbol Search A series of paired groups, each pair consisting of a target group and a search group, The examinee indicates, by marking the appropriate box, whether either target symbol appears in the search group.
13. Letter-Number Sequencing A series of orally presented sequences of letters and numbers that the examinee simultaneously tracks and orally repeats, with numbers in ascending order and the letters in alphabetical order.
14. Object Assembly A set of puzzles of common objects, each presented in a standardized configuration, that the examinee assembles to form a meaningful whole.
Source: D. Wechsler,
WAIS-III Administration and Scoring Manual
(San Antonio: The Psychological Corp. 1997)
As the Supreme Court observed, the WAIS-III is scored by adding together the number of points earned by a test subject on these different subtests, and then
using a mathematical formula to convert this raw score into a scaled score. The test measures an intelligence range from 45 to 155. The mean score of the test is 100,
[24]
which means that a person receiving a score of 100 is considered to have an average level of cognitive functioning. It is estimated that between 1 and 3 percent of the population has
an IQ between 70 and 75 or lower, which is typically considered the cutoff IQ score for the intellectual function prong of the mental retardation definition.
Atkins,
536 U.S. at 309 n. 5, 122 S.Ct. 2242 (emphasis supplied, citations omitted).
b.
The Stanfordr-Binet Intelligence Scales
The Stanford-Binet battery
of
intelligence assessment instruments is far older than the Wechsler series. The original test was devised by psychologist Alfred Binet, who was charged by the French government with developing a method of identifying intellectually-deficient school-age children for placement in special education programs. Research conducted by Binet and physician Theophilus Simon between 1905 and 1908 at a school for mentally-retarded boys led to the development of the Binet-Simon test, which employed questions of increasing difficulty to measure such attributes as attention, memory, and verbal skills. In 1916, Lewis Terman, a psychologist at Stanford University in California, released the “Stanford Revision of the Binet-Simon Scale.”
25
That test has been revised several times since its inception, and currently is in its fifth edi
*1268
tion — a version that is generally referred to as either the “Stanford-Binet 5” or “SB5.”
The SB5 consists of a battery of tests that assess a person’s intelligence across four areas of intellectual functioning: verbal reasoning; quantitative reasoning; abstract and visual reasoning; and short-term memory. These areas are covered by fifteen subtests, including vocabulary, comprehension, verbal absurdities, pattern analysis, matrices, paper folding and cutting, copying, quantitative, number series, equation building, memory for sentences, memory for digits, memory for objects, and bead memory. All test subjects take an initial vocabulary test that, together with the subject’s age, determines the number and level of sub-tests to be administered. Total testing time is 45 to 90 minutes, depending upon the subject’s age and the number of subtests administered. Raw scores are based on the number of items answered, and are converted into a “standard score” corresponding to the test subject’s age group, similar to an IQ measure.
26
According to the website of the SB5’s publisher, Riverside Publishing Company, “[njormative data for the SB5 were gathered from 4,800 individuals between the ages of
2
and 85 + years. The normative sample closely matches the 2000 U.S. Census. Bias reviews were conducted on all items for the following variables: gender, ethnicity, culture, religion, region, and socioeconomic status.”
27
The SB5 has a mean IQ score of 100,
28
and a standard deviation of 16: the “standard deviation” indicates how far a particular individual’s score falls above or below the mean score for the test subject’s age group.
29
For example, if an eight-year-old child achieved a score of 116 on the SB5, the child’s score would be one standard deviation above the mean performance score of all eight-year-olds in the representative sample.
2.
Measurement errors and cut scores
A key task for the
...
analyst applying a scientific method to conduct a particular analysis, is to identify as many sources of error as possible, to control or to eliminate as many as possible, and, to estimate the magnitude of remaining errors so that the conclusions drawn from the study are valid.
30
A critical question that must be addressed is:
“How much confidence can this court place in the IQ scores produced by the tests administered to petitioner?”
*1269
Even though most of the intelligence tests that will be discussed later in this opinion are generally considered to be reliable assessment instruments that produce valid IQ scores,
31
there still exists an inherent potential for “measurement error.” Measurement errors can be either random or systematic. “Random errors” are caused by any factors that randomly affect measurement of test variables. Examples include factors peculiar to the individual test-subject
(e.g.,
fatigue, poor health), the testing situation
(e.g.,
environmental distractions inhibiting concentration), the manner in which the test was administered
(e.g.,
the examiner’s failure to adequately explain each segment of the test, .or to strictly follow the test developer’s instructions), the examiner’s lack of training, or a multitude of other, unpredictable variables that artificially inflate or deflate a test subject’s performance. The important attribute of random errors is that they do not have consistent effects across the entire population of persons to whom the test instrument is administered.
“Systematic errors,” on the other hand, are test-specific sources of error that are caused by any factors that
systematically
affect IQ measurements across the entire population of test subjects. Systematic errors also can be generated by many variables, but usually they can be traced to inadequacies in the assessment instrument itself. Unlike random errors, systematic errors tend to have consistently positive or negative effects upon the performance scores generated by each individual to whom the test is administered. To use a pedestrian example, suppose “you recorded the temperature every day in your backyard. If your thermometer was incorrectly calibrated, so that it was always 4 degrees too high, the faulty thermometer would produce a systematic error (an upward bias) in your measurement.”
32
One such systematic inaccuracy in the intelligence assessment instruments administered to petitioner over the course of his life will be discussed in Part 11(A)(3)
infra,
addressing the so-called “Flynn effect.”
a.
The effect of “standard errors of measurement” on “true” IQ scores
A “true” IQ score is the hypothetical score a test subject would obtain if no measurement error influenced his or her performance during the administration of an intelligence assessment instrument. No clinician, much less this court, can state a test subject’s “true” score with
absolute
certainty, because error
always
is present in any testing situation. For well-designed test instruments, however, the random errors of individual test subjects are randomly distributed. Accordingly, the more scores that are grouped together, the more likely it is that an individual test subject’s errors will be cancelled out by the results obtained by administration of the assessment instrument to a very large sample of the population during the “normative research and survey process”
33
*1270
leading to the development of a “standardized test.”
34
That is the reason reputable IQ assessment instruments like the Wechsler and Stanford-Binet batteries were standardized (“normed”) on large, representative samples of the general population.
35
During the design process of developing normative standards, measurement errors are taken into account through use of a mathematical concept that statisticians and psychologists refer to as the “Standard Error of Measurement”
(SEM),
The
SEM
is an index of the variability of test scores produced by persons forming the normative sample. The
SEM
makes it possible to determine the reliability of a particular intelligence assessment instrument, and the level of confidence that can be placed in the scores produced by an administration of the test to an individual test subject.
Every intelligence test has a
SEM,
which is used to calculate a range of scores lying along a continuum (think of a yardstick), and evenly arranged on each side of the IQ score obtained during an individual administration of the test. The test subject’s “true” IQ most likely lies within that range above and below his or her actual test score.
For example, Kenneth Glenn Thomas was nearly 49 years of age on the date WAIS-III and SB5 IQ assessment instruments were administered to him by the parties’ expert witnesses.
36
The
SEM
for a full-scale IQ score produced by the administration of a WAIS-III test to a person between the ages of 45 to 54 years is 2.23 points.
37
The
SEM
for the full-scale IQ score produced by the administration of an SB5 assessment instrument to a person between the ages of 40 and 49 years is 2.12 points.
38
Taking those factors into account in connection with the IQ scores obtained from petitioner, this court can be confident that Thomas’s true “intellectual functioning ability” fell within a band of scores bordered on the high-end by
adding two SEMs
to the full-scale IQ score obtained by administration of each of the foregoing test instruments, and bordered on the low-end by
subtracting two SEMs
from his full-scale IQ score: in other words, and in the case of the WAIS-III, by
adding
4.46 points (2 x 2.23
SEMs =
4.46) to — and also
subtracting
4.46 points from — the full-scale IQ score produced by administration of that IQ assessment instrument. The same process applies to the administration of an SB5 assessment instrument, except that the number of points added to and subtracted from petitioner’s obtained full-scale IQ score would be 4.24 (2 x 2.12
*1271
SEMs
= 4.24 points). The 2002 AAMR manual explains this process as follows:
The assessment of intellectual functioning through the primary reliance on intelligence tests is fraught with the potential for misuse if consideration is not given to possible errors in measurement. An obtained standard IQ score must always be considered in terms of the accuracy of its measurement. Because all measurement, and particularly psychological measurement, has some potential for error, obtained scores may actually represent a range of several points.
This variation around a hypothetical
“true score”
may be hypothesized to be due to variations in test performance, examiner’s behavior, or other undetermined factors.
Variance in scores may or may not represent changes in the individual’s actual or true level of functioning. Errors of measurement as well as true changes in outcome must be considered in the interpretations of test results.
This process is facilitated by considering the concept of standard error of measurement (SEM), which has been estimated to be three to five points for well-standardized measures of general intellectual functioning.
This means that if an individual is retested with the same instrument, the second obtained score would be within one
SEM (i.e.,
± 3 to 4 IQ points) of the first estimates about two thirds of the time.
Thus an IQ standard score is best seen as bounded by a range that would be approximately three to four points above and below the obtained
score....
Therefore, an IQ of 70 is most accurately understood not as a precise score, but as a range of confidence
with parameters of at least one
SEM
(i.e., scores of about 66 to 74; 66% probability);
or parameters of two SEMs
(i.e., scores of 62 to 78; 95% probability).
This is a critical consideration that must be part of any decision concerning a diagnosis of mental retardation.
AAMR,
Mental Retardation
at 57-59 (all emphasis added, citation omitted).
b.
The stipulated
SEM
and its effect upon determination of petitioner’s IQ
The attorneys for both parties and their expert witnesses stipulated that a standard error of measurement in the neighborhood of approximately ± 5 points is proper for full-scale IQ test scores produced by the intelligence assessment instruments discussed in this opinion.
39
The American Psychiatric Association agrees: the most recent edition of its
Diagnostic and Statistical Manual of Mental Disorders
notes that “there is a measurement error of approximately 5 points in assessing IQ.” APA,
Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, Text Revision
41-42 (2000) (emphasis added) (hereafter “DSM-IV-TR”).
40
,
*1272
The parties disagree, however, as to how the “band of confidence” produced by the process of adding
and
subtracting that value from petitioner’s obtained, full-scale IQ scores should be interpreted. Petitioner argues, in combination with the “Flynn effect” discussed in the following section, that his IQ scores should be adjusted downward, because a “compelling showing of lifelong deficits in adaptive behavior ... yields scores that are perfectly consistent across [his] life.”
41
Petitioner points to several cases as support for this position, but none are particularly illustrative, much less precedential.
42
Respondent, on the other hand, argues that this court should refuse to adjust petitioner’s IQ scores downward using the standard error of measurement. Respondent also contends that he has been unable to find any authority to support a downward adjustment of petitioner’s full-scale IQ scores, “depending on ... assessment of Thomas’s adaptive functioning.”
43
Respondent asserts that, even though petitioner “argued during the evidentiary hearing [that] the range for mildly mentally retarded is from 70-75 ... under Alabama law, the cut-off for mild mental retardation is 70 or below.”
44
The resolution of these conflicting positions lies in a careful examination of the diagnostic criteria contained in the authoritative treatises published by the AAMR and APA,
45
and referenced in decisions of the Supreme Courts of the United States and the State of Alabama.
It is clear that neither of the professional organizations dedicated to the diagnosis and treatment of mental deficiencies advocates a fixed, finite IQ “cut score” as an impregnable barrier, separating persons who are mentally retarded from those who are not. The AAMR explicitly states that
“a fixed cutoff for diagnosing an individual as having mental retardation was not intended, and cannot be justified psycho-metrically.” Mental Retardation
at 58 (emphasis supplied). That manual also states — not just once, but at least eight
*1273
times — that the clinical standard for “significantly subaverage” intellectual functioning “is approximately two standard deviations below the mean,
considering the standard, error of measurement for the specific assessment instruments used
and the instruments’ strengths and limitations.”
Id.
at 13 (emphasis supplied);
see also id.
at 14, Table 1.2 (“Intelligence”) (same); 17
(same);
23, Table 2.1 (“IQ Cutoff’) (same); 27 (same); 37 (same); 58 (same); 198 (same).
In effect, this expands the operational definition of mental retardation to 75, and that score of 75 may still contain measurement error.
Any trained examiner is aware that all tests contain measurement error; many present scores as confidence bands rather than finite scores. Incorporating measurement error in the definition of mental retardation serves to remind test administrators (who should understand the concept) that an achieved Wechsler IQ score of 65 means that one can be about 95% confident that the true score is somewhere between 59 and 71.
Id.
at 59 (emphasis supplied).
In like manner, the American Psychiatric Association recognizes that measurement error must be taken into account when interpreting a full-scale IQ score obtained by assessment with any of the standardized, individually administered, intelligence assessment instruments discussed in this opinion.
Significantly subaverage intellectual functioning is defined as an IQ of about 70 or below (approximately two standard deviations below the mean). It should be noted that there is a measurement error of approximately 5 points in assessing IQ, although this may vary from instrument to instrument (e.g., a Wechsler IQ of 70 is considered to represent a range of 65-75). Thus, it is possible to diagnose Mental Retardation in individuals with IQs between 70 and 75 who exhibit significant deficits in adaptive behavior. Conversely, Mental Retardation would not be diagnosed in an individual with an IQ lower than 70 if there are no significant deficits or impairments in adaptive functioning....
APA, DSM-IV-TR at 41-42.
The Fourth Circuit endorsed these generally-accepted clinical standards when instructing a district court to consider whether a state statute defining mental retardation permitted measurement error to be taken into account when determining whether a capital murder habeas petitioner’s raw IQ score of 72 was “ ‘two standard deviations below the mean’ as set forth under that statute.”
Walker v. True,
399 F.3d 315, 323 (4th Cir.2005).
See also In re: Bowling,
422 F.3d 434, 442 (6th Cir.2005) (observing that “there appears to be considerable evidence that irrebuttable IQ ceilings are inconsistent with current generally-accepted clinical definitions of mental retardation and that any IQ thresholds that are used should take into account factors, such as a test’s margin of error, that impact the accuracy of a particular test score”) (Moore, J., concurring in part and dissenting in part) (footnote omitted).
In addition, many state courts have recognized that standard errors of measurement must be taken into account when interpreting IQ scores.
See, e.g., State v. Burke,
2005 Ohio 7020 , 2005 WL 3557641 , at *13 (10th Dist. Dec. 30, 2005) (“In accord with the AAMR’s standard, measurement error must be considered in determining an individual’s IQ score.”);
In re Hawthorne,
35 Cal.4th 40 , 24 Cal.Rptr.3d 189 , 105 P.3d 552, 557-58 (2005) (observing that IQ test scores are not precise due to measurement error, and holding that mental retardation should not be determined according to a fixed IQ cut score, but upon
*1274
an assessment of the defendant’s overall capacity based on a consideration of all relevant evidence);
State v. Williams,
831 So.2d 835 , 853 & n. 26 (La.2002) (observing that any IQ test must account for standard margins of error).
46
The California Supreme Court explained the point this way:
With respect to the intellectual prong of [California’s mental retardation statute], respondent Attorney General urges the court to adopt an IQ of 70 as the upper limit for making a prima facie showing. We decline to do so for several reasons: First, unlike some states,
the California Legislature has chosen not to include a numerical IQ score as part of the definition of mentally retarded.
[47]
...
Moreover, statutes referencing a numerical IQ generally provide that a defendant is presumptively mentally retarded at or below that level, rather than — as respondent impliedly argues— that a defendant is presumptively not mentally retarded above it. Second,
a fixed cutoff is inconsistent with established clinical definitions
and fails to recognize that significantly subaverage intellectual functioning may be established by means other than IQ testing. Experts also agree that an IQ score below 70 may be anomalous as to an individual’s intellectual functioning and not indicative of mental impairment. Finally,
IQ test scores are insufficiently precise to utilize a fixed cutoff in this context.
In re Hawthorne,
24 Cal.Rptr.3d 189 , 105 P.3d at 557 (emphasis added, citations and footnote omitted).
See also State v. Lott,
97 Ohio St.3d 303 , 779 N.E.2d 1011 , 1014 (2002) (“While IQ tests are one of the many factors that need to be considered, they alone are not sufficient to make a final determination on this issue. We hold that there is a rebuttable presumption that a defendant is not mentally retarded if his or her IQ is above 70.”).
One academician who surveyed state statutory developments following the
Atkins
decision has observed that
many states have incorporated a specific IQ cutoff score in their definitions of mental retardation, most often using an IQ of seventy as the cutoff for this component of the mental retardation definition. However, most of these definitions do not acknowledge that each assessment instrument has a standard measurement error, usually between three and five points, and that the standard measurement error is not the same for all instruments. Recognizing the impact of the standard measurement error, in the previous AAMR definitions and the current APA definition, the IQ cutoff for mental retardation has been quantified between seventy and seventy-five, as noted by the Court in
Atkins .
To avoid mistaken reliance on and potential misuse of a particular IQ score, especially if it does not include consideration of standard measurement error, the AAMR stated its current IQ cutoff in terms of being at least two standard deviations below the mean of the specific instruments used, considering their particular
*1275
standard measurement error, strengths, and limitations. The current APA definitional material also refers to the IQ cutoff as being approximately two standard deviations below the mean, with reference to measurement error of approximately five points.
Thus, any state’s use of a fixed IQ cutoff score, without reference to standard measurement error and other factors concerning the specific instrument used, risks an inaccurate assessment of the intellectual functioning component of the mental retardation definition.
Peggy M. Tobolowsky,
Atkins Aftermath: Identifying Mentally Retarded Offenders and Excluding Them From Execution,
30 J. Legis. 77 , 95-96 (2003) (emphasis added, footnote omitted);
see also id.
at 139 (“[SJtates that use a rigid IQ cutoff score of seventy for the intellectual functioning component may be excluding some individuals otherwise falling within the accepted clinical definition.”).
3.
The “Flynn effect,” IQ gains over time, and cut scores
The discussion in the preceding sections bears upon the following truism:
“An IQ score is only as valid as the test the person takes, and the test is only as valid as the standardization sample on which it is normed.”
James R. Flynn,
What is Intelligence?
Ill (Cambridge Univ. Press 2007) (emphasis supplied).
The “Flynn effect” is the name given in recognition of the central role played by Professor James R. Flynn in discovering and, in a series of fifteen or more publications between 1984 and today, documenting the fact that IQ scores have been increasing from one generation to the next in all fourteen nations for which IQ data is available.
48
Flynn recently explained the phenomenon this way:
For the Wechsler (WISC
[49]
and WAIS) and the Stanford-Binet IQ tests, the best rule of thumb is that Full Scale IQ gains have been proceeding at a rate of 0.30 points per year ever since 1947. This rate is based on comparisons of all of the Wechsler and Stanford-Binet tests used in recent years (see Box 11). It means that for every year that passes between when an IQ test was normed, that is, when its standardization sample was tested, obsolescence has inflated their IQs by 0.30 points. For example, if you took the WISC (normed in 1947-1948) in 1977-1978, you would get an unearned bonus of 9 IQ points (30 years x 0.30). Even though you might be dea,d average, you would be scored at 109 thanks to obsolete norms thirty years out of date.. After all, IQ gains over time mean that as we go back into the past, representative samples of Americans perform worse and worse. In this case, you are not being compared to your peers, the 14-year-olds of the late 1970s, but to a much lower-scoring
*1276
group, the 14-year-olds of the late 1940s. Your score of 109 against the old norms makes you appear above average, but you are actually no better than average and deserve an IQ of 100.
Flynn,
supra
at 112.
50
In other words, as an intelligence test ages — or moves farther from the date on which it was standardized (“normed”) — the mean score of the population as a whole on that assessment instrument increases, thereby artificially inflating the IQ scores of individual test subjects. Stated somewhat differently, IQ scores have been increasing over time for reasons that are totally unrelated to the actual, “true” intelligence of test subjects.
Even though the parties’ attorneys and expert witnesses uniformly agreed that the Flynn effect is an empirically proven statistical fact,
51
they disagreed on the extent to which an individual test subject’s IQ score should be adjusted to take that phenomenon into account. For example, petitioner’s psychological expert, Dr. Karen Salekin, testified that she always applies Dr. Flynn’s recommendation to reduce a full-scale IQ score by 0.30 points for each year elapsed beyond the date on which the test instrument was standardized (“normed”).
52
On the other hand, respondent’s expert, Dr. Harry McClaren, testified that the Flynn effect was something he would “take into consideration. But to slavishly say that this is definitely going to be right or a better estimate of this person’s true IQ, I don’t think that ... the profession [has come to a national consensus on] that point.”
53
Dr. McClaren’s opinion about the lack of professional acceptance of the validity of the Flynn effect is refuted by the AAMR’s 2002 manual, which explicitly states that, “as others have shown (e.g., Flynn, 1987), it is critically important to use standardized tests with the most updated norms.”
Mental Retardation
at 56 (emphasis supplied);
see also id.
at 59 (noting “variances in scores between successive revisions of intelligence measures”).
A manual published in 2007 under the present name of the organization formerly known as the “American Association on Mental Retardation” (AAMR) — i.e., the “American Association on Intellectual and
*1277
Developmental Disabilities” (AAIDD) — is even more explicit. It recommends that clinicians take into account
both
the Flynn effect
and
the standard error of measurement when performing retrospective diagnoses in less than optimal circumstances
(e.g.,
the legal and physical constraints of a maximum-security prison environment). Specifically, the
User’s Guide: Mental Retardation Definition, Classification and Systems of Supports
— 10th
Edition (“User’s Guide”)
directs diagnosticians to:
Recognize the “Flynn Effect.” In his study of IQ tests across populations, Flynn (1984, 1987, 1989) discovered that IQ scores have been increasing from one generation to the next in all 14 nations for which IQ data existed. This .increase in IQ scores over time has been dubbed the Flynn Effect. Flynn reported a greater increase in the Wechsler Performance IQ, which is more heavily loaded on fluid abilities than on the Wechsler Verbal IQs. On average, the Full-Scale IQ increases by approximately 0.33 points for every year elapsed since the test was normed (Flynn, 1999). The main recommendation resulting from this work is that all intellectual assessments must use a reliablé and appropriate individually administered intelligence test. In cases of tests with multiple versions, the most most recent version with the most current norms should be used at all times. In cases where a test with aging norms is used, a correction for the age of the norms is warranted. For example, if the Wechsler Adult Intelligence Scale (WAIS-III, 1997) was used to assess an individual’s IQ in July 2005, the population mean on the WAIS-III was set at 100 when it was originally normed.in 1995 (published in 1997). However, based on Flynn’s data, the population mean on the Full-Scale IQ raises roughly 0.33 points per year; thus, the population mean on the WAIS-III Full-Scale IQ corrected for the Flynn Effect would be 103 in 2005 (9 years X 0.33 = 2.9). ■ Hence, using the “at least two standard deviations below the mean” (Luckasson et al., 2002), the approximate Full-Scale IQ cutoff would be approximately 73
(plus or minus the standard error of measurement).
Thus the clinician needs to use the most current version of an individually administered test of intelligence and take into consideration the Flynn Effect
as well as the standard error of measurement
when estimating an individual’s true IQ score.
AAIDD,
User’s Guide
at 20-21 (emphasis supplied).
Respondent retorts that, with the
exception
of the AAMR/AAIDD, no other national organization or federal agency has officially endorsed the Flynn effect.
54
That may be so, but it does not justify ignoring th'e phenomenon in the face of its unchallenged existence.
55
The steady rise in IQ scores from year to year is a statisti
*1278
cally proven fact. Dr. McClaren admitted that he had no knowledge of any-study arguing that IQ scores in the general population of the United States have
not
increased at the average rate of 0.30 points each year after a test instrument was standardized.
56
Further, and even though “there is not a consensus among professionals as to why these gains -are occurring or what these gains actually mean (e.g., are we really getting smarter?),
all are in agreement that the gains
occur.”
57
“Since 1998, when the American Psychological Association issued
The Rising Curve: Long-Term Gains in IQ and Related Measures
(Neisser, 1998), no s'cholar published in a first-line journal has ignored the relevant data on IQ gains over time.”
58
It also is undisputed that Professor Flynn’s
recommendation
— i.e., “deduct 0.30 IQ points per year (3 points per decade) to cover the period between the year the test was normed and the year in which the subject took the test”
59
— is a generally accepted adjustment.
Outside the psychological and psychiatric communities, at least one Circuit Court of Appeals has held that the Flynn effect is relevant to the interpretation of IQ scores in capital cases. In
Walker v. True,
399 F.3d 315 (4th Cir.2005), the habeas petitioner, who was raising a mental retardation claim under
Atkins ,
argued that the district court committed reversible error by failing to -adjust his IQ scores to take into account both the Flynn effect and the test instrument’s standard error of measurement. The Fourth Circuit agreed, vacated the district court’s opinion dismissing the habeas petition, and remanded the case for consideration of “relevant evidence”:
.The district court, without much explanation, did not consider the Flynn Effect or the measurement error, stating that such evidence “does not provide a legal basis for ignoring Walker’s WAIS test scores.” J.A. 266. But, as the Virginia statute makes clear, the relevant question is whether Walker scored two standard deviations below the mean, a question which is directly addressed by Walker’s expert opinion as to the Flynn Effect.
Thus,
not only did
the district court
resolve a factual dispute against Walker — contrary to the claims in his petition and where the facts remained materially disputed — it also
refused to consider relevant evidence, namely the Flynn Effect evidence.
Therefore, on remand the district court should consider the persuasiveness of Walker’s Flynn Effect evidence. And if the district court does credit that evidence, it should then consider whether the Virginia statute permits consideration of measurement error in order to determine whether Walker’s purported score of 72 is “two standard deviations below the mean” as set forth under that statute.
*1279
Walker,
399 F.3d at 322-23 (emphasis added).
See also Walton v. Johnson,
440 F.3d 160 (4th Cir.2006)
(en
banc);
60
In re Hicks,
375 F.3d 1237, 1242 (11th Cir.2004) (Birch, J., dissenting from the denial of a stay of execution because the IQ scores generated by a 1985 administration of the Wechsler Adult Intelligence Scale to the habeas petitioner were “likely to have been artificially inflated by what has been labeled ‘The Flynn effect’ ”).
61
State appellate courts have reached the same conclusion. The Ohio Court of Appeals, for example, has held that “a trial court must consider evidence presented on the Flynn effect, but, consistent with its prerogative to determine the persuasiveness of the evidence, the trial court is not bound to, but may, conclude the Flynn effect is a factor in a defendant’s IQ score.”
State v. Burke,
2005 Ohio 7020 , 2005 WL 3557641 , at *13 (10th Dist. Dec. 30, 2005).
4.
Conclusions
Adherence to scientific principles is important for concrete reasons: they enable the reliable inference of knowledge from, uncertain information
....
62
George Bernard Shaw famously said that England and America are two countries separated by a common language. In an analogous manner, the discussion in the previous Parts of this opinion demonstrates that attorneys, judges, psychologists, and psychiatrists share a common language, but our specialized vocabularies often separate one professional group from the other. Legal and mental health professionals
differ in the acceptance of methods of analysis. In science in general and in statistics in particular, a method is evaluated according to well-established criteria. In the courts, theoretical justifications of a statistical method may be treated as if they are less important
*1280
than the general acceptance of the method by statisticians and other scientists.
The Evolving Role of Statistical
Assess
ment as Evidence in the Courts
145 (Stephen E. Fienberg ed., 1989). For such reasons, when assessing the role that full-scale IQ scores play in determining mental retardation, courts must be careful to distinguish between the language
{rules)
of law and the language
{diagnostic
criteria) of psychologists and psychiatrists. Stated differently, it is important for courts to guard against resolving the
factual question of
mental retardation as a
matter of law.
In finding the facts of a particular case, courts and juries untrained in science are sometimes called upon to resolve contested scientific issues, but such factual findings do not establish generally applicable rules of law.... [A]n appellate court cannot convert a disputed factual assertion into a rule of law simply by labeling it a “legal standard” ...
People v. Superior Court,
40 Cal.4th 999 , 56 Cal.Rptr.3d 851 , 155 P.3d 259, at 267 (2007).
Contrary to respondent’s argument that there is no diagnostic or legal basis by which this court may properly adjust petitioner’s raw IQ scores in answering the question of whether he suffers from significantly subaverage intellectual functioning,
63
the adjustments to raw IQ scores mandated by the “standard error of measurement” and the “Flynn effect” are well-supported by the accumulation of empirical data over many years. Both methodologies have been subjected to rigorous peer review and, while some psychologists may still ponder the precise cause(s) of the Flynn effect, no reputable member of the relevant professional communities denies that IQ scores have been increasing at the average rate of 0.30 points a year since the 1930s. General acceptance of both methodologies has come “as results and theories continue to hold, even under the
*1281
scrutiny of peers, in an environment that encourages healthy skepticism.”
64
Therefore, this court has taken both factors into account when evaluating the extent of petitioner’s intellectual functioning abilities. Stated differently, even though the legal cut-off score for a finding of “significantly subaverage intellectual functioning” is stated in opinions of the Alabama Supreme Court as “an IQ of 70 or below,” a court should not look at a raw IQ score as a
precise
measurement of intellectual functioning. A court must also consider the Flynn effect and the standard error of measurement in determining whether a petitioner’s IQ score falls within a
range
containing scores that are less than 70.
B. Assessment of Adaptive Behavior
Under
Atkins
and its progeny, a finding of “mental retardation” can only be made upon the basis of a conclusion that the petitioner’s significantly sub-average intellectual functioning is accompanied by substantial limitations in adaptive behavior.
See Atkins,
536 U.S. at 309 n. 3, 122 S.Ct. 2242 ;
Ex parte Perkins,
851 So.2d at 456 .
The term “adaptive behavior” is defined by the AAMR as “the collection of
conceptual, social,
and
practical skills
that have been learned by people in order to function in their everyday lives.”
Mental Retardation
at 41 (emphasis supplied). “Conceptual skills” include language, reading and writing,
money
concepts, and self-direction; in other words, a determination whether the test subject possesses a basic level of literacy and numeracy (so he can shop and make change), and remembers to do things on time. “Social skills” include interpersonal relationships, personal responsibility, self-esteem, gullibility and naiveté, following rules, obeying laws, and avoiding victimization. “Practical skills” include daily activities such as eating, personal hygiene, dressing, meal preparation, housekeeping, transportation, taking medication, money management, and telephone use, as well as occupational skills and maintaining a safe environment.
65
Dr. Salekin explained the concept as follows:
[Ajdaptive behaviors are everyday skills, such as walking, talking, grooming, cooking, cleaning, and participating in school or work. These abilities are learned over time in the context of one’s home and community, and they repre
*1282
sent skills that are necessary to function within that context. It is important to emphasize that adaptive behaviors develop over the course of time and with experience, and thus individuals are evaluated against their same age peers.
To measure adaptive skills, adaptive behavior scales have been developed and normed on individuals with and without intellectual disabilities. These scales require that an informant, typically a parent, teacher, or other individual who is very familiar with the individual’s daily level of functioning, rate the person of interest on a variety of skills.
For instance, the informant may rate the extent to which the individual follows directions or balances a checkbook along a continuum ranging from “Never Does or Can’t Do” to “Always Does or Can Do Without Assistance.”
There exist many misconceptions about how to conceptualize adaptive behavior for diagnostic purposes. First, some believe that adaptive behavior is measured by estimates of abilities or potential, but it is the individual’s actual performance that is important. Second, adaptive behavior is typical behavior that reflects an individual’s ability to function on a day-to-day basis. It is not measured by isolated successes or failures. Third, adaptive behavior is performance in one’s community, not in restricted settings, such as prison or therapeutic treatment programs. Fourth, the definition of mental retardation does not require that a cause of impairment be identified; diagnosis is determined by the presence of significant deficits in intellectual functioning and concomitant deficits in adaptive behavior that are evident before the age of 18 years.
Doc. no. 110-2 (Report of Karen Salekin, Ph.D.), at 15,16 (emphasis supplied).
1.
Standardized assessment instruments
In the Tenth (2002) edition of the AAMR’s manual, diagnosticians are, for the first time, instructed that
significant limitations in adaptive behavior should be established through the use of standardized measures normed on the general population, including people with disabilities and people without disabilities. On these standardized measures, significant limitations in adaptive behavior are operationally defined as performance that is at least two standard deviations below the mean of either (a) one of the following three types of adaptive behavior: conceptual, social, or practical or (b) an overall score on a standardized measure of conceptual, social, and practical skills.
Mental Retardation
at 76.
66
There are many scales for measuring adaptive behavior, but it is important to
*1283
recognize that “no one instrument can measure all of the relevant domains of adaptive behavior.”
Id.
at 84 (citation and internal quotation marks omitted).
See also, e.g.,
doc. no. 110-2 (Report of Karen Salekin, Ph.D.) at 16 (“There exist numerous scales of adaptive behavior that can be used for the purposes of diagnosis, classification, and planning for supports; [but]
no single measure is best for all three.”)
(alteration and emphasis added).
For that reason, review of school records, medical histories, public records, employment records, and personal observations of the test subject in face-to-face interviews, as well as consideration of other, collateral sources of information — such as interviews of “third-party informants”
(e.g.,
family members, friends, school teachers, and other persons who know the subject
well)
— are all used by clinicians to complement, but not replace, standardized assessment measures.
See
AAMR,
Mental Retardation at
84.
Those who use most current adaptive behavior scales to gather information about typical behavior rely primarily on the recording of information obtained from a third person who is familiar with the individual being assessed. Thus assessment typically takes the form of an interview process, with the respondent being a parent, teacher, or direct service provider [e.g., prison guards] rather than from direct observation of adaptive behavior or from self-report of typical behavior. It is critical that the interviewer and informant or rater fully understand the meaning of each question and response category in order to provide valid and reliable information to the clinician. It is also essential that people interviewed about someone’s adaptive behavior be well-acquainted with the typical behavior of the person over an extended period of time, preferably in multiple settings. In some cases it may be necessary to obtain information from more than one informant. The consequences of scores to the rater, informant, or individual being rated should also be taken into consideration, as well as the positive or negative nature of the relationship between the rater or informant and the person being assessed. Observations made outside the context of community environments typical of the individual’s age peers and culture warrant severely reduced weight.
Id.
at 85 (citations omitted).
2.
The importance of clinical judgment
As a result of the fact that no currently available assessment instrument can measure all domains of adaptive behavior, the importance of “clinical judgment” increases.
See id.
at 84. Clinical judgment is particularly important in cases such as this one, where the passage of many years limits the use and valid interpretation of standardized test instruments administered to third-party informants.
Id.
at 94 (“Clinical judgment is often required when ... difficulties arise in selecting informants and validating informant observations [or] ... direct observation of the individual’s actual performance has been
*1284
limited and additional direct observation is necessary.”).
67
Respondent’s expert, Dr. Harry McClaren, acknowledged the importance of clinical judgment. When he was asked if the fact that no currently-available assessment instrument completely measures all domains of adaptive behavior means that the question of whether a person is significantly limited in adaptive behavioral skills “essentially comes down to [a clinician’s] professional judgment,” Dr. McClaren answered: “absolutely correct,” “[t]hat’s right.”
68
However, Dr. McClaren’s unequivocal embrace of the importance of clinical judgment in the assessment of a test subject’s adaptive skills only begs a question:
What do psychologists mean when they use the term “clinical judgment”?
The AAIDD’s
User’s Guide
answers that query as follows:
Clinical Judgment is a special type of judgment that emerges directly from extensive data and is rooted in a high level of clinical expertise and experience. Its three characteristics are that it is (a) systematic (i.e., organized, sequential, logical), (b) formal (i.e., explicit and reasoned), and (c) transparent (i.e., apparent and communicated clearly). The result of competent clinical judgment is that its use enhances the precision, accuracy, and integrity of the clinician’s decisions and recommendations. It is important to point out that clinical judgment is not (a) a justification for abbreviated evaluations, (b) a vehicle for stereotypes or prejudices, (c) a substitute for insufficiently explained questions, (d) an excuse for incomplete or missing data, or (e) a way to solve political problems.
AAIDD,
User’s Guide
at 23.
69
Moreover, “it is crucial that clinicians conduct a thorough social history and align data and data collection to the critical question(s) at hand.”
Id.
3.
The assessment instrument used in this case
Petitioner’s expert, Dr. Karen Salekin, administered the Scales of Independent Behavior — Revised (SIB-R) test instrument to third-party informants for the purpose of probing their memories of petitioner’s adaptive skills prior to age eighteen. She also administered the SIB-R to prison guards for the purpose of exploring their observations of petitioner’s behavior, both currently and during the entire time
*1285
he has been incarcerated.
70
According to respondent’s expert, the SIB-R test instrument is designed for administration to third-party informants, to assess their knowledge of the subject’s
current
adaptive
skills
— not twenty- to thirty-year-old memories of the subject’s past behavior.
71
Respondent’s expert also asserts that the SIB-R is not normed for inmates who have been incarcerated in a highly-restrictive, death-row environment for long periods of time.
72
Dr. Salekin candidly confessed that the SIB-R was not created to be utilized in the manner she employed it. Nevertheless, she chose to make use of it in an effort to discern some sense of the petitioner’s adaptive skills, and also because the 2002 AAMR definition requires administration of a standardized assessment instrument for diagnostic purposes.
73
In short, any conclusions regarding the limitations of petitioner’s adaptive-behavioral skills prior to the age of eighteen, during adulthood, or as a result of his twenty-year stint on death row ultimately hinge upon the strength and credibility of each expert’s clinical judgment, as applied to the other sources of information she or he explored. As has been observed, how
*1286
ever, clinical judgments can “represent both the best and worst of assessment data. Judgments made by conscientious, capable, and objective individuals can be an invaluable aid in the assessment process. Inaccurate, biased, subjective judgment can be misleading at best and harmful at worst.”
74
III. EXPERT WITNESSES AND “BEST PRACTICES”
A. Petitioner’s Experts
1.
Dr. Karen Salekin
Dr. Salekin is a clinical and forensic psychologist. She earned a Bachelor of Science degree with double majors in Psychology and Kinesiology
75
from Simon Fraser University in Burnaby, British Columbia, during 1991, and a Master of Science degree and a Doctorate in Psychology from the University of North Texas in 1994 and 1997, respectively. Between 1995 and 1998, Dr. Salekin completed pre- and post-doctoral clinical training at a Federal penal facility and two University medical centers. She has maintained a private practice in clinical and forensic psychology in Tuscaloosa, Alabama, since 2001, but in 2003 she also was appointed an Assistant Professor of Psychology at the University of Alabama. She has various publications in print.
76
Dr. Salekin estimated that she has performed hundreds of forensic assessments throughout her career, and she has conducted eleven
Atkins
evaluations in Alabama death penalty cases — a number that includes petitioner, Kenneth Thomas.
77
All of those evaluations were conducted at the request of defense counsel.
78
Even so, Dr. Salekin has concluded that only three of those
*1287
eleven subjects were mentally retarded.
79
That fact indicates that Dr. Salekin is not a trained parrot, and enhances her credibility.
When conducting a forensic evaluation, Dr. Salekin follows recommendations contained in a treatise discussing the legal and clinical contexts of forensic assessments, including “best practice” guidelines for participating effectively and ethically in civil and criminal proceedings:
ie.,
Gary B. Melton, John Petrila, Norman G. Poythress
&
Christopher Slobogin,
Psychological Evaluations for the Courts: A Handbook for Mental Health Professionals and Lawyers
(New York: The Guilford Press 3d ed.2007).
80
One section of that treatise discusses the information that should be addressed in a written report of evaluation. For example, diagnosticians are encouraged to document, among other things: the reasons for conducting the mental assessment; the names of all contacts; the date, time, place, and length of interviews; what each informant said; and, the collateral sources and records that were reviewed.
81
2.
Dr. Daniel Marson
Dr. Marson has been a professor of neurology at the University of Alabama in Birmingham since 1990. He earned a Liberal Arts degree from Carlton College in Northfield, Minnesota during 1976. From 1977-78, he studied law at the University of Edinburgh in Scotland, and he earned a Juris Doctor degree from the University of Chicago School of Law in 1981. Dr. Mar-son pursued further graduate training in clinical psychology at Northwestern University Medical School in Chicago, and he obtained a Ph.D. in Clinical Psychology (with a specialization in geriatric issues) from that school in 1990.
82
As the term “neuropsychology” suggests, it is a field of study that combines neurology and psychology. It is concerned primarily with clinical and scientific aspects of the relationship between brain structure and human behavior. Neuropsychologists search for possible organic causes of mental disorders, including retardation.
83
Dr. Marson explained his discipline this way:
Well, to begin with, a good neuropsychologist is, first and foremost, a clinical psychologist. So I received that training at Northwestern.
While I was there I took classes in brain science within the medical school. And then as part of my Ph.D., I took a specialty internship in neuropsychology at Westside V.A. [Hospital] in Chicago.
*1288
Neuropsychology is a subspecialty of clinical psychology [that] focuses on disorders of higher cortical functioning or disorders of brain function. And so we’re psychologists who focus on brain behavior relationships.... [W]e use clinical interviews and specialized psychological tests to evaluate for brain dysfunction across a range of disorders.
84
With specific reference to the diagnosis of mental retardation, Dr. Marson said that neuropsychology allows a clinician to
conduct an evaluation [focusing upon] a wide range of specific or discreet cognitive functions, like memory, attention, perceptual abilities, problem solving.
And that can help the parties in a case, the Court understand how these specific deficits translate into specific kinds of problems in adaptive functioning .... A neuropsychologist can look at sort of the component cognitive abilities and relate them to specific deficits in everyday functioning.
85
Marson testified that, since 2000, he had “been involved in about ten forensic cases in which mental retardation was potentially an issue,” and all but one involved the death penalty.
86
Marson found mental retardation in the non-death-penalty case, and in only two of the nine death penalty cases.
87
Again, that statistic enhances Dr. Marson’s credibility.
B. Respondent’s Expert
Dr. Harry A. McClaren is a licensed clinical psychologist specializing in criminal forensic psychology. He earned a Bachelor of Science degree in Psychology from the University of Virginia in 1973, a Master’s degree in Clinical Psychology from Mississippi State University in 1974, and a Ph.D. in Clinical Psychology from Virginia Polytechnic Institute in 1981. He is licensed to practice psychology in Alabama and Florida, and he is a member of the American Psychological Association, as well as the American College of Forensic Examiners.
88
Dr. McClaren has maintained a private practice in Clinical and Forensic Psychology since 1996. He also has served as a staff psychologist at three institutions, including as Chief of Psychology at Alabama’s Taylor Hardin Secure Medical Facility in Tuscaloosa from November 1983 to May of 1985.
89
In that role, McClaren “worked with Norm Poythress to develop ... forensic examiner training for the State of Alabama.”
90
When asked by the court if he considered the
Handbook for Mental Health Professionals and Lawyers
referenced by Dr. Salekin to be an authoritative and learned treatise,
91
Dr. McClaren responded that he owned a copy of the book, and that he was “[a]wful proud of it,” because one of the four co-authors, Norman G. Poythress, “is a very fine psychologist. Knows a lot about the psychology and law.”
92
*1289
It is a book that talks about different kinds of forensic psychology, about applying psychology to the nexus of mental health and the law.
And it gives suggestions by people that are well known about how to go about doing evaluations, writing reports.
And it’s the kind of thing that is not like a guideline, but it’s a way to help people think about what they’re doing.
93
Dr. McClaren has been trained to conduct clinical forensic interviews and intelligence testing.
94
He has testified “[h]undreds of times” at the request of trial judges, defense attorneys, district attorneys, and private counsel.
95
He estimated that ten to fifteen of his cases have involved
Atkins
evaluations, and all were done “at the request of some division of the State of Florida or the State of Alabama.”
96
In those few instances in which Dr. McClaren found people to be retarded, the find had been “stipulated to” by the State.
97
C. “Best Practices”
Atkins v. Virginia
has required mental health professionals to hone diagnostic techniques originally designed for salutary
purposes
— e.g., supporting and promoting the education and welfare of persons afflicted with mental retardation — to the forensic evaluation of persons charged with or convicted of the most heinous offenses known to the law. Forensic assessments in such cases are more difficult than the ordinary case, at least in part because clinicians are asked to perform a retrospective diagnosis under less than optimal circumstances.
The AAIDD’s 2007
User’s Guide
was “developed to assist ... in understanding the 2002 [AAMR diagnostic] System fully and applying best practices based on that understanding.”
User’s Guide
at 2. “Best practices in [diagnosing] mental retardation are based on professional ethics, professional standards, research-based knowledge, and clinical judgment.”
Id.
at 1. The pertinent portions of the guidelines for clinicians performing “retrospective diagnoses” read as follows:
Considerations. A retrospective diagnosis ... may be required when clinicians are involved in determining ... sentencing eligibility questions such as those related to the recent Atkins (2002) case. As with all assessments for diagnosis, such situations require clinicians to act consistent with best practices— that is, to act consistent with professional ethics, professional standards, research-based knowledge, and clinical judgment.
[I]n reference to people in the criminal justice system, some criminal defendants fall at the upper end of the MR/ID [Mental Retardation / Intellectual Disability] severity continuum (i.e., people with mental retardation who have a higher IQ) and frequently present with a mixed competence profile. They typically have a history of academic failure and marginal social and vocational skills. Their previous and current situations frequently allowed formal assessment to be avoided or led to assessment that was less than optimal:
Guidelines. The following guidelines for clinicians are important in retrospective diagnoses....
*1290
1. Conduct a thorough social history that includes: (a) the investigation and organization of all relevant information about the person’s life including status, trajectory, development, functioning, relationships, and family; and (b) the exploration of possible reasons for absence of data or differences in data including poorly trained examiners, selection of inappropriate assessment instruments, improper interpretation of test scores, or lack of sensitivity or awareness of the impact of changing norms and practice effects (see Guidelines 4 & 5, below).
2. Conduct a thorough review of school records. Ideally, school records are available across the elementary-, middle-, and high-school years. Locate all those that are available and arrange them chronologically, numbering the pages. One also needs to make a summary table that lists placement on a year-by-year basis including any changes in the school or school system as well as special education or alternative placements. Citing page numbers from the school records is one way to provide documentation for each piece of relevant information identified. The thorough review should include:
(a) mapping out the grades earned across the school years, looking for consistency of low grades in the core academic areas such as reading, math, English, and science (only in the upper grades)
(b) indicating any grade levels failed or repeated
(c) summarizing teacher, social and behavior ratings
(d) identifying relevant teacher comments to student or parents and requests for, as well as actual, parent-teacher conferences
(e) identifying when periodic achievement assessment happened (instruments and results) and, if necessary, learn more about the psychometrics of instrument used and how to interpret the scores
(f) identifying results of hearing and vision and any other school-wide screening
(g) searching for failure or patterns that normally would trigger parent-teacher conferences, prereferral meetings, or referrals for special education consideration
(h) identifying the outcome of any eligibility assessment(s) and whether an individualized education plan was developed; if special education was provided, note the diagnosis (typically a developmental disability label is used until age 8), the years given, the type of placement (resource room, self-contained, separate school), and other supports
(i) noting any services that might be viewed as substitutes to special education, which could indicate difficulties in cognitive adaptive behavior (e.g., remedial reading, Chapter I services)
(j) looking for other evidence of difficulties in cognitive adaptive skills besides grades and test performance (e.g., student often late to class or confused about schedule, difficulty following classroom directions, poor record of handling homework, failing driver’s education, etc.)
(k) looking for difficulties in practical adaptive skills (e.g., poor grooming, unable to use money correctly, getting lost in school or on school grounds, unable to tell time, etc.)
© looking for evidence of difficulties in social adaptive skills (e.g. follows others, lack of self-direction, few friends,
*1291
gullible, does not understand social humor, etc.)
In addition to the above, having contact with the teacher is also valuable, if possible, as a means of clarifying questions in the records and getting specific comments on the student or explanations about gaps in the records. Getting peer comments from school years is another valuable source of anecdotal information.
3. In reference to the assessment of adaptive behavior: (a) use multiple informants and multiple contexts; (b) recognize that limitations in present functioning must be considered within the context of community environments typical of the individual’s peers and culture; (c) be aware that many important social behavioral skills, such as gullibility and naivete, are not measured on current adaptive behavior scales; (d) use an adaptive behavior scale that assesses behaviors that are currently viewed as developmentally and socially relevant; (e) understand that adaptive behavior and problem behavior are different constructs and not opposite poles of a continuum; and (f) realize that adaptive behavior refers to typical and actual functioning and not to capacity or maximum functioning.
4. Recognize the “Flynn Effect.”
[The text of this paragraph was quoted previously,
in Part 11(A)(3)
supra.]
5. Recognize the impact of practice effect. Practice effect refers to gains in IQ scores on tests of intelligence that result from a person being retested on the same test.... For example, ... [t]he WAIS-III manual reports an average increase of 5 points on the Full-Scale IQ between administrations with intervals of 2 to 12 weeks. Thus clinicians need to be sensitive to these practice effects and best practices in intellectual assessment recommendations against administering the same intelligence test to someone within the same year. Practice effects can apply as well to normal achievement tests and state tests to measure school and district performance.
6. Recognize that self-ratings have a high risk of error in determining “significant limitations in adaptive behavior.” However, consistent with the need for multiple informants or respondents, self-ratings can be used under the following cautions: (a) people with MR/ID are more likely to attempt to look more competent and “normal” than they actually are — which is sometimes interpreted as “faking”; (b) people with MR/ID typically have a strong acquiescence bias or inclination to say yes or agree with authority figures; and (c) MR/ID is a social status that is closely tied to how a person is perceived by peers, family members, and others in the community.
7. Conduct a longitudinal approach of adaptive behavior that involves multiple raters, very specific observations across community environments (especially in regard to social competence), school records, and ratings by peers in the development process. This longitudinal evaluation needs to be sensitive to the subtle issue of “the stigma of the label” and the concern that families and schools have about the label of “mental retardation” and over-representation of specific racial or ethnic groups within particular communities and/or schools.
8. Do not use past criminal behavior or verbal behavior to infer level of adaptive behavior or about having MR/ ID. Greenspan and Switzky (in press) discuss two reasons for this guideline. First, there is not enough available information; second, there is a lack of normative information.
*1292
AAIDD,
User’s Guide
at 17-20 (citations omitted, bracketed alterations added). None of the experts who testified in this case explicitly referenced the foregoing portions of the
User’s Guide,
but the reports tendered by petitioner’s experts, Dr. Karen Salekin and Dr. Daniel Marson, substantially adhere to the guidelines for conducting retrospective diagnoses.
In contrast, the report submitted by respondent’s expert, Dr. McClaren, is most notable for what it does
not
contain.
98
For example, Dr. McClaren did not evaluate petitioner’s adaptive behavior before the age of 18,
99
and he did not summarize interviews conducted with third-party informants for the purpose of obtaining some insight into petitioner’s adaptive skills during the developmental period.
100
Indeed, Dr. McClaren did not even identify the individuals to whom he spoke, and he devoted no attention to a discussion of diagnostic criteria.
101
Dr. McClaren acknowledged the absence of this information, but asserted that he did not include it in his reports “as a matter of practice.” He declared that, in his opinion, it was important to state his conclusions regarding petitioner’s pre-ageeighteen adaptive functioning only “if that is what the bottom line turned upon.”
102
That opinion is contradicted by the APA’s
Diagnostic and Statistical Manual,
which notes that, when a ± 5 point standard error of measurement is taken into account, “it is possible to diagnose Mental Retardation in individuals with IQs
between 70 and 75 who exhibit significant deficits in adaptive behavior.
Conversely, Mental Retardation would not be diagnosed in an individual with an IQ
lower than 70 if there are no significant deficits or impairments in adaptive functioning.”
DSM-IV-TR at 41-42 (emphasis supplied).
Dr. McClaren attempted to defend the paucity of substantive information included in his abbreviated report by answering ‘Tes” to the following, leading question posed by respondent’s counsel: “[T]hat report ... doesn’t include all of the work that you did in order to make your evaluation of Mr. Thomas; it’s just a summary, isn’t that correct?”
103
When asked by petitioner’s counsel if it was his “custom and practice ... to include all ... important conclusions and the supporting information in your report?,” Dr. McClaren answered: “I think it’s important to explain what the conclusion is based on. And I guess you could have a debate about broad-brush reports that may be three to six pages, seven pages long or fine-brush reports that may be 17, 20 pages long. And most of my reports are in the three-to six-page range.”
104
Regardless of Dr. McClaren’s rationalizations, this court finds that his approach to forensic report writing leaves a great deal to be desired, especially in eases such as this one, where important societal and
*1293
legal policies collide. Dr. McClaren’s report stands in stark contrast to the careful analysis and well-documented statements contained in the reports of petitioner’s witnesses. For such reasons, it is less persuasive.
Even so, Dr. McClaren reviewed the same historical records as Drs. Salekin and Marson; he conducted interviews (albeit, it is not clear how many, with whom, the length of each, and the questions that were asked); and he administered testing instruments. Thus, while the deficiencies of Dr. McClaren’s written report are painfully obvious, this court still must examine and consider his testimony at the evidentiary hearing to determine whether the paucity of substantive information contained in his written report simply reflects poor reporting practices, or whether it is indicative of an abbreviated, cursory evaluation.
IV. ASSESSMENTS OF PETITIONER’S INTELLECTUAL FUNCTIONING
As discussed in Part I
supra,
this court must determine whether petitioner’s intellectual functioning ability was significantly subaverage during three periods of his life: (1) before he reached the age of eighteen; (2) on the date the offense of conviction was committed; and (3) currently.
See Smith v. State,
No. 1060427, 2007 WL 1519869, at *8 (Ala. May 25, 2007).
A. Intelligence Assessments Conducted Prior to Age Eighteen
Intelligence tests were administered to petitioner at least seven times before his conviction for the capital murder that eventually brought him into this court. Four of those were conducted before March 7, 1977, the date upon which he attained the age of eighteen years.
1.
October Ip, 1968
— age
nine years and seven months
The first evaluation occurred on October 4, 1968, when Thomas was nine years and seven months of age. A psychologist on the staff of the North Central Alabama Mental Health Center in Decatur, Alabama, David Loiry, Ph. D., administered a short form version of the Wechsler Intelligence Scales for Children (WISC). The WISC was normed in 1949, and consisted of a series of twelve tests for minors between the ages of five and sixteen years. Verbal, performance, and full-scale intelligence quotients could be derived from computation of the subtest scales. According to Dr. Loiry’s report, Thomas’s estimated full-scale IQ score of 56 placed him
near the middle of the range of moderate mental deficiency according to the classification system used by the American Psychiatric Association. There is little variability in the sub-tests used to estimate his IQ and it is felt that this represents a valid estimate of his present intellectual functioning.
Kenneth was recommended for placement in [a] Special Class for the Educable Mentally Retarded.
105
Neither the standard error of measurement nor the Flynn effect can be taken into account, because the former is not known for this particular test, and no empirical studies of the latter phenomenon have been conducted using the short form of the WISC test instrument.
106
Even with
*1294
out such adjustments, however, Thomas’s full-scale IQ score of 56 clearly indicated significant limitations in his intellectual functioning at a young age.
2.
October 12, 1972
— age
thirteen years and seven months
The second evaluation occurred on October 12, 1972, when Thomas was thirteen years and seven months of age.
107
He was tested by the Athens, Alabama, School System with the California Test of Mental Maturity — Short Form (CTMM-SF). The scales of that test measured various aspects of intellectual functioning, such as memory and logical reasoning, at five different levels of difficulty. The test was administered to groups of students, as opposed to individual test subjects.
108
Thomas received a composite IQ score of 68. Again, even without taking into account either the standard error of measurement or the Flynn effect, Thomas’s composite IQ score of 68 indicated significantly sub-average intellectual functioning abilities.
109
3.
June 23, 1973
— age
fourteen years and two months
The third intelligence assessment occurred on June 23, 1973, when Thomas was fourteen years and two months of age. He was re-evaluated by David Loiry, Ph. D., at the request of Dr. Frank M. Cauthen, M.D. Dr. Loiry administered four tests: the full Wechsler Intelligence Scales for Children (WISC);
110
the Wide Range Achievement Test (WRAT);
111
the Bender Visual-Motor Gestalt Test (BVMGT);
112
a
*1295
Figure Drawing test;
113
and a Rorschach test.
114
Thomas’s performance on the WISC test instrument yielded a verbal IQ score of 69, a performance IQ score of 67, and a full-scale IQ score of 64 — clearly within the range of mental retardation. The WRAT was administered for the purpose of assessing Thomas’s level of academic achievement, and it revealed that he functioned at only a 3.7 grade level in spelling, a 3.9 grade level in reading, and a 2.9 grade level in arithmetic. Based upon these results, Dr. Loiry concluded:
“There appears to be little doubt about his mental
retardation.”
115
Without question, a full-scale IQ score of 64 on the WISC meets the intellectual functioning prong of a diagnosis of mental retardation. That score points even more strongly in favor of Dr. Loiry’s conclusion of mental retardation when the Flynn effect is taken into account. The WISC was standardized (“normed”) in 1949, but administered to Thomas 24 years later. Multiplying 24 by the Flynn factor of minus 0.30 points for each year elapsed after the date of standardization yields a correction of 7.2 points that must be deducted from Thomas’s full-scale IQ performance score, resulting in an adjusted IQ score of 57
(i.e.,
64-7.2 = 56.8, rounded-up to 57).
Not only is a score of 57 well below the cutoff for a diagnosis of mental retardation, it still is subject to adjustment by a ± 5 point standard error of measurement, resulting in a conclusion that can be stated with a 95% degree of confidence,
116
that Thomas’s “true” full-scale IQ score then lay within a band ranging from 52 on the low end to 62 on the high end.
Even if the Flynn effect is not taken into account, and the full-scale IQ performance score of 64 is adjusted only by the ± 5 point standard error of measurement, the high end of the resulting band (59 to 69) still is below the cutoff for mental retardation specified by the Alabama Supreme Court in
Ex parte Perkins,
851 So.2d at 456 .
*1296
In short, regardless of how the full-scale IQ performance score produced by this test administration is sliced or diced, the numbers
always
lead to only one conclusion: Thomas’s intellectual functioning was significantly below average at the age of fourteen years.
4.
March 18, 1975
— age
sixteen
The final intelligence assessment conducted prior to Thomas’s eighteenth birthday occurred on March 18, 1975, eleven days after his sixteenth birthday. He was tested on the Wechsler Adult Intelligence Scales (WAIS) by Jim Lenz, a psychometrist for the Limestone County School System,
117
for the purpose of determining whether Thomas should remain in the “special education” curriculum.
118
Thomas marked the highest, raw performance scores to that point in his life: verbal IQ 75, performance IQ 75, and full-scale IQ 74.
119
Lenz concluded, nevertheless, that Thomas was “functioning in the educable range of mental retardation,” and recommended that he remain in special education.
120
Thomas’s raw, full-scale IQ score of 74 does not bar the conclusion that he suffered from significantly subaverage intellectual functioning prior to the age of eighteen.
121
Adjustment for the Flynn effect alone reveals a diagnostically sound IQ score of less than 70. The Wechsler Adult Intelligence Scales were standardized (“normed”) in 1955, and Jim Lenz administered the assessment instrument to petitioner 20 years later. Multiplying 20 by the Flynn standard of minus 0.30 points for each year elapsed after the test was normed yields a correction of 6 points that must be deducted from Thomas’s raw performance score, resulting in an adjusted full-scale IQ of 68. Not only is that score below the cutoff for a diagnosis of mental retardation, but the score still is subject to adjustment by ± 5 points for measurement error, resulting in Thomas’s hypothetical, “true” IQ score lying within a 63 to 73 band of confidence.
122
Accordingly, this court concludes that petitioner did not present with a true IQ score that was greater than 70 based upon the results of the intelligence examination administered to him on March 18,1975.
5.
Findings
Based on the foregoing evidence, this court finds that petitioner has demonstrated, by a preponderance of the evidence, that he suffered from significantly subaverage intellectual functioning prior to the age of eighteen years:
i.e.,
during the so-called “developmental period.”
*1297
B. The Intelligence Assessment Performed on April 11,1977
The first assessment of Thomas’s intellectual functioning following the end of the developmental period was performed on April 11, 1977, just one month after his eighteenth birthday. A Wechsler Adult Intelligence Scales (WAIS) test instrument was administered by Joyce Raley, a counselor at the West Limestone School. The record is not clear on the questions of why Ms. Raley, as opposed to Jim Lenz, administered the test, or the purpose for which it was given.
123
In any event, Thomas’s raw test scores on that occasion were the highest he has ever recorded: verbal IQ 78, performance IQ 79, and full-scale IQ 77.
124
As noted in the previous section, however, the WAIS assessment instrument was standardized (“normed”) in 1955, and Ms. Raley’s administration of the test occurred 22 years later. Multiplying 22 by the Flynn corrective factor of minus 0.80 points for each year elapsed after the test was normed yields a correction of 6.6 points that must be deducted from Thomas’s raw score, resulting in an adjusted full-scale IQ of 70.4, which rounds down to 70.
125
Even that is not an absolute score, because a further correction of ± 5 points for the standard error of measurement must then be applied to the adjusted IQ score. In doing so, this court can be 95% confident
126
that Thomas’s “true” IQ on the date of Ms. Raley’s assessment fell within a band extending from 65 (65.4) on the low end to 75 (75.4) on the upper extreme.
This court cannot conclude its discussion of these test results, however, without addressing Dr. McClaren’s belief that Thomas’s raw, full-scale IQ score of 77 on the WAIS administered by Joyce Raley, a counselor, as opposed to the School System’s psychometrist, not only should be considered as an accurate measurement of his intellectual functioning abilities
prior to
age eighteen,
127
but also should be considered as proof that Thomas was not then, and is not now, mentally retarded. These contentions are deserving of at least some attention because it appears that there is a certain amount of arbitrariness in the selection of eighteen years as the end of the so-called “developmental period,”
128
and also because of Dr. McClaren’s opinion
*1298
that Thomas’s performance on this exam is
the
defining fact upon which his intellectual functioning ability during the developmental period should be determined.
Dr. McClaren made no effort to determine how many points, in his clinical judgment, should be deducted from Thomas’s IQ score for either measurement error or the Flynn effect. When asked to explain why he felt so strongly about the validity of Thomas’s raw, full-scale IQ score of 77, McClaren replied:
That appears to be the high point, and it happens at 18 years, one month. So I think it is highly likely that 30 days earlier before he was 18 that he would have scored in a like manner.
However, you do have the Flynn factor, which
if you gave every break and decided he should get the maximum discount, so to speak, from the Flynn effect, it
could reduce it about seven points.
So
— but
that’s not how IQ tests were interpreted in those days nor now when I see reports.
It’s with a score. And I often hear the psychologists say, well, probably the best in this confidence interval is the score you got.
And
— but
you know that things like the age of the test, the atmosphere of the testing situation, the environment in which the person is living all matters.
And when you report IQ scores, you report the score that you got and usually a confidence interval.
129
Several points in the foregoing portion of Dr. McClaren’s testimony require close scrutiny.
First — and addressing Dr. McClaren’s statement that “that’s not how IQ tests were interpreted
in those days
” — he is undeniably correct, but not for the reason he implies. Rather, that is a correct statement only because — as noted in Part 11(A)(3) of this opinion
supra
— the statistical phenomenon bearing Dr. Flynn’s name was not discovered until the middle of the following decade,
130
and clinicians did not begin to take it into account as a standard interpretative procedure until the 1990s, as the number of studies validating the accuracy of Flynn’s observations nudged his statistical data across the line dividing debatable propositions from empirically-proven, generally-accepted facts.
Second, Dr. McClaren’s assertion that he had not recently reviewed IQ reports interpreted in that
manner
— i.e., “nor now when I see reports” — conflicts with his acknowledgment that
“things like the age of the test,
the atmosphere of the testing situation, the environment in which the person is living
all matter
[ ].”
Third, if Thomas’s raw performance score was reduced by “seven points,” as Dr. McClaren conceded it could be, the resultant full-scale IQ score of 70 would indicate mental retardation.
Finally, Dr. McClaren simply ignored the generally-accepted practice of all competent professionals in the fields of statistics, test theory, and psychology of being concerned with how measurement errors affect the interpretation of an individual’s performance on a particular test. (See the discussion in Part 11(A)(2)
supra.)
In sum, Dr. McClaren glossed over the Flynn effect, ignored the standard error of measurement (which the parties’
stipulated
as being ± 5 points), and rational
*1299
ized his conclusion that Thomas did not meet the intellectual functioning prong of the clinical definition of mental retardation prior to the age of eighteen because the “trajectory” of Thomas’s IQ scores was “going up” during the developmental period,
131
and reached a peak one month after his eighteenth birthday.
132
This court does not find Dr. McClaren’s testimony persuasive, and concludes that the raw, full-scale IQ score discussed here was an aberration. The flaw in Dr. McClaren’s “trajectory” hypothesis was highlighted during further questioning by respondent’s counsel:
Q So do you have an opinion as to whether or not he meets the diagnostic criteria for mentally [sic] under the DSM-IV-TR?
A [McClaren] In my opinion, he does not meet the criteria for mental retardation.
Q So what would be his correct diagnosis as far as his intellectual ability is concerned?
A Borderline intellectual functioning. I believe that he probably does have some cerebral impairment. Whether that is related to his intellectual deficits or something distinct or some mixture is very hard to know.
But after 18 and the tests got re-normed, [and] he again, begins to drop off,
as
you would predict, using a Flynn
effect,
133
The last sentence in the foregoing quotation encapsulates the contradictions inherent in Dr. McClaren’s testimony. He is there referring to the 1981 standardization (“renorming”) of the first revised version of the Wechsler Adult Intelligence Scales, the so-called “WAIS-R.” (As will be discussed in the following section (Part IV(C)(1) of this opinion), the WAIS-R assessment instrument was administered to Thomas approximately eight years later, on March 21, 1985, and his raw, full-scale IQ score on that occasion was
71.)
In other words, Dr. McClaren was
tacitly conceding
that the only way to reconcile
(a)
Thomas’s raw, full-scale IQ score of 77 obtained at the age of eighteen years and one month with (6) his raw, full-scale IQ score of 71 obtained eight years later was by taking into account the “renorming” of the WAIS-R and,
thereby, the Flynn effect.
Stated differently — and considering the fact that during the eight years elapsed between April 11, 1977 (when Ms. Raley administered a 22-year-old WAIS assessment instrument) and March 21, 1985 (when, as discussed in the following Part of this opinion, a psychologist administered a four-year-old WAIS-R test), Thomas’s raw, full-scale IQ performance score dropped 6 points, from 77 to 71—
neither
those raw scores
nor
their “trajectory” can be rationally explained in the
*1300
absence of taking into account
both
the Flynn effect
and
the standard error of measurement. If that is done, then we can be 95% confident
134
that Thomas’s “true” IQ on the date he was tested by Ms. Joyce Raley was, in round numbers, 70 (70.4):
ie.,
virtually the same as when he was tested eight years later.
Another factor that causes the court to view Thomas’s performance on the April 11, 1977 test as an aberration is the fact that the record is silent on the question of what training or experience Ms. Raley had in the procedures for administering and scoring intellectual assessment instruments like the WAIS. The importance of that question is underscored by the facts summarized in the tables attached to the written report of Dr. Karen Salekin as Appendix A:
that is,
there is a consistency in the full-scale IQ scores recorded by Thomas both before and after Ms. Raley’s April 11, 1977 administration of the WAIS that causes his raw score on that date to stand out as not only the highest IQ score registered in Thomas’s lifetime, but also the
only
unadjusted, full-scale score above 74.
For all of these reasons, as well as those addressed in the following section, this court accepts Dr. Salekin’s well-supported, well-reasoned, and inherently-consistent analysis, and finds that the results of the April 11, 1977 IQ test are an aberration and should not be taken into account when determining Thomas’s intellectual functioning abilities during any of the relevant times:
ie.,
the developmental period, on the date of the offense, or currently.
C. Intelligence Assessments Performed Near the Date of the Offense
1.
March 21, 1985
— age
twenty-six years
The offense for which Thomas was convicted and sentenced to death occurred during the late-night or early-morning hours of December 15,1984.
135
Three months later, and prior to trial, he was evaluated by a “Dr. K. Hall” at the State of Alabama’s Taylor Hardin Secure Medical Facility in Tuscaloosa. Thomas then was twenty-six years of age. Dr. Hall administered the Wechsler Adult Intelligence Scales — Revised (WAIS-R), and Thomas’s raw scores were verbal IQ 70, performance IQ 74, and full-scale IQ 71. Dr. Hall concluded that he was “functioning within the borderline range of intellect.”
136
The WAIS-R was standardized (“normed”) in 1981, and Dr. Hall’s administration of it occurred four years later.
137
Multiplying 4 by the Flynn factor of minus 0.30 points for each year elapsed after the date of standardization yields a correction of 1.2 points that must be deducted from Thomas’s full-scale IQ score, resulting in an adjusted score of 69.8, rounded-up to 70.
138
Consideration of the ± 5 point standard error of measurement results in the conclusion, stated with a 95% degree of confidence,
139
that Thomas’s “true” full-scale IQ score lay within a band extending from 65 on the low end, to 75 at the upper extreme.
140
*1301
2.
January 2b, 1986
— age
twenty-six years and ten months
Thomas was evaluated once more prior to trial by a Ph.D. psychologist named James Crowder. The evaluation occurred on January 24, 1986, in the Limestone County Jail. Thomas then was twenty-six years and ten months of age. Dr. Crowder administered another WAIS-R assessment instrument, and Thomas’s raw scores were verbal IQ 65, performance IQ 69, and full-scale IQ 65. Even without taking the Flynn effect or standard error of measurement into account, these scores clearly fell within the range of mental retardation.
141
3.
Findings
In consideration of the evidence gleaned from the results of the March 21, 1985 and January 24, 1986 assessments, this court concludes that petitioner suffered from significantly subaverage intellectual functioning on the date of the offense.
See Smith v. State,
2007 WL 1519869, at *8 (noting that “subaverage intellectual functioning ... must be present at the time the crime was committed”).
D. Intelligence Assessments Performed in Preparation for Hearing
Dr. McClaren administered a WAIS-III intelligence assessment instrument to Thomas on September 6, 2007, and he obtained a verbal IQ score of 66, performance IQ score of 72, and a full-scale IQ score of 65.
142
Those results clearly support a finding of mental retardation, but the adjusted scores unquestionably do so. The WAIS-III assessment instrument was standardized (“normed”) in 1997, but it was administered to Thomas by Dr. McClaren some ten years later. Multiplying 10 by the Flynn factor of minus 0.30 points for each year elapsed after the date of standardization yields a correction of 3.0 points that must be deducted from Thomas’s raw, full-seale-IQ performance score, resulting in an adjusted IQ of 62 (65-3 = 62). Not only is that score well below the cutoff for a diagnosis of mental retardation, it still is subject to adjustment by a ± 5 point standard error of measurement, resulting in a conclusion that can be stated with a 95% degree of confidence that Thomas’s “true” full-scale IQ then lay within a band ranging from 57 to 67.
This court’s confidence that Thomas currently suffers from significantly subaver
*1302
age intellectual functioning abilities is bolstered by at least four additional facts. First, Dr. McClaren also administered a so-called “Test of Memory Malingering,”
143
and its results “did not suggest malingering” by Thomas during administration of the WAIS-III:
144
that is, deliberate feigning of mental retardation in order to avoid the death penalty.
145
Second, Dr. Salekin administered an SB5 intelligence assessment instrument to Thomas on November 11, 2007,
146
and it produced a full-scale IQ of 62 — identical to Thomas’s full-scale IQ score on the WAIS-III, when adjusted for the Flynn effect. Third, the “correlation coefficient” for the full-scale IQ scores produced by the WAIS-III and SB5 assessment instruments is + 0.82:
147
ie.,
very close to a perfect correlation coefficient of + l.O.
148
Finally, respondent
*1303
and Dr. McClaren concede this point.
149
Accordingly, the third requirement of
Smith ,
as that case construes
Atkins
and
Ex parte Perkins,
has been satisfied: petitioner has demonstrated, by a preponderance of the evidence, that he “currently” exhibits significantly subaverage intellectual functioning.
See Smith v. State,
2007 WL 1519869, at *8 ;
see also Holladay v. Allen,
555 F.3d at 1353 .
V. ASSESSMENTS OF PETITIONER’S ADAPTIVE BEHAVIOR
A. Petitioner’s Deficiencies Prior to Age Eighteen
No standardized, adaptive-behavior, assessment instruments were administered to Thomas prior to his eighteenth birthday. While consideration of a person’s limitations in adaptive skills became a component of the definition of mental retardation in 1959, “the notion that you had to test for it” did not become a diagnostic requirement until 2002, when the tenth edition of AAMR’s
Mental Retardation
manual was published.
150
Before then, psychologists relied primarily upon clinical judgment and the observations of third-party informants to assess a person’s adaptive skills. As Dr. Salekin put it, “we could talk to the school, we could talk to mom, we could talk to a bunch of people, [but we would] not hand over one of these [standardized tests], and be comfortable with giving a diagnosis that way.”
151
In order to gather information about Thomas’s adaptive skills during the period prior to his eighteenth birthday, therefore, the parties were relegated to a review of records maintained by public schools and social workers employed by the Alabama Department of Pensions and Security (“DPS”),
152
as well as the recollections of persons who were acquainted with petitioner during his youth.
153
Petitioner’s at
*1304
torneys argue that an objective assessment of such information establishes that Thomas’s adaptive behavioral skills during the developmental period were substantially below average in the areas of “functional academics, work, social and interpersonal skills, home living, and self-direction.”
154
Those assertions are evaluated below.
1.
The home environment
— birth
to age 12
DPS records establish that Kenneth Glenn Thomas was the third of nine children born into the marriage between William Thomas and Annie Ratcliff Thomas. The family was so socially and economically deprived that “even the poor people called them dirt poor.”
155
Dr. Salekin summarized the dysfunctional and abusive environment into which Thomas was born as follows:
With regard to his home of origin, records obtained from the Department of Pensions and Securities (DPS) are replete with information indicating that the parents of Mr. Thomas were repeatedly deemed to be unfit and that when in the home he was exposed to alcoholism (father), criminal activity (father), domestic violence (father toward mother), and there was evidence that Mr. Thomas Sr. had been molesting one or more of his daughters. Furthermore, there is information in the DPS records that suggests that on numerous occasions the Thomas family was without adequate food, clothing, and shelter (e.g., the home was unkempt and was too small for the amount of people residing there), and that Mrs. Thomas did not access appropriate medical services for the care and treatment of her children. In a notation dated 01/23/1963 (initials of writer IBM), the family was described as “destitute” and Mrs. Thomas was noted to have been wearing unseasonably thin and worn out clothing, and was described as “very cold and pitiful looking.”
Review of available documents provides ample information regarding DPS’s desire to reunify the family and the multitude of problems that prevented this from occurring. Throughout the records there are references made to the illegal, corrupt, violent, and unpredictable behavior of Mr. Thomas Sr., and the ineffective and dependent nature of Mrs. Thomas. During discussions with Ms. Carole Russell it became clear that the primary problem for the family lay in the relationship between the parents and the inability for Mrs. Thomas to extricate herself from this damaging union. According to Ms. Russell, it was common place for Mrs. Thomas to state that she would leave her husband permanently, but after a short period of separation (often times when he would be incarcerated or would otherwise abandon the family) she would return to the union. Over the course of the family’s involvement with DPS, Mrs. Thomas was informed multiple times that, should she make the decision to end the relationship with her husband, [she] and her children [could] be reunited. Mrs. Thomas was well aware of this fact, but routinely returned to their dysfunctional and violent relationship.
*1305
According to multiple informants (as well as information provided in court documents), Mr. Thomas was heavily influenced by his father who taught Mr. Thomas how to steal and to engage in other illegal activities. During Mr. Thomas’s testimony at trial, he stated that his dad would take him to bars when he was young, and that he witnessed him committing theft, arson, and assault. These accounts were corroborated by record review and interviews with numerous collateral sources (e.g., Connie Allen, Carole Russell, and Wayne Ridgeway). It is important to note that all nine of the children were at one point removed from the care of their natural parents, and in three of the nine cases the parental rights were terminated and the children adopted outside of the family.
According to records located in the DPS file, the Thomas family’s first contact with DPS occurred in 1957 at the time that Mr. Thomas’s father was first sentenced to prison (one year sentence). At this time Mrs. Thomas applied for assistance and did so again in 1963 when her husband was again incarcerated. In 1969 the family began to receive Social Security benefits due to Mr. Thomas Sr. becoming disabled secondary to a gunshot wound. In late 1971, Mr. Thomas was convicted of arson after having purposefully set fire to his family’s home and as a result, served two years in prison. Following his release Mr. Thomas Sr. once again behaved irresponsibly as evidenced by excessive drinking, “impulsive acts”, and he spent the majority of the family money on paying fines and court costs.
Doc. no. 110-2 (Report of Karen Salekin, Ph.D.) at 19-20 (footnote omitted).
Petitioner’s father, William Thomas, emphatically was a negative influence on his son’s life. Thomas did not recognize that fact, however. Instead, he perceived a “special relationship,” as Dr. Marson recorded:
His father apparently took a special interest in Mr. Thomas, and referred to his son as his “pet.” Mr. Thomas said he looked up to his father. Mr. Thomas reportedly often accompanied his father to the neighborhood bar where his father drank heavily, and also witnessed his father commit many criminal acts. For example, Mr. Thomas recalled an early childhood memory of watching his father break into a house, steal the contents of the home, and then set the house on fire. When Mr. Thomas got older, his father reportedly taught him how to strip cars and rob houses. Of note, when Mr. Thomas was 12 years old, he was arrested for starting a fire at the Capshaw Baptist Church in Athens. On February 22, 1974, at age 14, Mr. Thomas was reported arrested for stealing $5 from one of his teachers.
Ms. [Connie] Allen said that her brother learned everything bad from their father. According to Ms. Allen, Mr. Thomas was eager to please his father and did almost anything his father asked him. Ms. Allen stated that she did not think that Mr. Thomas ever learned to distinguish between right and wrong because his father used to tell him it was okay to do bad things.
Doc. no. 111-2 (Report of Dr. Daniel Marson & Dr. Kristen Triebel), at 4-5.
2.
Adolescence and foster care
— ages
12 through 18
Thomas resided with his family from birth until the age of twelve, when he was removed from the home and placed in DPS protective custody as a result of an arson. DPS records indicate that Thomas entered the Capshaw Baptist Church on a dare from two “older friends” and set fire to a
*1306
bulletin board and door.
156
He was arrested the following morning, after riding back to the church on his bicycle, and asking feigned questions about what had happened. Although the exact path of subsequent proceedings is not clear, it appears that Thomas was diverted from the juvenile court system into DPS-supervised foster care. He spent most of the following six years in three foster homes.
Thomas first resided for approximately one year (1971-72) with Mrs. Edith Austin and her family in Athens, Alabama, after which he returned to his family home for a short time, to live with his mother. Thomas reportedly was not happy living with his mother, however, because his father was in prison, and it was “not the same without him.”
Thomas then was placed in a second foster home, with Mr. and Mrs. Thomas Stevenson, for approximately a year and a half between the ages of thirteen and fourteen (1972 to 1974).
After repeatedly running away from the Stevensons’ home, Thomas was moved to a third foster placement, in the home of Mr. and Mrs. Wayne Ridgeway, who owned and operated a farm near Goodsprings, Alabama. Thomas described that placement as “the happiest home I knew.” No pleasure goes unpunished, however, and sadly, in September of 1977, Thomas was told that he no longer was eligible for foster-care or other DPS protective services as a result of becoming “an adult” on March 7, 1977 — his eighteenth birthday. Having no other options or resources, Thomas returned home to live with his mother.
Dr. Salekin summarized the circumstances of Thomas’s three foster-home placements in her written report as follows:
Mr. Thomas resided in foster care from the age of 12 years until he aged out of the system at 18 years of age. His first placement was with the Austin family ( — 08/23/1971) which lasted for a period of eight months after which time Mr. Thomas was removed from this placement and placed back in the family home. According to records, the change in placement was due to DPS’s observation that the placement was not longer of benefit to Mr. Thomas, but was instead hindering positive growth and development. In a progress note it was stated that DPS believed that, if possible, it would be much better for him to reside at home. After approximately one year, Mr. Thomas was again placed in foster care, but this time secondary to his mother being unable to manage and control his behavior.
On August 11, 1972, Mr. Thomas moved into the Stevenson Boarding home for a period of two years with termination of this placement occurring secondary to Mr. Thomas’s refusal to return to the home. This request was reportedly due to his growing dissatisfaction with the rules of the home, the reported alienation from his birth family, and the purported negative influence of a male foster child on Mr. Thomas. Mr. Thomas’s third and final placement was at the Ridgeway home (start date appears to be August, 1974). According to Ms. Carole Russell, Mr. Thomas was placed at this home because she believed that he would be able to function well in this environment. She further explained that there were few expectations
*1307
at the home and that Mr. Thomas would be able to carry out the simple tasks required of him. Ms. Russell indicated that the other children that were placed at the home were special needs children of low cognitive ability, a statement supported by Mr. Mitchell Rose, grandson of Ms. Ridgeway, and Mr. David Seibert, special education teacher.
According to Mr. Thomas, he was happiest when living at the Ridgeway home and he particularly enjoyed the outside chores that he and all of the children were required to do. He noted that during his time with the Ridgeways he did not get into trouble at school, did not use illegal drugs, played community baseball, and frequently attended church (these statements are supported by information available via record review and multiple collateral contacts). It is important to note that near the end of his residency at the Ridgeway home there were discussions between the Ridgeway family and DPS about continuing Mr. Thomas in foster care and assisting with his continued [vocational] education. It appears that at the time Mr. Thomas was about to “age out” of the system, DPS had either not made a firm commitment to providing further education or had not developed an appropriate transition plan. This lack of foresight and threatened abrupt discontinuation of services angered Mr. Wayne Ridgeway (as noted in DPS records and during personal communications with this examiner). It is noted that during discussion with Ms. Carole Russell he “became very upset, stating that he could not see how we could throw a child who could not take care of himself into the street.” The Ridgeways were clear that they could not keep Mr. Thomas unless the DPS payments continued.
In response to the above described situation, DPS looked into the vocational program that was offered by “the Junior College” (name unknown to this evaluator) at the time. Upon evaluation it was determined that “due to the fact that Kenneth is mildly mentally retarded, there was some concern as to whether or not he would be able to perform in the regular trade school program.” Apparently there was a relatively large amount of knowledge that would have been transmitted via text books and this was deemed to be unsuitable for the abilities of Mr. Thomas. Instead, DPS chose to look into the possibility of a good fit between Mr. Thomas’s abilities and the vocational program offered by “VPS” (exact name of the program unknown to this examiner). Although meetings with a VPS counselor were established, Mr. Thomas chose not to follow through with this recommendation. He reported that he no longer wanted to attend school, and he wanted to “be on his own.” Mr. Thomas’s contact with DPS was terminated on 09/19/1977.
Doc. no. 110-2 (Report of Karen Salekin, Ph.D.) at 20-21 (footnote omitted).
3.
Schooling
Thomas failed the first grade. During his second enrollment in that primary

[Text truncated at 120,000 characters. The full text is on the page linked above.]

---

Source: Frix Law Library, https://www.frixlaw.com/law-library/cases/2141960. Public record. Not legal advice.
