Citations
- 144 F. Supp. 3d 177
Full opinion text
FINDINGS OF FACT, RULINGS OF LAW, AND ORDER
YOUNG, DISTRICT JUDGE.
1. INTRODUCTION
In this action, ten black police sergeants (the “Plaintiffs”) employed by the Boston Police Department (the “Department”) brought suit against the City of Boston (“Boston” or the “City”) under Title VII of the Civil Rights Act of 1964, alleging that the multiple-choice examinations the Department administered in 2005 and 2008 to select which sergeants to promote to the rank of lieutenant had a racially disparate impact on minority candidates and were insufficiently job-related to pass muster under Title VII. The Plaintiffs also asserted a pendent claim under Massachusetts General Laws Chapter 151B (“Chapter 151B”). The City disputes that the exams had a disparate impact on minority candidates and claims that, even if they did, the exams were sufficiently job-related to survive a Title VII challenge.
This is a profoundly important case, one that evokes the finest of our nation’s aspirations to give everyone equal opportunity and a fair shot. In deciding this case, the Court first emphasizes what this case is not about: this is not a case about conscious racial prejudice. Rather, the Plaintiffs’ case is rooted in their allegation that the seemingly benign multiple-choice examination promotion process, while facially neutral, was slanted in favor of white candidates.
The parties engaged in a ten-day bench trial and submitted exhaustive post-trial briefs. The long trial involved substantial and dense discussions of statistical analysis. Consequently, the decision that follows is admittedly complex, but its conclusion is simple: the Department’s lieutenant-selection process — ranking candidates for promotion based on their scores on an exam administered in 2008 (“2008 exam”) — had a racially disparate impact and was not sufficiently job-related to survive Title VII scrutiny. Accordingly, the Court imposes liability on the City.
II. PROCEDURAL HISTORY
The Plaintiffs initiated this case in federal court in February 2012. Compl., ECF No. 1. Judge Tauro, to whom this case was originally assigned, dismissed without prejudice the claims of two of the Plaintiffs (John Johnson and Robert Tinker) for their failure to exhaust administrative remedies. Mem., ECF No. 28. Once this case was transferred to this Session on December 26, 2013, Mem., ECF No. 56, this Court denied the Plaintiffs’ motion to reconsider the dismissal, Elec. Order, ECF No. 67, and subsequently denied without prejudice the Plaintiffs’ motion to certify a class, Elec. Clerk’s Notes, ECF No. 70. Two years of discovery ensued, followed by the virtually inevitable cross-motions for summary judgment. Defi’s Mot. Summ. J., ECF No. 89; Pis.’ Mot. Summ. J., ECF No. 94. This Court denied summary judgment on all claims due to genuine disputes of material fact. Elec. Clerk’s Notes, ECF No. 120. In December 2014, the Court ruled that the remaining eight Plaintiffs had no viable disparate impact liability claim arising from their taking the 2005 lieutenant promotional exam (the “2005 exam”) due to their failure to exhaust administrative remedies. Elec. Order, ECF No. 150. Although no longer formally the subject of this litigation, the Court did consider evidence regarding the 2005 exam for background and context in evaluating the 2008 exam.
At the pre-trial conference, the Court bifurcated the case into separate liability and damages phases. Elec. Clerk’s Notes, ECF No. 98. The liability phase was tried before the Court between December 15, 2014 and January 7, 2015. See 12/15/14 Tr. 3:3-4, ECF No. 161; 01/07/15 Tr. 3:9-11, ECF No. 160. The following witnesses testified for the Plaintiffs: Dr. Joel Wiesen, PhD., industrial organizational psychologist (expert witness), 12/15/14 Tr. 3:6-12, Department Sergeant and Plaintiff Bruce Smith (fact witness), 12/17/14 Tr. 3:11-13, ECF No. 163, former Department Commissioner Edward Davis (fact witness), 01/05/15 Tr. 3:9-11, ECF No. 158, and Leatta M. Hough, PhD, industrial organizational psychologist (expert witness), 01/06/15 Tr. 3:5-7, ECF No. 159. The following witnesses testified for the City: Dr. Jacinto Silva, PhD, industrial organizational psychologist (expert witness), 12/17/14 Tr. 3:15-17, Dr. Michael Campion, PhD, industrial organizational psychologist (expert witness), 12/19/14 Tr. 3:5-7, ECF No. 166, Department Chief of the Bureau of Administration and Finance Edward P. Callahan (fact witness), 01/06/15 Tr. 3:9-11, and Department Commissioner William E. Evans (fact witness), 01/07/15 Tr. 3:5-7.
III. LEGAL CONTEXT
A. Title VII
It is the goal of Title VII “that the workplace be an environment free of discrimination, where race is not a barrier to opportunity.” Ricci v. DeStefano, 557 U.S. 557, 580, 129 S.Ct. 2658, 174 L.Ed.2d 490 (2009). The statute is designed to ‘“promote hiring on the basis of job qualifications, rather than on the basis of race or color.’ ” Id. at 582, 129 S.Ct. 2658 (quoting Griggs v. Duke Power Co., 401 U.S. 424, 434, 91 S.Ct. 849, 28 L.Ed.2d 158 (1971)).
Title VII, codified at 42 U.S.C. § 2000e, provides two theories of liability for discrimination in the employment context: disparate treatment and disparate impact. Ricci, 557 U.S. at 577-78, 129 S.Ct. 2658. A disparate treatment claim accuses an employer of intentionally basing employment decisions on an improper classification, such as race. See id. at 577,129 S.Ct. 2658. A disparate impact claim, by contrast, challenges an employment decision that is facially neutral, but which falls more harshly on those in a protected class. See id. at 577-78, 129 S.Ct. 2658.
This is a disparate impact case. Second Am. Compl., Compensatory, Injunctive & Declaratory Relief Requested (the “Complaint”) ¶ 1, ECF No. 14. Section 2000e-2(k)(l)(A) outlines the burden of proof in a disparate impact case:
An unlawful employment practice based on disparate impact is established under this subchapter only if—
(i) a complaining party demonstrates that a respondent uses a particular employment practice that causes a disparate impact on the basis of race, color, religion, sex, or national origin and the respondent fails to demonstrate that the challenged practice is job related for the position in question and consistent with business necessity; or
(ii) the complaining party makes the demonstration described in subpara-graph (C) with respect to an alternative employment practice and the respondent refuses to adopt such alternative employment practice.
42 U.S.C. § 2000e-2(k)(l)(A).
Under First Circuit case law, the plaintiff bears the burden of establishing a prima facie case of discrimination which consists of identification of an employment practice (in this case, the 2008 exam and promotions flowing therefrom), disparate impact, and causation. Bradley v. City of Lynn, 443 F.Supp.2d 145, 156 (D.Mass.2006) (Saris, J.) (quoting EEOC v. Steamship Clerks Union, Local 1066, 48 F.3d 594, 601-02 (1st Cir.1995)).
If the Plaintiff meets this burden, the employer may either debunk the Plaintiffs prima facie case, or alternatively, may demonstrate that the challenged practice is “job-related and consistent with business necessity.” Bradley, 443 F.Supp.2d at 157; see also Ricci, 557 U.S. at 578, 129 S.Ct. 2658. If the employer demonstrates the latter, the ball bounces back into the plaintiffs court to demonstrate that “some other practice, without a similarly undesirable side effect, was available and would have served the defendant’s legitimate interest equally well.” Bradley, 443 F.Supp.2d at 157.
The law of disparate impact has become a powerful tool for ensuring equal opportunity. It is both balanced and nuanced. Each step in its three-part doctrine serves a valuable function. Consider the following.
All testing, hiring, and promotion schemes are necessarily discriminatory. These programs exist because there are more applicants than there are jobs. Under the first prong, the Plaintiffs must make a significant showing of actual disparate impact upon an identified protected minority. This is as it should be: no one wants federal courts acting as super personnel agencies.
If the plaintiff can, however, make this showing, then under the second prong, the employer gets a chance to demonstrate that the test in question is both job-related and consistent with business necessity. Again, this step is sensible: courts ought not defer excessively to employers, but neither should they ignore the realities of particular jobs. The current debate over the exhaustive testing to determine the capabilities of women to engage in military ground combat comes immediately to mind as exemplifying the difficult issues encountered in prong two.
Even if the employer succeeds, however, the case is not over. Under the third prong, the plaintiff gets one more shot. If the plaintiff can demonstrate the availability of a testing program equally determinative of job performance, yet resulting in less disparate impact, the Court should fashion a remedy to secure the greatest degree of equal opportunity. In other words, to produce more equality of opportunity, Title VII empowers courts to impose on employers an equally effective means of evaluating applicants.
B. The Commonwealth’s Statutory and Administrative Framework
Under Massachusetts law, police sergeants seeking to be promoted to lieutenant are subject to the state civil service statutory promotion regime. See Mass. Gen. Laws ch. 31, § 51. The purpose of the examination regime is to “guard against political considerations, favoritism, and bias in governmental employment decisions ... and to protect efficient public employees from political control.” Cambridge v. Civil Serv. Comm’n, 43 Mass.App.Ct. 300, 304, 682 N.E.2d 923 (1997). Under this regime, to become a lieutenant, Boston police sergeants must first pass a competitive civil service examination. Mass. Gen. Laws ch. 31, § 59.
The Commonwealth of Massachusetts Personnel Administrator of the Human Resources Division (“HRD”), Mass. Gen. Laws ch. 31, § 1, is responsible for “conducting] examinations for purposes of establishing eligible lists” for promotion. Id. § 5(e). HRD is obligated by statute to “fairly test the knowledge, skills and abilities which can be practically and reliably measured and which are actually required” to perform the job. Id. § 16. To achieve this end, HRD develops the examination, id. §§ 5(e), 16, posts notices of the exams, id. §§ 18-19, and determines the passing requirements, id. § 22. Promotional examinations within the Department are typically administered every two or three years. 12/15/14 Tr. 64:4-10, ECF No. 161. Unlike other jurisdictions, Massachusetts requires candidates for each supervisory rank (sergeant, lieutenant, and captain) to take a test at each promotional level, even if there is substantial overlap in the questions that appear on the tests for each rank. Id. at 129:19-130:5.
The Department has the option of either using the exams developed by HRD, or seeking a delegation agreement with HRD by which HRD agrees to oversee the Department creating its own exam. Mass. Gen. Laws ch. 31, §§ 5(1), 59, 65; 01/06/15 Tr. at 91-93. Under this delegation regime, municipalities must still comply with civil service law and regulations, but may decide themselves how to satisfy these requirements. Lopez v. Massachusetts, 588 F.3d 69, 76 (1st Cir.2009). When the City enters into such a delegation agreement, it must bear the costs of developing and administering the tests. 12/17/14 Tr. 49:21-24. The City does not incur these costs when it uses an HRD test. 01/06/15 Tr. 112:16-21.
In a competitive examination, “an applicant shall be given credit for employment or experience in the position for which the examination is held.” Mass. Gen. Laws ch. 31, § 22. Such credit is known as an education and experience score (“E & E”), and is calculated using biographical information provided by the applicants. Id.; Ex. 85, Affidavit Edward P. Callahan (“Callahan Aff.”), ECF No. 177; Ex. 3, Education Experience Sheet Instructions (“E & E Instructions”) 1, ECF No. 177-3. Pursuant to Massachusetts statute, veterans and long-service employees receive preference points. Mass. Gen. Laws ch. 31, §§ 26, 59.
After combining the exam scores with the E & E scores, HRD issues an eligibility list for specific positions ranked in order of an applicant’s total score. Id. §§ 25, 27. An eligibility list remains in effect until replaced by a new eligibility list from a subsequent exam. Id. § 25; Callanan v. Pers. Adm’r for the Commonwealth, 400 Mass. 597, 601-02, 511 N.E.2d 525 (1987). The City has promoted police officers based on promotional examination results in strict rank order since at least 1977. 12/15/14 Tr. 63-64.
To hire for a vacancy, the Department submits a request to HRD, which certifies from the larger eligibility list a smaller list of names of persons for consideration in rank order. Mass. Gen. Laws ch. 31, § 6. Under HRD’s Personnel Administration Rules, the number of candidates appearing on the smaller list is determined by the formula 2n+l, with n representing the number of vacancies. 12/15/14 Tr. 63:5-16. For example, if there are two job vacancies corresponding to one applicable list, the list would be comprised of the candidates with the five highest scores (2x2+ 1=5).
The Department must then make selections from that list based on strict rank order, based on the candidates’ performance on the promotional exam. 01/06/15 Tr. 102:12-18. If the candidates have tied scores, the Commissioner may consider factors such as past work history and diversity. 01/07/15 Tr. 9:25-11:19. The statutory framework allows a municipal employer to' “bypass” a candidate on the list— that is, to step out of strict rank order— but the employer must have a defensible reason for the bypass, such as a history of disciplinary infractions. Mass. Gen. Laws ch. 31, § 27; City of Cambridge v. Civil Service Comm’n, 43 Mass.App.Ct. 300, 305, 682 N.E.2d 923 (1997). Former Commissioner Davis testified that as a practical matter, bypassing is difficult. 01/05/15 Tr. 100-101.
Dissatisfied candidates may challenge the results of the examination with the HRD by claiming that the examination was not a “fair test of the applicant’s fitness actually to perform the primary or dominant duties of the position for which the examination was held.” Mass. Gen. Laws ch. 31, § 22. The Massachusetts Civil Service Commission oversees the administrative appeals process for candidates to air their grievances with the hiring process and the initial review by the HRD. Id. § 24. Judicial review is available after the candidate has exhausted his or her administrative remedies. Id. § 44.
IV. FINDINGS OF FACT
This is not the first case in which Department employees or potential employees have challenged the Department’s hiring or promotional procedures. In fact, a case raising similar issues to this one was brought before Judge O’Toole in 2007. Judge O’Toole’s findings of fact and conclusions of law issued in September 2014. Findings Fact, Conclusions Law, Order J., Lopez v. City of Lawrence (“Lopez”), ECF No. 347 (D. Mass. Sept. 5, 2014). An appeal is pending. Lopez, appeal docketed, No. 14-01952 (1st Cir. Sept. 17, 2014). As there is some factual overlap, on the first trial day in this case, without objection, this Court admitted in evidence all the trial testimony and exhibits from Lopez. 12/15/14 Tr. 32:4-16.
More broadly, over the past few decades, candidates who are members of a racial minority and candidates who are not have challenged the hiring and promotional procedures employed by the Department as unlawfully discriminatory. The extensive litigation history is well documented elsewhere, see Lopez at 12-14, and the Court will not repeat it here. The Court mentions it only for the purpose of noting that, similarly to police departments in other jurisdictions, this pendulum of litigation has pressured the City constantly to re-tool its examination procedures in an effort to appease job applicants and courts alike. See Barnhill v. City of Chicago, Police Dep’t, 142 F.Supp.2d 948, 949 (N.D.Ill.2001).
A. The Role of a Boston Police Department Lieutenant
Boston Police Department lieutenants are second-line supervisors, meaning that they supervise sergeants, who themselves supervise police officers out in the field. 01/05/15 Tr. 102:1-10. Lieutenants are also in charge of station houses and are responsible for the proper arrest of suspects and for the safety of prisoners. Id. at 128:3-9, 131:4-8. The job involves a significant amount of desk work, 12/17/14 Tr. 104:9-12, as well as work outside of the station, including talking with citizens at community meetings, id. at 105:14-25, and taking control of scenes of major incidents, 01/05/15 Tr. 126:23-127:12. The job requires good management skills, including the ability to motivate employees, and to communicate information between ranks. Id. at 102:11-23.
The official Department job description for lieutenant has not changed since 1979, and current Boston Police Department Commissioner William Evans testified that it remains accurate today. Ex. 23, Boston Police Department Rules Procedures, Rule 105; 01/07/15 Tr. 16:15-17:9. Just prior to 2006, however, Boston began to shift its policing philosophy towards that of community policing, where police officers engage more directly with the community they serve. 01/05/15 Tr. 109:16-110:15. The skills required for a Boston Police Department lieutenant evolved with this shift, differing from what was needed in the early 1990s when the prevailing policing philosophy involved responding to, rather than preventing, emergencies. Id. at 110:9-12.
B. Job Analyses and Validation Studies Pre-Dating 2005
The first step in developing a valid civil service promotional exam is to create a job analysis, which identifies “important work behavior(s) required for successful performance and their relative importance.” 29 C.F.R. § 1607.4(C)(2) (“Uniform Guidelines”).
The development of the 2005 and 2008 exams began in 1991 with the creation of a job analysis and validity report, which HRD incorporated into the 2008 exam. 12/19/14 Tr. 14:5-15. HRD also incorporated into the 2008 exam a job analysis from 2000, which was in part based on the 1991 job analysis. Id. at 14:5-15:25, 32:13-16. The documents are somewhat dense, and the Court will discuss each in turn.
1. The 1991 Job Analysis
In 1991, the Massachusetts Department of Personnel Administration (“DPA”), the predecessor to HRD, prepared a statewide validation report for the ranks of sergeant, lieutenant, and captain. Ex. 71, Validation Report 1991 Police Promotional Selection Procedure (“1991 Validation Report”) at 00247. DPA relied on the Uniform Guidelines, as well as the Society for Industrial and Organizational Psychology Principles for the Validation and Use of Personnel Selection Procedures (“SIOP Principles”). Id. at 00247, 00250.
DPA began the job analysis by gathering information about the positions of police sergeant, lieutenant, and captain. It did so by surveying other jurisdictions and reviewing various documents including job analysis studies, articles, and the Uniform Guidelines. 1991 Validation Report at 00253-54. Based on this research, DPA created a list of 136 potentially critical tasks that sergeants, lieutenants, and captains perform. Id. at 00256. DPA sent surveys to municipal police departments across the Commonwealth asking incumbents of the positions to identify which tasks were critical to successful job performance and then rate them accordingly. Id. Police officers were asked to provide information regarding how often they performed the tasks, and to identify the fifteen tasks most critical to their jobs. 1991 Validation Report, App. H at 3638.
DPA then used these task ratings to create a list of 187 Knowledge, Skills, and Abilities (“KSAs”) necessary to carry out the critical tasks. 1991 Validation Report at 00257-58. In a survey administered to subject matter experts (“SMEs”) — superi- or officers serving in Massachusetts police departments — DPA asked the SMEs to evaluate two things: the importance of the KSAs, and whether the candidate needed the KSA at the time of appointment to the position or could acquire the KSA on the job. Id. Only KSAs that were determined to be important and necessary upon starting the job were considered for inclusion in the test for applicants. Id. at 00258.
In addition to these various surveys, DPA developed additional KSAs by holding “critical incident” technique discussions with the SMEs. Id. at 00252. These discussions involved SMEs providing narrative explanations or anecdotes about the tested positions. Id. The purpose was to allow DPA to design situational examination questions evaluating supervisory abilities. Id.
The next step was to link the KSAs with the critical tasks. Nine SMEs were tasked with this assignment. Id. at 00258. The SMEs identified, by group consensus, which KSAs had a direct relationship to identified clusters of critical tasks. Id. at 00259. Of the 187 KSAs, 60 were ultimately incorporated in the written test. 1991 Validation Report, App. EE.
The 1991 report acknowledged that some of the skills identified as important by SMEs could not be evaluated by a written test, such as the “ability to establish rapport with persons from different ethnic, cultural, and/or economic backgrounds.” Id. at 00265. The 1991 report also noted that “assessment of the performance of these skills and abilities would require the use of selection devices outside the scope of the written, multiple choice format.” Id. at 00265. Nevertheless, DPA decided to proceed exclusively with a written examination, offering eighty questions common to all three positions; an additional twenty questions for lieutenants and captains testing knowledge of police supervision, administration, and management; and another twenty-five questions to test captains on their knowledge of police administration. Id. at 00266-67. For the 1991 exam, the written portion accounted for 80% of an applicant’s final score, and the E & E portion for 20%. Id. at 00263.
2. 2000 Job Analysis and the Corresponding 2002 Exam
The 2000 job analysis (the “2000 report”) was prepared at the request of the City of Boston by Morris. & McDaniel, Inc., a consulting firm that specializes in the development of promotional systems. Ex. 39, Job Analysis Report Police Lieutenant City Boston (“2000 Job Analysis Report”). The 2000 report at issue in this case concerned Boston Police Department lieutenants only. Id. Morris & McDaniel came up with a list of 302 possibly relevant tasks that Boston police lieutenants perform, as well as KSAs necessary to carry out those tasks. Id. at 65; id. App. A, Task Inventory Police Lieutenant Boston Police Department.
Morris & McDaniel then had twelve SMEs, consisting of Department employees holding the rank of Lieutenant or higher, rate the tasks for frequency, importance, necessity of performing the task upon starting the job, and how correlative successful performance of the task was to successful job performance. 2000 Job Analysis Report at 10-14. If ten of the SMEs rated a task as “very important” or “important,” necessary upon entry to the job, and agreed that performance of that task clearly separated the best workers or better workers from inferior workers, then it satisfied Morris & McDaniel’s test criteria. Id. at 14. Of the initial 302 tasks, 281 fulfilled the criteria. Id. Morris & McDaniel also asked the SMEs to determine which of the following dimensions were required for each task: oral communication, interpersonal skills, problem identification and analysis, judgment, and planning and organizing. Id. at 15. Morris & McDaniel then composed a list of 149 KSAs potentially necessary to perform the 281 tasks. See id. at 48-49. Next, the SMEs were asked whether the KSAs related to the job of police lieutanant, when the KSA was learned (before or after assignment to the job), how long it took to learn the KSA, how the KSA differentiated performance, and whether the KSA was required to perform the job effectively. Id. For a KSA to be important enough to be tested, nine of the twelve SMEs must have rated the KSA as related to the job, learned before assignment to the job, requiring more training than a brief orientation period, capable of distinguishing performance to a high or moderate degree, and required or desirable to perform the job effectively. Id. at 49. Of the 149 KSAs rated by the SMEs, 145 were deemed sufficiently important to be tested. Id.
Based on its 2000 job analysis, Morris & McDaniel created a promotional test for the Department, to be administered in 2002. Ex. 80, Draft Police Lieutenant Written Examination Validity Report (‘Validity Report”) l. The 2002 exam consisted of three components: a written exam, an assessment center, and E & E. Id. at 2. Of a possible 100 points, Morris & McDaniel, after consulting with SMEs, assigned 30% to the written examination, 50% to the assessment center, and 20% to E & E. Id.; Ex. 81, Police Lieutenant Assessment Center Validity Report (“Assessment Center Report”) 17; 12/22/14 Tr. 46-47, ECF No. 167; 01/06/15 Tr. 110:7-13.
Morris & McDaniel composed the written test questions, which SMEs reviewed for accuracy and clarity. Validity Report 11-12. Based on the responses, Morris & McDaniel determined that the questions were internally consistent and reliable. Id. at 13.
The assessment center was designed to test oral communication skills, interpersonal skills, ability to quickly identify a problem and analyze it, ability to make sound decisions promptly, and ability to break work down into subtasks and prioritize them. Assessment Center Report 5-6. The assessment center consisted of an in-basket exercise (a simulated written exercise), and a situational exercise in which candidates were videotaped offering verbal responses to hypothetical scenarios that a lieutenant may encounter. Id. at 7-8. The assessment center exercises were evaluated by outside assessors. 01/06/15 Tr. 110:14-23.
After the 2002 exam had been administered, Morris & McDaniel prepared validity reports of the written examination and the assessment center. See Validity Report; Assessment Center Report. The validating report was written to comply with the Uniform Guidelines. Validity Report 1. Morris & McDaniel concluded that both portions of the examination were valid. Id. at 18.
This process cost $1,300,000, which included consulting fees and travel costs for outside assessors. 01/06/15 Tr. 112-113. HRD certified a list for promotion, from which one black sergeant was selected for promotion. Ex. 38; 01/06/15 Tr. 112:2-4.
C. Development and Administration of the 2008 Exam
Based in part on financial constraints and a perceived lack of improvement in diversity resulting from the 2002 exam, the Department elected to use HRD exams (conventional, written exams) in 2005 and 2008. Lopez 07/28/10 Tr. 17, 30; 01/06/15 Tr. 98:9-12, 118-19. HRD consulted with the outside firm EB Jacobs for the 2008 exam. 12/16/15 Tr. 55:17-56:3.
HRD did not create a comprehensive job analysis for the 2008 exam, but instead conducted “mini job analyses” for all three ranks, which were more or less updates or overlays to the 2000 report. HRD asked SMEs to rate tasks and KSAs. Ex. 55; Ex. 56; Tr. 01/07/15 at 20-36. The KSAs were pulled from the 2000 job report. 12/19/14 Tr. 38:9-19.
Commissioner Evans, who was a fact witness in this case, served as an SME for the 2008 exam. 01/06/15 Tr. 126:6-126:10. He testified that the SME review of the KSAs and tasks was “meticulous,” 01/07/15 Tr. 30:25-31:4, and that the purpose of the 2008 exam was not to test memorization of facts, but to test situational judgment, interpersonal relations, communication ability, and knowledge of rules and regulations. 01/07/15 Tr. 35-36.
Based on the mini job analyses, the Department’s consultant, EB Jacobs, created a test outline for the 2008 exam. Ex. 54; Ex. 60. The Civil Service then compiled a list of 100 questions for the exam. 01/07/15 Tr. 31:5-10. The SMEs reviewed the test questions for suitability for the different ranks, difficulty, and readability, and then indicated whether they recommended the question for the exam. Ex. 60; 01/07/15 Tr. 31:11-17.
HRD announced the 2008 exam and provided a corresponding reading list to members of the Department. Exs. 1, 5, 6, 8, 17. The Department provided “substantial amount[s] of tutorial information” for all candidates taking the promotional exams, including taped lectures and practice questions. 01/06/15 Tr. 128. The questions that appear on the written portion of the exam are taken directly from the reading list. 12/15/14 Tr. 51:19-52:12.
The 2008 exam consisted of two elements: a written, closed-book exam consisting of 100 multiple-choice questions, and an E & E rating. 12/15/14 Tr. 49:7-14. Out of 100 possible points on the written examination, a candidate heeded to score 70 to pass. Id. at 62:21-23. The E & E Score is calculated only for candidates who passed the written exam. Id. at 62:24-63:2. The written portion accounted for 80% of the final score; the E & E component for 20%. Id. at 50:9-50:15.
D. Development and Administration of the 2014 Exam
The Department typically offers promotional exams every two or three years. Id. at 64:4-7. The 2008 exam did not follow this trend: Commissioner Davis requested that the promotional certifications be extended due to the Lopez litigation. Id. at 64:11-18; Ex. 59; Ex. 61. Thus, it was not ■ until 2014 that the Department developed and administered a new promotional exam; the list from the 2008 exam was still in effect at the time of this trial. 01/06/14 Tr. 143:25-144:3.
The Department, in part due to an improved economic climate, and in part due to a desire to increase diversity, elected to go beyond a written examination and E & E for the 2014 promotional process. 01/05/2015 Tr. 139-140:8. The Department again retained the firm of EB Jacobs, this time to design and administer the 2014 exam. Callahan Aff. ¶ 1. At the firm’s recommendation, the Department in 2013 approved the development of a “very comprehensive” job analysis in anticipation of the 2014 exam (“2013 job analysis”). 01/06/15 Tr. 109:11-19. This was the first full job analysis performed since 2000. Id. at 131:5-19. Based on the 2013 job analysis, EB Jacobs recommended the use of an assessment center in addition to a written examination. Id. at 133:12-23. After securing funding from Boston for an assessment center, the Department secured a delegation from HRD to develop its own promotional exam. Id. at 134:5-17.
The Department posted an announcement of the exam, and indicated that the exam would be weighted as follows: technical knowledge written exam (36%); in-basket test (where candidates provide written essay-style responses to various job situations) (20%); oral board test (where candidates provide oral responses to hypothetical incidents and personnel issues) (24%); and E & E (20%). Callahan Aff. ¶¶ 4, 7; id., Ex. 2, Lieutenant Promotion Examination Candidate Preparation Guide In-Basket Oral Board Tests 3, 5, ECF No. 177-2.
Unlike the 2005 and 2008 exams, there was no cut-off score for the written portion of the 2014 exam. 12/17/14 Tr. 111:25-112:2. Much else was the same, however. Outside assessors evaluated each candidate’s performance in the assessment centers, id. at 114:22-115:4, the E & E component was based on self-reporting of candidates’ education, training and work experience, E & E Instructions 1, and veterans received an additional two points, Callahan Aff. ¶ 9. The full development of the promotional exam process (testing for promotion of sergeant, lieutenant, and captain) cost over $1,600,000. Id. ¶ 12.
E. Results of the 2005 and 2008 exams
While all of this history is informative and helpful to gain an understanding of the Department’s promotional process, it is the results of the 2005 and 2008 exams that form the crux of this dispute. By and large, the parties agree on all of the numbers in this section. The crux of the dispute is which of these numbers are important, and which methodology is the most appropriate for analyzing these numbers.
One hundred and twenty seven candidates reported for the 2005 lieutenant promotional exam: 104 were white, 22 black, and 1 Hispanic. Ex. 47, Adverse Impact Evaluation: 2008 and 2005 Exams Promotion Lieutenant, BPD (“Wiesen Report”) 8; Ex. 72, Report Jacinto M. Silva (“Silva Report”) 2. The passing rate of the written exam for minorities was 50%, and for whites, 88%. Wiesen Report 12. The mean score for minorities on the 2005 exam was 69.9, and for whites, 78.7. Id. at 16. Of the 127 sergeants who sat for the exam, 27 were promoted: 25 were white (out of 104 white applicants), one was black (out of 22), and one was Hispanic (who was the only Hispanic candidate). Id. at 8; Silva Report 2.
Ninety-one sergeants sat for the 2008 promotional exam: 65 were white, 25 were black, and one was Hispanic. Wiesen Report 7; Silva Report 7. Of the 91 candidates who took the exam, the passing rate for minorities was 69%, and for whites was 94%. Wiesen Report 11. The mean score for minorities was 76.6, and for whites was 83.2. Id. at 14. Of the 91 candidates, 33 were promoted: 28 were white (out of 65), and 5 were black (out of 25). Silva Report 7.
After reviewing the final scores from the 2008 exam, the Department’s consultant, EB Jacobs, recommended “banding” the results in nine-point increments (meaning that scores within nine-point ranges would be deemed equivalent). Lopez Exs. 70, 71.
The background and raw numbers for this case are straightforward and self-explanatory. Their statistical analysis and corresponding legal conclusions, less so.
Y. CONCLUSIONS OF LAW
A. Disparate Impact (Prong 1)
The use of the 2008 exam is the employment practice subject to challenge under Title VII. The parties agree that minority test-takers passed the 2008 exam and were promoted to lieutenant at a lower rate when compared to white candidates. But a lower rate of passage or promotion is not, by itself, sufficient to establish a prima facie case of discrimination. Statistical disparities must be significant enough to “raise an inference of causation.” Bradley, 443 F.Supp.2d at 157 (citation omitted). The Supreme Court has stated:
[T]he plaintiff must offer statistical evidence of a kind and degree sufficient to show that the practice in question has caused the exclusion of applicants for jobs or promotions because of their membership in a protected group. Our formulations, which have never been framed in terms of any rigid mathematical formula, have consistently stressed that statistical disparities must be sufficiently substantial that they raise such an inference of causation.
Watson v. Fort Worth Bank & Trust, 487 U.S. 977, 994-95, 108 S.Ct. 2777, 101 L.Ed.2d 827 (1988) (plurality opinion); see also Texas Dep’t of Hous. & Cmty. Affairs v. Inclusive Communities Project, Inc., — U.S. -, 135 S.Ct. 2507, 2523, 192 L.Ed.2d 514 (2015) (noting that “[a] robust causality requirement ... protects defendants from being held liable for racial disparities they did not create”). In other words, .the Plaintiffs must show that any disparity between races is not the result of mere chance. See Jones v. City of Boston, 752 F.3d 38, 43 (1st Cir.2014).
There is no “single test” to demonstrate disparate impact. Langlois v. Abington Hous. Auth., 207 F.3d 43, 50 (1st Cir.2000). Plaintiffs in Title VII disparate impact cases often demonstrate causation by presenting evidence that the disparity in outcomes between white and minority candidates is “statistically significant,” meaning that statistical analysis reflects that the odds of the disparity occurring by mere coincidence are less than 5%. Jones, 752 F.3d at 43-44. This is demonstrated when a statistician determines that the “p-value,” which stands for “probability,” is less than .05, meaning that the probability of the result occurring by chance is less than 5%. Jones, 752 F.3d at 46-47; 12/15/14 Tr. 74:19-75:7; Wiesen Report 5.
Another way to express this same mathematical calculation is to utilize the statistical measure of “standard deviation,” which measures how dispersed a set of data is (the more dispersed the data, the higher the standard deviation). See 12/15/14 Tr. 75:19-25. One “can calculate the standard deviation ... for any [p-value].” Id. at 77:1-9. For a so-called one-tailed test, the relevant standard deviation is 1.645. 12/15/14 Tr. 77:1-9. For a two-tailed test, it is 1.96. Id. In other words, if one utilized a two-tailed test and found that the mean test scores of minority candidates were located 1.96 standard deviations away from the overall mean, there would be only a 5% probability that such difference was due to chance. (And if their mean score was more than 1.96 standard deviations from the mean, the probability that it was due to chance would be even lower.)
Parties alleging disparate impact also sometimes rely on what is known as the “four-fifths rule,” articulated in the Uniform Guidelines, a non-binding set of guidelines authored by the Equal Employment Opportunity Commission to help employers comply with Title VII. The thrust of the “four-fifths rule” is that a selection rate for any racial group that is less than four-fifths (or 80%) of the rate of the group with the highest rate is evidence of adverse impact. 29 C.F.R. § 1607.4(D). The four-fifths rule is “widely used” among organizational psychologists. 12/15/14 Tr. 72:25-73:6. The First Circuit has acknowledged, however, that it is' not decisive. Jones, 752 F.3d at 51.
Reports and testimony from experts play a crucial role in evaluating a disparate impact claim. The Plaintiffs’ expert witness regarding disparate impact was Dr. Joel Peter Wiesen, an industrial organizational psychologist. 12/15/14 Tr. 34:3. Dr. Wiesen holds his PhD in psychology. Id. at 34:11. He worked for HRD from 1977-1992, during which time he was the chief of test development and validation. Id. at 36:13-25. Since leaving HRD, Dr. Wiesen has been consulting in the area of test development, and has served as an expert witness. Id. at 38:10-16. The City rebutted Dr. Wiesen’s testimony with that of Dr. Jacin-to M. Silva, who holds a PhD in industrial organizational psychology. 12/17/14 Tr. 141:19-142:11; Silva Report 1. Dr. Silva is currently a senior managing consultant at EB Jacobs, which is the firm that developed the 2014 lieutenants’ exam for the Department and consulted on the 2008 exam. Id. at 143:25-144:11.
Drs. Wiesen and Silva agreed on many issues: the raw data underlying each other’s analysis (although there were some minor discrepancies based on the timing of their reports); that each other’s mathematical calculations were correct; and that the Fisher Exact Test was the appropriate test for this analysis. Def. City Boston’s Post-Trial Proposed Findings Fact & Conclusions Law (“Post-Trial City’s Proposed Findings”) 38 n.21, ECF No. 190. The experts disagreed as to which statistics were relevant (among promotion rates, mean scores, pass-fail rates, or delay in promotion); whether the results from the 2005 and 2008 tests should be aggregated; and whether a one-tailed or a two-tailed Fisher Exact Test was the appropriate methodology. The Court will address these issues in turn.
1. The Relevant Data Points
Before beginning any analysis, statistical or otherwise, the Court must determine the proper scope of its inquiry. The Plaintiffs argue the Court should cast a wide net, looking to several statistics regarding the 2008 exam. The City, however, suggests that one number, promotion rates, provides all the necessary information. As explained more fully below, the Court largely agrees with the Plaintiffs.
Dr. Wiesen examined various aspects of the lieutenant promotional procedure employed by the City in an effort to determine whether there was disparate impact. Specifically, he compared the numbers between minority and non-minority candidates for: (1) promotion rates; (2) passing rates, (3) average exam scores; and (4) delays in promotion. 12/15/14 Tr. 81:23-82:11; Wiesen Report 3.
Dr. Wiesen, acknowledging that the first measurement, promotion rates, is the “most important[ ],” Wiesen Report 5, nevertheless argued that the other measurements were also important for various reasons. He argued that the passing rates and average scores were relevant because of the current system in which candidates are promoted in strict rank order. Id. He posited that passing rates and average scores were more “sensitive” or “statistically] power[ful]” than promotion rates. Id. at 5-6. Dr. Wiesen opined that delays in promotion in the Department — meaning the time between becoming eligible for a promotion and actually receiving that promotion — were important because the timing of promotion from sergeant to lieutenant affects how job tasks are assigned. Id. at 6. Dr. Wiesen ultimately concluded that the 2005 and 2008 exams “had fairly severe adverse impact on minority candidates, black and Hispanic.” 12/15/14 Tr. 48:21-24.
In contrast to the Plaintiffs’ consideration of four sets of data, Dr. Silva and the City argued that promotion rates were the only appropriate measurement for determining adverse impact. Post-Trial City’s Proposed Findings 44-45; see Silva Report 13. They argued essentially that the Court should ignore average exam scores and pass-fail rates because the only value they have is in predicting promotion rates, which can measured directly. See Silva Report 13; Post-Trial City’s Proposed Findings 45. Dr. Silva argued that a delay in promotion is an inappropriate measuring device because “the number of days between promotions is not a function of the test, it is a function of when the positions open up. The test is only responsible for the order in which the promotions are made.” Silva Report 6. Analyzing only the promotion rates in 2005 and 2008 using a two-tailed test, the City argues, one cannot conclude that the 2008 exam resulted in a statistically significant adverse impact, and thus, judgment should enter in the City’s favor. Silva Report 7-10; Post-Trial City’s Proposed Findings 41.
The Court agrees with the Plaintiffs that statistics other than promotion rates are relevant in evaluating disparate impact. The City’s argument that promotion rates are the only relevant factor is a “bottom line” defense which the Supreme Court has rejected. In the seminal case Connecticut v. Teal, the Supreme Court stated:
In. considering claims of disparate impact under [Title VII], this Court has consistently focused on employment and promotion requirements that create a discriminatory bar to opportunities. This Court has never read § 703(a)(2) as requiring the focus to be placed instead on the overall number of minority or female applicants actually hired or promoted.
Teal, 457 U.S. 440, 450, 102 S.Ct. 2525, 73 L.Ed.2d 130 (1982). In Teal, a pass-fail test had an adverse impact on minorities but, due to an affirmative action program, no adverse impact on promotion rates. Id. at 443-44, 102 S.Ct. 2525. The Supreme Court rejected the agency’s “bottom-line” defense, admonishing that “[t]he suggestion that disparate impact should be measured only at the bottom line ignores the fact that Title VII guarantees these individual respondents the opportunity to compete equally with white workers on the basis of job-related criteria.” Id. at 451, 102 S.Ct. 2525. In other words, “individual components of a hiring process may constitute separate and independent employment practices subject to Title VII even if the overall decision-making process does not disparately impact the ultimate employment decisions involving a protected group.” Bradley, 443 F.Supp.2d at 158-59. Under the progeny of Teal, even in the absence of adverse impact on promotion rates, an exam can lead to liability for an employer if it functions as “a gateway that has a disparate impact on minority hiring.” Id. at 159. Promotion rates — the “bottom line” — are thus not the only relevant inquiry in this Court’s disparate impact analysis.
The 2005 and 2008 exams served two functions: they were used as pass-fail hurdles and they accounted for 80% of the final score that determined candidates’ rank on an eligibility list from which they were promoted in rank order. Under such a scheme, this Court cannot rule that passing rates and average scores are irrelevant. Indeed, the Second Circuit has ruled that where an exam is used as both a pass-fail hurdle as well as a mechanism for ranking candidates for a promotion (as is the case here), courts should consider the disparate impact in the pass-fail rates, as well as the placement of ethnic groups on the ranking list. Waisome v. Port Auth. of New York & New Jersey, 948 F.2d 1370, 1377 (2d Cir.1991). The average scores and passing rates are relevant to the Court’s determination of whether the Plaintiffs have met their burden on prong one.
The City argued that the timing of promotions is largely determined by the timing of vacancies, so delays are more relevant to the damages phase of the litigation than to the liability phase. See Post-Trial City’s Proposed Findings 78. Neither argument persuades the Court to exclude delays in promotion from its analysis. The first argument falls flat because promotion rates themselves are determined, at least in part, by vacancies, and the City nowhere argues that promotion rates are irrelevant. Regarding the damages argument, the Court acknowledges that the delay in promotion will be relevant at the damages phase of the litigation, but it can also constitute disparate impact in the form of loss of pay, benefits, and seniority. See Bradley, 443 F.Supp.2d at 168 (ranking by exam score disproportionately precluded minority candidates from earning an earlier promotion, thus constituting an adverse impact); Guinyard v. City of New York, 800 F.Supp. 1083, 1089 (E.D.N.Y.1992) (same). This comports with common sense, with case law, and with the spirit of Teal — that employers may not circumvent Title VII “by merely showing that eventually they may hire some members of the disadvantaged minority group.” Bridgeport Guardians, Inc. v. City of Bridgeport, 933 F.2d 1140, 1147-48 (2d Cir.1991) (citing Teal, 457 U.S. at 455-56, 102 S.Ct. 2525).
The Court will therefore consider all of the factors that Dr. Wiesen statistically analyzed: promotion rates, pass-fail rates, average scores, and delays in promotion.
2. To Aggregate or Not to Aggregate?
For two of the four aspects of the promotional procedure, promotional rates and passing rates, Dr. Wiesen aggregated the data for the 2005 and 2008 exams, Wiesen Report 10,13, properly taking care to account for people who took both exams. Id. at 5. Dr. Wiesen stated that aggregation yields a “more powerful statistical test [because] you have a larger sample size and you see ... the big picture.” 12/15/14 Tr. 83:10-17. See also Wiesen Report 5.
Dr. Jacobs, a colleague of Dr. Silva’s, argued in a pretrial affidavit that aggregation was inappropriate. He posited that aggregation “only increases the sample size without a strong underlying logic as to the appropriateness of treating candidates from 2005 and 2008 as competing for the same jobs.” Aff. Rick R. Jacobs, PhD (“Jacobs Aff.”) ¶ 15, ECF No. 93. Dr. Silva opined in his expert report that aggregating the data between 2005 and 2008 is inappropriate because of Simpson’s Paradox, a phenomenon by which an “effect exists in two separate data sets but disappears when the data sets are combined or vice versa.” Silva Report 10. During trial, Dr. Silva testified that aggregation creates a risk of distortion, even if Simpson’s Paradox is not present. 12/18/14 Tr. 33:6-19, ECF No. 165.
Dr. Wiesen does not offer a sound basis for aggregating the data, other than its favoring the case of the party who hired him. He implies in his report that he only aggregated data if his original analysis did not produce statistically significant and practically important numbers for individual exam years. See Wiesen Report 16. In other words, he aggregated when the results using individual exams were not strong enough to help the Plaintiffs. The Court is not persuaded that this provides a sufficient basis to aggregate two data sets.
This Court rules that aggregation is inappropriate in this case. The 2005 and 2008 exams presented different questions, had different mean scores, and tested different candidates. 12/16/14 Tr. 129-131, ECF No. 162. Moreover, the Court has already ruled that the 2005 exam is not actionable; combining the 2005 scores with the 2008 scores would allow the Plaintiffs to circumvent this ruling. It would also raise so-called slippery slope concerns: why not aggregate with exams from the 1990s? Why not the 1980s? The Court sees no legitimate basis for aggregating the statistics on the facts of this case and therefore declines to do so.
3. One-Tailed vs. Two-Tailed
The next dispute the Court must resolve involves statistical methodology. It is a familiar one in disparate impact cases: whether to use a one- or two-tailed test when testing for statistical significance.
Drs. Wiesen and Silva both used “Fisher Exact Tests” to compare the exam results for candidates who are members of a minority group and white candidates. 12/17/14 Tr. 148:8-20. Fisher Exact Tests produce a bell curve with a tail on either end representing the lowest probability events. Id. at 149:24-150:9. When conducting a Fisher Exact Test, one can use a one-tailed or a two-tailed test. The terms “one-tailed” and “two-tailed” reflect whether statistical significance is determined from one or both the tails of the sampling distribution.
A two-tailed test assumes that any result could come from the test: in this case, in determining whether a given result was due to random chance, a two-tailed test would entertain three possibilities: that minorities would outperform non-minorities, non-minorities would outperform minorities, or that their performances would be equal. 12/17/14 Tr: 150:10-19. In contrast, a one-tailed test assumes only two possibilities: in this case, the performance between the groups was equal, or minorities performed worse on the test than non-minorities. Id. at 150:20-151:1; see Palmer v. Shultz, 815 F.2d 84, 94-95 (D.C.Cir.1987).
In Dr. Wiesen’s original report, he conducted all of his analyses using a two-tailed test, stating that although the one-tailed approach is more logically defensible, the two — tailed approach is more conservative. Wiesen Report 7 n.4; 12/16/14 Tr. 119:25-120:2. Subsequent to his initial report, and before the City offered its expert report, two additional sergeants were promoted to lieutenant, one white and one black. 12/15/14 Tr. 108:19-109:3.
When analyzing the data with the two new hires using a two-tailed test, Dr. Jacobs, the City’s expert, concluded that the adverse impact for the promotion rates for the 2008 exam was not statistically significant. Jacobs Aff. ¶ 14. Dr. Jacobs found a p-value of .052, a hair above the .05 threshold. Id. ¶ 10. Dr. Silva arrived at the same statistical conclusion. Silva Report 8.
Dr. Wiesen acknowledged this lack of statistical significance using a two-tailed test to examine the new data set. Second Aff. Joel P. Wiesen (“Wiesen Second Aff.”) ¶ 6, ECF No. 103. He subsequently switched his analysis from a two-tailed test to a one-tailed test, defending this approach by explaining that the question asked in this litigation is whether there was adverse impact on minorities, and thus, the “one-tailed test is more appropriate.” 12/16/14 Tr. 122; Wiesen Second Aff. ¶¶ 1, 7. Dr. Wiesen concluded that, even accounting for the two new hires, the “proper statistical conclusion is that there was adverse impact in promotions, both for the 2008 and the 2005 exams.” Id. ¶ 10.
The City bristled at Dr. Wiesen’s flip-flop from the two-tailed test to the one-tailed test, arguing that the two-tailed test is the appropriate one because it “accepts it is possible on a promotional examination that minorities could outscore non-minorities or that non-minorities could outscore minorities[J” Post-Trial City’s Proposed Findings 41. Dr. Silva argued that a one-tailed approach was inappropriate because it “implies that there is no chance that the direction of a promotion rate difference will ever favor minorities,” which contradicted Dr. Silva’s professional experience. Silva Report 2-3.
Whether to use a one-tailed or a two-tailed test is a common point of contention in disparate impact cases. Title VII defendants often argue that both minority and majority groups are protected from discrimination, and “it is therefore inequitable to disregard the probability of outcomes that may favor either group.” Palmer, 815 F.2d at 95. Defendants are also aware, of course, that it is easier for plaintiffs to prove significance and thus disparate impact with a one-tailed test. See, e.g., Brown v. Delta Airlines, Inc., 522 F.Supp. 1218, 1228 n. 14 (S.D.Texas 1980).
The weight of the case law appears to favor two-tailed tests. In Palmer, the D.C. Circuit Court of Appeals favored the two-tailed test for Title VII cases, 815 F.2d at 95-96, and that court continues to do so today. Csicseri v. Bowsher, 862 F.Supp. 547, 564 (D.D.C.1994) aff'd, 67 F.3d 972 (D.C.Cir.1995). Other courts have agreed. See, e.g., Dicker v. Allstate Life Ins. Co., No. 89 C 4982, 1997 WL 182290, at *41 (N.D.Ill. Apr. 9, 1997) (two-tailed is especially appropriate in disparate impact claims involving facially neutral employment selection procedures).
Courts recognize, however, that a one-tailed test can be appropriate in some Title VII circumstances, such as when “one population is consistently over-selected over another.” Stender v. Lucky Stores, Inc., 803 F.Supp. 259, 323 (N.D.Cal.1992); see Brunet v. City of Columbus, 642 F.Supp. 1214, 1230 (S.D.Ohio 1986). rev’d on other grounds, 1 F.3d 390 (6th Cir.1993) (stating that the one-tailed test is appropriate where “the raw numbers indicate that women are selected at a lesser rate than men. In these circumstances, the question being asked is whether this apparent difference is real or a statistical artifact.”). A one-tailed test can also be appropriate if “[t]here is little chance, from a facial review of the evidence, that applicants” in the plaintiff class were “treated statistically better” than those in the other group. Csicseri, 862 F.Supp. at 564-65. The First Circuit has overtly avoided choosing between the two. Jones, 752 F.3d at 43 n. 5.
There are two good arguments for using a one-tailed test in this case. The first is broader, relying on current inequalities in our society to put this case in context, and the second is narrower, implicating the history between this particular defendant, and this particular class of plaintiffs.
Experts for both sides testified to the phenomenon of written multiple-choice tests producing high levels of adverse impact on minority candidates. 12/15/14 Tr. 113-114; 12/18/14 Tr. 45-46, 70; Lopez, 07/13/10 Tr. 82-85; Lopez, 07/14/10 Tr. 43-48, 55, 59-60; Lopez, 07/26/10 Tr. 30; L^ pez, 09/15/10 Tr. 58-59; Lopez, 09/16/10 Tr. 110. Judge O’Toole in Lopez also recognized this phenomenon. Lopez, slip op. at 14. Experts in the seminal case of Ricci v. DeStefano similarly testified. 557 U.S. at 570, 572, 129 S.Ct. 2658 (2009). Why the difference in performance between members of minority groups and white applicants on these tests? Neither expert addressed this question, other than Dr. Wiesen suggesting that “there are ... dozens of reasons, each of which account[ ] for just a small amount of that difference.” 12/15/14 Tr. 114:10-13. Without wading into social-scientific debates, the Court agrees with Dr. Wiesen that there are likely several factors driving this disparity (e.g., legacies from historical discrimination, economic inequality, current explicit and implicit biases). Whatever the causes, the so-called “achievement gap” is real, and might recommend adopting a one-tailed test as more rooted in reality.
In addition, Boston has a history of discrimination against minority police officers. See Boston Police Superior Officers Fed’n v. City of Boston, 147 F.3d 13, 20 (1st Cir.1998). This history might suggest that viewing white police-officer-applicants as equally likely to over- and under-perform minority applicants on an exam developed by Boston is overly idealistic.
The Court is hesitant, however, to analyze a disparate impact case, a case in which no one has accused the Department of any malfeasance or conscious desire to discriminate, under the assumption that minorities could only underperform and not also possibly outperform their white peers on a promotional exam. Cf. Parents Involved in Cmty. Sch. v. Seattle Sch. Dist. No. 1, 551 U.S. 701, 748, 127 S.Ct. 2738, 168 L.Ed.2d 508 (2007) (Roberts, C.J.) (plurality opinion) (“The way to stop discrimination on the basis of race is to stop discriminating on the basis of race.”); id. at 789, 127 S.Ct. 2738 (Kennedy, J., concurring) (suggesting that “facially race-neutral means” or considering race as part of “a more nuanced, individual” evaluation of applicants is permissible, whereas broad-based classifications based on it are subject to strict scrutiny). Moreover, although the First Circuit and the Supreme Court have remained relatively quiet on the issue of one versus two tails, the Supreme Court has previously suggested that a protected class’s treatment that falls two or three standard deviations beyond the mean is evidence of disparate impact. See Castaneda v. Partida, 430 U.S. 482, 496 n. 17, 97 S.Ct. 1272, 51 L.Ed.2d 498 (1977). This requirement is closer to that of a two-tailed test (which requires a result be more than 1.96 standard deviations from the mean to reach statistical significance) than of a one-tailed test (1.65). The Court is inclined to agree with the City that a two-tailed test is the more appropriate methodology for evaluating statistical significance in this case.
The debate over the one versus two-tailed test, while important and fascinating, is not dispositive in this case. As the Court explains in the next section, the Plaintiffs have met their burden of demonstrating disparate impact, regardless of whether the Court accepts the one or two-tailed approach for the 2008 promotion rates.
4. Conclusions Regarding Disparate Impact (Prong 1)
In his expert report, Dr. Wiesen concluded that:
• The adverse impact ratio for promotions from the 2008 exam was .45 (meaning that minorities were promoted at a rate of .45 as compared to the rate at which whites were promoted, well below the % rule, which would be satisfied by a rate of up to .80), Wiesen Second Aff. 7, and .38 for the 2005 exam. Wiesen Report 8. He concluded that the p-value for the 2008 promotion rates was .052 for a two-tailed test, and .027 for a one-tailed test. Wiesen Second Aff. 7. The ratio for the 2005 exam was not statistically significant. Wiesen Report 9.
• The adverse impact ratio for passing scores was .74 for the 2008 exam and .57 for the 2005 exam, both of again satisfy the % rule, and both of which were statistically significant at a p- . value of .004 for the 2008 exam and .00005 for the 2005 exam, using a two-tailed test. Wiesen Report 11-12; 12/15/14 Tr. 88:1-4.
• The adverse impact for average scores was 6.6 points on the 2008 exam (meaning that minority applicants scored, on average, 6.6 points lower than white applicants) and 8.8 points on the 2005 exam, both of which were “highly statistically significant” at a p-value for the 2008 exam of .0015 and for the 2005 exam of .00003 (both using a two-tailed test). Wiesen Report 3,15-16.
• The adverse impact for delay in promotions for the