Citations

Full opinion text

MEMORANDUM AND OPINION

LEE H. ROSENTHAL, District Judge.

This Title VII disparate-impact suit challenges the City of Houston’s system for promoting firefighters to the positions of captain and senior captain. Historically, the City has promoted firefighters based on their years of service with the Houston Fire Department (“HFD”) and their scores on a multiple-choice exam. The format and general content of that exam are set out in the Texas Local Government Code (“TLGC”) and in the collective bargaining agreement (“CBA”) between the City and the firefighters’ union, the Houston Professional Fire Fighters Association (“HPFFA”). Seven black firefighters sued the City, alleging that the promotional exams for the captain and senior-captain positions were racially discriminatory, in violation of the Fourteenth Amendment, 42 U.S.C. § 1981, and Title VII, 42 U.S.C. § 2000e-2. After mediation in February and March 2010, the City and the seven firefighters reached a settlement that included a proposed consent decree. The decree would require the City to implement changes to the captain and senior-captain promotion exams in two phases. The first phase required minor changes to the November 2010 captain exam. The second phase required more significant changes to the May 2011 senior-captain exam that would apply to future captain and senior-captain exams. The HPFFA intervened in the lawsuit and objected because the proposed consent decree changed the exams in ways inconsistent with the TLGC and the CBA. The HPFFA contended that the City and the plaintiffs had not shown discrimination that would permit this court to approve the consent decree.

This court bifurcated the proceedings to resolve the HPFFA’s objections. The first stage addressed the narrow set of changes proposed for the November 2010 captain exam. The second stage addressed the broader changes proposed for subsequent captain and senior-captain exams. In the first stage, this court, with the HPFFA’s agreement, approved changes to the November exam. This opinion addresses the HPFFA’s objections to the proposed changes to the future captain and senior-captain exams.

The HPFFA vigorously objects that the proposed changes involve “a far-reaching and wholesale restructuring of the entire promotional process that goes beyond anything plaintiffs have even alleged in this lawsuit” and “bypass both the long-established protections of state law and the union’s protected role in being the sole, collective voice for the city’s firefighters.” (Docket Entry No. 89, at 8). The City and the seven individual plaintiffs acknowledge that the changes are far-reaching but argue that they are needed to comply with federal antidiscrimination law. They argue that the current exam system is not “job related for the position[s] in question and consistent with business necessity” as required by 42 U.S.C. § 2000e-2(k)(1)(A)(i).

This court held an evidentiary hearing to consider the proposed changes to the May 2011 senior-captain exam and to future exams. The summary of the evidence shows the welter of expert opinions the parties presented on whether the existing format and content of the City’s promotion exams for the captain and senior-captain positions have a disparate impact on African-American candidates; whether the existing exams are reliable and valid measures of the knowledge and qualities relevant to the promotion decisions; whether the existing exams are reliable and valid ways to compare candidates; and whether the proposed changes to the exams will provide reliable and valid exams and address disparate impact. The experts’ testimony and submissions left the court with a sense of disquiet about the opinions expressed. The science of testing to measure and compare promotion-worthiness is admittedly imperfect. The expert witnesses, particularly for the City, acknowledged some errors and some incomplete aspects of their work in designing and administering the promotion exams. At best, all the witnesses’ opinions amount to uncertain efforts to gauge how well different exam approaches measure, compare, and predict job performance. The analytical steps required by the applicable legal standards must be approached with a recognition of the limits of the expert testimony.

At the same time, courts clearly lack expertise in the area of testing validity. ‘“The study of employment testing, although it has necessarily been adopted by the law as a result of Title VII and related statutes, is not primarily a legal subject.’ Because of the substantive difficulty of test validation, courts must take into account the expertise of test validation professionals.” Gulino v. N.Y. State Educ. Dep’t, 460 F.3d 361, 383 (2d Cir.2006) (quoting Guardians Ass’n of N.Y.C. Police Dep’t, Inc. v. Civil Serv. Comm’n of City of N.Y., 630 F.2d 79, 89 (2d Cir.1980)). The combination of the lack of judicial expertise in this area and the limits of the expertise of those who do have training and experience support a cautious and careful judicial approach.

Based on the parties’ filings, the evidence, and the applicable law, this court finds that the City and the seven individual plaintiffs have shown that the captain and senior-captain exams violate Title VII. But this court also finds that some of the changes in the proposed consent decree violate the CBA and TLGC and that the City and the plaintiffs have not shown that all these changes are necessary to comply with Title VII. Based on these findings and conclusions, the proposed consent decree is accepted in part and denied in part. The use of situational-judgment questions and an assessment center are justified by the record evidence and are job-related and consistent with business necessity. But other parts of the proposed modified consent decree violate the TLGC and CBA, and the City and the plaintiffs have not shown that they are tailored to respond to the disparate impact alleged. Using the parties’ descriptions of the proposed changes, the provisions that violate the TLGC and CBA without the necessary justification in the record, and which this court does not accept, are as follows:

2. Job-Knowledge Written Test

— Pass/fail test (test designer to determine cut-off score)

— No rank-order list

— Test designer may elect not to use any written job-knowledge cognitive test

3. Scenario-Based Computer-Objective Test

— Rank-order list from which all intended promotions to assessment center will be made

4. Assessment Center

— Rank-order list

•— Sliding bands based on test accuracy as determined by consultant

— Fire Chief will document reasons for the selection of each candidate within bands

■— No Rule of Three

(Docket Entry No. 69-2, at 29).

The reasons for finding these aspects of the proposed changes to the promotion examinations invalid, and the remaining aspects supported by the record and the applicable law, are explained below. This opinion first describes the promotion system in place before any changes; reviews the expert and other evidence relevant to assessing disparate impact; and analyzes whether, under the applicable law, the proposed settlement is tailored to remedying the disparate impact that is shown. A hearing is set for February 21, 2012, at 1:30 p.m. to address the issues that remain to be resolved and a timetable for doing so.

Finally, because the terminology used by the HFD and the industrial psychologists who served as experts in this case produced a number of acronyms and abbreviations, a list of the most commonly used is attached to this Memorandum and Opinion.

I. Background

A. The Houston Fire Department

The HFD has approximately 4,000 employees involved in firefighting. Ninety percent are in the Emergency Operations Division (“EOD”). Half of the EOD employees are at the “firefighter” level and perform “task-level jobs” such as retrieving and using fire hoses. The next rank above firefighter is “engineer operator” (“EO”). In addition to performing firefighters’ tasks, EOs drive fire trucks and HFD ambulances and operate ladders and pumps. Firefighters outnumber EOs two-to-one. (Evidentiary Hr’g Tr. 115, Docket Entry No. 130).

Captains are ranked immediately above EOs. HFD captains are the “first line of supervisor position[s] in the fire department.” A captain supervises the operation of fire engines, which are smaller fire trucks that carry hoses and pump water. Each HFD fire station has at least one fire engine and one captain. A captain supervises an EO and two firefighters assigned to an engine. When a captain misses a day of work, an EO may “ride up” and perform the absent captain’s job duties.

Senior captains are ranked immediately above captains. A senior captain supervises the operations of “ladder trucks,” which are large fire trucks with aerial ladders. Only half of the City’s fire stations have a ladder truck with a senior captain in addition to a fire engine and captain. A senior captain may supervise up to eight firefighters, including EOs. When a senior captain misses a day of work, a captain may “ride up.”

During a fire emergency, a district chief — ranked above senior captain — is responsible for developing the firefighting strategy. The district chief may decide, for example, whether firefighters will enter a burning building and address a fire directly or instead contain it by protecting adjacent buildings. Senior captains may participate in the strategy development, but the district chief bears ultimate responsibility. Once a strategy is set, the senior captain and captain are responsible for implementing it. Usually a senior captain and the ladder-truck crew are responsible for forcible entries into a building to ventilate it, for attempting rescues, and for creating ways for other firefighters to enter. The captain and the fire-engine crew are usually responsible for locating, confining, and extinguishing fires. (Id. at 116—20).

To summarize the promotional system that is discussed in detail below, promotion from EO to captain and from captain to senior captain depends largely on a candidate’s score on a multiple-choice test. Any person meeting the experience requirement can take the test. An EO can apply for captain after four years in the fire department. A captain can apply for senior captain after two additional years of service as a captain. Tex. Loc. Gov’t Code § 143.028(a). A candidate’s length of service with the HFD will add some points to the test score, but the test score largely determines promotion.

The City makes promotion decisions based on a rank-order list of the candidates’ test points added to their length-of-service points. For each captain or senior-captain position available during the three years after the exam, the top three candidates’ names and scores are submitted to the HFD fire chief. The presumption is that the fire chief will select the candidate with the highest test score. If the fire chief selects the second or third highest scoring candidate, the chief must explain his reasons in writing. If a candidate is not selected for promotion within the three-year period, the candidate must retake the exam. These promotional procedures for the captain and senior-captain positions are based on the TLGC and the CBA.

1. The Texas Local Government Code

The City of Houston adopted the Fire Fighter and Police Civil Service Act (“CSA”), codified as Chapter 143 of the TLGC, on January 31, 1948. The CSA’s “fundamental principle” is ensuring that public-service appointments and promotions are made “according to merit and fitness, ascertained by competitive examinations.” Klinger v. City of San Angelo, 902 S.W.2d 669, 671 (Tex.App.-Austin 1995, writ denied). The Texas legislature passed the CSA “to secure efficient fire and police departments composed of capable personnel who are free from political influence.” Tex. Loc. Gov’t Code § 143.001(a). The TLGC requires a test-based promotional system for firefighters. Section 143.021(c) states that positions within fire departments must be filled “from an eligibility list that results from an examination held in accordance with [the CSA].” The TLGC contains detailed rules describing the eligibility list, exam, and procedure for selecting firefighters for promotion.

The promotional process begins when a city posts notice of an upcoming examination. Municipalities like the City of Houston, with populations greater than 1.5 million, must post notice in plain view on a bulletin board located in’City Hall’s main lobby and in the Firefighters’ and Police Officers’ Civil Service Commission office by the 90th day before the date a promotional exam is scheduled. This 90-day notice must show the positions to be filled and the date, time, and place of the exam. Tex. Loc. Gov’t Code § 143.107(a). The 90-day notice must also list the sources from which the exam questions are taken. Id. § 143.029(a). By the 30th day before the date a promotional exam is scheduled, another notice must be posted in the same locations. Id. § 143.107(b). The 30-day notice must state the number of newly created positions and may “include the name of each source used for the examination, the number of questions taken from each source, and the chapter used in each source.” Id. § 143.029(c).

The TLGC requires that the test be in writing and forbids tests that “in any part consist of an oral interview.” Id. § 143.032(c). The questions must “test the knowledge of the eligible promotional candidates about information and facts.” Id. § 143.032(d). The information-and-fact questions “must” be based on:

(1) the duties of the position for which the examination is held;

(2) material that is of reasonably current publication and that has been made reasonably available to each member of the fire or police department involved in the examination; and

(3) any study course given by the departmental schools of instruction.

Id. The questions must also be taken from the sources identified in the posted notices. Id. § 143.032(e). Finally, the “examination questions must be prepared and composed so that the grading of the examination can be promptly completed immediately after the examination is over.” Id. § 143.032(f).

The exam grade determines whether the candidate will be placed on a promotion-eligibility list. Grading begins as soon as an individual candidate completes the exam. The candidate may remain present during the grading. Id. § 143.033(a). The multiple-choice exam score is based on a maximum grade of 100 points and is determined by the correctness of the answers to the questions. Id. § 143.033(c). Each candidate also receives one point for each year of seniority, with a maximum of 10 points. Id. § 143.033(b). In municipalities like Houston, a candidate must score at least 70 points on the exam to be eligible for promotion. Id. § 143.108(a).

All scores must be posted within 24 hours of the exam. Id. § 143.033(d), Each candidate may see the answers, grading, and source materials after the exam and can appeal a score within 5 days. Id. § 143.034(a). The City has 60 days to decide the appeal. Id. § 143.1015(a). A candidate who appeals is entitled to a hearing. Id. § 143.1015(b).

Once the scores are finalized, all candidates who pass are listed in rank order on a promotion-eligibility list. See id. § 143.021(c); id. § 143.108(f). When vacancies occur, the names of the three persons with the highest scores for the position are certified and provided to the head of the department with the vacancy. Id. § 143.036(b). This is known as the “Rule of Three.” The TLGC provides that “[u]n-less the department head has a valid reason” for not doing so, “the department head shall appoint the eligible promotional candidate having the highest grade on the eligibility list.” Id. § 143.036(f). If the candidate with the highest grade is not selected, the department head must personally discuss the reason with that candidate and file a written explanation. Id.

2. The Collective Bargaining Agreement

Texas law establishes firefighters’ right to collective bargaining. Tex. Loc. Gov’t Code § 174.002(b) (“The policy of this state is that fire fighters and police officers, like employees in the private sector, should have the right to organize for collective bargaining, as collective bargaining is a fair and practical method for determining compensation and other conditions of employment. Denying fire fighters and police officers the right to organize and bargain collectively would lead to strife and unrest, consequently injuring the health, safety, and welfare of the public.”); id. § 143.204(a) (stating that a firefighter association submitting a petition signed by the majority of the paid firefighters in the municipality “may be recognized ... as the sole and exclusive bargaining agent for all of the covered fire fighters”). The HPFFA is the sole and exclusive bargaining agent for the City’s firefighters.

The TLGC allows the City and the HPFFA to enter into a written agreement binding when ratified by both. Id. § 143.206(a). Such an agreement can supersede the TLGC’s provisions “concerning wages, salaries, rates of pay, hours of work, and other terms and conditions of employment to the extent of any conflict with the [written agreement].” Id. § 143.207(a). The agreement “preempts all contrary local ordinances, executive orders, legislation, or rules adopted by the state.” Id. § 143.207(b).

The 2009-2010 CBA between the City of Houston and the HPFFA made few departures from the TLGC’s exam provisions. Like the TLGC, the CBA required a grade of at least 70% for promotion eligibility. The CBA specified that the test must consist of “not less than 100 and not more than 150 questions.” (Docket Entry No. 69-6, at 20). Unlike the TLGC, the CBA allowed only a .5-point increase in the score for each year of service, with a maximum of 10 points. The CBA also allowed a .5-point increase for each year of service for certain ranks. For example, an engineer or operator applying to be a captain is awarded .5 points for each year of service as an engineer. (Id.). Aside from these changes, the 2009-2010 CBA provided that the TLGC “remain[s] in full force in the same manner as on the date [the CBA] became effective.” (Id. at 13).

B. Title VII

“Congress enacted Title VII of the Civil Rights Act of 1964, 42 U.S.C. § 2000e et seq., to assure equality of employment opportunities by eliminating those practices and devices that discriminate on the basis of race, color, religion, sex, or national origin.” Alexander v. Gardner-Denver Co., 415 U.S. 36, 44, 94 S.Ct. 1011, 39 L.Ed.2d 147 (1974). Title VIPs prohibitions include using “a particular employment practice that causes a disparate impact on the basis of race, color, religion, sex, or national origin” unless the employment practice “is job related for the position in question and consistent with business necessity.” 42 U.S.C. § 2000e-2(k)(1)(A)(i). The plaintiffs alleged that the City’s promotional procedures for captain and senior captain violated Title VTI’s disparate-impact provision.

“Congress intended voluntary compliance to be the preferred means of achieving the objectives of Title VII.” Local No. 93, Int’l Ass’n of Firefighters, AFL-CIO v. City of Cleveland, 478 U.S. 501, 515, 106 S.Ct. 3063, 92 L.Ed.2d 405 (1986). To help employers comply with Title VII, Congress authorized the Equal Employment Opportunity Commission (“EEOC”) to issue compliance guidelines (the “Guidelines”). The Guidelines “are not administrative regulations promulgated pursuant to formal procedures established by the Congress. But ... they do constitute ‘(t)he administrative interpretation of the Act by the enforcing agency,’ and consequently they are ‘entitled to great deference.’ ” Albemarle Paper Co. v. Moody, 422 U.S. 405, 431, 95 S.Ct. 2362, 45 L.Ed.2d 280 (1975) (citing Griggs v. Duke Power Co., 401 U.S. 424, 433-34, 91 S.Ct. 849, 28 L.Ed.2d 158 (1971)).

The Guidelines require employers who make promotional decisions based on test scores to maintain records of tests and test results. 29 C.F.R. § 1607.4(A). The Guidelines’ rule of thumb for determining disparate impact is the “4/5 Rule.” Under this Rule:

A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact, while a greater than four-fifths rate will generally not be regarded by Federal enforcement agencies as evidence of disparate impact.

Id. § 1607.4(D). There are exceptions to the 4/5 Rule. The Guidelines state:

Smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms or where a' user’s actions have discouraged applicants disproportionately on grounds of race, sex, or ethnic group. Greater differences in selection rate may not constitute adverse impact where the differences are based on small numbers and are not statistically significant, or where special recruiting or other programs cause the pool of minority or female candidates to be atypical of the normal pool of applicants from that group. Where the user’s evidence concerning the impact of a selection procedure indicates adverse impact but is based upon numbers which are. too small to be reliable, evidence concerning the impact of the procedure over a longer period of time and/or evidence concerning the impact which the selection procedure had when used in the same manner in similar circumstances elsewhere may be considered in determining adverse impact. Where the user has not maintained data on adverse impact as required by the documentation section of applicable guidelines, the Federal enforcement agencies may draw an inference of adverse impact of the selection process from the failure of the user to maintain such data, if the user has an underutilization of a group in the job category, as compared to the group’s representation in the relevant labor market or, in the case of jobs filled from within, the applicable work force.

Id.

If analyzing an employer’s test results under the 4/5 Rule shows “that the total selection process for a job has an adverse impact, the individual components of the selection process should be evaluated for adverse impact.” Id. § 1607.4(C). The method for evaluating individual components is a “validity study.” Id. § 1607.3(A). The Guidelines describe three types of validity studies: criterion-related-validity studies; content-validity studies; and construct-validity studies. Id. § 1607.5(A). A criterion-related-validity study analyzes whether test results correlate to “criteria that [are] predictive of job performance.” Mark R. Bandsuch, Ten Troubles with Title VII and Trait Discrimination Plus One Simple Solution (A Totality of the Circumstances Framework), 37 Cap. U.L. Rev. 965, 1089 (2009). A content-validity study analyzes whether test results correlate to “the knowledge, skills, and abilities related to that job.” Id. A construct-validity study examines whether test results correlate to “general characteristics important to job performance.” Id.

The Guidelines also describe the evidence each type of validity study requires. A criterion-related-validity study requires “empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance.” 29 C.F.R. § 1607.5(B). A content-validity study requires “data showing that the content of the selection procedure is representative of important aspects of performance on the job for which the candidates are to be evaluated.” Id. A construct-validity study requires “data showing that the procedure measures the degree to which candidates have identifiable characteristics which have been determined to be important in successful performance in the job for which the candidates are to be evaluated.” Id.

One court has summarized the content-validity and criterion-validity methods for evaluating a promotion or other employment test, as follows:

[Ejmployers can establish job-relatedness by one of three methods, including “content validity,” which entails showing that the test measures the job or adequately reflects the skills or knowledge required by the job. A typing test for secretaries exemplifies this kind of approach. This method does not require empirical evidence, but instead “should consist of data showing that the content of the selection procedure is representative of important aspects of performance on the job.” 29 C.F.R. § 1607.5(B). In contrast, the “criterion related” approach evaluates whether a test is adequately correlated with future job performance and is constructed to measure traits thought to be relevant to future job performance. An IQ test is a typical criterion-related method. Unlike content validity, this ' method requires “empirical data demonstrating that the selection procedure is predictive of or significantly correlated with important elements of job performance.” 29 C.F.R. § 1607.5(B).

Banos v. City of Chicago, 398 F.3d 889, 893 (7th Cir.2005) (citations and internal quotations marks omitted).

Before conducting a validity study, an employer should conduct a “job analysis.” 29 C.F.R. § 1607.14(A). Each type of validity study requires a different type of job analysis. Criterion-related-validity studies require “reviewing job information to determine measures of work behavior(s) or performance that are relevant to the job or group of jobs in question.” Id. § 1607.14(B)(2). “These measures or criteria are relevant to the extent that they represent critical or important job duties, work behaviors or work outcomes as developed from the review of job information”; “[b]ias should be considered.” Id. Content-validity studies should include “an analysis of the important work behavior(s) required for successful performance and their relative importance and, if the behavior results in work product(s), an analysis of the work product(s). Any job analysis should focus on the work behavior(s) and the tasks associated with them.” Id. § 1607.14(C)(2). Construct-validity studies “should show the work behavior(s) required for successful performance of the job, or the groups of jobs being studied, the critical or important work behavior(s) in the job or group of jobs being studied, and an identification of the construct(s) believed to underlie successful performance of these critical or important work behaviors in the job or jobs in question.” Id. § 1607.14(D)(2). “Each construct should be named and defined, so as to distinguish it from other constructs.” Id.

If one or more validity studies produces evidence “sufficient to warrant use of the procedure for the intended purpose under the standard of these guidelines,” the promotional procedure is “properly validated.” Id. § 1607.16(X). But if no validity study produces sufficient evidence, an employer “should initiate affirmative steps to remedy the situation.” Id. § 1607.17(3). These steps, “which in design and execution may be race, color, sex, or ethnic ‘conscious,’ include, but are not limited to,” the following:

(a) The establishment of a long-term goal, and short-range, interim goals and timetables for the specific job classifications, all of which should take into account the availability of basically qualified persons in the relevant job market;

(b) A recruitment program designed to attract qualified members of the group in question;

(c) A systematic effort to organize work and redesign jobs in ways that provide opportunities for persons lacking “journeyman” level knowledge or skills to enter and, with appropriate training, to progress in a career field;

(d) Revamping selection instruments or procedures which have not yet been validated in order to reduce or eliminate exclusionary effects on particular groups in particular job classifications;

(e) The initiation of measures designed to assure that members of the affected group who are qualified to perform the job are included within the pool of persons from which the selecting official makes the selection;

(f) A systematic effort to provide career advancement training, both classroom and on-the-job, to employees locked into dead end jobs; and

(g) The establishment of a system for regularly monitoring the effectiveness of the particular affirmative action program, and procedures for making timely adjustments in this program where effectiveness is not demonstrated.

Id.

C. The Procedural History of this Case

On August 4, 2008, seven firefighters sued the City of Houston, alleging that the 2006 captain and senior-captain exams had a discriminatory effect on their promotion opportunities, in violation of § 1981 and 42 U.S.C. § 2000e-2. The seven individual plaintiffs contended that the 2006 exams had a disparate impact on the promotion of black firefighters to captain and senior-captain positions compared to white firefighters. Four plaintiffs — Dwight Bazile, Johnny Garrett, Trevin Hines, and Mundo Olford — were lieutenants denied promotion to captain. Three plaintiffs — George Runnels, Dwight Allen, and Thomas Ward— were captains denied promotion to senior captain. (Docket Entry No. 1).

The City and the plaintiffs settled. (Docket Entry No. 64). The HPFFA was not a party to the negotiations or settlement. The City agreed to promote Bazile, Olford, and Hines to captain; to promote Allen to senior captain; to allow Garrett to retire as a captain; and to allow Runnels and Ward to retire as senior captains. The City also agreed to pay each plaintiff backpay in amounts ranging from $376.80 to $23,075.46. (Docket Entry No. 69-2, at 2-8, 17-22, 26-27).

The settlement also contained a proposed consent decree to be submitted to the court for approval. (Id. at 9-10, 28-30). The decree required the City to implement changes to the captain and senior-captain exams in two phases. In the first phase, the City agreed to implement “modest” changes to the November 2010 captain exam. In the second phase, the City agreed to implement broader changes, beginning with the May 2011 senior-captain exam and applying to all future captain and senior-captain exams. The settlement agreement required the parties to give notice to the HPFFA of “this conceptual agreement” and to “meet in person or conference call [with the HPFFA] to explore potential adjustments of union suggestions prior to final settlement meeting.” (Id. at 10). The settlement agreement also required approval by the Houston City Council and by this court.

The parties notified this court of the settlement and their intent to file the proposed consent decree. Before filing the decree, the parties moved to join the HPFFA to the suit because the proposed changes to the promotion exams conflicted with the TLGC and the CBA. (Docket Entry No. 69). The HPFFA moved to intervene and asked this court to bifurcate review of the proposed consent decree. The first step would be to consider the HPFFA’s objections to the proposed changes to the November 2010 captain exam. The second stage would be to consider the HPFFA’s objections to the proposed changes to the subsequent senior-captain and later captain and senior-captain exams. This court granted the motion and entered a scheduling order. (Docket Entry Nos. 70 & 71).

1. The November 2010 Captain Exam

The HPFFA objected to certain proposed changes to the November 2010 captain exam. (Docket Entry No. 75). This court heard arguments and evidence on the objections on September 16, 2010. On the same date, and with the HPFFA’s agreement, this court found that the existing captain exam disparately impacted black firefighters and entered an order allowing the City to implement the consent decree provisions changing the November 2010 captain exam. (Docket Entry Nos. 82 & 85). The proposed consent decree described those changes to the 2010 captain exam, as follows:

Hybrid written examination

Content-validated job-knowledge portion of exam

— weighted to Houston Departmental material, plus

— carefully selected directly relevant test material,

— supported by job analysis and incumbent/supervisor feedback in direct interviews

— job analysis

Content-validated multiple-choice situational-judgment exam

— designed by an industrial/organization psychologist

— HFD departmental scenario based

— zero, partial, and full-credit options

— additional points to stay the same as under the Collective Bargaining Agreement in an effort to minimize adverse impact potential

— rank order for selection process by Chief

— Rule of Three for promotion

(Docket Entry No. 69-2, at 28-29).

The most significant change to the November 2010 captain exam was the inclusion of multiple-choice “situational-judgment” questions in addition to the “job-knowledge” questions used on previous exams. Situational-judgment questions present hypothetical situations encountered on the job and ask candidates how they would respond. By contrast, the job-knowledge questions that made up the previous captain exams, mandated by the TLGC, see Tex. Loc. Gov’t Code § 143.032(d), ask about facts related to the job, such as the content of applicable regulations or specific HFD policies and procedures.

This court’s order approved the inclusion of situational-judgment multiple-choice questions for the November 2010 captain exam based on a finding that “[t]he continued exclusive use of questions based on ‘fact’ and ‘information’ as stated in Local Government Code § 143.032(d) is likely to continue to result in adverse impact.” (Docket Entry No. 85, at 2). The situational-judgment questions included in the November 2010 captain exam were developed by industrial-psychology consultants selected by the City, the plaintiffs, and the HPFFA (the “consultants”). These consultants developed the questions by creating a job analysis for the captain position. See 29 C.F.R. § 1607.14(A) (describing a job analysis). To create the job analysis, the consultants interviewed “subject-matter experts” (“SMEs”); analyzed HFD materials related to the captain position, such as internal job descriptions and policies and procedures; and analyzed external source materials such as published industry standards and firefighting textbooks. The consultants interviewed both internal SMEs — incumbent HFD captains and their supervisors — and external SMEs — individuals with similar experience who did not work for HFD. Through the job analysis, the consultants identified the “knowledge, skills, abilities and other characteristics” (“KSAOs”) required for successful performance in the captain position and designed the situational-judgment questions to measure the identified KSAOs.

This court also approved an additional consent-decree provision inconsistent with the TLGC and CBA. The TLGC and CBA allow a promotional candidate to be present while the candidate’s exam is scored and require that scores be posted within 24 hours of the exam. Tex. Loc. Gov’t Code § 143.033(a), (d). The consent decree required an “item analysis” of the score before it was finalized. Item analysis requires the consultants to aggregate data related to each question — or “item”— to eliminate questions that did not reliably measure an individual candidate’s exam performance. Because item analysis requires collecting data from all promotional candidates’ exams and time to evaluate this data, the parties agreed to post the raw scores from the exam within 24 hours to meet the TLGC and the CBA requirements, but these raw scores would not be the final scores until the item analysis was completed. The promotional candidates would not remain throughout the item analysis.

Many of the TLGC and CBA requirements remained in place under the consent decree for the November 2010 captain exam. The consent decree still required job-knowledge questions. The exam that was administered contained 75 job-knowledge questions. The job analysis was included as a “source material” in the posted notices and made available before the exam. See Tex. Loc. Gov’t Code § 143.029(a) (requiring posting of source material). The consent decree also required a rank order of the candidates based on their exam scores and required that promotions be made according to the Rule of Three set out in the TLGC. See id. § 143.036(b), (f) (describing the Rule of Three).

The City administered the captain exam on November 17, 2010. The consultants conducted an item analysis after the exam. A panel consisting of representatives for the City, the individual plaintiffs, and the HPFFA met to review scoring. Initially, based on the item analysis, the consultants recommended giving candidates full credit for seven job-knowledge questions and for fourteen situational-judgment questions, effectively eliminating those questions as a way to differentiate among the candidates. In addition, the City’s internal SMEs recommended giving full credit for one job-knowledge question and for six situational-judgment questions. The panel agreed with the SMEs’ recommendation. (Docket Entry No. 94, at 2; Docket Entry No. 94-1 at 2-3). The panel also agreed that a candidate’s score on the job-knowledge portion and the situational-judgment portion would be weighted equally in calculating the final score. (Docket Entry No. 94, at 2). Based on these decisions, a rank-order results list for the exam was created. Some discrepancies emerged in the statistical calculations and the consultants recommended a credit adjustment for additional questions. On January 10, 2011, the City submitted a revised rank-order list of candidates who passed the exam. (Docket Entry No. 104, at 2).

On January 12, 2011, the City advised this court that there were more than 200 appeals by the promotional candidates. On January 14, the City moved for additional time to finalize the scores and rank-order list. (Docket Entry Nos. 110 & 112). The City sought more time- than the 60 days the TLGC allowed to decide whether to sustain the appeals. The HPFFA did not object to the request, and this court granted the motion. (Docket Entry No. 117). This court has not been updated on the status of the appeals or on promotions to captain under the November 2010 exam.

2. The May 2011 Senior-Captain Exam and Future Exams

The HPFFA filed its objections to the proposed changes to the May 2011 senior-captain exam and to subsequent captain and senior-captain exams. (Docket Entry No. 89). The proposed changes are described as follows:

1. Officer Development Program

— 2 years in grade

— Educational Courses

• Officer Development I, II

• Available online at stations for all to participate

2. Job-Knowledge Written Test

— New job analysis

— Designed by an industrial/organizational psychologist

— HFD departmental based

— Pass/Fail Exam (test designer to determine cut off score)

— No rank-order list

— Test designer may elect not to use a written job-knowledge cognitive test

3. Scenario-Based Computer-Objective Test

— Situational-judgment exam -with HFD departmental scenarios

— Computer simulations, such as in-basket exercises or incident-scenario judgment

— Zero, partial, and full credit answers may be used

— Same responses receive same points

— Scored test

— Rank-order list from which all of intended promotions will proceed to assessment center

4. Assessment Center

— Rank-order list

— Sliding bands based on test accuracy as determined by consultant

— Fire Chief will document reasons for selection of each candidate within bands

— No Rule of Three

— Two-year eligibility list

— Will be applied to the May 2011 Senior-Captain exam

4a. [Blank]

— Additional points in effort to minimize disparate impact potential

■ All points stay the same as the Collective Bargaining Agreement until May 1, 2011

■ The City will propose and will bargain for a point system which does not exceed the following points:

• 10 points for seniority

• 5 points for time in rank

• Education/Certification

O 1 point-intermediate certification

O 2 points-Advanced certification

O 3 points-Masters certification or Associates degree

O 4 points-Bachelor’s degree

O 5 points-Masters degree

(Docket Entry No. 69-2, at 29-30).

The parties agree that many of these proposed changes violate the TLGC and the CBA. Under the consent decree, the test designer “may” elect to use a “written job-knowledge test” depending on the job analyses for the captain and senior-captain positions. Whether such questions are included depends on the importance of the “knowledge” component compared to the skills, abilities, and other characteristics identified for the positions. If the test designer elects to use written job-knowledge questions, they are scored on a pass/ fail basis. Only candidates who get a passing score remain promotion-eligible. A candidate’s specific score on the job-knowledge test is otherwise irrelevant. The score is not used to produce a rank-order list and the Rule of Three is abandoned as to this part of the promotional process.

The remaining two parts of the captain and senior-captain exam are not questions based exclusively on facts and information. One part is a scenario-based eomputer objective test. The second part uses an assessment center.

The scenario-based computer objective test is based on situational-judgment concepts, using a computer to present hypothetical situations that captains and senior captains would likely encounter on the job. One type of situational-judgment question identified in the consent decree is an “in-basket exercise.” In such an exercise, a candidate is given documents or other information creating a hypothetical fact pattern and is asked to analyze or describe a response. An in-basket exercise testing training abilities might ask the candidate to review a firefighter’s performance evaluations and identify what training that firefighter needs to improve. (Dr. Brink Report 48). The consent decree allows for scoring these situational-judgment questions on a full-credit, partial-credit, and zero-credit basis, provided that the “same responses” receive the same credit. The consent decree requires ranking the candidates according to their scores on this part. In the initial settlement agreement, the candidates’ scores on this situational-judgment component determined whether the candidate could proceed to the final phase of the exam, but the modified settlement agreement provides that all candidates advance. (Compare Docket Entry No. 69-2, at 29, with Docket Entry No. 86-1, at 3).

The final exam component is an assessment center. “An assessment center consists of multiple exercises simulating job activities that are designed to allow trained observers, or assessors, to make judgments about candidates’ behaviors as related to job performance.” (Dr. Brink Report 47). One type of simulation used in assessment centers is a “role play.” A role play “is a simulation of a face-to-face meeting between the candidate (playing the role of a job incumbent) and a trained role player acting as a person incumbents frequently encounter on the job (such as a subordinate or citizen).” (Id.). Assessment-center activities such as role play violate the CBA and TLGC. See Tex. Loc. Gov’t Code § 143.032(c) (forbidding tests that “in any part consist of an oral interview”). “Assessors” score promotional candidates’ performance on the assessment-center activities. Although there are preset criteria distinguishing better from worse performance, the scoring system is subjective and violates the TLGC and the CBA.

Another inconsistency between the TLGC and the CBA on the one hand and the consent decree provisions on the other is the requirement in the consent decree to “band” the promotional candidates’ assessment center scores. “Banding” scores means adjusting the individual test scores based on statistical analyses showing the likelihood that: (1) a candidate could score higher or lower on the same exam; and (2) the individual assessor could have given the candidate a higher or lower score for the same performance. Banding tends to convert individualized score differences into homogenized “bands” of more uniform scores. For example, three candidates’ scores of 85, 86, and 87 might be “banded” as one score of 86, depending on the results of the statistical analysis. Banding is like converting individual scores of 95%, 97%, and 100% on a 100-question multiple choice test into three “As” that are viewed as identical. The conversion is based on statistical analysis showing that an individual scoring 95% on the exam has the same chance of scoring 100% on the exam as the person scoring 100% on the exam has of scoring 95%. Banding is based on the assumption that small differences in scores do not reliably demonstrate superiority in the KSAOs the exam is supposed to measure. One of the City’s expert witnesses, Dr. Morris, summarized banding as follows:

“[A] band is ... saying if I made 87 and someone else made 85, is it possible that the next day I could have made 85 and they could have made 87? So a band— the band we’re trying to calculate the standard error of measurement is simply saying that certain number of times that band is going to fall within a standard error of measurement that we calculate. So, it’s a reasonable thing. And most people in our field accept using bands as a way to minimize the error that could be assumed in the minds of the decision-makers.

(Evidentiary Hr’g Tr. 100, Docket Entry No. 130).

Banding is inconsistent with the Rule of Three. Under the proposed consent decree, the final promotion decision is based on the banded assessment-center scores. Names are submitted by score “bands,” not subject to the Rule of Three that would have applied under the TLGC and the CBA. Under the Rule of Three, if the three highest scores were 85, 86, and 87, the names of those applicants would be submitted. The person who scored the 87 would be selected unless the decision-maker provided a written reason for selecting the person who scored the 86 or 85. Under the banding system, the three individuals would be treated by the decision-maker as having the same score. The “band” might also be larger than three persons; its size would be determined by statistical analyses rather than a preset number. The consent decree requires the decision-maker to select one within the band and to provide a written explanation for the selection.

The consent decree does not state the role of a candidate’s race or how the decision-maker may consider race in choosing who to promote within a band. There was testimony that using race as a factor to select a candidate within a band could reduce the exam’s disparate impact on African-American applicants. Within a band, all applicants are viewed as equal. (Evidentiary Hr’g Tr. 107-08, Docket Entry No. 130). But the consent decree does not explicitly authorize race-based promotional decisions.

II. The Evidence in the Record

At an evidentiary hearing, the parties presented evidence as to (1) whether the senior-captain exam disparately impacted black firefighters, and (2) whether the proposed changes to the captain and senior-captain exams were justified by business necessity.

A. The Evidence as to Disparate Impact of Past Exams

On February 8, 2006, the City of Houston administered the senior-captain exam to 221 promotional candidates. Of the 221 candidates taking the exam, 172 were white, 15 were black, 33 were Hispanic, and 1 was “other.” The 212 candidates who passed by scoring above 70 consisted of 166 Caucasians, 13 African-Americans, 32 Hispanics, and 1 “other.” The City promoted 70 candidates based on the rank-order list of those who passed the exam. Of those promoted, 59 were Caucasian, 2 were African-American, 8 were Hispanic, and 1 was in the “other” category. (Evidentiary Hr’g Ex. 7, Dr. McPhail Report, at 5).

The following experts submitted reports or testified as to whether the 2006 senior-captain exam disparately impacted black firefighters:

• Dr. S. Morton McPhail, an industrial-psychology consultant, on behalf of the City. Dr. McPhail is licensed by the Texas State Board of Examiners of Psychologists and is a Fellow of the Society for Industrial and Organizational Psychology (“SIOP”). He has served as an adjunct faculty member in the psychology departments of Rice University and the University of Houston. Dr. McPhail has “authored scholarly articles and [has] given many symposia, presentations, and continuing education workshops for peers on issues relating to employment, and in several instances, on the topics of job analysis.” (Docket Entry No. 37-1, at 3).

• Dr. Kyle Brink, an industrial-psychology consultant and a tenure-track assistant professor in the management department of the Bittner School of Business at St. John Fisher College, testified on the plaintiffs’ behalf. Dr. Brink’s doctorate is in industrial psychology. He has experience developing and validating promotion procedures at both private companies and governmental organizations. Recently, he worked with the Personnel Board of Jefferson County, Alabama, helping end a federally imposed consent decree. (Dr. Brink Report 5).

• Dr. Kathleen Lundquist, an industrial-psychology consultant, testified for the City. Dr. Lundquist is the president and CEO of APT, Inc. She has “extensively researched, designed and conducted statistical analyses and provided consultation in the areas of job analysis, test validation, performance appraisal and research design” for “major corporations in the banking, financial services, retail, electronics, aerospace, pharmaceutical,, telecommunications, and electric utility industries, as well as for federal, state, and local agencies.” (Dr. Lundquist Aff. 1, Docket Entry No. 93-2). She has a Ph.D. in psychometrics from Fordham University.

• Dr. Winfred Arthur, a full professor of psychology and management at Texas A & M University, testified for the HPFFA. Dr. Arthur has a Ph.D. in industrial/organizational psychology from the University of Akron and “over 20 years of practical experience in the areas of test development, selection, public safety testing, and training.” He is a SIOP fellow. (Dr. Arthur Aff. 1, Docket Entry No. 89-1).

• Dr. David M. Morris, an industrial-psychology consultant, testified for the City. Dr. Morris is the president of Morris & McDaniel, Inc., an industrial-psychology consulting firm he started in 1976. He received his Ph.D. in psychology, with a specialization in industrial/organizational psychology, from the University of Southern Mississippi. Dr. Morris has authored numerous scholarly articles and is a member of the industrial/organizational division of the American Psychological Association and also a member of SIOP. (Evidentiary Hr’g Ex. 14).

All the experts agreed that the “total selection process” for promoting HFD captains to senior captain showed disparate impact under the 4/5 Rule. The Guidelines require that “[a]dverse impact is determined first for the overall selection process for each job.” Adoption of Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures, 44 Fed.Reg. 11966, 11998 (1979) [hereinafter Guidelines Questions & Answers], “The ‘total selection process’ refers to the combined effect of all selection procedures leading to the final employment decision such as hiring or promoting.” Id. The experts agreed that the rate of blacks promoted to senior captain— 13.3% — is less than 4/5ths the selection rate for whites — 34.3%. (See, e.g., Dr. Brink Report 9-10). But only Dr. Brink and Dr. Lundquist found disparate impact for the senior-captain exam.

Both parties’ experts testified that a 4/5 Rule violation is an unreliable basis to find disparate impact when the population size of one group is small. Only 17 black firefighters were eligible for promotion to senior captain. The experts agreed that this is too small a number to make a 4/5 Rule violation sufficient to find disparate impact. (Dr. Brink Report 10; Dr. Lundquist Aff. 6; Dr. Arthur Aff. 2-3; Dr. McPhail Report 6-7; Evidentiary Hr’g Ex. 12, Dr. Morris Report, at 0010447). There was general agreement among the experts that when group populations are small, statistical analyses should be used to determine whether the 4/5 Rule violation is the product of “chance.” This requires determining the statistical significance of the 4/5 Rule violation. Dr. Morris’s report noted that the 4/5 Rule risks “sampling error,” which statistical-significance analysis mitigates. Dr. Morris’s report stated:

[T]he 4/5ths Rule has two major limitations, precision and sampling error. The 4/5ths Rule provides a descriptive ratio; it is not a statistical test. As such, the 4/5ths Rule cannot determine if an observed disparity is the result of mere chance or an indication of underlying bias. Use of the 4/5ths Rule is limited further by sample error. Unlike statistical tests, the 4/5ths Rule does not make adjustments for sampling error and, in cases where sample sizes are small, may fail to detect disparities. More problematic, the 4/5ths Rule has proven to falsely show adverse impact when no adverse impact exists.

When sample sizes are small, results from the 4/5ths Rule will vary, often dramatically, because the composition of candidates vary each time a personnel decision is made. Statistically speaking, this variation is referred to as sampling error. The 4/5ths Rule is insensitive to sampling error. Boardman (1979) and Greenberg (1979) demonstrated that the 4/5ths Rule is susceptible to either falsely identifying adverse impact when none exists or failing to identify adverse impact when it does exist. At their very heart, statistical tests directly address sampling error.

(Dr. Morris Report 0010446-47). Dr. Brink’s report discussed how the 4/5 Rule risks “Type I” error by leading to the conclusion “that adverse impact exists, when in reality the difference in selection rates is a result of sampling error (or chance).” (Dr. Brink Report 12). Dr. Brink’s report explained how statistical tests can “control the potential amount of Type I error”:

Type I error in this context is defined as concluding that adverse impact exists, when in reality the difference in selection rates is a result of sampling error (or chance). A statistically significant result is one in which the probability of incorrectly concluding that adverse impact exists (i.e., a Type I error) is less than a specified level; this specified level is referred to as an alpha level.... Statistical tests produce a probability value ... that determines or estimates the probability of obtaining the sample result assuming there were no differences in the population. If the [probability value] resulting from the statistical test is less than the specified alpha level, we say the result is statistically significant and would decide, based on the test, that there is adverse impact. For example, if an alpha level of .05 is chosen and the [probability value] resulting from the statistical test is less than .05, then there is less than a 5% probability that the difference is due to chance (i.e., there is less than a 5% probability of making a Type I error) and we say the result is statistically significant. Conversely, you can conclude that there is a 95% probability that the difference is not due to chance.

(Id.).

The experts also identified peer-reviewed journal articles critical of the 4/5 Rule. In 1979, Anthony Boardman and Irwin Greenberg authored analyses showing that the 4/5 Rule could lead to both Type I (falsely identifying disparate impact when none exists) and Type II (failing to identify disparate impact when it does exist) statistical errors. See Irwin Greenberg, An Analysis of the EEOCC ‘Four-Fifths’ Rule, 25 Mgmt. Sci. 762 (1979); Anthony E. Boardman, Another Analysis of the EEOCC ‘Four-Fifths’ Rule, 25 Mgmt. Sci. 770 (1979). A recent article similarly concluded that “there is a fairly high false-positive rate for the 4/5ths rule used by itself.” Phillip L. Roth et al., Modeling the Behavior of the I/5ths Rule for Determining Adverse Impact: Reasons for Caution, 91 J. Applied Psychol. 507, 519 (2006). The article’s authors cautioned that “other factors (e.g., sample size) were quite important” and recommended using “a test such as Fisher’s exact test or a chi-square test to mitigate false-positives.” Id. Also noting the 4/5 Rule’s shortcomings, Scott Morris and Russell Lobsenz recently proposed a “more complex” statistical technique, the Zir test, for evaluating disparate impact. See Scott B. Morris & Russell Lobsenz, Significance Tests and Confidence Intervals for the Adverse Impact Ratio, 53 Personnel Psychol. 89 (2000).

The experts applied three tests to measure statistical significance: the Fisher exact; the Pearson chi-square; and the Zd. The Fisher exact test provides “the exact probability of obtaining the observed frequency table (or one more extreme) under the null hypothesis [that no disparate impact exists]” and is particularly suited to analyses involving small sample sizes. Michael W. Collins & Scott B. Morris, Testing for Adverse Impact When Sample Size Is Small, 93 J. Applied Psychol. 463, 464 (2008). Dr. Brink reported the Fisher exact test probability to be .15 — that is, a 15% chance that the observed results were due to pure chance. (Dr. Brink Report 15). Social scientists generally require the p-value — the probability that the observed results are due to pure chance — to be less than .05 for results to be considered statistically significant.

The Pearson chi-square test estimates the probability of obtaining the observed frequency table under the null hypothesis that no disparate impact exists. (Id. at 13). Dr. Brink reported the Pearson chi-square probability to be .10 — that is, a 10% chance that the observed results were due to pure chance. (Id. at 15). Like the Fisher exact test’s p-value, the Pearson chi-square p-value is greater than .05 and is viewed as statistically insignificant.

Unlike the Pearson chi-square and Fisher exact tests, the Zd test does not provide a p-score, the probability that the observed results were due to pure chance. Instead, the Zd test yields a Z statistic. The difference between the selection rates of the two groups being compared is statistically significant if the absolute value of the Z statistic is greater than 1.96. See Collins & Morris, supra, at 464. Using the Zd test, Dr. Brink reported a statistically insignificant Z statistic of — 1.66. (Dr. Brink Report 15).

Neither the Fisher exact, the Pearson chi-square, nor the Zd test demonstrated that the 4/5 Rule violation for the total selection process for senior captain was statistically significant. (Dr. Brink Report 12; Dr. Arthur Aff. 3; Dr. McPhail Report 8; Dr. Morris Report 0010448). Only one test of statistical significance, the Zir test, showed that the 4/5 Rule violation was statistically significant. Unlike the Zd test, which evaluates the observed difference in selection rates, the Zir test evaluates the difference in selection-rate ratios. Dr. Brink argued for the use of the Zir test in these circumstances because the test uses the same comparison as the 4/5 Rule and is “slightly more powerful than the Zd or chi-square tests, especially as the proportion of minorities is smaller.” (Dr. Brink Report 13-14). But the use of the Zir test is not as well supported in the literature as the other tests, especially in the context of very small sample sizes. One peer-reviewed journal article explained that while the Zir test was “interesting and deserve[s] greater thought,” further research was needed on the test’s ability to evaluate disparate impact. Roth et al., supra, at 520. The test’s creators acknowledged that the Fisher exact test will “provide a more accurate evaluation of statistical significance” than the Zir test when the smallest expected value in the analysis is less than five.