Citations

Full opinion text

FRANK A. KAUFMAN, Chief Judge:

In an opinion filed March 29, 1979 , this Court held, inter alia, that the sergeant’s promotional examinations used by the Police Department (“Department”) of Baltimore City (“City”) in 1972, 1973, 1974, 1976 and 1977 violated Title VII of the Civil Rights Act of 1964, 42 U.S.C. § 2000e et seq., (“Title VII”) because those exams had a racially adverse impact upon blacks and because defendants had not shown that the exams were job-related. In that opinion, certain questions relating to relief were held sub curia pending further presentation of evidence and legal argument. Those outstanding issues are now ripe for decision.

Plaintiffs herein challenge the validity of the 1982 written sergeant’s promotional exam. That written exam was designed by Baltimore Civil Service Commission (“Commission”) personnel and was administered to 605 candidates for sergeant on May 8, 1982. During a nonjury trial held on January 5-6 and January 17-20,1984, a number of expert and lay witnesses testified. Subsequently, pre- and post-trial memoranda, along with other documents, were filed. After careful review of the entire record in this litigation, this Court holds that the 1982 written exam is invalid and that appropriate relief as set forth infra is required. Findings of fact and conclusions of law, in accordance with Federal Civil Rule 52(a), are set forth below.

I. FACTS

The written exam in question, designated Exam No. 820508, was one part of a three-step promotional procedure employed in 1982 to promote police officers to the rank of sergeant. The other two parts consisted of a promotional appraisal and an oral examination. All three components of the promotional procedure were designed and administered by the Commission, under the supervision of Robert G. Wendland, Deputy Personnel Director. The 605 candidates for sergeant, in addition to sitting for the 115-question written exam, were evaluated by their supervisors on the basis of a supervisory appraisal, called the promotional appraisal. The promotional appraisal was designed to test seven skills, deemed by the Commission to be “significant elements of a sergeant’s job.” A candidate’s scores on the written exam and the promotional appraisal were scaled, multiplied by the weight assigned to the written exam (40%) and to the promotional appraisal (30%) and added together to produce an overall score for each candidate on those two components. Then, each candidate’s weighted, composite score was ranked. Only the top 95 candidates, of the original 605, were given the oral exam. The oral exam (weighted 30%) consisted of problem analysis exercises and was designed to test for six skills which the Commission deemed “essential” to the sergeant’s job. After all three component scores were computed, the Commission published an eligibility list, ranking the top 95 candidates.

The eligibility list is designed to be used until the list is exhausted, a new selection procedure is developed, or the list expires. Individuals are promoted, in accordance with their ranking, from the eligibility list as vacancies occur. At the time of trial on January 20, 1984, 15 persons — 12 whites and 3 blacks — had been promoted from the 1982 eligibility list. The eligibility list is scheduled to expire in accordance with applicable law in October, 1984.

Plaintiffs in this litigation are presently challenging only the written exam. Plaintiffs concede that any racially adverse impact of the overall promotional procedure is attributable solely to the written exam. The 115-question, multiple-choice written exam was developed by a two-phase process: first, a job analysis was devised and second, the test itself was constructed. A thorough job analysis was prepared by Management Scientists, Inc. (“MSI”) experts hired by the City in connection with the 1981 sergeant’s promotional procedure. No independent job analysis was performed for the 1982 promotional procedure. Rather, in the Commission’s 1982 Validation Report, “the reader is referred to the Ford report and volume one of the MSI report for a comprehensive discussion of the task analysis and the linkage of measured knowledges, skills and abilities to the individual tasks required on the job.” In other words, while an expert psychometric firm, MSI, was consulted and did prepare a job analysis for the 1981 promotional procedure, no new job analysis was prepared in 1982. Instead, the City, acting without expert assistance, relied on the 1981 job analysis, the 1981 Ford report, and earlier MSI reports, in devising its 1982 promotional procedure and, in particular, the written exam. Mr. Wendland testified at trial that the 1982 Validation Report “piggybacked” upon the 1981 job analysis, although some modifications were made to the job analysis in 1982. Thus, in considering the validity of the job analysis used in the 1982 sergeant’s promotional procedure, review of both the 1981 MSI job analysis and the modifications made thereto in 1982 is required.

The job analysis prepared by MSI in 1981 consisted of five separate steps. First, based on interviews, full shift observation and questionnaires completed by 73 randomly selected sergeants, MSI developed a list of “tasks” related to various “job components” which sergeants generally perform. As part of that step, the incumbent sergeants assigned point values to each task and to each job component based on frequency of occurrence, importance, complexity and eriticality, and also based on the amount of time normally spent performing that task or job component. Second, SKAP lists (lists of S' kills, K now-ledges, A bilities and P ersonal Characteristics) were generated for each job component found to be a part of the sergeant’s job.

In the third step, the MSI Project Director, David Wagner, prepared a SKAP rating questionnaire to measure the importance, or to determine the relative weight, of each SKAP with respect to the successful performance of a sergeant’s job. The SKAP rating questionnaire was distributed to fifty randomly selected incumbent sergeants who assigned points to each SKAP to reflect the importance of that SKAP to successful performance of each job component. The sergeants also indicated whether they were of the opinion that a sergeant, on the first day he was on the job, needed to perform a particular SKAP.

In the fourth step, MSI determined the importance of each SKAP in relation to the job of police sergeant. To so do, MSI multiplied the average SKAP weight by each job component value. That produced a total value for a SKAP within a given job component. The values for SKAPs which appeared in more than one job component were then added together to arrive at a total SKAP value. In its simplest form, that procedure assigned a numerical value to each SKAP which each sergeant needed in order to perform the job of police sergeant. The relevant SKAPs and their values are set forth in Table 1, “Description of SKAP Clusters with Associated Job Analysis Points.”

As a final step in the 1981 job analysis, MSI determined which of three measurement components — the written test, the performance appraisal or the assessment center — would be most effective for testing each SKAP cluster. MSI determined that the Knowledges SKAP, assigned 4182 job analysis points, was to be measured on the written exam, while 27 other SKAPs were to be measured by the other two test components. MSI calculated the appropriate weights for the written exam (25%), the performance appraisal (30%) and the assessment center (45%) by determining the percentage of job analysis points allocated to each test component.

As noted earlier, the 1981 MSI job analysis was modified substantially by Commission personnel in 1982. Only 12 of the 28 SKAPs identified in the 1981 job analysis were included in the 1982 job analysis and were the subject of testing in the 1982 promotional procedure. Among those SKAPs deleted in 1982 was the second most-heavily weighted SKAP, that is, supervisory ability. Of the 12 SKAPs included in the 1982 job analysis, two, knowledges and reading comprehension, were tested for in the 1982 written exam. The 10 remaining SKAPs were measured by the promotional appraisal and/or the oral exam. Additionally, the weightings of the three test components were substantially altered in 1982. In explaining the re-weighting, the Commission stated:

Some of the factors [SKAPs] rated or .assessed in 1981 were deleted from the 1982 examination. Based on the new set of factors, the relative weights for the written test, supervisory rating and oral examination are as follows:

Written Test: 40% Supervisory Promotional Appraisal: 30%

Oral Examination: 30%

The reweighting resulted in the written exam, which is the only test component alleged to have a racially adverse impact, being weighted most heavily in the 1982 promotional procedure.

The second major phase in developing the 1982 promotional procedure was the construction of the three testing components. Because the only component of the 1982 promotional procedure challenged in this litigation is the written exam, discussion in this opinion is limited to the construction of that exam. It bears emphasizing that the procedures used in constructing the 1982 written exam differed significantly from those employed by MSI in constructing the 1981 written exam. In 1981, for example, MSI first conducted a series of workshops for Commission personnel to train them in writing multiple choice test items. MSI then reviewed the job analysis to decide what major knowledge areas were to be tested. The number of test items for each knowledge area reflected the relative number of job analysis points associated with that knowledge area.

The next step involved in construction of the 1981 written exam was the identification of viable primary and secondary sources. Thereafter, a test outline was prepared based on an 100-item test. Finally, construction of the 1981 written exam involved drafting the 100 multiple-choice test items. In so doing, the Commission retrieved items from previously administered sergeant, lieutenant and captain examinations from the Civil Service item pools. The items were then classified, if possible, into one of the knowledge areas and were reviewed to determine whether they possessed current vitality. Each previously used item was also classified according to its level of difficulty, level of discrimination, and level of racial impact. All previously used items which had questionable or unacceptable impact levels were deemed nonusable, as were all items which were poor discriminators.

After determining the number of previously used items available for pretesting in 1981, the Commission staff drafted additional items in order to produce three times the number of items for each knowledge area as were in the end considered to be needed. In other words, more than 300 usable items were developed for the 100-item test. The 300-question test was then pilot-tested in parts in three cities, St. Louis, Detroit and Denver. Each pilot-tested item was analyzed for difficulty, disparate racial impact, internal consistency, validity and relevance. Based on those analyses, “the MSI Project Director selected for inclusion in the [final 100-item multiple-choice] test those items which possessed the best psychometric information.” An item was considered “desirable” if, among other things, it demonstrated “no difference between minorities and non-minorities on [sic] Difficulty level.”

In constructing the 1982 written exam neither MSI nor any other expert in psychometrics played any active role. Instead, Commission personnel determined the 1982 written test content and wrote test items “in-house.” In its 1982 Validation Report, the Commission described the basis for construction of the 1982 written exam as follows:

The specific areas of knowledge listed in the job analysis were evaluated based on their statistical ratings and trivial knowledges were eliminated. The remaining knowledges were grouped into broader categories of related knowledge, and the categories were then statistically analyzed to determine approximately what proportion of the test should be related to each category. Reference materials and source documents for each area were identified by Police Department personnel.

The major change in test content was the inclusion of reading comprehension items. Since this skill was rated so heavily in the job analysis, since it is appropriate to test reading comprehension in a written test and since Dr. Barrett made unsubstantiated references to reading level differences between black and white populations, the Commission felt it was important to measure reading comprehension. In addition, areas measured by fewer than 5% of the items were dropped because of problems in sampling, reliability and importance. The test was therefore lengthened to 115 items in 1982 compared to 100 items in 1981.

1982 Validation Report, supra, at 2. The Commission’s Report does not indicate which Commission personnel were item-writers in 1982, nor whether such personnel received any training in item-writing. Moreover, the Report states: “The test items were not reviewed by experts within the Police Department prior to the examination for test security purposes, and the CSC [Civil Service Commission] did not have the expertise to spot the highly technical flaws in [some] questions.” Further, the Report indicates that “in 1982, pretesting multiple choice items was not possible.” As an alternative to pilot testing, the Commission ‘.‘elected to administer the test and then delete those items which might later prove to be defective” based on candidates’ appeals. Thus, at the time the written test was administered, “candidates were given a form on which they could comment on specific parts or questions in the test.” Based on a. review of those comments by Department and Commission staff, eighteen items were deleted from the written test prior to final scoring.

The test was then scored from 0 to 97 (115 original questions minus 18 deleted questions) with one point given for each correct answer. The raw scores were converted to standard scores based upon standard deviations.

As noted earlier, the combined scores on the written test (weighted 40%) and the promotional appraisal (weighted 30%) were then computed for all 605 candidates. The top 95 candidates, with their respective rankings achieved on the basis of their written exams and promotional appraisal scores, were then given the oral examination (weighted 30%). Only those 95 candidates were provided the opportunity to proceed through all three stages of the promotional procedure because the Department estimated that that number would meet the Department’s needs for 40 sergeants during the two-year life of the eligibility list. In that regard, the Department stated in its 1982 Validation Report that the cutoff point of 95 was “based on the number of anticipated vacancies,” not on “the level of performance below which a candidate could not function as a Sergeant.”

Of the 605 applicants who participated in the first two components of the 1982 promotional procedure, 460 were white, 128 were black, and 8 were members of other minority groups. The whites constituted 76% of the total applicants, while blacks constituted 22% of the total applicant pool.

Of the white candidates, 17.4% passed the promotional procedure, i.e., were given the chance to proceed to the oral examination and were placed on the eligibility list. By contrast, only 10.9% of the black candidates passed the procedure. Thus, the black pass rate was about three-fifths of the white pass rate, substantially below the four-fifths minority-to-white pass ratio deemed acceptable by the Uniform Guidelines § 4(D). Because the results of the promotional appraisal were almost identical for blacks and whites, any adverse impact is attributable to the written exam.

II. LAW — TITLE VII

Title VII, and more particularly 42 U.S.C. § 2000e-2(h), forbids the use of employment “practices, procedures or tests, neutral on their face, and even neutral in terms of intent” if they are discriminatory in effect. Griggs v. Duke Power Co., 401 U.S. 424, 430, 91 S.Ct. 849, 853, 28 L.Ed.2d 158 (1971). See also Dothard v. Rawlinson, 433 U.S. 321, 329, 97 S.Ct. 2720, 2726, 53 L.Ed.2d 786 (1977); Albemarle Paper Co. v. Moody, 422 U.S. 405, 430, 95 S.Ct. 2362, 2377, 45 L.Ed.2d 280 (1975). In Griggs, Chief Justice Burger wrote:

Nothing in the Act precludes the use of testing or measuring procedures; obviously they are useful. What Congress has forbidden is giving these devices and mechanisms controlling force unless they are demonstrably a reasonable measure of job performance. Congress has not commanded that the less qualified be preferred over the better qualified simply because of minority origins. Far from disparaging job qualifications as such, Congress has made such qualifications the controlling factor, so that race, religion, nationality, and sex become irrelevant. What Congress has commanded is that any tests used must measure the person for the job and not the person in the abstract.

Id. 401 U.S. at 436, 91 S.Ct. at 856.

In a disparate impact case, such as the one at bar, a tripártate analysis must be applied. First, a Title VII plaintiff bears the initial burden of making out a prima facie case of discrimination. To so do, such a plaintiff need only show “that the tests in question select applicants for hire or promotion in a racial pattern significantly different from that of the pool of applicants.” Albemarle Paper Co. v. Moody, 422 U.S. at 425, 95 S.Ct. at 2375 (1973). A prima facie case may be established by evidence of statistical disparities alone. Dothard v. Rawlinson, supra, 433 U.S. at 329, 97 S.Ct. at 2726; International Brotherhood of Teamsters v. United States, 431 U.S. 324, 329, 97 S.Ct. 1843, 1851, 52 L.Ed.2d 396 (1977). Once a plaintiff establishes a prima facie case, the burden shifts to the employer to demonstrate that the selection process which has produced the disparate racial impact is job related or, to state it more precisely, that the selection process has a “manifest relationship to the employment in question.” Griggs v. Duke Power Co., supra, 401 U.S. at 432, 91 S.Ct. at 854. Finally, if the employer meets that burden, the plaintiff may still prevail by persuading the trier of fact that the challenged selection process was a mere pretext for discrimination in hiring or promotion. Connecticut v. Teal, 457 U.S. 440, 446-47, 102 S.Ct. 2525, 2530-31, 73 L.Ed.2d 130 (1982). In a Title VII cause of action, discriminatory purpose need not be proven. Albemarle Paper Co. v. Moody, supra, 422 U.S. at 422, 95 S.Ct. at 2373; Griggs v. Duke Power Co., supra, 401 U.S. at 432, 91 S.Ct. at 854.

A. Adverse Impact

The parties have asked this Court to assume, on a binding basis for all purposes of this litigation, that the 1982 written exam had a disparate impact as defined by the Equal Employment Opportunity Commission (“EEOC”) Uniform Guidelines on Employee Selection Procedures, 29 C.F.R. § 1607 (1983) (“Guidelines”). Under the Guidelines, “[a] selection rate for any race, sex, or ethnic group which is less than four-fifths (4/s) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact ____” Section 4(D). See Connecticut v. Teal, supra, 457 U.S. at 443 n. 4, 102 S.Ct. at 2529 n. 4. Plaintiffs here have proffered evidence that the written exam had an adverse impact ratio of 62.9% and that the promotional procedure as a whole had an adverse impact that fluctuated between an acceptable ratio of 89.9% to an unacceptable ratio of between 50 and 75%. However, plaintiffs’ expert, Dr. Richard S. Barrett, testified that if relatively few of the 95 candidates on the eligibility list were promoted, e.g., 23 or fewer, the 80% rule of thumb established in the Guidelines would not be violated because several black candidates scored disproportionately well on the written exam. Additionally, Dr. Barrett noted that if more than 23 candidates were promoted from the eligibility list, adverse impact would become apparent, fluctuating between unacceptable ratios of 50% and 75%, depending upon the number promoted.

The City has informed the Court that as of the time of trial only 15 eligible candidates — 12 whites and 3 blacks — had been promoted to sergeant. Thus, in terms of actual promotions made, the 80% selection ratio has not been violated.

On the basis of this initial testimony, the parties agreed not to litigate the issue of adverse racial impact but to assume that such adverse impact existed for the 1982 sergeant promotional procedure as a whole and for the written exam in particular. The parties further agreed that no additional promotions would be made from the eligibility list pending this Court’s ruling with regard to job-relatedness and validity of the 1982 written exam. That agreement not to promote further from the 1982 eligibility list was designed to preserve a promotion ratio of black to white sergeants which did not, in actual practice, offend the 80% Guidelines rule.

B. Job Relatedness

In view of defendants’ concession of adverse impact, defendants must rebut plaintiffs’ prima facie case of discrimination by demonstrating that the 1982 written exam is job-related. See, e.g., Connecticut v. Teal, supra, 457 U.S. at 446-47, 102 S.Ct. at 2530-31; Albemarle Power Co. v. Moody, supra, 422 U.S. at 425, 95 S.Ct. at 2375; Griggs v. Duke Power Co., supra, 401 U.S. at 432, 91 S.Ct. at 854; Vanguard I, 471 F.Supp. at 698-99. To so do, defendants must show that the written exam is “demonstrably a reasonable measure of job performance.” Griggs, supra, 401 U.S. at 436, 91 S.Ct. at 856.

“The threshold task in determining the validity or job-relatedness of a challenged examination is to select the appropriate measure for assessing its job-relatedness.” Guardians Ass’n of New York City v. Civil Service Comm’n, 630 F.2d 79, 91 (2d Cir.1980), cert. denied, 452 U.S. 940, 101 S.Ct. 3083, 69 L.Ed.2d 954 (1981) (“Guardians IV”). The Guidelines, which draw heavily upon professional standards of test validation established by the American Psychological Association recognize three validation techniques: content validation, construct validation and criterion-related validation. Guidelines §§ 5(B), 14. The Guidelines specify when each technique is appropriate and also specify the requirements for successfully validating an examination by use of each technique. Justice White, in Washington v. Davis, 426 U.S. 229, 247 n. 13, 96 S.Ct. 2040, 2051 n. 13, 48 L.Ed.2d 597 (1976), recognized those three methods of test validation as set forth in the Guidelines and described them as follows:

Professional standards developed by the American Psychological Association in its Standards for Educational and Psychological Tests and Manuals (1966), accept three basic methods of validation: “empirical” or “criterion” validity (demonstrated by identifying criteria that indicate successful job performance and then correlating test scores and the criteria so identified); “construct” validity (demonstrated by examinations structured to measure the degree to which job applicants have identifiable characteristics that have been determined to be important in successful job performance); and “content” validity (demonstrated by tests whose content closely approximates tasks to be performed on the job by the applicant).

Further, Chief Justice Burger in Griggs v. Duke Power Co., 401 U.S. 424, 433-34, 91 S.Ct. 849, 854-55, 28 L.Ed.2d 158 (1971), and Justice Stewart in Albemarle Paper v. Moody, 422 U.S. 405, 431, 95 S.Ct. 2362, 2378, 45 L.Ed.2d 280 (1975), concluded that the EEOC Guidelines are entitled to great deference. This Court earlier in this litigation surveyed the substantial authority favoring deference to the Guidelines and observed that “a determination as to whether any test is job-related requires consideration of the EEOC Guidelines,” Vanguard I, supra, 471 F.Supp. at 728. Other courts, too, have expressed similar viewpoints. In Guardians IV, supra, the Second Circuit discussed extensively the proper weight to be given the EEOC Uniform Guidelines. Judge Newman, in cautioning against viewing every deviation from the Guidelines as an automatic violation of Title VII, wrote:

To the extent that the Guidelines reflect expert, but non-judicial opinion, they must be applied by courts with the same combination of deference and wariness that characterizes the proper use of expert opinion in general ____ Thus, the Guidelines should always be considered, but they should not be regarded as conclusive unless reason and statutory interpretation support their conclusions. As this Court has previously stated: “If the EEOC’s interpretations go beyond congressional intent, the Guidelines must give way.”

Id. at 91 (citations omitted). In order for an examination to be found job-related, it should usually comply, not 100%, but nevertheless in substantial measure, with the reasonably attainable requirements set forth in the Guidelines. Thus, in this ease, where defendants contend that the written exam is both content valid and criterion valid, defendants will be required to demonstrate substantial compliance with the Guidelines requirements governing content and/or criterion validation.

C. Content Validity

A test’s degree of content validity has been defined as “the degree to which the subject matter of the test relates to those significant parts of the curriculum for which the test is intended to predict the applicants’ capability.” Rivera v. City of Wichita Falls, 665 F.2d 531, 537 (5th Cir. 1982). Another oft-quoted definition of content validity is set forth by Judge Weinfeld in Vulcan Society of the New York City Fire Dep’t, Inc. v. Civil Service Comm’n, 360 F.Supp. 1265 (S.D.N.Y.), aff’d in part, remanded in part, 490 F.2d 387 (2d Cir.1973):

An examination has content validity if the content of the examination matches the content of the job. For a test to be content valid, the aptitudes and skills required for successful examination performance must be those aptitudes and skills required for successful job performance. It is essential that the examination test these attributes both in proportion to their relative importance on the job and at the level of difficulty-demanded by the job.

Id. at 1274 (footnotes omitted). In short, an examination is content valid if it tests knowledges, skills and abilities critical to a job and thereby rates applicants on the basis of their ability to perform that job.

While the Guidelines describe various aspects of content validation, they “do not neatly list ingredients of an adequate exam.” Guardians IV, supra, at 95. Yet, courts have been able to distill certain factors from the Guidelines for consideration in determining whether an examination possesses sufficient content validity to justify its use, notwithstanding its disparate racial impact. To begin with, constructing a content valid exam requires proof of a thorough job analysis. Thus, in United States v. County of Fairfax, Va., 629 F.2d 932, 943 (4th Cir.1980), cert. denied, 449 U.S. 1078, 101 S.Ct. 858, 66 L.Ed.2d 801 (1981), Judge Winter wrote: “Usually the starting point in proof of validity is evidence of a thorough job analysis.” Second, the test-makers must have used “reasonable competence” in constructing the examination itself. Third, the designers of a test must develop a test whose content has a direct relationship with the content of the job. Fourth, the content of the test must be representative of the content of the job. Finally, the test must be used or scored in such a manner as to assure selection, with some precision, of those applicants best able to perform the job.

1. Job Analysis

Judge Weinfeld has described a job analysis as “a thorough survey of the relative importance of the various skills involved in the job in question and the degree of competency required in regard to each skill.” Vulcan Society, supra, 360 F.Supp. at 1274. According to the Guidelines, a job analysis consists of an assessment “of the important work behavior[s] required for successful performance and their relative importance.” § 14(C)(2). In the case at bar, the Commission did not devise an independent job analysis for the 1982 promotional procedure. Rather, as noted supra, the Commission relied largely upon the 1981 job analysis which was prepared by MSI based upon information gathered in 1979 and 1980.

The 1981 job analysis seemingly comported with the Guidelines requirements. First, the 1981 job analysis identified the important work behaviors required for successful performance of the sergeant’s job. As discussed in more detail, supra, the sergeant’s job was divided into “job components” and associated “tasks” based upon questionnaires distributed to 73 incumbent sergeants. Those job components and tasks were defined with a relatively high degree of precision. Each of the 73 sergeants completing the task questionnaire was asked to indicate, inter alia, whether each component and each task was a part of his job and the relative importance of each component and of each task. The record herein calls for the conclusion that the first part of the Guidelines standard, i.e., “identification of important job behaviors,” was satisfied by the 1981 job analysis.

Second, the MSI job analysis met the requirement of determining the relative importance of the identified work behaviors. The City assessed the importance of each job component and task by means of the task questionnaires referred to above. Further, the City identified a list of SKAPs which a sergeant needed to possess in order to perform each component of his job, and established the relative weight of each SKAP by virtue of SKAP rating questionnaires distributed to 50 randomly selected sergeants. Finally, in order to determine the overall importance of each SKAP to the job of sergeant, the City calculated an overall SKAP value by multiplying the SKAP weight by each job component weight. The care taken by MSI first to determine the relative importance of work behaviors and then to identify the relative importance of critical SKAPs associated with those work behaviors was in full accord with the Guidelines. Moreover, not only was the job analysis for the 1981 promotional procedure as a whole adequate, but the three components of that procedure — the written exam, the performance appraisal and the assessment center — were properly weighted in accordance with the number of job analysis points tested by each component. In 1981, the written exam was weighted 25%, the performance appraisal, 30% and the assessment center, 45%. Finally, the distribution of questions on the 1981 written exam, which tested only for the SKAP of knowledges, reflected the number of job analysis points assigned to each identified knowledge area.

The piggybacking of the 1982 promotional procedure onto the 1981 job analysis presents difficulties. To begin with, there exists some question as to whether work behaviors (job components and tasks), identified as early as 1979, retained vitality as late as 1982. The 1982 Validation Report makes no mention of any inquiry as to whether the sergeant’s job, and the work behaviors or job components associated therewith, changed somewhat between 1979 and 1982. Cf. Vanguard I, supra at 740. Further, and of much greater concern, are the unexplained modifications made by the Commission in 1982 in connection with the 1982 use of the 1981 job analysis. The 1982 Validation Report states that “[sjome of the factors rated or assessed in 1981 were deleted from the 1982 examination.” 1982 Validation Report, supra, at 2. In fact, of the 28 SKAPs measured in 1981, only twelve were tested for in the 1982 promotional procedure. Among the SKAPs eliminated in 1982 was supervisory ability, which was deemed to be the second most important SKAP in the 1981 job analysis, receiving 968 job analysis points out of a total of 15,160 job analysis points. Indeed, the failure to test in 1982 for such a critical work behavior might alone be sufficient to defeat defendants’ claim of content validity. In Firefighters Institute for Racial Equality v. United States, 549 F.2d 506, 511-14 (8th Cir.), cert. denied, 434 U.S. 819, 98 S.Ct. 60, 54 L.Ed.2d 76, the Court concluded that a fire captain’s examination, which did not test for supervisory ability, was fatally deficient. The Court there noted that that supervision was the fourth most important task of a fire captain’s job, and that'supervision was “the only major job attribute that separates a firefighter from a fire captain.” 549 F.2d at 511. Here, too, supervision is one of the major factors separating the job of a police officer from that of a police sergeant. In 1981, the City did test for supervisory ability by way of the assessment center. It did not so do in 1982 because the assessment center had been eliminated and replaced with a more limited oral examination. The Court in Firefighters did not accept the contention that the City of St. Louis could delete a major job attribute because “an Assessment Center is too expensive.” 549 F.2d at 512. This Court concurs.

Other significant SKAPs no longer tested for by the 1982 promotional procedure included written communication skills (725 job analysis points in 1981), honesty (612 job analysis points in 1981), objectivity (43 job analysis points in 1981) and logical reasoning (422 job analysis points in 1981). It is also to be noted that while the Commission in 1982 eliminated some SKAPs with significant job analysis point values, it continued to test for SKAPs with relatively low job analysis point values, including Ability to Act under Stress (190 job analysis points in 1981) and Report Preparation Ability (239 job analysis points in 1981). Those changes, in and of themselves, impaired the integrity of the 1981 MSI job analysis. Accordingly, it cannot be said that the 1982 job analysis, which tested for only 12 SKAPs, accurately identified important work behaviors of a sergeant’s job. See Guidelines § 14(C)(2).

Further, the relative importance of each SKAP was incorrectly assessed in 1982 because the Commission re-weighted the three components of the promotional procedure (written examination — 40%; promotional appraisal — 30%; oral examination— 30%). The Commission’s seemingly random inclusion and exclusion of SKAPs from the 1981 job analysis resulted in an end product in 1982 which omitted important work behaviors and overemphasized the relative weights of certain work behaviors which were tested for.

2. The Test Construction Process

With a job analysis of dubious accuracy, defendants must shoulder an unusually heavy burden to demonstrate that the written exam, as constructed, is content valid. “Because of the unlikelihood that an examination prepared without benefit of a probing job analysis will be content valid, ... in the absence of such an analysis the proponent of the examination carries a greater burden of persuasion on the issue of job-relatedness.” Guardians Ass’n of New York City v. Civil Service Comm’n (Guardians V), 633 F.2d 232, 242-43 (2d Cir.1980), cert. denied, — U.S. —, 103 S.Ct. 3568, 77 L.Ed.2d 1410 (1983). In a similar vein, Judge Weinfeld observed that a showing of a substandard job analysis must be met by “the most convincing testimony as to job-relatedness,” Vulcan Society, supra, 360 F.Supp. at 1276. In affirming Vulcan on appeal, Judge Friendly characterized with seeming approval Judge Weinfeld’s approach as follows: “[T]he poorer the quality of the test preparation, the greater must be the showing that the examination was properly job-related, and vice versa ”, Vulcan Society of New York City Fire Dept., Inc. v. Civil Service Comm’n, 490 F.2d 387, 396 (2d Cir.1973). It is in this context, then, that the written test itself must be analyzed.

The 1982 written exam was developed “in-house” by staff members of the Commission. MSI, which had played an instrumental role in developing the 1981 sergeant’s promotional procedure, did not participate in designing the 1982 written exam. Rather, MSI, at most, answered questions propounded to it by the City and assisted the City in post-exam administration validation analyses. While, as Judge Newman has written, “the law should not be designed to subsidize specialists, ... employment testing is a task of sufficient difficulty to suggest that an employer dispenses with expert assistance at his peril____ [T]he decision to forgo such assistance should require a Court to give the resulting test careful scrutiny.” Guardians IV, supra, 630 F.2d at 96. Moreover, it is further worthy of note that many “in-house” examinations scrutinized by the courts have failed to pass muster.

Exam No. 820508 (i.e., the 1982 written exam) does not meet the Guidelines standard required for test construction. That written exam was designed to measure 18 different knowledges, and reading comprehension. Reading comprehension had not been tested for in the 1981 written exam because MSI personnel believed that it could best be measured through the assessment center. Because of the addition of reading comprehension to the written exam, the 1982 written test was expanded from 100 items in 1981 to 115 items in 1982.

In 1982 the Commission determined how many questions to allocate to each identified knowledge by reference to the number of job analysis points associated with each knowledge. Yet despite the Commission’s initial care in 1982 in testing for each knowledge in proportion to its relative importance to the sergeant’s job, the test, as finally scored, was not proportionately representative of each knowledge area. The deletion of 18 questions by the Commission prior to final scoring of the written exam resulted in a written exam which, as scored, placed far too little emphasis on the knowledge areas of “Reports” and “Search and Seizure.”

Further, and of far greater concern, is the fact that the 115 questions were apparently drafted in an haphazard manner. The questions were written by unidentified Commission personnel, who have not been shown to be expert in the art of test-writing. Indeed, according to the Commission’s 1982 Validation Report, the item-writers lacked the expertise to spot “highly technical flaws in the questions.” Additionally, the 1982 Validation Report does not state whether, and to what extent, the item-writers relied on the job analysis materials. Nor, as far as the record herein reflects, were the 1982 questions ever reviewed by persons within or without the Department for accuracy or reliability. The 1982 Validation Report explains that the items were not reviewed by Department personnel for “security reasons.” Thus, incumbent sergeants had no input into the test-construction process and were not provided an opportunity to comment on whether the items tested for the knowledges which they purported to test for. Equally disturbing is the fact that there was no pre-test administration review to assure that the questions were not ambiguous, overly complex, overly specialized, dependent on prior knowledge, or based on information which would be acquired during a training period. Nor were the questions tested on a sample population, even though such pilot testing had been employed by the City in selecting the items for inclusion in the 1981 written exam. Moreover, no item-analysis was performed on the questions prior to their use, despite the fact that such was considered necessary by MSI in constructing the 1981 written exam. Specifically, prior to administering the 1981 written exam, item-impact and item-discrimination analyses were performed. Those analyses, which were designed to select the items most predictive of job performance and with the least racially adverse impact, were not explored prior to test administration in 1982. In short, the Commission in 1982 made far too little effort to ensure that the test items were reliable, comprehensible, ambiguous or racially-neutral in impact.

Not surprisingly, the 1982 test construction process resulted in a written exam of questionable validity. The Commission itself found 18, or 15.7%, of the items to be defective and removed them before scoring the exam. Two other items were rekeyed on the grounds that they had been miskeyed initially. Further, as Dr. Barrett noted in his March 23, 1983 Critique, 17 or 14.8% of the items were answered correctly by 90% or more of the candidates and, therefore, contributed little to the value of the exam. Additionally, Dr. Barrett expressed the view that eight items were “inappropriate distractors,” in that an incorrect alternative was selected more often than the keyed answer.” In his October 13, 1983 affidavit, Dr. Barrett also expressed the view that of the 80 “knowledge” items reviewed by the panelists, only 56%, or 45 questions, were part of the sergeant’s job and that 26% of the questions were the responsibility of someone else. Illustrative of this type of question is item 25 which reads: “Who makes the final decision on a complaint of excessive force filed against a member of the Dept.? The (a) Police Commissioner; (b) Complaint Evaluation Board; (c) Director of the Internal Investigation Division; (d) Commanding Officer of the Person charged.” That question, in addition to being the responsibility of someone other than a sergeant, is ambiguously worded. That is true because although the Police Commissioner takes final action on a complaint of excessive force, the Complaint Evaluation Board is charged, by statute, with making final decisions on such complaints. Moreover, question 25, like so many others, does not test for critical knowledge. Indeed, of the 45 questions which were deemed to be job-related, many were, at best, marginally so. An example of a question identified as marginally related to the sergeant’s job is item 38, which states:

According to the Digest of Laws, concerning domestic violence which of the following acts between family members would be considered “abuse”?

(a) Verbally assaulting another;

(b) Putting another in constant fear of bodily harm;

(c) Malnutrition of children;

(d) None of the above.

It is of marginal importance that a sargeant knows the Digest of Laws definition of domestic abuse. A prosecutor will be responsible for making an official charge of abuse. The sergeant’s role is more centrally what action should be taken if he is confronted with any one of the alternatives delineated in question 38(a)-(c).

At trial, Dr. Barrett testified that “a big problem with the test is ambiguity,” specifically observing that items could not be deemed job-related if a test-taker could not understand what was being asked, Dr. Barrett identified at least 15 items, out of the 80 knowledge items, which he considered to be ambiguous. This Court agrees with Dr. Barrett that far too many of the 80 knowledge items are ambiguous.

Moreover, Dr. Barrett was not the only witness who testified to the fact that a significant number of questions were either ambiguous or unrelated to satisfactory performance of the sergeant’s job. The City’s own witness, Major Norris, testified to such. In particular, Major Norris identified 17 of the first 80 questions on the written exam which were not job-related. Of those 17 questions, Major Norris stated that five items either had more than one correct answer or no correct answer. Additionally, the Major expressed the view that at least seven of those questions covered information wholly or substantially outside the scope of a sergeant’s job and that other questions challenged were ambiguous.

While this Court is mindful of Judge Newman’s warning that the “ ‘burden of judicial examination-reading’ ... need not inevitably be assumed,” Bridgeport Guardians v. Bridgeport Police Dep’t, 432 F.Supp. 931, 937 (D.Conn.1977), the assumption of that burden is appropriately undertaken by this Court because it is important herein to determine the approximate number of questions which are subject to attack as non-job-related. Further, the subject matter of the 1982 written exam, namely, the knowledge required to be a good sergeant, is not too far removed from judicial competence.

Herein, defendants have the burden of showing, in the face of conceded disparate racial impact, that the written exam tested for those knowledges which “are critical and not merely peripherally related to successful job performance,” Kirkland, supra, 374 F.Supp. at 1372. In the words of the Uniform Guidelines, the employer “should show that (a) the selection procedure measures and is a representative sample of that knowledge, skill, or ability; and (b) that knowledge, skill, or ability is used in and is a necessary prerequisite to performance of critical or important work behaviors).” Guidelines § 1607.14(C)(4). In this case, this Court is unconvinced that the 1982 written exam adequately tested for the knowledge areas which it was designed to test for. And, the examination also seemingly failed to test for those knowledges which are a necessary prerequisite to the performance of critical or important aspects of a sergeant’s job. Approximately 17 of the first 80 questions of the 1982 written exam relate to information which is only of marginal importance to a sergeant’s job. Further, an additional 12 of the first 80 questions are so ambiguous or incomprehensible that they cannot be said to measure the knowledge areas which they were intended to test for.

To sum up, 14 of the first 80 questions were deemed deficient by the Commission and were deleted. An additional 29 questions out of the first 80 are substandard for reasons set forth above in this opinion. Thus, considering only the first 80 items, arguably there remain at the most 37 “good” questions. That finding alone warrants a conclusion that the 1982 written exam is fatally deficient. In addition, it is to be noted that with respect to questions 101 through 115, they dealt with Department procedures and have not been vigorously analyzed or challenged by plaintiffs. Yet, defendants’ witness, Major Norris, when asked to comment on questions 101 through 115, expressed the view that 6 of those 15 items were not job-related. Thus, of the final 97-item test (115 items minus 18 items deleted by the Commission), 35 (29 out of the first 80 and 6 out of the last 15) are not job-related. Those numbers speak for themselves and strongly suggest that Exam No. 820508 is not content valid. However, while the substantial shortcomings of the job analysis and test construction process might alone compel the conclusion that the written exam is not content valid, it would appear helpful, at least in terms of the future, for this Court to reach and discuss other factors related to content validity.

3. The Direct Relationship Requirement

As the Second Circuit has observed, the central requirement of Title VII is “relationship of test content to job content.” Guardians IV, supra, 630 F.2d at 97-98. In this case, it is questionable, at best, whether Exam No. 820508 tests for those knowledges which are requisite to successful performance of a police sergeant’s job. While the 1981 job analysis adequately identified the relevant knowledge areas necessary to performance of the sergeant’s job and while those same knowledge areas ostensibly formed the basis for the 1982 written exam, there is considerable doubt as to whether the 1982 test, as constructed, measured such knowledge areas. In other words, although the knowledge areas sought to be tested, including “evidence collection,” “warrants,” “discipline,” “crime-related matters,” “reports,” and “court-related matters,” appear to be related to information which a “good” sergeant should know, the questions relating to each knowledge area may in fact test for unimportant rather than important information.

For example, the knowledge area of “crime related matters” received the most job analysis points in 1981 and was most heavily tested for in 1982. Yet, of the 15 questions devoted to that knowledge area, most did not test for crime related information which is central to a sergeant’s successful job performance. Thus, question 35 asked: “Which of the following states does not participate in the Non-Resident Violator Compact?” There were some testimonial differences during trial concerning whether such a compact in fact exists. Even assuming one does exist, it is doubtful, at best, whether a sergeant needs to memorize the information. Rather, it would seem likely that a sergeant could refer to reference materials when the need to know arises. Moreover, the knowledge tested for does not appear to go to the core of crime-related matters. The same can be said of question 37 which asked: “How the word wilful, used in the term first degree murder, was defined according to the Digest of Laws.” Again, it is hard to see how knowledge of such a definition in the Digest of Laws is critical to a sergeant’s job. Someone other than a sergeant will presumably determine whether a murder is chargeable as first degree murder. Questions of this sort predominate throughout the 1982 written exam to far too great an extent, and militate against a finding that the 1982 written exam satisfies the direct relationship requirement of content validation.

4. The Representative Requirement

The Guidelines require that a test be “a representative sample of the job.” Guidelines § 14(C)(4). That requirement is construed primarily to mean that the content of the test must be proportionately representative of the content of the job. Many courts have construed that requirement strictly. For example, in Kirkland v. New York State Dep’t, supra, 374 F.Supp. at 1372, Judge Weinfeld observed that the test must be shown “to examine all or substantially all the critical attributes of the sergeant’s position in proportion to their relative importance to the job and at the level of difficulty which the job demands.” Similarly, Judge Ready commented: “A content valid examination must not only test required skills and aptitudes, it must test them in proportion to their relative importance on the job.” Walls v. Mississippi State Dep’t of Public Welfare, 542 F.Supp. 281, 312 (N.D.Ill. 1982). The Standards for Educational and Psychological Tests of the American Psychological Association state: “[A]n employer cannot justify an employment test on the grounds of content validity if he cannot demonstrate that the content universe includes all, or nearly all, important parts of the job ” (at 29), cited favorably in United States v. City of Chicago, 573 F.2d 416, 425 (7th Cir.1978) (emphasis added by Chief Judge Fairchild). In Guardians TV, supra, the Second Circuit rejected the contention that an examination must test for all abilities of a job and for each in its proper proportion, and adopted the following more relaxed standard:

The reason for a requirement that the content of the exam be representative is to prevent either the use of some minor aspect of the job as the basis for the selection procedure or the needless elimination of some significant part of the job’s requirements from the selection process entirely; this adds a quantitative element to the qualitative requirement— that the content of the test be related to the content of the job. Thus, it is reasonable to insist that the test measure important aspects of the job, at least those for which appropriate measurement is feasible, but not that it measure all aspects, regardless of significance, in their exact proportions.

630 F.2d at 99 (emphasis in original).

Precise proportionality of each knowledge in relation to its importance need not be demonstrated. Similarly, measurement of each and every knowledge area is not required. Yet significant knowledge areas, comprising up to 30% of the job, may not be omitted from a written promotional exam simply because testing for such knowledge areas is difficult. Rather, in order for a written exam to pass muster under the representativeness requirement, the exam must measure all, or nearly all, of the significant knowledge areas of a job in approximate proportion to each knowledge area’s relative importance to the job.

Assessed in the light of that standard, the 1982 written exam is not adequately representative. While the written exam purported to measure the relative importance of 18 knowledge areas, the exam failed to measure those knowledge areas because of faulty test construction. Moreover, the deletion of 18 test items skewed the actual weight attributed to each knowledge area. Finally, knowledge areas and reading comprehension, the two SKAPs ostensibly measured by the written exam, were weighted too heavily in 1982 because of the reweighting of the written exam component.

5. The Scoring Requirement

As described supra, the 1982 promotional procedure was scored as follows: A candidate’s scaled written exam score (weighted 40%) was combined with his promotional appraisal score (weighted 30%). On the basis of that composite score, the top 95 candidates were administered the oral examination (weighted 30%). Thus, a cutoff score was established after two parts of the procedure were administered. That cutoff score was not a score which indicated a candidate’s ability to perform the job but was simply the composite score of the 95th candidate at that stage of the procedure. After the oral exam was scored, the 95 candidates were placed on the eligibility list in order of their rank. The rankings were derived by adding a candidate’s weighted, scaled score on each of the three promotional procedure components. Although each candidate was not ranked solely on the basis of his written exam score, that score was a significant factor because of the heavy weighting given to it, i.e., 40% of the total promotional procedure score. Similarly, the written exam score was the most important factor in the cutoff score (57% of the composite score before administration of the oral exam).

The Guidelines provide that rank ordering should be employed only if it can be shown that “a higher score ... is likely to result in better job performance.” Guidelines § 14(C)(9). Further, the Guidelines state that “evidence which may be sufficient to support, the use of a selection procedure on a pass/fail (screening) basis may be insufficient to support the use of the same procedure on a ranking basis ____” Guidelines § 1607.5(G).

One court has summarized those Guidelines requirements by stating that “[i]f test scores do not vary directly with job performance, ranking the candidates on the basis of their scores will not select better employees.” Guardians IV, supra, 630 F.2d at 100. Further, the inference that higher scores closely correlate with better job performance must be “closely scrutinized.” Id. This close scrutiny is required because “[a] test may have enough validity for making gross distinctions between those qualified and unqualified for a job, yet may be totally inadequate to yield passing grades that show positive correlation with job performance.” Id. at 100.

Courts have consistently disapproved of ranking where a test’s content validity is suspect. In Guardians IV, supra, the Second Circuit stated that “the defects we noted in the job analysis and the test construction are substantial enough to preclude an inference that passing scores will correlate with job performance closely enough to justify rank-ordered selections.” 630 F.2d at 101. Similarly, in Firefighters Institute v. City of St. Louis, 616 F.2d 350, 358 (8th Cir.1980), cert. denied sub nom. City of St. Louis v. United States, 452 U.S. 938, 101 S.Ct. 3079, 69 L.Ed.2d 951 (1981), the Court noted that the EEOC’s “Questions and Answers ... specifically require empirical evidence that mastery of more knowledge is linked with better performance on the job.” (Emphasis in original). Finding that no empirical evidence had been adduced demonstrating an association between levels of performance on the multiple choice exam and the job, the Court there held the rank-ordering was invalid. Id. at 357-60. Finally, in Berkman v. City of New York, supra, the Court criticized the City’s rank-ordering of a firefighter’s exam, stating, “neither the job analysis instrument, the test instrument, or the validation instruments appear able to perform their tasks with the precision necessary to justify the rank-ordering used here.” 536 F.Supp. at 212. The Court went on to note that:

This conclusion does not necessarily mean that a future exam must be administered on a pass/fail basis. It does suggest, however, that some larger use of random selection within fewer, more rationally grounded ranks will have to be substituted for the present system, unless finer test instruments can be found.

Id. See also EEOC Questions and Answers, supra n. 76, Q. 62 which states that it is “easier” to make the inference of a relationship between higher scores and better job performance “the more closely and completely the selection procedure approximates the important work behaviors.”

In the case at bar, the rank-ordering of the 1982 promotional procedure cannot be justified because the written exam, which was a substantial ingredient of a candidate’s ranking, possessed insufficient validity and reliability. Here, as described in detail, supra, there exist substantial defects in the written exam which dictate against any use of the test, particularly a scaled usage, which forms the basis for ranking. The 1982 written exam fails to test adequately for important knowledges, and there is no indication that there exists any association between levels of performance on the written exam and on the job. The written exam lacks another essential feature required before rank-ordering can be justified, namely, reliability. Specifically, an exam is said to be reliable if there is some likelihood that the exam will produce consistent results among applicants who repeatedly take it or a similar exam. “Like content validity, reliability is not an all or nothing matter. It too comes in degrees. What is required is not perfect reliability, but rather a sufficient degree of reliability to justify the use being made of the test results.” Guardians IV, supra, 630 F.2d at 101.

In the case at bar, no evidence has been produced to indicate that the 1982 written exam is reliable. Indeed, important indicia of reliability are absent. The exam questions are not of high quality. Thus, there is no reason to believe success on one question will correlate with success on other questions. Further, no reliability analyses were performed by the test designers prior to test administration. No pilot testing of questions was undertaken on a sample population, despite the fact that such was done in 1981. Nor was a “reliability estimate” computed. See, e.g., Rivera v. City of Wichita Falls, 665 F.2d 531, 537 (5th Cir.1982). While a post-test administration, upper-group/lower-group item analysis was performed after the 1982 test was administered, the results of that analysis were not relied upon by the City to demonstrate reliability and, indeed, the results seemingly do not indicate such. In addition, the reliability of the questions is suspect because of the high number of easy questions (17) and of the high number of inappropriate distractors (8). Easy questions — ones answered correctly by 90% or more of the candidates — contribute little to the value of a test because they do not differentiate among candidates based on their respective capabilities and “magniffy] effects that may make scoring arrangements unjustified.” Guardians IV, supra, 630 F.2d at 103 n. 19. Inappropriate dis-tractors — “wrong” answers which attract more response