Citations
- 815 F.2d 84
Full opinion text
Opinion for the Court filed by Chief Judge WALD.
WALD, Chief Judge:
In this action, a class of women plaintiffs allege various forms of unlawful employment discrimination in the Foreign Service from 1976 to 1983. After a trial, the District Court found that no unlawful discrimination had occurred. See 616 F.Supp. 1540 (D.D.C.1985). This appeal followed. The record, however, discloses that the District Court’s decision was premised on errors of law and that several of its critical findings of fact were clearly erroneous. Consequently, we reverse, and remand for further proceedings in accordance with this opinion.
I. Background Information
A. The Foreign Service and Its Employment Practices
The Foreign Service is our nation’s professional diplomatic corps. Members of the Service represent the interests of this nation abroad and assist the Secretary of State in the formulation of foreign policy at home. See 22 U.S.C. § 3904(1)-(1)2). The organization of Foreign Service personnel draws on the model of the United States military as well as the United States civil service. See S.Rep. No. 913, 96th Cong., 2d Sess. 2 (1980), U.S.Code Cong. & Admin. News 1980, P. 4419. For example, the Foreign Service is a “rank-in-person” system: members of the Service have an individualized rank which is independent of the rank of the particular job they happen to hold at any given time. H.R.Rep. No. 992, pt. 1, 96th Cong., 2d Sess. 3 (1980).
The Foreign Service also copies the military in its “up or out” personnel system. Individuals must serve a probationary period of up to five years before they can receive a career appointment in the Service. 22 U.S.C. § 3946. If at the end of that period an individual has not received a career appointment, he or she must leave the Service. Id. § 3949. (Although according to the Foreign Service Act of 1980, the term “Foreign Service Officer” refers only to members of the Service with career appointments, and those serving under a limited, probationary appointment are called “career candidates,” the parties to this lawsuit use the term “Foreign Service Officer,” or “FSO,” to refer to those serving under both career and limited appointments. To avoid confusion, we will do likewise.)
The Foreign Service assigns its officers to one of four areas of functional specialization, known as “cones”: political, economic, administrative, and consular. Officers in the political and economic cones deal with, respectively, political and economic dimensions to foreign relations and foreign policy. Officers in the administrative cone “are responsible for the support operations of U.S. embassies and consulates.” 616 F.Supp. at 1544 (115). Officers in the consular cone “work closely with the public providing assistance to American travelers and residents abroad, issuing visas [and dealing with] other immigration related issues.” Id. (116). As the District Court expressly found, the State Department does not encourage FSOs to change cones, and “[officers are expected to serve the major portion of their time in the Service” in the cones to which they were initially assigned. Id. (111110,14). Some officers, however, do switch cones. Senior FSOs who have demonstrated leadership ability may transfer into a “prestigious” program direction cone. Id. at 1554 (11104). Other FSOs are occasionally given temporary assignments to other cones or to some “inter-functional” positions. Id. at 1550 (H 70).
Most FSOs applying to the Foreign Service at junior entry levels must take a written examination. Beginning in 1975, the examinations have tested applicants for aptitude in all four functional areas, and the Foreign Service has used the results of these examinations to determine a new FSO’s initial cone assignment. Id. at 1545 (1115.) A relatively small number of individuals have entered the Service laterally as mid-level FSOs. These lateral entrants bypassed the examination process and “selected, in advance, the functional field in which they wished to compete and were evaluated only for that specific cone.” Id. (H 17).
Once in the Foreign Service, individuals change specific jobs frequently; the State Department has a policy of assigning individuals to positions for a set period of time, generally two to three years. See id. at 1550 (1171); H.Rep. No. 96-992, pt. 1, 96th Cong., 1st Sess. 3 (1980). Since 1975, job assignments in the Foreign Service have been made pursuant to an Open Assignment Policy, in which all members of the Service receive a list of vacant positions and submit “a bid list” indicating their preferences. These bid lists are compiled into a “bid book” from which assignment panels make their selections, after considering the interests and preferences of the bureau in which each position is located. Id. at 1550 (H1Í 73, 74). As previously indicated, some FSOs receive “out-of-cone” assignments pursuant to this process but in the main, job transfers are made inside the cones of initial assignment. In addition, FSOs do not necessarily receive a job position with a rank corresponding to the individual’s personal rank. Positions that have a higher rank than the individual are known as “stretch” assignments. Positions with a lower rank than the individual’s are “down-stretch” assignments. Pursuant to the Open Assignment Policy, individuals do not receive stretch or down-stretch assignments unless they bid for them, but as with any other assignment, individuals do not receive these assignments simply because they bid for them. Id. at 1551 (1177).
The Foreign Service prepares annual written evaluations of its officers’ job performance. In addition to rating the actual past performances of FSO’s, the evaluations rate the potential of the FSOs future job performance. 616 F.Supp. at 1549. The State Department also gives out Honor Awards in recognition of outstanding achievement. In descending order of prestige are the Distinguished Honor Award, the Superior Honor Award, and the Meritorious Honor Award. See Plaintiffs’ Post-Trial Brief at 112-13.
Except for Senior members, salaries in the Foreign Service are based on a schedule established by the President which consists of nine salary classes. 22 U.S.C. § 3963. The Secretary of State assigns all Foreign Service Officers to a particular salary class. Id. § 3964. By statute, except in limited circumstances, a career candidate for appointment as a Foreign Service Officer may not be initially assigned to a salary class higher than class 4 (class 1 being the highest). Id. § 3947. Usually career candidates are placed initially in class 7 or class 8. Promotions from one salary class to another are made by the Secretary of State after receiving recommendations and rankings submitted by selection boards which evaluate the members of each class. Foreign Service Officers do not compete for promotions until the transition from class 6 to class 5; until then, they are promoted at the end of an established time period if they perform their duties satisfactorily. See Joint Appendix (“J.A.”) at 117-121; Defendant’s Post-Trial Brief at 96.
B. The History of This Litigation
This class action began over ten years ago when appellants filed their complaint alleging that widespread discrimination against women in the Foreign Service violated Title VII of the Civil Rights Act of 1964, as amended in 1972 to cover employment discrimination in the federal government. See 42 U.S.C. § 2000e-16. The parties subsequently resolved by consent decree all claims relating to admission into the Foreign Service. The appellants’ claims of discriminatory personnel actions against women already in the Foreign Service proceeded to trial in the District Court. The parties agreed to try initially only the issue of liability, leaving appropriate remedies to a subsequent phase of the proceedings, if necessary. After trial on the liability issue, the District Court concluded that appellants “failed to show by a preponderance of the evidence any sexual discrimination by the State Department.” 616 F. Supp. at 1561. The court entered a final judgment for the Secretary of State, dismissing the complaint. Id.
This appeal followed from the District Court’s failure to find sex discrimination in seven different types of personnel practices. First, the appellants claim that from 1976 to 1983, the Foreign Service discriminated against women in the initial cone assignments of entering FSOs; the State Department assigned proportionally fewer women than men to the political cone and proportionately more women than men to the consular cone. This disparity was allegedly caused by the differing scores of women and men on the Foreign Service entrance examinations, producing a disparate impact on women and men candidates in violation of Title VII. Second, women were given proportionally fewer out-of-cone assignments to the program direction cone and proportionally more out-of-cone assignments to the consular cone. Third, women were given proportionally fewer “stretch” assignments and proportionally more “downstretch” assignments than men in the same class. Fourth, women received a disproportionately low number of appointments as Deputy Chief of Mission, the position just below that of Ambassador. Fifth, in its evaluation reports, the State Department gave lower future potential ratings to women than men despite equivalent ratings for their past performance. Sixth, women received a disproportionately low number of Foreign Service Honor Awards. And seventh, the State Department promoted women from class 5 to class 4 at a lower rate than it promoted men.
With respect to each of these seven personnel practices, the appellants offered data showing a disparity between men and women, along with a statistical analysis designed to demonstrate the improbability that a disparity of that scale could result from chance. The data and analysis, they allege, provide a strong basis for inferring that this disparity was the product of unlawful discrimination. In addition, the appellants introduced nonstatistical evidence pertaining generally to the existence of a prejudicial attitude towards women in the Foreign Service from 1976 to 1983. The District Court, however, rejected the inference of unlawful discrimination in each of the seven areas.
In discounting the probative force of appellants’ statistics, the District Court said that their statistical studies rested on faulty data, or flawed methodology, or omitted a crucial variable that would explain the disparity between men and women in a nondiscriminatory way. The District Court also said that some of the statistical evidence focused on too narrow a segment of Foreign Service personnel practices. As we shall explain, the District Court’s treatment of the appellants’ evidence was in some instances contrary to law and in other respects clearly erroneous as a matter of fact.
II. Title VII Claims: Two Different Theories
Under Title VII a plaintiff can rely on either of two different theories to support a claim of unlawful sex discrimination. A “disparate treatment” claim alleges that the defendant intentionally based an employment decision on the sex of the plaintiffs. See, e.g., International Brotherhood of Teamsters v. United States, 431 U.S. 324, 335 & n. 15, 97 S.Ct. 1843, 1854 & n. 15, 52 L.Ed.2d 396 (1977). Disparate treatment claims can involve an isolated incident of discrimination against a single individual, or, as in this case, allegations of a “pattern or practice” of discrimination affecting an entire class of individuals. Id. A “disparate impact” claim alleges that the defendant based an employment decision on a criterion that although “facially neutral” nevertheless impermissibly disadvantaged individuals of one sex more than the other. Id. at 336 n. 15, 97 S.Ct. at 1854 n. 15. This case is a “classic” example of a disparate impact claim in which plaintiffs allege that the defendant based employment decisions on the results of a test for which members of one sex on average received lower scores than members of the other sex. See B. Schlei & P. Grossman, Employment Discrimination Law at 13 (1983-84 Supp.); see also Griggs v. Duke Power Co., 401 U.S. 424, 91 S.Ct. 849, 28 L.Ed.2d 158 (1971) (the original disparate impact case).
Because these two theories are distinct, we must consider them separately. Appellants’ only disparate impact claim concerns the initial cone assignments; the other six claims involve disparate treatment and we will consider them first.
III. Legal Principles Applying to Pattern or Practice Disparate Treatment Claims
In a typical sex discrimination pattern or practice disparate treatment case, plaintiffs allege the existence of a disparity between men and women in selection rates for a particular job or job benefit and further allege that this disparity was caused by an unlawful bias against members of the disadvantaged sex, usually women. To prevail in their claim, plaintiffs must prove, by a preponderance of the evidence, that these allegations are true. Proof of the disparity itself is based upon a comparison of the proportion of those women eligible for selection who were actually selected with the corresponding proportion of eligible men who were actually selected. Plaintiffs establish a disparity disfavoring women if the evidence demonstrates that the selection rate for eligible women was less than the selection rate for eligible men. Sometimes, the disparity is expressed as the difference between the number of women actually selected and the number of women one would expect to have been selected, assuming equality in the selection rates for men and women. (If one knows the number of women eligible and the selection rate for men, one can determine, using algebra, the expected number of successful women.)
Proof that the observed disparity was caused by an unlawful bias against women need not be direct. Circumstantial evidence that the disparity, more likely than not, was a product of unlawful discrimination will suffice to prove a pattern or practice disparate treatment case. See Teamsters, 431 U.S. at 335 n. 15, 97 S.Ct. at 1854 n. 15. Indeed, this circumstantial evidence may itself be entirely statistical in nature. See, e.g., Segar v. Smith, 738 F.2d 1249, 1278-79 (D.C.Cir.1984), cert. denied sub. nom. Meese v. Segar, 471 U.S. 1115, 105 S.Ct. 2357, 86 L.Ed.2d 258 (1985). In this case, appellants rely to a great extent on statistical evidence to prove their claims of disparate treatment. We find it necessary, therefore, to discuss how statistical analysis of an observed disparity can raise an inference of unlawful discrimination.
A. Raising An Inference of Discrimination With Statistical Evidence
A disparity between the selection rates of men and women for a particular job or job benefit has one of three possible causes. See D. Baldus & J. Cole, Statistical Proof of Discrimination 291 (1980). First, the disparity may be a product of an unlawful discriminatory animus; this is what plaintiffs are attempting to prove. Second, the disparity may have a legitimate and nondiscriminatory cause. For example, prior experience of a certain type may be an important factor in making certain employment decisions, and if it happened to be true that women on the average have less of this experience than men, one would expect that women could be selected less frequently. Third, the disparity may simply be a product of chance. Even if we may properly assume that, as a general rule, women and men on average are equally qualified to be selected for a particular job or job benefit, for any particular group of men and women who happen to constitute the actual pool of eligible candidates at the time the selections are made, there may be some deviation from this general rule because the actual qualifications of men and women differ from individual to individual and any particular pool of eligible candidates constitutes an inherently random collection of individuals. Thus, even if selections were made entirely on the basis of qualification, without a trace of discriminatory bias, random deviations in the selection rates for men and women may result.
A statistical analysis of a disparity in selection rates can reveal the probability that the disparity is merely a random deviation from perfectly equal selection rates. Statistics, however, cannot entirely rule out the possibility that chance caused the disparity. Nor can statistics determine, if chance is an unlikely explanation, whether the more probable cause was intentional discrimination or a legitimate nondiscriminatory factor in the selection process. See id. at 290-92.
Title VII nevertheless provides that if the disparity between selection rates for men and women is sufficiently large so that the probability that the disparities resulted from chance is sufficiently small, then a court will infer from the numbers alone that, more likely than not, the disparity was a product of unlawful discrimination — unless the defendant can introduce evidence of a nondiscriminatory explanation for the disparity or can rebut the inference of discrimination in some other way. See Hazelwood School District v. United States, 433 U.S. 299, 307-08, 97 S.Ct. 2736, 2741, 53 L.Ed.2d 768 (1977) (“Where gross statistical disparities can be shown, they alone in a proper case constitute prima facie proof of a pattern or practice of discrimination.”); see also Segar, 738 F.2d at 1278 (“[W]hen a plaintiffs methodology focuses on the appropriate labor pool and generates evidence of [a disparity] at a statistically significant level,” this evidence alone will be “sufficient to support an inference of discrimination.”).
The preliminary question for a court, then, is at what point is the disparity in selection rates is sufficiently large, or the probability that chance was the cause sufficiently low, for the numbers alone to establish a legitimate inference of discrimination. Although this question is crucial in Title VII litigation, the answers given by courts have been regrettably imprecise. The Supreme Court has twice stated that “[a]s a general rule for ... large samples, if the difference between the expected value and the observed number is greater than two or three standard deviations, then the hypothesis that [the disparity] was random would be suspect to a social scientist.” Castaneda v. Partida, 430 U.S. 482, 497 n. 17, 97 S.Ct. 1272, 1281, n. 17, 51 L.Ed.2d 498 (1977); see also Hazelwood, 433 U.S. at 309 n. 14, 97 S.Ct. at 2742 n. 14 (quoting Castaneda). But many lower courts and commentators have noted that the difference between two and three standard deviations is considerable and that, therefore, the Supreme Court’s statement falls short of establishing an exact legal threshold at which statistical evidence, standing alone, establishes an inference of discrimination. See, e.g., Segar, 738 F.2d at 1283 n. 28.
This court, using different terminology, has stated that statistical evidence meeting “the .05 level of significance ... [is] certainly sufficient to support an inference of discrimination.” Segar, 738 F.2d at 1283. “[T]he .05 level,” the Segar opinion explained, “indicates that the odds are one in 20 that the result could have occurred by chance.” Id. at 1282. (This statement is somewhat imprecise and has predictably led to confusion, as we discuss infra.) The Segar court justified the consistency of its statement with the statements of the Supreme Court by observing that “[a] level of two standard deviations corresponds to statistical significance at the .05 level.” Id. at 1283 n. 28. In this case, the District Court cited Segar in its Conclusions of Law, stating: “The Court adopts the .05 level for establishing that a [statistical] study is statistically significant.” 616 F.Supp. at 1559 (if 14). But the District Court then went on to say that “[t]he .05 level generally corresponds to 1.65 standard deviations.” Id.
How can a 5% probability of randomness correspond both to a measurement of two standard deviations and a measurement of 1.65 standard deviations, one may reasonably ask? There is a legitimate answer: it depends on whether one is using a “one-tailed” or a “two-tailed” test of statistical significance. A disparity measuring 1.65 standard deviations corresponds to a 5% probability of randomness under a one-tailed test. A disparity measuring two standard deviations (to be more precise, 1.96 standard deviations) corresponds to a 5% probability of randomness under a two-tailed test.
This difference between one-tailed and two-tailed tests obviously requires further explanation. It also presages the obvious question, given the substantial differences in result, of which test is the more appropriate one to use in Title VII cases. Neither this court’s opinion in Segar nor the District Court's opinion in this case discusses the difference between “one-tailed” or “two-tailed” approaches. The Supreme Court has given us no explicit guidance on this issue. And, unfortunately, neither side to this litigation has devoted more than a single footnote each to this difficult but important issue. See Appellants’ Reply Brief at 32 n. 38; Appellee’s Brief at 62 n. 73. For obvious reasons we, too, confront this issue with some trepidation. But appellants’ and appellee’s evidence on the un-derpromotion of women from FSO class 5 to class 4 measures 1.88 and 1.76 standard deviations, respectively. (The difference results from the use of some different data. See 616 F.Supp. at 1557 (H130).) Whether one adopts the appellants’ or the appellees’ number as the better evidence, it falls between 1.65 and 1.95 standard deviations. Therefore, if one tests the statistical significance of this number using the Se-gar standard of a 5% probability of randomness, the outcome turns on whether one uses a one-tailed or two-tailed test. Under a one-tailed test, the number is statistically significant (because it is larger than 1.65 standard deviations, which correspondents to a 5% probability of randomness under a one-tailed test) and therefore by itself establishes a prima facie case of disparate treatment. Under a two-tailed test, the number does not quite reach the statistically significant threshold (because it is smaller than 1.96 standard deviations, which corresponds to a 5% probability of randomness using a two-tailed test) and therefore by itself does not raise an inference of discrimination.
Given the unavoidability of embarking upon a journey into the statistical maze, we begin with the terms “one-tailed” and “two-tailed”; they refer to the “tails” or ends of the bell-shape curve, which represents in graph form a “random normal distribution.” E.g., W. Curtis, Statistical Concepts for Attorneys 72-73 (1983); see Diagram 1 copied from id. In these random distributions, the area under any segment of the bell curve measures the probability of that range of results occurring randomly. Id. Furthermore, the percentage area underneath the bell curve within one standard deviation (