How Do We Measure Merit in Foreign Policy?

Dan Spokojny | August 12, 2026

Promotion Process

The old-timey baseball scouts portrayed in the movie Moneyball are gathered around a conference table assessing free agent prospects for their team. “He’s got a baseball body,” one scout asserts. “Good looking ball player. Good jaw,” adds another. But the General Manager of the team, Billy Beane, looks skeptical. He asks, “Can he hit?” A scout chimes in with confidence: “He’s got a classy swing, it’s a real clean stroke.” Another agrees: “A lot of pop coming off the bat.” Beane, growing frustrated, asks again, “If he’s a good hitter why doesn’t he hit good?”

Conversations about merit in foreign policy are often reminiscent of that scene. The grist used to form our opinions about merit often seems ambiguous and subjective. Overly-confident assertions about who is (and who is not) an expert are laid upon questionable foundations.

The concept of merit has never been more essential in foreign policy. Those who are judged to possess the most merit rise to positions of great power and influence, while those lacking merit are left behind. And ideas about merit drive the design of our institutions of foreign policy. It’s time to get clear on what merit really means, and how we identify it.

Defining Merit

The concept of merit must be anchored in an objective understanding of performance within a specified domain. The more a person is judged to have the capabilities to achieve success, the more merit they have.

A prerequisite for evaluating performance is therefore possessing a clear understanding of the goals at hand. Defining goals in a democracy is a difficult and contested process. Most modern policy environments face competing demands and priorities: Is the goal to prevent war? Defend our borders? Project American values? Fight for human rights? These are complex, highly political decisions. But if the objectives are unclear, there can be no meaningful assessment of merit. At its worst, foreign policy is the art of setting ambiguous goals so one can always claim success. But merit is only a useful concept if it captures the ability of some actor to achieve certain goals.

A healthy organization will design its personnel processes to reward the qualities most likely to lead success. The salesman may be judged on total revenue; the software engineer will be encouraged to ship error-free code, etc. In contrast, an unhealthy organization may cling to ambiguous understandings of merit. This may result in poor decisions about who to hire, who to promote, and what to train. This introduces distorting incentives. Staff may learn, for example, that the best way to get ahead is to kiss up to the most powerful official in the room.

It is worth noting that merit need not be entirely about performance. In addition to the objective qualities that comprise merit, the concept of merit may also include some more subjective facets, such as the qualities of being courageous, ethical, or friendly. For instance, an organization might decide that those who have triumphed over great obstacles to achieve their position possess more merit than those weaned with a silver spoon, even if their performance is exactly the same.

The Validation Challenge

To have a more constructive conversation about merit, it helps to understand the concept of internal validity. There are two categories: construct validity and inter-rater reliability.

Construct validity is the extent to which a metric accurately captures what it claims to measure. For instance, if a coach is trying to identify the best chess players to recruit, measuring shoe sizes will offer poor construct validity. Measuring SAT scores would offer better construct validity. Best yet, looking at applicants’ official chess rankings would offer high construct validity. But notice that the coach needs to be clear on the precise goal in recruitment: Does she care only about demonstrated skill or long-term prospects? Does it matter if the recruit is a supportive teammate?

Inter-rater reliability evaluates the uniformity between scores from different judges or measurements. If you ask five people, are you going to get the same five answers? For instance, imagine a team of magazine editors judging a poetry contest. If the question presented to the reviewers is “Which poem is best?” there will likely be wide divergence between the judges. Low inter-rater reliability doesn’t mean that none of the poems are beautiful, just that “best” isn’t a concept that’s satisfactorily measured in this way. Recognizing this challenge, the judges might agree to use more precise metrics to evaluate their contest. They may, for instance, evaluate “which poem is the most technically complicated?” or “which poem is the most original?” They could even average these two scores.

Ideally, measurements have both high construct validity and high inter-rater reliability. Then we can really trust that the identified metric measures what it claims to measure, and does so reliably.

Evaluating the Foreign Service Entrance Exam

The Department of State sometimes makes strong claims about its ability to measure merit and performance. In 2006, then-Director General of the Foreign Service George Staples was exploring some major changes to the exam. Before his decision was inked, the Department commissioned McKinsey & Company to review the evaluation process, which concluded that the exam was “the best screen, and people who do well there are going to do well in the Foreign Service.” Many at the State Department continue to parrot McKinsey’s claim that the exam represents the “gold standard” for hiring processes.

Unfortunately, the claim lacked any supporting evidence (at least that I’m aware of). It did not appear that McKinsey conducted the sort of evaluation tests that would have validated this claim. But it is useful to imagine what kind of evidence would be convincing.

First, it would help to set up clear definitions: the better someone does on the entrance exam, the more likely they will be to succeed in the Foreign Service (exam score → success). But, a) How exactly do we measure test scores? and, b) How do we measure success in the Foreign Service? For (a), we might just acquire each officer’s original test scores. A discerning observer might raise some concerns here. What do we do with parts of the test that are scored more subjectively (such as the essay, biographical information, and in-person interviews)? And how do we account for changes in the test over time? Regarding question (b), What does it mean to do well in the Foreign Service? I have chosen to use the metric “Years to promotion into the Senior Foreign Service.” Again, there are weaknesses in this measure: The promotion process has been heavily criticized as somewhat arbitrary. We might worry, for instance, that those most quickly promoted aren’t necessarily the “best” employees for a variety of reasons. Good science demands we carefully answer these questions. But let’s ignore these wrinkles for now.

Now that we’ve carefully defined our terms, we can collect some data and run some analysis. Below, I produced two graphs to demonstrate hypothetical relationships between entrance exam scores and promotion speed (to be clear, this data is entirely made up). The first graph shows that officials who scored best on the entrance exam tend to do better in the Foreign Service. According to this data, those who score in the top 10 percent are expected to get promoted into the Senior Foreign Service two years faster than those in the bottom 10 percent. This would be a meaningful finding; the test seems to be capturing merit.

Next, I produce an alternative dataset. This one suggests that there is no correlation between scores on the entrance exam and years to promotion. While scoring 100% on the exam seems better than 60%, the scores demonstrate no relationship with future promotion prospects. This finding would seriously call into question the construct validity of the entrance exam.

Evaluating the Promotion Process

Let’s turn to another example focused on the Foreign Service’s promotion process. I’ll briefly describe the process so this makes sense to outsiders: the State Department dictates that every official gets a performance evaluation every year. When an employee’s time arrives to be considered for promotion, five years’ worth of evaluations will be stapled together into a file for each official. This file will be read and rated by multiple judges. The judge’s scores are averaged and, after some discussion, the highest-rated officials get promoted. (It’s a bit more complicated than this, but you get the idea.)

The State Department’s Bureau of Examinations (BEX) boasted in 2025 that “candidates could demonstrate all dimensions on virtual platforms and assessors could assess candidate performance accurately.” The office also explained that subsequent training of assessors and program assistants had ensured the ability to “administer and score the exercises accurately and without any potential for bias.” In short, BEX claimed that their evaluation process has both high construct validity and high inter-rater reliability. But, again, they’ve given us no data to validate their claim. Let’s visualize the sort of evidence would support their claim.

Below, I visualize some (entirely artificial) data to exhibit what high inter-rater reliability looks like. I imagine an environment in which there were 40 officers being evaluated (listed across the X-axis), and each of their files is read by the same 5 judges. Individual scores for each file are represented by colored dots, each representing a different judge. The average score (Y-axis) for each file is indicated with a black bar. You can see from the distribution of dots and scores that clear preferences emerge about the best and worst employees. The reviewers’ scores clump together tightly; there is almost perfect agreement about the quality of each file. We might still have important questions about the construct validity of this procedure, but inter-rater reliability is pretty good. If this is what BEX’s data looks like, it would lend confidence to their process.

Let’s look at an alternative situation. Below, I produced another dataset that still demonstrates clear preferences for promotions. But you can see that the inter-rater reliability in this environment is much lower. In fact, when we measure the underlying agreement between raters (using, in this case, a statistical procedure called intraclass correlation coefficient), there is virtually no consistency between reviewers. While the process produces the appearance of preferences, promotions given on this basis would be entirely arbitrary. They may as well have been rolling dice.

Which of these situations most accurately resembles the State Department’s recruitment and promotion processes? We don’t know. And that’s a problem.

It is possible that State Department officials have a reasonably accurate sense of merit. Certainly, there is knowledge within the system about what makes a great foreign policy official. And the aggregated judgment of many managers producing many evaluation reports for their staff likely adds up to something better than completely random. Yet, the experience of most observers of the Employee Evaluation Process is that it is rather arbitrary and unpredictable, often dismissed as little more than a “creative writing contest.”

Further, an institution genuinely committed to evidence-based personnel selection would likely describe its work very differently. It would publish its validation studies and acknowledge the limits of its findings. Then it would learn from the results and constantly innovate to improve and refine its methods of detecting merit. But the divergence between the Department’s confident public characterization of its assessment process and the absence of the underlying validation evidence is, in a sense, a form of evidence. The institution in charge of evaluating the quality of candidates’ critical thinking does not appear to be demonstrating sophisticated evaluation itself.

Indeed, the State Department invests few resources in studying merit. It neglects to study success and failure in our foreign policy. And it presents little agreed-upon measurement, data, or clarity about what it looks like for a staffer to be effective. This leaves us with little ability to comment on the construct validity of any performance management systems. Nor are we able to check to see that there is at least inter-rater reliability in the measurement of performance files.

This is the natural result of a culture of foreign policy that treats itself as “more art than science.” As we look to improve our institutions in the coming years, I hope we can do better.

Next
Next

A New Foreign Service Act: What Does This Mean?