Back to All Articles

How To Rank Jeopardy Contestants

I’ve written a decent amount of sports analytics content over the years, and the subject that inevitably produces the most arguments and discussion are ratings. Sports produces ratings debates of all kind: we argue over who’s the true GOAT, we compare players across eras, and we try and figure out who plays the best and why (and for that matter, what it means to be the best). The advent of analytics in sports has produced rating systems for just about every sport, and Board Control is no different with its BUTTREY ratings that ranks Jeopardy players over the last 25 years. While rating systems can often be incredibly divisive (I’m rating some of my friends with BUTTREY- it’s weird!), I think they’re still a valuable forcing function in forcing someone to explain how and why their ratings are what they are. Rating Jeopardy players is no different on that front, so we’re going to unpack a little bit on what BUTTREY does and does not include.

All rating systems are full of choices, assumptions, and methodologies. The rating systems that approximate a one-size-fits-all number, like Bill Conelly’s SP+ or Ken Pomeroy’s KenPom, make a ton of choices on what goes into their ratings, but their intent is still to capture all of the components of the teams they’re rating: independent factors that contribute to their overall strength, the ability of certain data points and metrics to capture and represent those effects, and a choice in how those metrics are weighted and adjusted to be as predictive as possible. Each of those components have several properties associated with them: what particular skill or ability they’re trying to measure, how statistically reliable those metrics are in terms of measuring their underlying skill, and how they need to be adjusted in order to be as predictive as possible. All good rating systems at some point have a long list of things they’d like to measure, or are at least trying to measure, and have some editorial choices on what they do and do not include and why. Before we get into the nuts and bolts of BUTTREY, which reflect my specific editorial choices, I think it’s more helpful to walk through what I think are the distinct components of Jeopardy skill regardless of our ability to measure them accurately. The structure of the gameplay naturally lends itself to a series of discrete and independent skills, each of which can be measured with varying degrees of accuracy and have different natural levels of importance in rating someone’s Jeopardy ability. We’ll go through each of them here and describe what these distinct skills are, how they can be measured, estimate their importance to a Jeopardy player’s overall ability, and make a judgment call on if the ability should be factored into a rating system.

Clue Selection

Description: The ability to select clues that maximize a contestant’s chance of winning when they have control of the board. Typically, this manifests itself as demonstrating some knowledge and understanding of where Daily Doubles are distributed and selecting clues accordingly, given the outsized role that Daily Doubles play in determining the outcome of a game. There is some nuance here: automatically hunting for Daily Doubles isn’t necessarily the optimal strategy, as James Holzhauer demonstrated in his run when he chose to go for bottom row clues to start to build up a stack and then looking for DDs to maximize the value he got out of them. There are some more minor applications of this skill, such as not selecting from a category in Double Jeopardy that has already had a DD revealed since a category can’t have two DDs, but more broadly, this skill typically manifests as “don’t pick a top row clue until the very end”.

Metric: Mistake rate: number of clear clue selection mistakes divided by total number of clue selection opportunities. This is a little more subjective than it may appear, too: even with picking top-row clues early, it’s debatable on if doing so is a clear cut mistake, and how much win equity these mistakes would cost a player. Maybe picking a top-row clue early in a category you like eases you into the game, or maybe you’re also willing to skip DD hunting to try and max out on a category you think you have an edge in, so there might not even be such a thing as a “clear” mistake, even if it’s likely that it is. This is also a highly variable metric, as your ability to select a clue depends on your ability to maintain control of the board, which isn’t 100% within your control. You have to out-buzz your other contestants to maintain board control, which is a completely separate skill, so the sample size for demonstrating this skill is pretty variable.

Importance: Low, and arguably getting lower. DD hunting used to be an exploitable edge that contestants who studied tactics were able to utilize more successfully than others, but beginning in 2021, this strategy has become commonplace, and this isn’t a skill that can be improved to a variable degree, like knowledge bases or buzzer speed. It can be more or less reduced to “you have clue selection awareness or you don’t”, and these days, just about everyone has it.

Rating suitability: Low. The combination of the subjectivity of the metric, the variance in its components, and its diminishing importance in the modern game make the effort not worth it.

Buzz-In Rate

Description: The ability to digest the question, come up with the answer, and attempt to buzz. The heart of the game: being confident that you think you know the answer and trying to ring in.

Metric: Buzzer attempts. Jeopardy started tracking this metric in early 2022 as part of their Jeopardata initiative in order to isolate how often a player tries to ring in and answer the question, which is a much closer proxy for their knowledge base than how often they actually buzz in, since the latter is also dependent not only on their own buzzer speed, but their opponents’ speed as well.

Importance: High. This is almost redundant, but coming up with the answers in a fast-paced environment is what Jeopardy is all about at its core. It’s also what separates Jeopardy from your local bar trivia game: you don’t have minutes to deliberate on what the answer is, the pace is rapid-fire and the questions are filled with clues and nuggets you have to decipher on the fly.

Rating suitability: Medium. The biggest problem with including buzzer attempts in a rating system is we only have data on this stat going back to 2022. This is a common shortcoming of many stats in many sports: not all stats are equally historically available, so your period of analysis is limited to when the stat started being measured. If I were ranking players from 2022 onward, buzzer attempts would be factored in much more heavily, but its absence from earlier games make it difficult to incorporate into a rating system that covers players before then.

Buzzer Speed

Description: How fast a player can buzz in as soon as they are eligible to. (A reminder: contestants do not buzz in as soon as they know the answer. They are eligible to ring in right as the question is finished being read, as indicated by a visual light that you cannot see while watching at home.)

Metric: Milliseconds between when a player is eligible to buzz in and when they actually buzz in. This raw metric needs some adjustment to account for the quarter-second lockout penalty for buzzing in early, but

Importance: Medium. Buzzer training for prospective contestants is becoming more commonplace, and for most appearances on the show, this skill can be reduced to “have you practiced this or not before coming on the show”. Some contestants probably do have better reaction time than others, which will absolutely help them, but for my money, if I were in something like create-a-player mode for Jeopardy, I’d still put my points into knowledge base over reaction time.

Rating suitability: Low. Buzzer speed is not tracked anywhere as part of any official statistics, so it literally can’t be used for anything. If I had a single wishlist for data that I wish the show started officially tracking, it would be this. The relative importance of knowledge base versus buzzer speed is still debated among fans in terms of its importance, and having this data available would go a long way in beginning to answer this question.

Answer Rate on Buzzed Questions

Description: How often you get the question right after you buzz in. This might seem redundant at first, since if you buzz in, of course you know the answer, but plenty of things can happen between buzzing and answering. Sometimes your guess is wrong, sometimes you buzz in without fully knowing the answer and think you’ll get there but you don’t, and sometimes you just straight up buzz when you didn’t mean to. Most of the time though, if you buzz in, you’re going to get it right.

Metric: Ratio of successful answers to overall answer attempts.

Importance: Low. The naive assumption of “if you buzz in, you’ll probably get it right” generally holds as a baseline, and there’s a bit of a selection bias going on here. If you overachieve relative to baseline, it’s not going to be a huge difference since the baseline percentage is already pretty high, and if you’re noticeably bad at answering questions after you buzz in, you’re probably not going to last long enough to be properly rated anyway.

Rating suitability: Low. Buzzer attempts and speed are far more informative about a player’s ability compared to their correct answer percentage.

Daily Double Wagering

Description: Knowing how much risk to take once you have found a Daily Double to maximize your win equity. Covered in much greater detail here.

Metric: Win equity gained or lost from wagers. Even with using something like a DD wager analysis tool to measure win probability added, there is still some subjectivity in this evaluation, because the calculations around win equity from wagers rely on a lot of assumptions on what a player’s get rate will be for a given question. There are player-wide averages to provide estimates, but get rate is incredibly variable between players, to say nothing of how the specific category influences a player’s get rate. There are ways to account for that natural variation and still come up with at least an estimate of how much a player wins or loses on their wagering decisions over the long run with things like sensitivity analysis and stress testing over a range of get rates, though, so it’s not a completely fruitless endeavor to try and quantify something like this.

Importance: Low. Even if you train extensively on the math behind DD wagering, there’s no guarantee you’ll have an opportunity to put those skills to use, as many a contestant will tell you. There are only 3 DDs to be found, and the DD in the first round is far less important than the ones in the second round, so arguably there are only two that matter in the whole game, and even if you’re actively looking for them, there’s no guarantee that you’ll find them. Daily Doubles play an outsized role in determining the outcome of a game given how much they can swing a score, and at the same time, the randomness from their sparsity and their hidden placement means that a player’s skill in utilizing them isn’t guaranteed to show up in a given game.

Rating suitability: Low. If there were something like 2 DDs in single Jeopardy and 4 in Double Jeopardy, wagering skill might play a larger role in a player’s overall ability. But there just aren’t enough DDs that actually happen during a game for a player’s DD skill to influence a game in a predictable manner.

Final Jeopardy Wagering

Description: The ability to know what to wager in Final Jeopardy in order to maximize a player’s chance of winning.

Metric: I’m not sure there actually is one, because Final Jeopardy wagering is ultimately a matter of game theory in the long run. You might have seen some wagering guidelines that describe how to wager in general. These are good guidelines, and when I have helped prospective contestants before they go on the show, I tell them to use the same guidelines. But in the postseason, we’re starting to see players deviate from these guidelines, and they’re rational to do so: what happens when all the players all know the guidelines, and they know that you know the guidelines? I think we’re going to see a lot more adjustments of players’ Final Jeopardy wagers, especially if we have more games where players have played against each other multiple times and know their tendencies, so it will be very difficult to determine if players are making objective mistakes in their FJ wagering.

Importance: Low to medium. I have seen enough games where contestants with no knowledge about how to think about wagering scenarios has cost them a game, but for regular Jeopardy, this can be effectively reduced to “you know the strategy or you don’t” similar to clue selection. For repeat players who are frequently in the postseason, the skill level becomes more variable, but at the same time more difficult to measure.

Rating suitability: Low. FJ wagering is absolutely a distinct skill, but at the higher levels, it’s difficult to know whether a player actually made a mistake in wagering, or if their adjustments were defensible.

Final Jeopardy Get Rate

Description: How often a player gets Final Jeopardy right. This is distinct enough from regular Jeopardy play since the questions are more engineered towards 30 second deliberation, and coming up with the answer and not having to worry about buzzing in is a distinct skill from regular Jeopardy play.

Metric: Percentage of times a player gets Final Jeopardy right. This is only meaningful for players that have

Importance: Low. It’s hard to imagine training for answering FJ questions specifically compared to regular questions. The core of both types of questions is the same: there’s a question in front of you and you need to get it right, and they have the same path to getting it right: just knowing stuff. On top of that, even if you’re specifically better on FJ questions than regular questions, there’s no guarantee that you’ll have a meaningful chance to use that ability to swing a game. You might be in a lock game where it doesn’t matter, your get rate might only be a little higher than your opponents, and the scores going into FJ might be the dominant force that dictate the outcome, not your ability to answer the question. There are too many other skills and components that drive a Jeopardy result for this particular skill to matter too much.

Rating suitability: Low. There’s just not that much separation between answering regular Jeopardy questions versus FJ questions: all else being equal, if you’re good at answering questions in regular gameplay, you’re probably going to be good at answering FJ questions as well. The overlapping signal from the other metrics more than cover the concepts that this

This covers all the distinct skills that go into playing Jeopardy, regardless of their overlap, measurability, or importance to provide the foundation of what you could rank a player on. We’ll cover how BUTTREY actually works next time to show what the editorial process looks like for trying to get a rating system that’s actually practical and functional next time.

Enjoy this article?

Subscribe to get new articles on Jeopardy strategy and analytics delivered to your inbox.