On a recent Championship Sunday at USTA Texas 18+ League Sectionals, someone passed an interesting observation to me. After a particular team won the first three lines, clinching advancement to Nationals, the level of play on the remaining two courts dropped precipitously.
There are any number of innocent explanations for that. Once the championship was decided, some of the competitive intensity naturally disappeared. Players knew that the outcome of their individual matches could no longer change the team result. Fatigue may also have been a factor in the blistering heat. However, those were not the explanations that tennis insiders immediately gravitate toward. Instead, the perceived effort is often viewed through the lens of ratings management.
Sometimes players seem to modulate their performance to intentionally influence their NTRP rating. The less charitable way to put it is that sometimes people deliberately lose games or even a complete match they would otherwise win. I wasn’t present on that Championship Sunday, and even if I were, there is no way to know whether anything like that happened in those particular final two matches. Tennis is a game of high variance, after all. However, the fact that other tennis players noted it and immediately speculated on the possibility says everything you need to know about the culture surrounding NTRP ratings.
Yesterday, I argued that the NTRP ratings system is complex rather than merely complicated. One reason is that the people being measured understand how the system works, or at least believe that they do, and adjust their behavior in response to the incentives it creates. Ratings management is a clear example of that phenomenon.
When people weren’t complaining about Self-Rated players over the weekend at Sectionals, the other dominant theme involved long-time players perceived to have carefully curated their ratings to stay at a particular NTRP level. The accusation was that established computer-rated players had learned how to remain within a desirable ratings band by selectively modulating their performance over time.
Every individual data point cited to support that accusation is technically impossible for an outside observer to prove. An unexpected score can result from a bad day, an unfavorable matchup, an injury, fatigue, or another player just having the match of their lifetime. Tennis has enough natural variability that almost any individual result can have a benign and legitimate explanation.
What we can determine with greater confidence, however, is why players might be motivated to manage their ratings. That requires examination of how the competitive system incentivizes them.
NTRP is intended to create competitive opportunities by grouping players of reasonably similar ability for leveled play. That is a terrific goal. Unfortunately, a player’s rating also gates the competitive opportunities that are available to them. Moving up an NTRP level may reflect improved performance, but it can also mean losing eligibility for their team or even league play at all. A player who gets to play all the time at one NTRP level may find considerably fewer opportunities after moving to the next.
Consequently, a higher rating is sometimes viewed as a punishment rather than a reward.
The NTRP system gives tennis an unusual incentive structure when compared to other competitive systems. The general ethos of sports is that better performance is rewarded. People practice because they want to improve, and once a match begins, the sporting objective is to win. Ideally, the player’s interests and the competitive system’s objectives all point in the same direction.
NTRP creates a curious phenomenon where those interests diverge. A player may genuinely want to improve their performance level, while also preferring that the algorithm not fully recognize that improvement. Winning remains desirable, but winning by too much may not be. Performing well enough to help the team advance can be advantageous, while producing the same results in matches that no longer matter to the team outcome may carry considerably less benefit.
A concept from economics and management known as Goodhart’s Law helps explain the problem. When a measure becomes a target, it stops being a good measure. I would not assert that the statement applies literally to every aspect of NTRP, but the underlying idea is useful. A metric behaves differently once the people being measured have reasons to care about the value of the metric itself.
The purpose of the NTRP rating system is to be a measure of tennis ability. However, once that number determines eligibility and playing opportunities, players have reasons to prefer particular outcomes from the measurement process. The rating no longer merely describes their tennis. It gates what tennis they will be permitted to play.
That does not mean all or even most players manipulate their ratings. I want to believe that the overwhelming majority of people compete and allow the computer to do whatever it does with the results. Additionally, the existence of an incentive does not absolve anyone who intentionally underperforms in an effort to influence a rating. Players remain responsible for their own conduct.
However, the system itself bears some responsibility. Competitive frameworks also have to be evaluated according to the behavior they incentivize rather than only the behavior their designers would prefer. If remaining at a particular NTRP level offers substantial benefits, it should not be surprising that at least some participants will attempt to remain there. The more consequential the rating becomes, the stronger that incentive can become.
This is where yesterday’s discussion about complexity becomes practical. A complicated ratings algorithm could theoretically be improved by making the mathematics more accurate. A complex ratings system has a harder problem because changing the calculation can also change how participants behave. The people producing the data are capable of understanding what the system values and adjusting the data they generate through their on-court performance.
That feedback loop creates another problem even when nobody is manipulating anything. Let’s return to the scenario that opened this post. Perhaps the players really did lower their effort after their team clinched the championship. If so, their performance still changed because the incentives changed. The matches mattered differently once the team outcome was determined. No nefarious intent is required for human behavior to respond to changing circumstances. Also, it was bloody hot.
That highlights a damaging secondary effect. The perception of ratings management can erode confidence in the system even when the perception is wrong. Players stop viewing an opponent’s strong performance simply as good tennis and begin wondering how that person managed to stay at that level. At the same time, poor performances can spark speculation about whether the score represented legitimate effort.
None of this means that level-based competition is a bad idea. NTRP creates an enormous amount of competitive tennis that would not exist without some mechanism for grouping players by approximate ability. Any classification system that determines eligibility will also create boundaries, and boundaries inevitably produce incentives around which side of them people would prefer to occupy.
The uncomfortable part is that a system intended to facilitate fair competition can sometimes create incentives that conflict with one of the most basic expectations of sportsmanship: players should compete to the best of their ability. That conflict does not have to be widespread to matter. The perception that it occurs can be enough to affect how participants regard both their opponents and the legitimacy of the competition.
Yesterday, I suggested that understanding the nature of a problem is necessary before attempting to solve it. Ratings management illustrates why NTRP cannot be treated solely as a mathematical equation. Improving the algorithm may improve the measurement, but the people being measured will continue to respond to whatever incentives the system creates.
That leaves us with a considerably harder question for tomorrow. If every change to a complex system has the potential to alter the behavior of the people inside it, what does meaningful improvement actually look like?
The objective cannot realistically be to devise one perfect version of NTRP and declare the problem solved. Managing complexity requires something more iterative: understanding what a change was intended to accomplish, observing how people actually respond, identifying unintended consequences, and being willing to adjust when reality does not match the assumptions that produced the decision.