Over the past two days, I have been exploring the NTRP ratings system through the lens of complexity. Friday started with the distinction between complicated and complex problems and the recognition that NTRP belongs firmly in the latter category. Yesterday, we looked at a specific reason why. The people being measured understand that their rating carries consequences, and at least some of them will respond to the incentives those consequences create.
That leaves us with a practical question. If NTRP is a complex system that cannot be permanently solved through a better algorithm, self-ratings questionnaire, or rules and regulations, what does improving it actually look like?
The answer is certainly not to accept whatever happens and shrug our shoulders because “it’s complex.” In many respects, recognizing complexity raises the standard for managing a system. If we know in advance that interventions may produce consequences that cannot be completely predicted, then monitoring what happens after those interventions becomes every bit as important as the analysis that preceded them.
That requires a different way of thinking about policy changes. When the USTA modifies some aspect of the competitive framework, it is presumably addressing an identified problem. The rationale for the decision, the assumptions underlying it, the expected outcome, and some method of determining whether that outcome actually occurred should all be part of the process. Then, after the change has operated long enough to generate useful information, someone needs to go back and look.
That last step is easy to neglect, particularly for the USTA, which is culturally wired to regard every change and initiative as a smashing success. They’re not alone. Organizations are generally much better at making decisions than systematically revisiting them. A rule gets changed, everyone adjusts to the new reality, institutional memory gradually fades, and eventually people know what the rule says without necessarily remembering what problem it was originally intended to solve. Once that happens, determining whether the rule remains useful becomes much harder.
This is particularly important in a complex system because the original intervention changes the environment in which it operates. Some responses will be exactly what the rule intended to produce, while others may never have been anticipated. The existence of an unintended consequence does not necessarily mean that the original decision was poor. What it does mean is that the consequence becomes another piece of information to consider when making decisions about what happens next.
There is an additional complication when collecting that information. Complex systems generate a tremendous amount of noisy feedback, and any individual data point can be almost impossible to interpret with certainty.
Yesterday’s discussion of ratings management provides a useful example. An unexpected match outcome cannot prove that a player intentionally modulated their performance. Tennis has too many variables and too much natural performance variance to establish intent from any one observation.
However, that does not mean the observations in aggregate are useless. Any individual data point may be impossible to prove or explain with certainty, but the pattern tells the story. If the same behavior, complaint, or unintended consequence repeatedly surfaces across different players, teams, leagues, and Sections, eventually the pattern itself becomes the meaningful information. We do not need to establish exactly what happened in every individual instance before recognizing that the aggregate may be telling us something about how the system behaves.
Consider the Championship Sunday example that opened yesterday’s post. The level of play on two courts was perceived to decline after a team clinched the championship by winning the first three lines. That observation proves essentially nothing about ratings management. Several completely benign explanations could explain what happened.
Now imagine that the same performance pattern repeatedly appeared after team matches were decided. Suppose it occurred across different teams, championship events, and Sections. At some point, the question would no longer be whether we could prove what happened on one particular court. The repeated pattern would provide a reason to investigate whether some characteristic of the competitive system was influencing player behavior.
That distinction between individual events and broader patterns is particularly important when organizations evaluate feedback. It would be foolish to modify the NTRP system every time someone complains about an opponent at Sectionals. Tennis players who have just lost a match are not always the most dispassionate evaluators of the person who beat them. At the same time, it would be equally foolish to dismiss recurring complaints simply because no individual allegation can be conclusively established.
The challenge is separating signal from noise to identify patterns, not to validate or refute any individual data point.
That requires collecting feedback rather than merely reacting to it. Local league administrators see implementation problems that may not be visible at the Sectional or National level. Captains discover incentives and edge cases because they are constantly constructing rosters and lineups within the rules. Players experience the consequences directly. None of those perspectives provides an objective description of the entire system, but taken together over time, they provide information that cannot be generated by the rating algorithm alone.
Anecdotes alone tell us nothing. Aggregated versions of the same story, however, can be significant. A complex system requires collection of enough information to determine whether anecdotes are isolated events or manifestations of a larger theme. When similar observations begin accumulating from independent sources, they deserve examination even when each one remains individually inconclusive.
The same principle applies when evaluating changes intended to improve NTRP. The relevant question is rarely whether one metric moved in the desired direction. Rating accuracy matters, but so do competitive opportunities, player experience, administrative burden, and confidence that the competition is reasonably fair. Improving one aspect of the system can create problems somewhere else.
The changes to dynamic disqualification coming to USTA League in 2027 provide a useful example. Beginning next year, Mixed Exclusive players will become subject to dynamic disqualification, and disqualifications will be possible through the National Championships. Those changes presumably address concerns about rating integrity. However, as I discussed a few weekends ago, they can also create significant consequences for teammates and captains, particularly in Mixed competition where losing one player can make it difficult or impossible to construct legal pairings.
We cannot know in advance exactly how those competing effects will balance. What we can do is decide what we expect the changes to accomplish and then collect enough information to determine what actually happens. How frequently are players disqualified at Nationals? At what point in the competition does that occur? What happens to the remaining players on those teams? Are Mixed teams disproportionately affected? Do complaints about clearly above-level competition decline? Do we observe changes in self-rating behavior?
Those questions turn a rule change into something that can be evaluated rather than merely implemented.
This is also why feedback mechanisms matter so much in complex systems. People closest to day-to-day operation will frequently recognize consequences that were difficult to anticipate when a policy was written. Their observations will not all be correct, and some will directly contradict one another. That is the nature of feedback from a system populated by people with different interests and experiences. The objective is not to treat every complaint as evidence that something needs to change. It is to remain receptive enough to recognize when independent observations begin forming a pattern.
No version of the league competitive framework or the NTRP system will eliminate every questionable rating, disputed self-rate, unfavorable matchup, or incentive to influence the measurement. Expecting that degree of precision from a complex system sets an impossible standard. However, complexity should not become an excuse for accepting systemic problems simply because individual examples are difficult to prove.
I have long asserted that you have to understand a problem before you can solve it. This weekend has forced me to refine that philosophy. On Friday, we added the requirement to understand the nature of the problem and recognize that some problems may not have a permanent solution. Yesterday demonstrated that the people inside a complex system respond to its incentives, sometimes in ways that make the original problem even harder to address.
The final addition is that managing complexity requires becoming comfortable with evidence that is individually imperfect. We observe what happens, collect feedback, look for patterns, compare those outcomes against what we expected, and make another decision when the evidence warrants it. Then we start observing again.
Perhaps that is a more realistic standard by which to evaluate the impacts of leveled play under NTRP. The measure of success is not whether the system eventually reaches some perfected state where every player receives an indisputably correct rating. It is whether the organization responsible for the system can recognize meaningful patterns, learn from what they reveal, and continue adapting as the people inside the system adapt in response.
Sometimes understanding the problem means recognizing that solving it was never really the objective. The mission is to manage it well.