The Observer Effect
Half a season into MLB's robot-umpire era, the machine's biggest effect on baseball has come from the calls it never made. Umpires shade toward the machine's zone, hitters have learned to wait, walks hit a rate unseen since 1950 — and knowing the strike zone better than the umpire is now a measurable, tradable skill.
Twenty-one and four.
That was Carson Kelly’s record on challenged ball-strike calls as the Cubs’ catcher when ESPN checked in this May — twenty-one calls overturned, four lost. Baseball Savant scored the work at 2.3 runs above expectation. Not runs from hitting. Not runs from framing, blocking, or throwing out runners. Runs from knowing the strike zone better than the umpire standing behind him, and proving it, in public, twenty-one times out of twenty-five.
A year ago that skill did not exist as a measurable thing. There was no column for it, no run value, no leaderboard. Now there is a public dashboard.
Half a season into Major League Baseball’s robot-umpire era, this is what the machine has produced: a new market, a new skill, a new anxiety — and, the part I keep coming back to, a changed game on the thousands of pitches where the machine did nothing at all.
The Tap
The ABS challenge system went live on Opening Day this season, and its design is worth pausing on, because the design is most of the story. MLB did not hand ball-strike calls to the machine. The umpire still calls every pitch. What the league added is an appeal court. Each team gets two challenges per game and keeps any challenge it wins. Only the pitcher, the catcher, or the batter can invoke it — no dugout signals, no manager theatrics, no video coordinator whispering into a headset. The player taps the top of his cap within roughly two seconds of the call, the Hawk-Eye tracking layer that already powers Statcast renders the verdict over a T-Mobile 5G link, an animation goes up on the board, and the game moves on. The whole exchange takes seconds.
Notice what the design refuses to do. It refuses to make the machine the default. It refuses to give the decision to anyone who was not on the field feeling the pitch. It refuses to allow unlimited appeals, which means every challenge is a bet placed with a scarce resource. This is a league that has spent the decade making deliberate product decisions — the pitch clock, two games every Friday night sold to Apple as a habit rather than a broadcast — and the challenge system carries the same fingerprints. The machine was installed the way you would install a guest you were not sure about. Powerful, but on probation, and only allowed to speak when spoken to.
Half a season in, the guest has rearranged the furniture anyway.
Half a Season of Receipts
Start with the totals, because they undercut the loudest prediction about this system before we get to the interesting part.
By mid-April, when CBS Sports ran the numbers, players had issued 1,050 challenges — 4.05 per game, covering roughly 1.4% of all pitches — and won 54% of them.
Sit with that success rate for a moment. These were not random appeals. Challenges are scarce, losing one costs you the resource, and the only people allowed to burn one are the three players with the best view in the building. These were the calls that professional hitters and catchers felt surest the umpire had missed. And the umpires were still right almost half the time.
The pre-season fear — that the machine would expose umpires as frauds and hollow out the human element — inverted on contact with data. The half-season record reads closer to a vindication: on the most contested one-and-a-half percent of pitches, the humans split the difference with a tracking system accurate to fractions of an inch. Half a season of this data has, if anything, increased my appreciation for the umpires. The machine did not dethrone the umpire. It gave the umpire a grade, and the grade was better than the folklore.
The distribution of who wins challenges is where the texture starts. In spring training, the defense won 60% of its challenges while batters won only 45%. That gap makes intuitive sense once you say it out loud — the catcher watches the pitch into his glove from directly behind it, while the hitter is processing a ball moving at close to a hundred miles an hour with an at-bat’s worth of adrenaline in his bloodstream. Conviction and accuracy are different skills. The batter feels certain. The catcher knows.
If the story ended there — a well-designed appeals system, a modest error-correction rate, a shrug from the fans — this would be a footnote in the history of sports officiating.
It does not end there.
The Calls It Never Made
In late April, the Associated Press published a piece I have been thinking about ever since. Players around the league, it reported, believe the challenge system is shrinking the effective strike zone. Umpires now get instant feedback on every challenged call, and — in the words of Diamondbacks catcher James McCann — the zone is “tighter in general,” the calls “much more uniform.” Hitters, sensing the shift, have adjusted — laying off the borderline pitches they used to defend against. And walks spiked: 9.8 percent of plate appearances through the season’s first month, a rate the league had not seen since 1950. The AP was careful with the causality — early-season walk rates always run high, and nobody can prove ABS is the reason — but even adjusted for the calendar, the jump from last season was, in its word, massive.
Hold the two numbers side by side. The machine directly touches about 1.4% of pitches. The behavioral shadow it casts touches all of them.
The machine's biggest effect on baseball has not come from the calls it overturned. It has come from the calls it never made.
The mechanism is not mysterious; it is incentives. For a century, an umpire’s borderline strike call was an assertion. It could be argued with, but it could not be falsified — not in the moment, not in front of the crowd, not on the scoreboard. The challenge system changed the epistemic status of every close call. Any borderline strike is now, potentially, two seconds and a cap-tap away from being rendered as an animation that shows the ball missing the zone by half an inch, in front of thirty thousand people and a broadcast audience. Umpires are professionals with reputations, and professionals with reputations respond to falsifiability the way anyone does. They retreat toward the calls that cannot embarrass them.
This is the observer effect, and it is not new. Physicists have known for a century that measuring a system disturbs it. Management theorists have known nearly as long — the famous studies at Western Electric’s Hawthorne plant found that workers’ output changed when researchers watched them, regardless of what the researchers actually changed. Being observed is itself an intervention. Baseball has now run the cleanest version of this experiment that professional sports has ever produced: install a perfect observer, let it act on 1.4% of events, and watch the other 98.6% move.
I wrote earlier this year about the measurement problem in creator economics — the way measurement frameworks quietly reshape the businesses they claim to merely describe. That process usually takes years, because dashboards work on humans slowly, through budget meetings and performance reviews. Baseball compressed it into weeks. The umpires did not need a quarterly review to internalize the new incentive. They felt it the first time a cap-tap turned one of their strikes into a ball on the videoboard.
And here is the detail that makes it art: nobody decided this. MLB did redraw the zone’s formal geometry for the machine — a precise slab, seventeen inches wide, pegged to a percentage of each batter’s height — but the shrink the players describe is not in the rulebook. The humans shrank the effective zone, in anticipation of the machine, one flinched borderline call at a time. The walk spike is not an output of the technology. It is an output of the psychology the technology created.
The Market That Formed Overnight
Every new measurement creates a market in whatever it measures, and this one formed with startling speed.
Carson Kelly’s 21-4 is the headline case — an 84% success rate against a league average of 54%, worth 2.3 runs above expectation by Baseball Savant’s accounting. Two-point-three runs is not a rounding error; it is real value, produced by a skill nobody could price a year ago. And because successful challenges are retained, the skill compounds: a catcher who challenges well does not just win calls, he preserves his team’s optionality deep into games, which is worth something extra in every close ninth inning. The system’s design accidentally invented a portfolio-management job and handed it to the catcher.
The market noticed immediately. By early July the Associated Press was writing up Hunter Goodman’s challenge record the way wire services used to cover hitting streaks. When challenge skill gets the feature-story treatment inside one season, it has already crossed from curiosity to attribute — the kind of thing that shows up in scouting reports and, eventually, in arbitration hearings. Because the dashboard is public, every front office can see which of its players spend challenges well and which ones torch them on wishcasting. Somewhere in every analytics department, there is now a spreadsheet ranking players by challenge discipline, and I would bet the spread between the best and worst organizations on that spreadsheet is wider than anyone expected.
Every market has a short side, and the short side of this one is pitch framing. For a decade, framing was the sabermetric community’s favorite hidden skill — catchers who could receive a borderline pitch so smoothly that the umpire called it a strike were quietly worth wins, and teams paid for the craft accordingly. The challenge system attacks framing from both ends at once. The stolen strike can now be challenged and un-stolen. And the shrinking zone means there are fewer borderline calls being awarded in the first place. The early ESPN and FanGraphs comparisons point the same direction: the elite framers are among the system’s first losers. A skill that took the analytics movement a decade to price is being repriced in a single season.
Though not everyone agrees on the direction. Will Smith — the Dodgers catcher, a man whose living sits directly on this fault line — offered the counterpoint in the Los Angeles Times in May: under a challenge format, framing becomes “more important, in a way.” I read his logic like this: the machine only speaks when invoked, which means the human umpire still calls more than 98% of pitches, and a catcher who can still win those borderline calls forces the other team to spend scarce challenges disputing them. Framing used to steal strikes. Now it can also tax the opponent’s challenge budget. Whether that nets out positive is exactly the kind of question the next half-season of data will answer, but the fact that the league’s best-paid catchers are openly theorizing about it tells you how fast the ground is moving under them.
The umpires, meanwhile, are living the other side of the observer effect. Former umpires have described to reporters the particular anxiety of working under the system — every close call now carrying the possibility of instant, animated, public contradiction. It is worth saying that the umpires’ union agreed to all of this inside a routine collective bargaining agreement. There was no labor war over the robot. The fight everyone predicted simply did not happen; the anxiety got absorbed as a working condition, the way surveillance usually does.
The Machine’s Confession
There is one more number in this dataset, and it is the one the enthusiasts skip.
When MLB detailed the system’s accuracy — figures given to The Athletic and reported widely in April — the league said the tracking is accurate to within 0.39 inches at a 95% confidence level, and within 0.48 inches at 99%.
Read that as an engineer and it is a triumph: sub-half-inch precision on a ball moving a hundred miles an hour. Read it as an epistemologist and it is a confession. The machine has an error bar. Roughly one call in twenty could be off by more than four-tenths of an inch — and some of the calls the machine overturns are calls the umpire missed by less than that.
Which means that on the closest pitches, the challenge system is not necessarily replacing a wrong answer with a right one. It is replacing one uncertain measurement with another uncertain measurement that happens to carry more authority. The umpire’s zone came with visible fallibility — a face, a stance, a history of blown calls you could yell about. The machine’s zone comes with an error distribution nobody can see and an animation rendered in the crisp, confident graphics of certainty. The animation never wobbles. The confidence interval never makes the broadcast.
I do not raise this to argue the machine is bad at its job. It is astonishingly good at its job, and materially better than any alternative on the closest calls. I raise it because the difference between accurate and authoritative is exactly the gap every league adopting these systems is about to fall into. We did not install truth behind home plate this season. We installed a measurement — a very good one — and agreed, collectively, to treat its output as truth because arguing with an error bar is unsatisfying. The strike zone did not become objective in 2026. It became settled, which is a different thing, and the fact that the sport is happier with settled than it ever was with human is one of the more revealing findings of the whole experiment.
Who the Humans Become
The diffusion has already started. The SEC ran ABS-style challenges at its 2026 baseball tournament in May — which means there are now college players building challenge discipline as a skill before they ever sign a professional contract, the way previous generations built framing. The attribute is propagating downward faster than the infrastructure is.
And it will not stop at baseball. Adam Silver said in May, on the record, that NBA out-of-bounds calls “will be done by an AI-automated system with cameras lined around the court” — taking, in his words, the whole category of objective calls out of the referees’ hands — and the league has had Sony’s Hawk-Eye tracking installed league-wide since 2023 to do it. Gary Bettman wants AI to resolve “where exactly is the puck.” Cathy Engelbert says the WNBA is looking at the same AI playbook as its “big brother,” the NBA. Every commissioner in American sports is now, in effect, promising to hire the observer.
Here is what I would tell each of them, from baseball’s half season of data.
The system will work. The calls will be more accurate, the appeals will be fast, the fans will absorb it within a month, and the labor war you are budgeting for will probably never arrive. All of that is the easy part, and none of it is the point.
The point is that you are not installing a tool. You are installing a witness. And the consistent lesson of every watched workplace since the Hawthorne plant — now confirmed at the scale of a major professional sport — is that the watched do not stay the same. Your officials will shade their judgment toward whatever the machine cannot contradict. Your athletes will discover, within weeks, which behaviors the machine rewards, and a new skill market will form around gaming the appeal before your competition committee has scheduled its first review meeting. The second-order effects will be larger than the first-order ones, they will arrive faster, and they will not be the effects you designed.
Baseball’s machine was given the narrowest possible mandate — speak only when spoken to, rule on one pitch at a time, touch 1.4% of the game. It still moved the walk rate of the entire sport, repriced a decade-old catching skill, created a new one from nothing, and changed how every umpire in the league experiences a two-strike count.
Every measurement system changes the thing it measures. The leagues lining up behind Silver’s promise are not adopting a technology; they are agreeing to be observed, and observation is never neutral. The machine behind home plate had made its first thousand calls by the middle of April. Its real work is everything the humans did because it was watching.
Published 31 July 2026, revised 31 July 2026. Narendra Nag is a founder and media executive writing on attention, streaming, and the economics of live sports.