Wittgenstein’s Stuff+

TL;DR

  • Wittgenstein’s Stuff+: Unless you have confidence in the reliability of a stuff model, if you use stuff to measure a pitcher you may also be using the pitcher to measure the stuff model. The less you trust a stuff model’s reliability, the more information you are getting about the stuff model and the less about the pitcher.

  • botStf is a steadier ruler than Stuff+, agreeing with itself year-to-year at r = .83 to Stuff+’s .77.

  • Stuff+ predicts rest-of-season better on most pitches - sinkers and sweepers are a coin flip and botStf wins on curveballs.

  • Both models are within noise on predicting next-year ERA (.41 and .39). Location+ and botCmd aren’t predictive year-to-year. Pitching+ doesn’t notably outperform Stuff+ on its own.

  • Stuff+ and botStf disagree most often on cutters and sweepers with the primary reason purportedly being release extension.

In Nassim Nicholas Taleb’s book, Fooled by Randomness, he outlines a concept, Wittgenstein’s ruler, that states, “Unless you have confidence in the ruler’s reliability, if you use a ruler to measure a table you may also be using the table to measure the ruler. The less you trust a ruler’s reliability (in probability called the prior), the more information you are getting about the ruler and the less about the table.” Taleb’s aphorism can be directly translated to interpreting stuff models, a discourse that’s often littered with directionally-similar models and rarely involves pitch shape discussion. 

This piece aims to deep dive into FanGraphs’ two stuff models - Cameron Grove’s botStf and Stuff+, Max Bay, Owen McGrattan, and Eno Sarris’ model. For the purpose of investigation, Wittgenstein’s ruler gets turned into Wittgenstein’s Stuff+: unless you have confidence in the reliability of a Stuff+ model, if you use Stuff+ to measure a pitcher you may also be using the pitcher to measure the Stuff+ model. The less you trust a Stuff+ model’s reliability, the more information you are getting about the Stuff+ model and the less about the pitcher.

THE DATA

To compare the two models, we pull Stuff+ and botStf FanGraphs leaderboards for the full 2020 through 2026 regular seasons and FanGraphs game logs for 2024 through 2026. The game logs carry both models’ grades by outing and by pitch type. We pair our FanGraphs data with Statcast to judge outcomes and postseason data is omitted.

Stuff+ and botStf are on different scales with Stuff+ being centered at 100 and botStf mimicking the traditional 20 to 80 scouting grade scale. In order to compare models, we make every grade a percentile, rather than try to pick the 100 average or 20 to 80 scale. The percentiles are within-season, an important distinction given that the spread of arsenal-level botStf grade in 2026 is half of what it was in every prior year so raw numbers aren’t comparable across seasons. 

FanGraphs publishes only one slider grade which includes all of sweepers, sliders, and slurves. In an effort to not muddy the data, I split the family using Statcast labels and if 90% or more of a pitcher’s slider-family offering carried a single label, the SL grade became that pitch. For example, Jesus Luzardo threw 1,000+ slider-family pitches in 2026 and Statcast deemed 99% of them sweepers, so Luzardo’s Stuff+ and botStf for his slider become a sweeper grade. On the flipside, pitchers like Max Meyer and Eury Perez who throw 47 and 59 percent sliders and 53 and 41 percent sweepers (within the family), respectively, have their slider grades dropped from our evaluation pool because a grade averaged over two shapes doesn’t allow us to drill down on model performance.

HOW FAST DO THE RULERS SETTLE?

FanGraphs’ Stuff+ primer claims that the model’s output becomes reliable around 80 pitches while no such claim is made by PitchingBot’s primer. Our research finds that a pitcher’s first 100 pitches predict his rest-of-season arsenal Stuff+ at r = .81 and r = .84 for arsenal botStf. By pitch type half-reliability for each model reads as the following:

How a grade becomes reliable, by pitch type

Share of a pitcher's grade that is signal rather than outing-to-outing noise, as pitches of the type accumulate. Dots mark the value at FanGraphs' 80 pitches.

Stuff+botStfFanGraphs' 80-pitch figure
Knuckle curven=9650%90%050100150Knuckle curve · Stuff+ · half point 91 pitches · 47% signal at 8047%Knuckle curve · botStf · half point 6 pitches · 93% signal at 8093%Splittern=19450%90%050100150Splitter · Stuff+ · half point 50 pitches · 62% signal at 8062%Splitter · botStf · half point 4 pitches · 95% signal at 8095%Slidern=72350%90%050100150Slider · Stuff+ · half point 43 pitches · 65% signal at 8065%Slider · botStf · half point 8 pitches · 91% signal at 8091%Curveballn=48350%90%050100150Curveball · Stuff+ · half point 40 pitches · 67% signal at 8067%Curveball · botStf · half point 6 pitches · 93% signal at 8093%Changeupn=68750%90%050100150Changeup · Stuff+ · half point 33 pitches · 71% signal at 8071%Changeup · botStf · half point 7 pitches · 92% signal at 8092%Cuttern=52550%90%050100150Cutter · Stuff+ · half point 26 pitches · 75% signal at 8075%Cutter · botStf · half point 6 pitches · 93% signal at 8093%Four-seamn=147550%90%050100150Four-seam · Stuff+ · half point 23 pitches · 78% signal at 8078%Four-seam · botStf · half point 9 pitches · 90% signal at 8090%Sweepern=29550%90%050100150Sweeper · Stuff+ · half point 15 pitches · 84% signal at 8084%Sweeper · botStf · half point 5 pitches · 94% signal at 8094%Sinkern=90950%90%050100150Sinker · Stuff+ · half point 12 pitches · 87% signal at 8087%Sinker · botStf · half point 8 pitches · 91% signal at 8091%Whole arsenaln=216150%90%050100150Whole arsenal · Stuff+ · half point 8 pitches · 91% signal at 8091%Whole arsenal · botStf · half point 13 pitches · 86% signal at 8086%

Curves are n ÷ (n + half point) from the noise decomposition on 2024 to 2026 outing-level grades; the half point is where each curve crosses 50%. Sweepers split from sliders by Statcast labels. n = pitcher-seasons.

The number for botStf and Stuff+ above is a ratio of outing-to-outing noise to the true spread between pitchers, so it’s sensitive to both. We know that Stuff+ is noisier outing-to-outing on every pitch type, which is the aforementioned reliability edge we mentioned for botStf. The startling number here is the knuckle curve, an extreme case that looks drastic because Stuff+’s single-outing grades for it swing more than any other pitch by a wide margin. On top of that, the pitch is rare enough (96 pitcher-seasons, median of seven per outing) that a single outing’s grade entirely depends on a small sample.

The one place the stabilization order flips in Stuff+’s favor is at the arsenal level. Arsenal grades are naturally larger samples and pitchers differ far more in overall arsenal grade than in the grade of a single pitch, therefore creating more spread for the grade to find. A ruler settles fastest when the things it’s trying to measure are further apart, and whole arsenals are further apart than a single offering:

Why the arsenal settles fastest

Each point is a pitch type. Right means pitchers truly differ more on that grade; up means a single outing's grade bounces more. The half point is the ratio, so points on the same line from the origin settle at the same speed, and steeper is slower.

Stuff+ (labeled)botStflines of equal half point
02468010203040true spread between pitchers (Stuff+ points, season level)outing-to-outing noise (Stuff+ points)5 pitches10 pitches25 pitches50 pitches100 pitchesWhole arsenal · Stuff+ · spread 7.6, noise 21.0 · half point 8 pitchesWhole arsenalWhole arsenal · botStf · spread 4.4, noise 15.7 · half point 13 pitchesFour-seam · Stuff+ · spread 6.1, noise 29.2 · half point 23 pitchesFour-seamFour-seam · botStf · spread 4.4, noise 13.4 · half point 9 pitchesSinker · Stuff+ · spread 7.2, noise 24.5 · half point 12 pitchesSinkerSinker · botStf · spread 6.3, noise 18.0 · half point 8 pitchesCutter · Stuff+ · spread 4.1, noise 21.0 · half point 26 pitchesCutterCutter · botStf · spread 5.4, noise 13.6 · half point 6 pitchesSlider · Stuff+ · spread 4.5, noise 29.6 · half point 43 pitchesSliderSlider · botStf · spread 4.3, noise 12.2 · half point 8 pitchesSweeper · Stuff+ · spread 7.2, noise 27.8 · half point 15 pitchesSweeperSweeper · botStf · spread 4.7, noise 10.8 · half point 5 pitchesCurveball · Stuff+ · spread 4.8, noise 30.2 · half point 40 pitchesCurveballCurveball · botStf · spread 3.6, noise 8.7 · half point 6 pitchesKnuckle curve · Stuff+ · spread 3.7, noise 35.3 · half point 91 pitchesKnuckle curveKnuckle curve · botStf · spread 4.4, noise 10.9 · half point 6 pitchesChangeup · Stuff+ · spread 4.6, noise 26.5 · half point 33 pitchesChangeupChangeup · botStf · spread 5.1, noise 13.2 · half point 7 pitchesSplitter · Stuff+ · spread 4.4, noise 31.4 · half point 50 pitchesSplitterSplitter · botStf · spread 5.8, noise 11.5 · half point 4 pitches

Spread = square root of the between-pitcher variance (τ²), noise = square root of the outing variance (σ²), from the noise decomposition on 2024 to 2026 outing-level grades. For Stuff+ the whole arsenal sits furthest right and lowest: the widest true spread of any grade and the least outing noise, so it settles fastest. botStf's points cluster near the bottom: it settles faster on single pitches because its grades barely move outing to outing, not because its pitchers are further apart.

DOES THE RULER AGREE WITH ITSELF NEXT YEAR?

In our dataset, there are 2,262 instances of a pitcher throwing 300+ pitches in back-to-back seasons. We find that Stuff+ correlates to next-year Stuff+ at r= .77 and botStf correlates year-to-year at .83. botStf proves to be steadier on nine of ten pitch types with the one exception being the splitter, a smaller sample pitch at just 56 year-to-year instances. The disparity between the two models on splitters is .82 to .78, which is within noise. The findings on year-to-year correlation are illustrative in this context given lack of public information on how much the model’s inner workings changed and whether those inner workings were applied to past seasons. Regardless, this finding is a notch in botStf’s belt to a degree, but a model agreeing with itself isn’t necessarily the same as being more accurate. 

DOES THE RULER PREDICT THE TABLE?

We ran studies on next-year ERA according to the same logic Jonathan Judge used in his 2023 Baseball Prospectus article; our evaluation pool used pitchers with 40 innings in back-to-back seasons (coming out to 1,404 instances) and we weighted by innings. Innings weights are imperative so that a 42-inning reliever whose ERA went from 2.5 to 5.8 counts differently than a 200-inning starter whose ERA moved by .1. In this example, we can conclude that the reliever’s ERA swing can be more attributed to noise than that of the starter, so the noise in this case dilutes the correlation we are seeking to find. 

Stuff+ predicted next-year ERA at r = .41 and botStf did so at r = .39, essentially a wash. Although this is a stuff model evaluation, it’s worth noting that the corresponding Location+ and botCmd aren’t additive in a predictive sense. Location+ is r= -.03 and botCmd is r= .04. This leads Pitching+ to post r = .39, holding the level of Stuff+ and botOvr sits at r = .32, giving up decent ground to its botStf component. 

As expected, the performance of both models are vastly different for starters and relievers, owing to the importance of innings pitched. 

Starting pitcher next-year ERA prediction:

  • Stuff+: .43

  • botStf: .40

  • Pitching+: .44

  • botOvr: .36

Reliever next-year ERA prediction:

  • Stuff+: .25

  • botStf: .22

  • Pitching+: .28

  • botOvr: .21

Reliever ERA is mostly noise because of the innings sample, but stuff models far exceed predictive power for reliever year-to-year ERA with reliever ERA predicting itself at r= .07. Put more concretely, Stuff+ ERA predictiveness goes from .26 for pitchers throwing 20 to 40 innings, compared to .46 for those that threw 160+.

PITCH-BY-PITCH

Pitch type analysis paints a clearer picture. Here’s how a pitcher’s first 100 pitches of a certain pitch type predicts rest-of-season RV/100:

First 100 pitches vs. rest-of-season run value

Rank correlation between a pitcher's grade on his first 100 pitches of a type and that pitch's run value per 100 over the rest of the season. Higher is more predictive. 2024 to 2026.

Stuff+botStf
.00.10.20.30Spearman r with rest-of-season RV/100Four-seamFour-seam · Stuff+ · .28 (n=965).28Four-seam · botStf · .20 (n=965).20SliderSlider · Stuff+ · .28 (n=367).28Slider · botStf · .12 (n=367).12CutterCutter · Stuff+ · .22 (n=236).22Cutter · botStf · .06 (n=236).06SplitterSplitter · Stuff+ · .24 (n=93).24Splitter · botStf · .15 (n=93).15ChangeupChangeup · Stuff+ · .21 (n=309).21Changeup · botStf · .12 (n=309).12SweeperSweeper · Stuff+ · .21 (n=142).21Sweeper · botStf · .12 (n=142).12SinkerSinker · Stuff+ · .19 (n=488).19Sinker · botStf · .17 (n=488).17CurveballCurveball · Stuff+ · .23 (n=205).23Curveball · botStf · .27 (n=205).27

Sweepers and sinkers are inside the noise; curveballs are the one pitch botStf reads better. n = pitcher-seasons with 100 pitches of the type and 100 more after.

Here’s what it looks like, but swapping in whiff rate for RV/100:

First 100 pitches vs. rest-of-season whiff rate

Same test, with whiff rate as the outcome.

Stuff+botStf
.00.10.20.30.40.50Spearman r with rest-of-season whiff rateFour-seamFour-seam · Stuff+ · .48 (n=965).48Four-seam · botStf · .36 (n=965).36CutterCutter · Stuff+ · .41 (n=236).41Cutter · botStf · .27 (n=236).27SplitterSplitter · Stuff+ · .41 (n=93).41Splitter · botStf · .11 (n=93).11ChangeupChangeup · Stuff+ · .40 (n=309).40Changeup · botStf · .19 (n=309).19SliderSlider · Stuff+ · .39 (n=367).39 (tie)Slider · botStf · .39 (n=367)CurveballCurveball · Stuff+ · .44 (n=205).44Curveball · botStf · .48 (n=205).48SinkerSinker · Stuff+ · .24 (n=488).24Sinker · botStf · .25 (n=488).25SweeperSweeper · Stuff+ · .21 (n=142).21Sweeper · botStf · .25 (n=142).25

Sliders, sinkers and sweepers are ties; curveballs lean botStf.

At the arsenal level, botStf is the better model for predicting strikeouts. Rest-of-season whiff rate has botStf greatly ahead (.43 to .31) and next-year strikeout rate paints a similar picture (.54 to .49). A blend of the two models does not help at the pitch level. The best weight we found was 88% Stuff+ and it ties Stuff+ alone within noise. However, for next-year strikeout rate a ⅓ Stuff+ and ⅔ botStf model yields a r = .555 against a .541 or .492 using each model on their own. 

WHERE THE RULERS DISAGREE

How often do the two models agree on pitch grades?

How much the two models agree, by pitch type

Rank correlation between a pitch's Stuff+ percentile and its botStf percentile, pitcher-seasons with 100 or more pitches of the type, 2020 to 2026.

.00.20.40.60.80Spearman r between the two modelsWhole arsenalWhole arsenal · .73 (n=4902).73SplitterSplitter · .70 (n=319).70SinkerSinker · .70 (n=1881).70Four-seamFour-seam · .69 (n=3288).69ChangeupChangeup · .64 (n=1507).64Knuckle curveKnuckle curve · .60 (n=257).60CurveballCurveball · .57 (n=1043).57SliderSlider · .57 (n=2153).57CutterCutter · .42 (n=1034).42SweeperSweeper · .32 (n=293).32

The in-between shapes, cutters and sweepers, are where the rulers disagree; the rare splitter is where they agree most.

What’s surprising here is that it’s not the rare pitches that have the least correlation, but rather the in-between shapes. This may partially be attributable to pitch classification. PitchingBot’s model pools by fastballs, breaking balls, and offspeed pitches while we lack information on how Stuff+ does it. This may cause methodology differences that create divergences in grades for pitches that are different in classification, but similar in movement.

An especially strange trend is splitters - a pitch that has seen its correlation between models precipitously drop over time. In 2020, r=.94 between the two models on splitters and that number has subsequently dropped, going to .84, .84, .79, 74, .58, and now .55 in 2026. 

We attempted to find what drove the gap between pitch types and our LightGBM model posited that how release extension was incorporated may be an ostensible culprit. The model was fitted on 29 shape features and explained 35% of the gap’s variance out-of-sample with the single biggest driver showing to be release extension. Using a ridge regression, we found one standard deviation of release extension moved the gap 15 percentile points on sliders, 17 on sweepers, 14 on curveballs, and 15 on cutters, and the gap moved toward botStf. It would be naive to say that PitchingBot doesn’t utilize extension as an input, but it’s fair to say how extension is being used likely differs. 

THE PITCHER MEASURES THE MODEL

If you look at the 50 pitches where the models disagree most, they average a staggering 81 percentile point difference. Stuff+ won 27 of these instances and botStf won 23, showing that each model is correct in certain instances. Widening the evaluation pool to 200, Stuff+ wins 111 and botStf 89. Unsurprisingly, the pitch type analysis matches our previous takeaways where Stuff+ dominates on changeups (15 wins to 5) and four-seamers (14 to 6) while botStf outclasses Stuff+ on curveballs (15 to 5). The aforementioned divergent splitter is nearly a dead heat at 11 to 9.

The difference in grade and outcome between Paul Skenes, Foster Griffin, and Yuki Matsui (in 2026) illustrate why the models split results cleanly. Skenes threw 291 splitters with Stuff+ rating it in the 11th percentile and botStf having it in the 57th. Skenes’ shape is a unique one, throwing it nearly harder than anyone with a whiff rate ranking in the 6th-percentile. Griffin’s ranks in the 72nd-percentile by Stuff+ and dead last in botStf. Another distinct shape, Griffin’s has 1st-percentile active spin, 99th-percentile seam-shifted movement, and an 18th-percentile whiff rate. Matsui has an 82nd-percentile splitter by Stuff+ and 8th by botStf on a 19th-percentile velocity, posting a 92nd-percentile whiff rate. In these three cases, we see three different shapes under the same classification, differing whiff rates, Stuff+ winning on wildly different velocity splitters in Skenes and Matsui, and botStf winning on Griffin.

HOW TO READ THE MODELS

If you’re trying to analyze a single pitch, outside of curveballs, use Stuff+. The best method for predicting next-year strikeout rate, use a ⅓ Stuff+ and ⅔ botStf blend. If you’re trying to predict next-year ERA, either model is acceptable. If the pitcher throws from a unique slot, gets unusual extension, or has an outlier shape, check both models before figuring out which to trust. The same goes for analyzing any pitch with a massive difference in grade - don’t average them. Instead, look into the pitch and why the models may grade them differently. These models aren’t the end of the discussion, but rather the start of it.

Next
Next

3 Cubs Offseason Acquisition Targets