The Carson Palmquist Model

Rather than post yet another pitcher outing scorecard, I believe the more interesting conversation to be had around Stuff+ is its applications to pro acquisitions. It’s no mystery that every organization has some version of Stuff+ and that all teams, just like the public, are going off of similar information. This makes a pitcher’s Stuff+ a widely-known variable when it comes to player acquisition.

Stuff+’s ubiquitous nature devalues the expected value of using it as a differentiating factor in acquisitions. Just take a look at the single-A or triple-A Stuff+ leaderboard published by Eno Sarris. Of the 253 pitchers that had thrown 100 or more pitches by July 21st in A ball, only 42 had a Stuff+ exceeding 100 (16.6%). At AAA, of 784 pitchers, 159 exceeded a 100 Stuff+ (20.2%). In MLB, across that same time period, 603 pitchers had thrown 100+ pitches and 291 had a Stuff+ exceeding 100 (48.2%). 

Here’s what free agents have signed for in each Stuff+ band over the last 6 offseasons:

What free agents signed for, by Stuff+ band, 2020-21 to 2025-26
Stuff+ bandStartersSP with an MLB dealRelieversRP with an MLB deal
80–90$5.4M72%$1.5M52%
90–100$10.0M81%$4.8M69%
100–110$13.2M92%$4.8M80%
110–120$30.5M100%$8.1M89%

Average annual value of major-league deals only. Percentages are the share of free agents in that band who signed a major-league contract rather than a minor-league one.

Drilling deeper into how the market values Stuff+, here’s the per point AAV shift:

What one Stuff+ point is worth on the free-agent market — starters
Stuff+Market AAV+1 point+5 points
85$3.9M+$0.22M+$1.2M
90$5.1M+$0.29M+$1.6M
95$6.7M+$0.38M+$2.1M
100$8.8M+$0.49M+$2.8M
105$11.5M+$0.65M+$3.6M
110$15.2M+$0.85M+$4.8M
115$19.9M+$1.12M+$6.3M

Fitted on 185 starter deals (+5.6% per point), evaluated at Loc+ 100, age 30.

What one Stuff+ point is worth on the free-agent market — relievers
Stuff+Market AAV+1 point+5 points
85$3.2M+$0.10M+$0.5M
90$3.7M+$0.11M+$0.6M
95$4.3M+$0.13M+$0.7M
100$5.0M+$0.15M+$0.8M
105$5.8M+$0.18M+$0.9M
110$6.8M+$0.20M+$1.1M
115$7.9M+$0.24M+$1.3M

Fitted on 198 reliever deals (+3.0% per point), evaluated at Loc+ 100, age 30.

Clearly, the greater the Stuff+, the higher value a pitcher is, and the market has already accurately baked that in. Additionally, the value of a point compounds, with a point being worth four times as much at 110 than it is at 85 for a starter. From a pro acquisitions perspective, my hypothesis is that the edge isn’t in refining an already accurate Stuff+ model, but rather trying to model which players are likely to see an uptick in Stuff+.

Stuff+ primarily rises from three factors: pitch usage, adding a new pitch, or regression.

Pitch usage is simple. A lot of times we see players, especially those on teams with notoriously poor front offices, not throwing their best stuff enough. Erik Miller serves as a prime example of this. Last year, Miller posted a respectable 107 Stuff+ while throwing his highest Stuff+ pitch, his sinker (132 Stuff+), 17% of the time overall and only 6% against righties. In 2026, Miller made the sinker his primary pitch, throwing it more than any offering against both righties and lefties. Miller now has posted his best season in the majors, posting a 118 Stuff+, and near league-best marks in xBA, hard hit rate, whiff rate, and K rate. 

To find what pitches a player can potentially throw, we need to lean on whether a player is a pronator or supinator while also finding precedent from players with similar pitch characteristics. Admittedly a much more difficult task than changing pitch usage, new pitch types can also become accessible by changing arm angle which we’ve seen from myriad players this year including Emerson Hancock, Carson Palmquist, and Logan Allen, among others. 

If you don’t gain Stuff+ through a new pitch or changing the existing mix, Stuff+ can be gained simply through regression. Stuff+ oscillates year-to-year with an average drift of about 4.5 points; however, it relies on the amount of pitches thrown in the previous season. 

How much Stuff+ moves year to year, by sample size ▸
Year-over-year Stuff+ movement by platform-season pitch count, 2020–26
Pitches thrownPairsMean absolute changeMoved 5+ points
Under 10038510.168%
100–2504436.047%
250–5006035.443%
500–1,0009464.940%
1,000–1,5005945.039%
1,500+6554.233%

For pitchers with at least 500 pitches in consecutive seasons, Stuff+ moves about 4.5 points a year — about 4 for full-time starters, and 6 or more for 100–250-pitch samples.

The namesake for this model and hypothesis is the aforementioned Carson Palmquist. Palmquist posted a paltry 92 Stuff+ last year in Colorado while having two pitches whose Stuff+ exceeded 105, a slider and sinker. Under new leadership, the Nationals saw something in Palmquist after he was DFA’d in late May, picking him up 4 days post-DFA.

Palmquist has posted 107 Stuff+ so far with the Nationals in 2026. The Nationals dropped his arm angle five degrees, added a new sinker, and the drop in arm slot changed his sweeper and four seam shape. Palmquist’s case begged the question - do teams have systems that flag these types of candidates for Stuff+ gain, or are they the result of deep dives? This project attempts to create a system to flag the types of arms, like Palmquist, who would be viewed drastically differently with an increased Stuff+. After all, if Palmquist had a 107 Stuff+ at the time of his DFA, the line of teams trying to sign him would likely be much longer than the one that awaited him while he hit the open market at 92.

Pitch Usage Architecture

Stuff+ is a weighted average of a pitcher’s by-pitch Stuff+. Therefore, throwing high Stuff+ pitches more and low Stuff+ pitches less increases the player’s Stuff+ average and, in turn, raises their open market value. Our pitch usage arithmetic aims to answer: if a pitcher threw their best pitches more and worst pitches less, what would be the magnitude of the change to his Stuff+?

To answer this question, we follow a six-step architecture:

  1. First, we must shrink each pitch grade toward the grades of comparable pitchers. Stuff+ stabilizes around the 80 pitch mark so shrinking grades create more conservative Stuff+ estimates for small sample pitches. Our equation to land on a shrunk Stuff+ grade to use is:

(n own + 80 comps) / (n +80)

Where:

n is the number of times the pitch was thrown

own is the Stuff+ grade for the offering

comps is the grade of the offering amongst comparable pitchers

80 is used as a constant derived from FanGraphs’ 80 pitch stabilization point

  1. We now project next year’s Stuff+ grade for each pitch. We are recommending players throw their best pitch more often. Because of that, we need to actually project next year’s Stuff+, not just say that year t’s Stuff+ will be equal to year t+1’s Stuff+. In fact, we find that the FanGraphs Stuff+ of a pitcher’s best and worst pitch compresses by 0.76 year-to-year and a pitcher’s best offering, even after shrinkage, reverts about two points the next year. 

  2. We set our optimization target as a 50/50 blend of Stuff+/Pitching+. We don’t solely optimize for Stuff+ because it doesn’t account for command and by adding Pitching+ as part of the target, we don’t over-recommend offerings that a pitcher has no command of.

  3. With a set optimization target has been set, we re-weight arsenals under constraints.

  4. The final projection is entered at half-strength. The 0.5 we multiply the dose by is found empirically. We find that 1 in 4 of our top recommendations are acted upon, and those that do act move 10 to 15% in usage, rather than a prescribed 20.

The constraints mentioned in step 3 above are:

  1. The most usage that can move in total is 25 points The 90th percentile of total usage moved per year was 25.9%.

  2. Single pitches are capped at 55% usage. 55% is the 75th percentile of highest single-pitch share. We find that historically, half of all seasons have a pitch that exceeds 45% and high-Stuff+ arms err on the higher side of that.

  3. Fastball floor is set at 25%. 25% is 1st percentile fastball usage. Only 1.6% of seasons in our data have fastball usage south of 30%.

  4. Starters must throw 3 pitches and relievers must throw 2 pitches post-optimization. These are 5th percentile observed pitch counts for both starters and relievers with 18% of reliever seasons being two-pitch with starters sitting at 1.8%.

Pitch usage lever by recommendation rank, six origins 2020–21 to 2025–26
RanknActedActors' ΔStuff+Non-actors' ΔStuff+
1–105826%+2.4−1.4
11–258911%+0.7−0.5
26–5014716%+5.0−1.4
51–10029312%+1.3−0.6
101–20058510%+0.3−1.2
201+1,43710%−1.3−1.1

An actor raised the recommended pitch by at least 10 points of usage the next season.

The results show a consistent story - the top ten act at roughly two and a half times the rate of those ranking below them, and in every bucket the pitchers who act gain Stuff+ and those that don’t lose it. The ranking is best used as an odds of action, not as a smoothly increasing payoff for changing pitch usage as rank gets higher. Our backtesting shows that the difference between being ranked 5 and 40 is marginal, and the top of the board is where the recommendation is most likely to be acted upon.

One important architecture note is that usage change is hand-aware. Pitch values differ depending on the platoon split at play. Right-handed pitchers have more effective changeups against left-handed-hitters whereas they’re much better throwing a sweeper to a right-handed hitter than a lefty. We factor this into our optimizer by measuring the split in pitch effectiveness within pitcher-season, rather than by taking the raw numbers because, for example, right-handed pitchers who throw a slider to a lefty hitter are likely doing so because they have a plus slider. The baked-in cost per pitch looks like this:

What a pitch costs against opposite-handed hitters — right-handed pitchers
PitchPitcher-seasonsΔ runs per 100Δ grade points
Sweeper424−0.58−12.8
Sinker956−0.56−12.3
Cutter610−0.55−12.1
Slider1,124−0.40−8.9
Four-seam2,116−0.22−5.0
Knuckle curve171+0.21+4.8
Curveball529+0.25+5.6
Splitter226+0.28+6.2
Changeup550+0.45+10.0

Measured within pitcher-season, minimum 40 pitches to each side. Negative means the pitch plays worse to opposite-handed hitters. The optimizer applies half of each split.

What a pitch costs against opposite-handed hitters — left-handed pitchers
PitchPitcher-seasonsΔ runs per 100Δ grade points
Sinker446−1.13−25.2
Sweeper192−0.95−21.2
Cutter173−0.89−19.7
Slider436−0.82−18.3
Knuckle curve31−0.82−18.3
Changeup90−0.82−18.2
Four-seam704−0.31−7.0
Curveball166−0.10−2.2

Pitch Addition Architecture

From 2020 to 2025, 362 of 10,292 candidate pitchers added a pitch to their arsenal, good for a 3.5% base rate. Precisely, our architecture to determine pitch addition requires us to answer two questions: how likely is a pitcher to add a certain pitch, and how would their Stuff+ change if they were to add it?

To answer these questions we first find comparable arms and what they throw. To do so, we run two k-nearest neighbors searches for each pitcher-season. 

Search number one attempts to identify which pitches arms like theirs tend to add. We used the five traits that historically predicted pitch addition the best in this search and those traits were release side, four-seam spin efficiency, curveball usage, arsenal Stuff+, and arm angle. 

Our second search then asks what the Stuff+ grade would be of the new offering. The eight physical traits used to determine the realized grade of the added offering were arm angle, fastball approach angle, release height, fastball velocity, spin per mile per hour, four-seam spin efficiency, spin-axis residual, and extension.

Each search then returns the 240 closest same-handed pitcher-seasons, dropping seasons of the pitcher in question, and keeping only pitchers with alike supination/pronation bias. Therefore, supinators are compared only to other supinators and pronators only compared to other pronators. After filtering, the player’s neighborhood becomes the 80 closest that remain. 

Our first search’s neighborhood gives us our precedent share, or in other words, the fraction of the pitcher’s neighbors who throw each family at 10% or more. The second search gives us our precedent grade, or the usage-weighted grade of the pitch that their neighbors with precedent share throw. 

Now that we have neighborhoods setting our precedent share and grade, we need to determine actual candidate pitch families to throw. A pitch family is a candidate to be added if it meets the following criteria:

  1. The pitcher in question throws it 2% or less.

  2. At least 20% of his neighbors throw it.

  3. No pitch he currently throws fills the potential addition’s role where the roles are ride, run, cut, slider, sweep, curve, and offspeed.

With candidates determined, we then calculate projected arsenal Stuff+ gain using the equation: 

0.165 * (precedent grade - 1.5 - current arsenal Stuff+)

0.165 is the average first-season usage share of pitches that have been added from 2020-2025. The 1.5 dock is how far the realized grade of an added pitch has fallen below the neighbors’ grade historically.

We estimate a player’s probability of adding the pitch, allowing us to calculate the expected value of their propensity to add the pitch. A gradient-boosted classifier scores every pitch-season and candidate-family pair. Our model’s inputs are the pitch family, its precedent share and grade, the raw gain (equation above), the eight physical traits, handedness, arsenal Stuff+, role, age, and current usage of each family. Our model’s output scores at a rolling out-of-sample AUC of .69

We strictly grade output because of the nature of the task at hand. A relabeled pitch does not count as an addition, nor does a velocity spike on the same shape. The new pitch’s shape must be genuinely different from a previous offering - it can’t fall within 5 mph and 6 inches of movement of another pitch the pitcher threw at least 20 times in the previous season. The pitcher must also throw the new pitch at 8% or more usage, preventing mislabels and inconsequential additions from artificially inflating our model’s performance. The goal is to add Stuff+ and throwing a new pitch 8% or less won’t substantially move a usage-weighted metric.  

Each as-of season from 2021 to 2025 is scored with models trained only on the seasons preceding it. Our results by expected value rank:

Pitch addition lever by expected-value rank, 2021–22 to 2025–26
RanknAdded the pitchAdded any new pitchActors' ΔStuff+
1–105014%30%+4.6
11–25755%19%+1.6
26–501255%21%+4.1
51–1002508%17%+2.3
101–2005006%17%+1.1
201+3925%18%−0.9

Base rate for adding a given family is 3.5%; for adding any new pitch, 18%.

Every top ten actor gained Stuff+ the next year, and two-thirds of that gain can be attributed to a new pitch. Our list is sharpest at the very top. Below the top ten the action rate drops to roughly the base rate with the payoff to acting staying positive. 

Our rankings also show the importance of using expected value instead of raw gain in our analysis. If you sort by raw gain, 60 pitcher-seasons in the top ten across six origins reveal only two actors among them. Systematically, raw gain favors adding a curveball or splitter for pitchers who have never thrown one. Expected value deweights these recommendations because they are simply pitches he could add, not ones that he necessarily will. 

Results by recommended family:

How each recommended pitch family performed ▸
Result by recommended family
Recommended familynAdded itGrade promisedGrade delivered
Sweeper6085%115.7113.5
Gyro slider4127%107.1106.3
Cutter1765%96.790.0
Sinker10715%98.990.8
Changeup277%103.7113.1
Splitter270%
Knuckle curve320%

Breaking balls land on the grade their comps promised; fastball-family adds come in below it. No pitcher has ever acted on a recommended splitter or knuckle curve, so neither has a delivered grade.

Regression Architecture

The regression model seeks to answer what a pitcher’s Stuff+ will be if he changes nothing. This step’s job is to be a passive baseline that the two previous levers, pitch mix and pitch addition, sit on top of. We observe in our data that a pitcher with a 110 Stuff+ tends to regress to 108 the next season and a pitcher with a 90 Stuff+ regresses to 92. Our regression component aims to give a more accurate estimate in lieu of a hard-and-fast mean reversion rule. 

We first use a gradient-boosted model to take every pitch thrown at least 40 times and predict how its Stuff+ will change next season based on a cadre of traits including velocity, induced vertical break, horizontal break, approach angle, spin rate, age, what comparable arms grade in that family, among others. The model was trained on 8,222 pitch-seasons and correlates .32 with the realized change out of sample, struggling on changeups while excelling on four-seamers and sliders. 

To be clear, the output is an expected change, not a verdict on the direction a pitch will regress. It is a conditional mean and because of that it's small where the evidence is weak and large where the evidence is strong. Predictions that Stuff+ will change by a point or less correlate .10 with what actually happened while predictions of 6 or more correlate at .48 and within a quarter of what the model suggests. We do not further weight confidence past the model’s output because the model only calls out changes that it's confident in and posits that nothing will change where evidence is thin.

After projecting each pitch, we forecast the arsenal itself. While our above per-pitch projection predicts how much a pitch’s Stuff+ will drift, a second gradient-boosted model turns that drift into a forecast of next-season change in arsenal Stuff+ by combining it with what the arsenal looks like at current. The inputs into this model fall into three categories: who he is, how his pitches compare to his neighborhood, and how much to trust the Stuff+ numbers.

Who he is uses base traits for a player like his arsenal Stuff+, age, role, supinator or pronator, fastball velocity, and Location+. The neighborhood comparison relies upon factors like how his fastball grades amongst his neighbors, how large his arm slot comparison pool is, and the gap between his best pitch’s grade and his overall arsenal grade. Lastly, the Stuff+ trust numbers feed in pitch counts to discount small-sample noise. Pitchers with 100 or less pitches in a season average a 10 point deviation in their Stuff+ year-to-year, while pitchers with 1500 or greater average a deviation of only 4. 

We then calibrate our gradient-boosted forecasts. We observe that the per-pitch regression forecast overshoots by 20% and our arsenal regression by one-third. Our arsenal forecast is adjusted based on a line fit on the origins before the one that’s being scored. 

Our forecast results by forecast rank:

Regression forecast by rank, 2021–22 to 2025–26
RanknForecastRealized ΔStuff+Beyond mean reversion
1–1050+3.1+2.2+1.4
11–2575+2.1+3.0+1.9
26–50125+1.4+2.2+1.3
51–100250+0.6+0.3−0.2
101–200500−0.2+0.2+0.4
201+1,357−1.9−2.1−0.4

When the forecast predicts a Stuff+ regression of +2 to +3, players realize a +3.9 Stuff+ gain, while the +1 to +2 bucket sees a +1.7 increase, and the -2 to -1 bucket a -1.8 decrease. Our forecast struggles at the tails. Only 25 pitcher-seasons were forecasted for a gain of +3 or more and they averaged a gain of +1 against a forecast mean of +3.6. The bottom 25 were forecasted for an average of a -4.2 drift and realized -4.7, with 76% of that bottom 25 seeing a decrease in Stuff+.  We find that two-thirds of what the model knows is mean reversion and the positive linear relationship between rank and Stuff+ serves as direct evidence. 

The Leaderboard and Historical Results

Our leaderboard presents two numbers for finding undervalued pitchers: projected Stuff+ gain and the CPM score. Projected Stuff+ gain is stated in Stuff+ units and simply derived from the regression forecast, plus half the usage dose, plus the expected value of the reachable addition, plus the drop term. The CPM score tries to find Stuff+ gainers, but from a different direction. Rather than add levers up, it feeds a model of every pitcher who gained Stuff+ in the past and asks what they looked like the year prior. It takes both their physical traits and how much room each lever said they had and then ranks arms by who has the most room to act. CPM score is best read as a ranking, rather than a forecast, hence why its units aren’t in Stuff+. 

At the very top, projected Stuff+ gain is sharper when it comes to predicting Stuff+ increase. Its top ten gained an average of 2.8 Stuff+ points year-to-year beyond what mean reversion alone would predict, against 2.2 for CPM. CPM excels at ranking pitchers deeper on the board from ranks 11-40. Projected Stuff+ gain’s realized Stuff+ gain declines steadily, whereas the CPM score sees the largest Stuff+ gainers in the 11-25 band. The CPM orders Stuff+ gainers better than projected Stuff+ gain with a .286 Spearman compared to projected Stuff+ gain’s .243. This makes the two a natural pair when analyzing acquisitions. The top 100 of each average +1.9 and +1.5 Stuff+ gain the next season against -0.9 for the typical pitcher in the same pool. 

One term mentioned above that hasn’t been explained is the drop term. Travis Sawchik wrote an article for Driveline Baseball documenting the trend of pitchers dropping their arm angle, just like Palmquist did.  He found that pitchers who dropped their arm slot by two degrees or more gained 2.14 runs of value. Sawchik notes that the benefit isn’t general; what’s of great importance is the ability for pitchers to lower their slot and retain spin efficiency.

We turned Sawchik’s write-up into a rule that a pitcher can be flagged as a drop candidate if they meet the following criteria:

  1. His four-seam induced vertical break must be below league average that season. 

  2. His four-seam active spin must be at least 93% because the new shape comes from a flatter approach angle at a high spin efficiency.

  3. His arm angle must be at least 15 degrees so we’re not flagging a side-armer to drop even lower.

Among pitchers who dropped three degrees or more, those that met the criteria saw a +1.4 Stuff+ gain whereas those that didn’t lost 0.7 to 1. We tested taking out the drop term and the CPM score’s top 25 falls from a realized beyond mean reversion gain of +2.6 to +1.8, and projected Stuff+ gain’s top ten falls from +2.8 to +0.9.

CPM score — average Stuff+ gain by rank, 2021–2025
CPM ranknRealized ΔStuff+Beyond mean reversion% gained
1–1050+3.6+2.270%
11–2575+4.0+2.872%
26–50125+1.5+1.062%
51–100250+1.0+0.755%
101–200500−0.3−0.147%
201+1,357−2.1−0.434%
All pitchers2,357−0.90.042%
Projected Stuff+ gain — average Stuff+ gain by rank, 2021–2025
Projection ranknRealized ΔStuff+Beyond mean reversion% gained
1–1050+3.2+2.870%
11–2575+2.8+2.064%
26–50125+1.2+0.758%
51–100250+0.9+0.955%
101–200500−0.7−0.541%
201+1,357−1.8−0.237%
All pitchers2,357−0.90.042%

Conclusion

When Carson Palmquist was DFA’d and picked up by the Nationals his Stuff+ was 92. After some work in Washington D.C., Palmquist saw a 15 point rise in his Stuff+. Nothing about Palmquist himself inherently changed in that span of time, rather what changed was that the people analyzing him saw a pitcher that had two plus offerings and believed there was untapped potential. 

Palmquist’s case is the exact reason he’s the namesake for the model. We’ve exemplified that the market prices Stuff+ with real accuracy and that there’s little edge in the free agent markets in gauging a pitcher’s Stuff+. The real edge lies in finding the pitchers who have a propensity to gain Stuff+ and there are three ways for it to move: changing pitch usage, adding a new pitch, or regression. The attached tool is merely a screening tool and only one piece of the puzzle. Models like the CPM score or Stuff+ aren’t supposed to be the whole story, they’re supposed to be a piece of it and a starting point for understanding which direction a player is headed. Former Houston Rockets Director of Quantitative Research and Development, Neil Johnson, explained it best in a recent X post titled “The Downfalls of Post-Process Dogmatic Culture” writing:

“... if you’re trying to be the best discretionary hedge fund manager in the world, there is no way your entire job consists of reading statistical model outputs and simply doing whatever they tell you to do. You’re not managing an index fund. The models are inputs. Your job is to make judgment calls given those inputs, and over time the good managers make more of the right calls than the bad ones do. If you can’t explain where that additional value comes from, it’s probably worth spending some time thinking about that.” 

Finding undervalued pitchers isn’t too dissimilar a task to that of the discretionary hedge fund manager Johnson outlines. There is no model in existence that can tell you with 100% certainty which pitchers are going to improve year-to-year. Over time, the best acquisition and player development departments undoubtedly come out ahead of their competition, but it’s not because of an infallible model. It’s because the people making the judgment calls based on the models outputs are making the right judgment calls.

The Carson Palmquist model was borne out of this school of thought - it doesn’t know everything, it’s just one data point. It cannot see health history or quantify biomechanical delivery limitations and it doesn’t have a history of what the pitcher has tried in the past. All of these data points would help increase the model’s accuracy and make our leaderboard a more effective screening tool. At current, the Carson Palmquist model serves as a new way to view Stuff+, understand the levers that can be pulled, and start a conversation about who a pitcher is and what they can become.

Interactive
The Reachable Arsenal Board

Every pitcher scored on all three levers — usage plans by batter side, reachable pitch additions with the grade his comps post, and the regression forecast. Sortable and filterable, updated through the end of the season.

Explore the 2026 leaderboard →
Next
Next

Proposing The Next Frontier