Case Study
StatPacks
Making predictive models transparent and usable
StatPacks is a product we designed and built, not a client project. It predicts pitcher strikeout performance in Major League Baseball using a proprietary analytical framework. The design challenge was making the model's reasoning visible, honest, and verifiable.
The Problem
What predictive tools typically hide
Predictive analytics tools produce outputs that users are expected to act on: a model analyzes data, generates a recommendation, and presents it as a conclusion. The user sees what the model predicts but not why it predicts it.
This creates a binary relationship with trust. The user either believes the tool or doesn't, because there is nothing between those two positions to evaluate. The methodology is proprietary. Confidence levels are unstated or uniform across all outputs. The track record, when it exists at all, tends to be selective.
We designed StatPacks against that default. The question we kept returning to was why anyone should trust a given prediction. Answering it required transparency at four levels: the reasoning behind individual predictions, the methodology underlying the model, the model's confidence in its own outputs, and the historical accuracy of its results.
The Thinking
Decisions and the reasoning behind them
The card interface
Predictive tools typically present their outputs as finished conclusions. The user sees a recommendation and decides whether to follow it, but the factors that produced the recommendation are hidden behind the interface.
We used a card-based layout where each prediction is a self-contained object. The front of a card shows the output: which pitcher, which game, what the model expects. Flipping the card reveals the reasoning. SHAP factor analysis breaks down the variables that pushed the predicted strikeout count higher or lower, with each factor's contribution visible and proportional.
The product's visual register is different from anything else we've built because the context of use is different. Analytical data reviewed in focused sessions calls for a darker palette with denser information and more typographic distinction than operational software checked between job sites.
Daily Picks
Today's Recommendations
Picks generated from the production model. Tap a card to see model reasoning.
Authenticated snapshot · Updated August 6, 2026 at 09:59 · Poisson V6
★ Pick of the Day
RS
LHP
Ranger Suárez
Chicago White Sox · LHP · Home
/ /
tap to flip
2026 Season Stats
| Avg IP | PSI+ | Stuff+ | Loc+ | Pitch+ |
| 4.9 | 90 | 95 | 106 | 101 |
#2
BM
RHP
Bryce Miller
Detroit Tigers · RHP · Home
/ /
tap to flip
2026 Season Stats
| Avg IP | PSI+ | Stuff+ | Loc+ | Pitch+ |
| 5.6 | 102 | 108 | 110 | 116 |
#3
BA
RHP
Braxton Ashcraft
Milwaukee Brewers · RHP · Away
/ /
tap to flip
2026 Season Stats
| Avg IP | PSI+ | Stuff+ | Loc+ | Pitch+ |
| 5.3 | 116 | 106 | 107 | 109 |
The front of each card presents the prediction. The back reveals the model's reasoning through SHAP factor analysis, showing which variables pushed the predicted strikeout count higher or lower.
The framework
Even when individual predictions show their reasoning, the system producing them usually remains opaque. Users can see what the model predicted and sometimes why, but not what the model is designed to look for in the first place.
PSI+ is the analytical framework underlying every StatPacks prediction. It evaluates pitchers across four weighted components built from pitch-level data spanning five seasons. Rather than keeping the methodology proprietary, we designed a dedicated page that explains what each component measures, how the components are weighted, and how individual pitchers rank within the framework.
The explanations describe what each component captures in practical terms rather than mathematical ones. A leaderboard shows how pitchers score across all four components, giving users a reference point for evaluating individual predictions. The validation section publishes the full results of a blind test against a season the model had never seen, including year-over-year consistency metrics and predictive accuracy comparisons against standard industry measures.
The page also publishes case studies where PSI+ disagreed with a pitcher's raw strikeout rate and was subsequently validated, two failed experiments that were tested and abandoned, and the technical specifications used to build the framework.
Pitcher Strikeout Index · 2026
PSI+
A new framework for evaluating strikeout ability in modern baseball.
Tells you which pitchers are built to strikeout hitters, not just which ones have recently.
See Who's Leading PSI+ →
Validation Results
Built on 2020–2024 data · Blind-tested on 2025 season
Count Leverage · Velocity · Pitch Angle · SLWR
Last updated: August 4, 2026
PSI+ Framework
Selected Production Sections
Components, leaderboard, and blind-test validation
Authenticated snapshot · August 4, 2026
Four Components · Starters and Relievers Scored Separately
What Goes Into PSI+
Hover each card to reveal the formula.
Count-Leveraged Whiff Rate
Measures how often a pitcher misses bats when it matters most. Two-strike whiffs count double. First-pitch misses count half. The heaviest component in PSI+.
Hover to see the formula ↺
Count-Leveraged Whiff Rate (CLW)
Pitches are weighted by count leverage before calculating whiff rate.
Count Weights
Two-strike2.0×
First-pitch0.5×
All others1.0×
YoY correlationr = 0.5818
Fastball Velocity Ceiling
Captures the top-end speed a pitcher can reach when the moment demands it. Not average velocity — the high gear they can access in big counts.
Hover to see the formula ↺
Fastball Velocity Ceiling (Velo P95)
95th percentile of release speed across four-seam fastballs, sinkers, and cutters — measuring the top-end speed a pitcher can reach, not their average.
YoY correlationr = 0.4815
Fastball Vertical Approach Angle
Measures how flat or steep a fastball enters the strike zone. Flatter angles are harder for hitters to square up. A smaller component, but stable year over year.
Hover to see the formula ↺
Fastball Vertical Approach Angle (VAA)
Mean vertical approach angle of fastballs at the front of home plate. More negative values indicate a flatter plane into the zone.
YoY correlationr = 0.4159
Secondary Leverage Whiff Rate
The same count-leverage logic as CLW. Applied only to secondary pitches — breaking balls, changeups, and off-speed offerings. Only included when a pitcher has thrown at least 50 secondary pitches.
Hover to see the formula ↺
Secondary Leverage Whiff Rate (SLWR)
Applies count-leverage multipliers to whiffs on breaking balls, changeups, and off-speed pitches. New in PSI+ v2.
Fallback weights · <50 secondary pitches
CLW57.89%
Velo36.84%
VAA5.26%
When SLWR is excluded the remaining three weights are rescaled proportionally.
2026 Season · Through August 4
Leaderboard
Minimum 200 pitches · Scored within role · 100 = league average · Showing top 10 starters
| # | Pitcher | PSI+ | K% | CLW | Velo p95 | VAA | SLWR |
| 1 | Jacob Misiorowski | 124.8 | 40.6% | 0.191 | 103.3 | −4.10° | 0.187 |
| 2 | Tarik Skubal | 124.7 | 30.6% | 0.177 | 98.6 | −5.08° | 0.254 |
| 3 | Jacob deGrom | 123.8 | 28.9% | 0.172 | 98.8 | −4.42° | 0.195 |
| 4 | Dylan Cease | 121.4 | 36.3% | 0.159 | 99.1 | −4.82° | 0.193 |
| 5 | Chase Burns | 121.3 | 28.7% | 0.154 | 99.5 | −4.98° | 0.247 |
| 6 | Cristopher Sánchez | 121.0 | 27.8% | 0.164 | 96.7 | −6.83° | 0.238 |
| 7 | Gavin Williams | 120.4 | 32.1% | 0.155 | 99.3 | −5.09° | 0.185 |
| 8 | Shohei Ohtani | 120.0 | 27.9% | 0.154 | 100.2 | −4.96° | 0.164 |
| 9 | Nathan Eovaldi | 118.8 | 25.3% | 0.170 | 95.7 | −5.44° | 0.200 |
| 10 | Hunter Greene | 118.8 | 27.3% | 0.146 | 100.3 | −4.92° | 0.228 |
≥ 120
110–119
90–109
80–89
< 80
Blind Test · 2025 Season
Does It Actually Work?
PSI+ was built entirely on pre-2025 data. Then tested on 2025 pitchers the model had never seen. Every result below is from that blind test.
Year-over-Year Consistency
Is PSI+ Consistent Year to Year?
Starters · Higher = more consistent from year to year
2024 Metric Predicting 2025 K%
Predictive Accuracy
306 pitchers with back-to-back seasons
r measures predictive accuracy — closer to 1.0 means better.
| Metric | All | Starters | Relievers |
| PSI+ | 0.5815 | 0.6799 | 0.5136 |
| CSW% | 0.5416 | 0.5930 | 0.4285 |
| SwStr% | 0.6049 | 0.6423 | 0.5227 |
SwStr% edges PSI+ in the overall numbers. PSI+ pulls ahead when you look at starters specifically, and holds up better from year to year.
Overall accuracy when PSI+ disagreed with K%
69.0%
236 of 342 cases
Accuracy flagging underrated pitchers
58.8%
PSI+ high, K% rose the next year
Accuracy flagging overrated pitchers
79.1%
PSI+ low, K% fell the next year
Case Studies · 2020–2024
Where K% Missed, PSI+ Didn't
Pitchers where PSI+ disagreed with their raw strikeout rate. The following season showed who was right.
UNDERRATED
2021
Jesús Luzardo
Change:+7.4pp
"One of the first cases where I went back and double-checked the number. A 22.5% K rate doesn't look like a pitcher worth flagging, and 114.2 felt too high. But the two-strike whiff data was clean. It wasn't noise."
OVERRATED
2022
Adam Wainwright
Change:−6.4pp
"Velocity had dropped. Two-strike whiff rate had weakened for two seasons in a row. The 17.8% K rate in 2022 masked the decline until the numbers aligned in 2023."
UNDERRATED
2020
Zack Wheeler
Change:+10.6pp
"A 10.6 point jump the following year is the kind of movement that makes you want to go find the next Wheeler. That's the whole point of building this."
What We Tried · What Failed
What Didn't Work
Two approaches that looked promising on paper and failed in testing. Showing them here because credibility means showing what didn't work, not just what did.
Geometric Pitch Tunneling (AIS)
Measured how similar two pitches look to a hitter before they break in different directions.
YoY r = 0.15 · Too weak.
Outcome-Based Sequencing (OBAI)
Whether throwing one type of pitch made the next pitch harder to hit. We analyzed over 2 million consecutive pitch pairs with adjusted baselines.
YoY r ≈ 0 · Essentially zero.
Technical Specification
How It's Built
Data Source
Baseball Savant
3.6M pitches, 2020–2026
Training Period
2020–2024
2,060 pitcher-seasons
Holdout
2025 Season
473 pitcher-seasons, untouched
Role Detection
≥ 45 P / app
avg pitches per appearance
Normalization
p2 / p98 clip
within role, scaled 0–1
Scaling
100 = avg
SD = 10 points
Rolling Window
1,000 pitches
min 200, strictly pre-game
Qualifier
500 pitches
season min / 200 rolling
Both had intuitive appeal. Both failed validation. What the data consistently rewarded was simpler: count leverage + stuff quality.
Selected production sections show the four weighted components, the 2026 leaderboard, and blind-test validation. Authenticated framework snapshot dated August 4, 2026.
Confidence and uncertainty
Most predictive tools present all their outputs with the same level of implied confidence. Every recommendation looks the same regardless of how certain the model actually is, and a prediction the model is highly confident about appears identical to one where the edge is marginal.
StatPacks includes a threshold breakdown that stratifies the model's historical performance by strikeout line. Some lines show a clear edge; others fall below breakeven. Both appear in the same visualization with equal prominence, not hidden in fine print or filtered from the interface.
A user looking at the breakdown can see exactly where the model's predictions carry the most weight and where the edge narrows or disappears.
Threshold
Total
Win %
vs −110
Overs
Unders
3.5K
114–91n=205
+3.2pp
76–6354.7%
38–2857.6%
4.5K
223–160n=383
+5.8pp
107–7957.5%
116–8158.9%
5.5K
146–135n=281
−0.4pp
76–6653.5%
70–6950.4%
6.5K
55–54n=109
−1.9pp
21–3338.9%
34–2161.8%
7.0K+
32–14n=46
+17.2pp
6–366.7%
26–1170.3%
Performance stratified by strikeout line from authenticated historical data. The model's strongest and weakest ranges appear in the same visualization with equal prominence. Snapshot dated August 6, 2026.
The performance record
A tool that publishes its methodology, explains its reasoning, and communicates its confidence levels is transparent about what it intends to do. Whether it actually delivers is a separate question, and most predictive tools leave that question unanswered or answer it selectively.
StatPacks publishes its complete performance record. Through August 5, 1,124 predictions had been settled and recorded. The performance page shows season statistics, a rolling win rate across 128 days, segment analysis, and a daily calendar. Losing days appear alongside winning ones. A day where the model went 0-4 is given the same format and the same prominence as a day where it went 5-0.
The performance page is a continuation of the same principle that produced the card flip and the threshold breakdown. A product that asks users to trust its predictions should give them a way to verify that trust against actual results.
A bounded production-structure excerpt showing season statistics, the full 128-day rolling record, August daily results, and top-performing segments. All values preserve the authenticated August 6, 2026 snapshot.