StatPacks

Making predictive models transparent and usable

StatPacks is a product we designed and built, not a client project. It predicts pitcher strikeout performance in Major League Baseball using a proprietary analytical framework. The design challenge was making the model's reasoning visible, honest, and verifiable.


What predictive tools typically hide

Predictive analytics tools produce outputs that users are expected to act on: a model analyzes data, generates a recommendation, and presents it as a conclusion. The user sees what the model predicts but not why it predicts it.

This creates a binary relationship with trust. The user either believes the tool or doesn't, because there is nothing between those two positions to evaluate. The methodology is proprietary. Confidence levels are unstated or uniform across all outputs. The track record, when it exists at all, tends to be selective.

We designed StatPacks against that default. The question we kept returning to was why anyone should trust a given prediction. Answering it required transparency at four levels: the reasoning behind individual predictions, the methodology underlying the model, the model's confidence in its own outputs, and the historical accuracy of its results.


Decisions and the reasoning behind them

The card interface

Predictive tools typically present their outputs as finished conclusions. The user sees a recommendation and decides whether to follow it, but the factors that produced the recommendation are hidden behind the interface.

We used a card-based layout where each prediction is a self-contained object. The front of a card shows the output: which pitcher, which game, what the model expects. Flipping the card reveals the reasoning. SHAP factor analysis breaks down the variables that pushed the predicted strikeout count higher or lower, with each factor's contribution visible and proportional.

The product's visual register is different from anything else we've built because the context of use is different. Analytical data reviewed in focused sessions calls for a darker palette with denser information and more typographic distinction than operational software checked between job sites.

Daily Picks
Today's Recommendations
Picks generated from the production model. Tap a card to see model reasoning.
Authenticated snapshot · Updated August 6, 2026 at 09:59 · Poisson V6
The front of each card presents the prediction. The back reveals the model's reasoning through SHAP factor analysis, showing which variables pushed the predicted strikeout count higher or lower.

The framework

Even when individual predictions show their reasoning, the system producing them usually remains opaque. Users can see what the model predicted and sometimes why, but not what the model is designed to look for in the first place.

PSI+ is the analytical framework underlying every StatPacks prediction. It evaluates pitchers across four weighted components built from pitch-level data spanning five seasons. Rather than keeping the methodology proprietary, we designed a dedicated page that explains what each component measures, how the components are weighted, and how individual pitchers rank within the framework.

The explanations describe what each component captures in practical terms rather than mathematical ones. A leaderboard shows how pitchers score across all four components, giving users a reference point for evaluating individual predictions. The validation section publishes the full results of a blind test against a season the model had never seen, including year-over-year consistency metrics and predictive accuracy comparisons against standard industry measures.

The page also publishes case studies where PSI+ disagreed with a pitcher's raw strikeout rate and was subsequently validated, two failed experiments that were tested and abandoned, and the technical specifications used to build the framework.

Pitcher Strikeout Index · 2026
PSI+

A new framework for evaluating strikeout ability in modern baseball.

Tells you which pitchers are built to strikeout hitters, not just which ones have recently.

See Who's Leading PSI+  →
Validation Results
Built on 2020–2024 data  ·  Blind-tested on 2025 season
Count Leverage  ·  Velocity  ·  Pitch Angle  ·  SLWR
Last updated: August 4, 2026
PSI+ Framework
Selected Production Sections
Components, leaderboard, and blind-test validation
Authenticated snapshot · August 4, 2026
Four Components · Starters and Relievers Scored Separately

What Goes Into PSI+

Hover each card to reveal the formula.

2026 Season · Through August 4

Leaderboard

Minimum 200 pitches · Scored within role · 100 = league average · Showing top 10 starters

#PitcherPSI+K%CLWVelo p95VAASLWR
1Jacob Misiorowski124.840.6%0.191103.3−4.10°0.187
2Tarik Skubal124.730.6%0.17798.6−5.08°0.254
3Jacob deGrom123.828.9%0.17298.8−4.42°0.195
4Dylan Cease121.436.3%0.15999.1−4.82°0.193
5Chase Burns121.328.7%0.15499.5−4.98°0.247
6Cristopher Sánchez121.027.8%0.16496.7−6.83°0.238
7Gavin Williams120.432.1%0.15599.3−5.09°0.185
8Shohei Ohtani120.027.9%0.154100.2−4.96°0.164
9Nathan Eovaldi118.825.3%0.17095.7−5.44°0.200
10Hunter Greene118.827.3%0.146100.3−4.92°0.228
≥ 120 110–119 90–109 80–89 < 80
Blind Test · 2025 Season

Does It Actually Work?

PSI+ was built entirely on pre-2025 data. Then tested on 2025 pitchers the model had never seen. Every result below is from that blind test.

Year-over-Year Consistency
Is PSI+ Consistent Year to Year?
Starters · Higher = more consistent from year to year
PSI+
0.769
K%
0.669
CSW%
0.590
2024 Metric Predicting 2025 K%
Predictive Accuracy
306 pitchers with back-to-back seasons
r measures predictive accuracy — closer to 1.0 means better.
MetricAllStartersRelievers
PSI+0.58150.67990.5136
CSW%0.54160.59300.4285
SwStr%0.60490.64230.5227

SwStr% edges PSI+ in the overall numbers. PSI+ pulls ahead when you look at starters specifically, and holds up better from year to year.

Overall accuracy when PSI+ disagreed with K%
69.0%
236 of 342 cases
Accuracy flagging underrated pitchers
58.8%
PSI+ high, K% rose the next year
Accuracy flagging overrated pitchers
79.1%
PSI+ low, K% fell the next year
Case Studies · 2020–2024
Where K% Missed, PSI+ Didn't

Pitchers where PSI+ disagreed with their raw strikeout rate. The following season showed who was right.

UNDERRATED
2021
Jesús Luzardo
K%
22.5%
PSI+
114.2
Next K%
29.9%
Change:+7.4pp

"One of the first cases where I went back and double-checked the number. A 22.5% K rate doesn't look like a pitcher worth flagging, and 114.2 felt too high. But the two-strike whiff data was clean. It wasn't noise."

OVERRATED
2022
Adam Wainwright
K%
17.8%
PSI+
80.0
Next K%
11.4%
Change:−6.4pp

"Velocity had dropped. Two-strike whiff rate had weakened for two seasons in a row. The 17.8% K rate in 2022 masked the decline until the numbers aligned in 2023."

UNDERRATED
2020
Zack Wheeler
K%
18.5%
PSI+
105.8
Next K%
29.1%
Change:+10.6pp

"A 10.6 point jump the following year is the kind of movement that makes you want to go find the next Wheeler. That's the whole point of building this."

What We Tried · What Failed
What Didn't Work

Two approaches that looked promising on paper and failed in testing. Showing them here because credibility means showing what didn't work, not just what did.

Geometric Pitch Tunneling (AIS)

Measured how similar two pitches look to a hitter before they break in different directions.

YoY r = 0.15 · Too weak.
Outcome-Based Sequencing (OBAI)

Whether throwing one type of pitch made the next pitch harder to hit. We analyzed over 2 million consecutive pitch pairs with adjusted baselines.

YoY r ≈ 0 · Essentially zero.
Technical Specification
How It's Built
Data Source
Baseball Savant
3.6M pitches, 2020–2026
Training Period
2020–2024
2,060 pitcher-seasons
Holdout
2025 Season
473 pitcher-seasons, untouched
Role Detection
≥ 45 P / app
avg pitches per appearance
Normalization
p2 / p98 clip
within role, scaled 0–1
Scaling
100 = avg
SD = 10 points
Rolling Window
1,000 pitches
min 200, strictly pre-game
Qualifier
500 pitches
season min / 200 rolling

Both had intuitive appeal. Both failed validation. What the data consistently rewarded was simpler: count leverage + stuff quality.

Selected production sections show the four weighted components, the 2026 leaderboard, and blind-test validation. Authenticated framework snapshot dated August 4, 2026.

Confidence and uncertainty

Most predictive tools present all their outputs with the same level of implied confidence. Every recommendation looks the same regardless of how certain the model actually is, and a prediction the model is highly confident about appears identical to one where the edge is marginal.

StatPacks includes a threshold breakdown that stratifies the model's historical performance by strikeout line. Some lines show a clear edge; others fall below breakeven. Both appear in the same visualization with equal prominence, not hidden in fine print or filtered from the interface.

A user looking at the breakdown can see exactly where the model's predictions carry the most weight and where the edge narrows or disappears.

Breakdown

Performance by K Line Threshold

Where the model is strongest and where to be cautious.

Threshold
Total
Win %
vs −110
Overs
Unders
3.5K
114–91n=205
55.6%
+3.2pp
76–6354.7%
38–2857.6%
4.5K
223–160n=383
58.2%
+5.8pp
107–7957.5%
116–8158.9%
5.5K
146–135n=281
52.0%
−0.4pp
76–6653.5%
70–6950.4%
6.5K
55–54n=109
50.5%
−1.9pp
21–3338.9%
34–2161.8%
7.0K+
32–14n=46
69.6%
+17.2pp
6–366.7%
26–1170.3%
Performance stratified by strikeout line from authenticated historical data. The model's strongest and weakest ranges appear in the same visualization with equal prominence. Snapshot dated August 6, 2026.

The performance record

A tool that publishes its methodology, explains its reasoning, and communicates its confidence levels is transparent about what it intends to do. Whether it actually delivers is a separate question, and most predictive tools leave that question unanswered or answer it selectively.

StatPacks publishes its complete performance record. Through August 5, 1,124 predictions had been settled and recorded. The performance page shows season statistics, a rolling win rate across 128 days, segment analysis, and a daily calendar. Losing days appear alongside winning ones. A day where the model went 0-4 is given the same format and the same prominence as a day where it went 5-0.

The performance page is a continuation of the same principle that produced the card flip and the threshold breakdown. A product that asks users to trust its predictions should give them a way to verify that trust against actual results.

MLB · PSI+ V2

Performance

3/26–8/5 · 1,124 picks

Authenticated snapshot · Updated August 6, 2026 at 09:59
Season Record
625–499
1,124 settled picks
Win Rate
55.6%
+3.2pp vs −110 breakeven
Overs
308–264
53.8% win rate
Unders
317–235
57.4% win rate
Rolling Win Rate
7-Day
30-Day
PSI+ Added
80% 75% 70% 65% 60% 55% 50% 45% 40% 35% PSI+ 6/11 3/26 4/10 4/25 5/10 5/25 6/9 6/24 7/9 7/28 8/5
← scroll to view earlier dates
Segments

Top Performing Segments

Top 5 segments by win rate, each side.

Overs
RHP Home ≤4.5
74–47
61.2%
+3.4pp
RHP Home ≤5.5
109–71
60.6%
+2.8pp
RHP Home (all)
123–87
58.6%
+0.8pp
P(Line) 80–85%
21–15
58.3%
+0.5pp
Line 4.5
107–79
57.5%
−0.3pp
Unders
Line 7.0+
26–11
70.3%
+12.5pp
P(Line) 25–30%
40–20
66.7%
+8.9pp
Line 6.0–6.5
43–23
65.2%
+7.4pp
LHP Home
50–27
64.9%
+7.1pp
P(Line) 20–25%
32–19
62.7%
+4.9pp
Daily Results

August 2026

Results through August 5

53.8%
14–12
8/1
3–2
60%
8/2
2–2
50%
8/3
0–4
0%
8/4
5–0
100%
8/5
4–4
50%
A bounded production-structure excerpt showing season statistics, the full 128-day rolling record, August daily results, and top-performing segments. All values preserve the authenticated August 6, 2026 snapshot.

What the product creates

A user who has seen the reasoning, understood the methodology, reviewed the performance thresholds, and examined the historical record can decide for themselves how much weight to give any individual prediction.


Continue