Skip to main content

Accountability

We keep score.

Every day we record a verdict for the whole universe: a PAS (Pythia Academic Score), a Guru Pyramid tier, a best-match guru. This page measures how those cohorts actually performed against the S&P 500 afterward. It is computed mechanically, never hand-picked, and it is a forward, out-of-sample record, not a backtest.

Scorecard vintage October 7, 2026 · cohorts May 16, 2026 → Sep 7, 2026 (93 verdict days) · live since May 16, 2026

Median top-decile name vs SPY · 90d

−2.5 pp

n = 12,142

Top-decile win rate vs SPY

42%

share of names beating SPY

Decile 10 − decile 1, median

+7.8 pp

Too early to tell

Does the score rank the future?

Pooled excess return vs SPY per Pythia decile at the 90-day horizon. A working score slopes upward: low deciles underperform, high deciles outperform.

Current read: the median top-decile name finished 7.8 pp ahead of the median bottom-decile name. Too early to tell: Only 0.4 independent 90-day windows so far (38 cohort dates, but consecutive ones overlap by 89 of 90 days). No interval is credible yet.

Cohort by cohort

Each point is one verdict day's 90-day forward return: the top decile, the bottom decile, and SPY over the identical windows. No pooling, no smoothing.

The full scorecard

By PAS decile

Decile of each cohort's own PAS distribution (10 = highest).

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By PAS decile — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
Decile 10 (highest)+2.5%+1.0%-1.2%mean +1.6%42.8%46.1%14% pooled50,48078
Decile 9+2.9%+0.9%-1.1%mean +2.0%42.4%47.6%14% pooled50,77178
Decile 8+1.2%+1.0%-1.0%mean +0.1%42.5%48.4%14% pooled53,45578
Decile 7+6.5%+0.8%-0.8%mean +5.7%43.4%48.3%14% pooled41,11178
Decile 6+2.8%+1.1%-1.3%mean +1.7%41.0%47.3%14% pooled58,97778
Decile 5+4.4%+1.0%-1.5%mean +3.4%41.7%46.0%14% pooled50,15278
Decile 4+10.0%+0.9%-1.8%mean +9.1%41.0%45.4%14% pooled47,97678
Decile 3+7.9%+0.9%-1.8%mean +7.0%41.0%44.2%14% pooled50,85478
Decile 2+29.5%+1.0%-2.3%mean +28.5%40.1%41.6%14% pooled49,04678
Decile 1 (lowest)+31.6%+0.9%-3.6%mean +30.7%37.9%38.2%14% pooled50,35278

By pyramid tier

Canonical pass-count band at the time of the verdict.

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By pyramid tier — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
DEEP VALUE+1.2%+1.1%-1.1%mean +0.1%43.6%50.2%19% pooled5,87076
QUALITY VALUE+1.1%+1.0%-0.9%mean +0.1%43.2%51.0%19% pooled57,72678
FAIR VALUE+2.7%+0.9%-0.8%mean +1.8%43.8%48.9%19% pooled158,53778
OVERVALUED+15.6%+0.8%-1.7%mean +14.8%41.7%43.7%19% pooled373,09778

By best-match guru

The guru archetype each company most resembled that day.

n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature

By best-match guru — forward return and win-rate vs SPY. n is names pooled; cohorts are verdict days.
BucketAvg returnSPYExcess (median)Win rateP(beat SPY)n, names pooledCohorts, verdict days
Marks Cycle+16.4%+0.8%-1.3%mean +15.6%43.2%44.5%23% pooled155,37478
Pabrai Value+20.2%+1.0%-1.8%mean +19.2%40.6%43.1%23% pooled93,11678
Buffett Moat+9.0%+0.5%-1.3%mean +8.5%41.7%43.3%23% pooled84,55478
Safety First+1.6%+0.6%-1.4%mean +0.9%40.7%46.3%23% pooled59,44178
Greenblatt Magic+2.1%+0.7%-0.5%mean +1.4%45.7%47.9%23% pooled41,96178
Drucker Efficiency-0.5%+1.2%-2.4%mean -1.7%35.1%42.5%27% pooled37,97661
CAPEX Efficiency+4.3%+1.7%-2.5%mean +2.6%36.7%38.9%40% pooled18,64734
Thorndike Outsiders+2.0%+2.0%-1.6%mean -0.0%42.0%42.3%27% pooled13,27263
Buffettology Growth+0.6%+1.6%-1.6%mean -1.0%39.6%47.5%28% pooled6,57059
Value Line Cash+2.1%+2.9%-2.2%mean -0.9%38.3%44.6%46% pooled6,20427
Owner Earnings+3.5%+2.0%-2.2%mean +1.4%41.4%43.0%28% pooled5,71559
Lynch Growth-0.8%+1.2%-1.9%mean -2.0%41.5%44.4%23% pooled74378
Methodology & caveatsShow

Ledger status

as of October 7, 2026 · cohorts May 16, 2026 → Sep 7, 2026 (93 verdict days) · live since May 16, 2026

How a number gets here

Each daily verdict anchors every company to its adjusted-close price on that date (no look-ahead). Once a horizon of 30, 90, 180, or 365 days has fully elapsed, we measure each company's total return to the next available bar and compare it to SPY over the exact same window. Companies without an anchor price or a matured end bar are excluded — never imputed. Deciles and quintiles are assigned over the full anchored cohort on the verdict day (not just the names that later survive to a matured bar), and companies with identical scores always share a bucket — assignment is deterministic and identical across horizons.

How the table is pooled

Across all matured cohorts we report the n-weighted mean return, SPY return, excess (return − SPY), and win-rate (share of names that beat SPY) for each PAS decile, pyramid tier, best-match guru, and PGS / PCI quintile. A bucket is shown only with at least 5 pooled observations — thinner samples are noise and are dropped. Tables also report cohorts (verdict days that contributed) beside n (names pooled).

Caveats — read before trusting

  • Survivorship. Names that delist or stop trading lose their end bar and leave the cohort, so returns tilt slightly toward survivors. These are price returns of names that kept trading, not a tradable portfolio P&L.
  • Short history. Snapshots began May 16, 2026; samples are small and noisy until cohorts accumulate. This is a forward record that grows in real time — not a backtest.
  • New score families start at zero. PGS and PCI quintile cohorts accrue only from the day their snapshot columns shipped — earlier verdicts have no point-in-time record of those scores, so their history is never reconstructed after the fact.
  • No look-ahead, but adjusted-close mixing. Base anchors and end bars share the adjusted-close basis, so there is no look-ahead. One residual: the base anchor is a true adjusted close while ordinary daily bars store the raw close as adjusted — at most a dividend-yield-scale (≈1–2%/yr) drift on individual names. The SPY benchmark is refreshed on a fully-adjusted basis, so the comparison line is unaffected.
  • Not a recommendation. This measures historical price behavior of scored cohorts. Past performance does not predict future results, and nothing here is investment advice.

Why we grade the process, not just the outcome

Markets are a wicked learning environment: over short windows a sound process loses often and an unsound one wins often, so judging a method purely by its recent outcomes systematically rewards luck. That is why every claim on this page carries its sample size and a confidence interval, why verdicts refuse when the interval cannot carry them, and why a negative spread is published as plainly as a positive one.

The benchmark is SPY (S&P 500) on a dividend-adjusted basis, a deliberately unforgiving bar.

The full record lives in the app

Sign in for the interactive explorer: every horizon, per-bucket cohort histories, the PGS and PCI score-family tables, and each company's own since-first-verdict track record.