Track record · we keep score in public
Every thesis, on the clock.
Anyone can sound confident. We publish the exact condition that would prove each thesis wrong, then resolve it in public — played out, or invalidated. Strictly non-monetary; just whether the falsifiable claim was falsified. Below: every live commitment and the clock on it.
Calibration · 122 scored
meaningful sampleBrier score
0.25
0 perfect · 0.25 coin-flip
Played-out rate
41%
50 of 122
Calibration error
0.020
lower = conviction tracks reality
Skill vs base rate
−0.02
Brier skill score
Read these as directional, not a verdict: with 122 resolved, a single-tier row (n=1) can show 100% off one outcome and the skill score sits near zero on a young record. The numbers are published as they stand and sharpen as more theses resolve.
What a Brier score of 0.25 means
A Brier score is the mean squared error between the probability a forecast stakes and what actually happened. Each conviction tier carries an implied probability of playing out — SUPREME around 90%, LOW around a coin-flip 50% — and every resolved thesis is graded against the odds its tier published in advance. A score of 0 would be perfect foresight; 0.25 is what you'd get by forecasting 50% on everything.
That gives the number a reference band. Expert human forecasters in Tetlock's Good Judgment work are commonly cited around 0.08, and the strongest LLM forecasters in 2024–25 evaluations around 0.10; an uninformed coin-flip sits at 0.25. Those benchmarks score different questions than equity theses, so read them as scale markers rather than a like-for-like comparison. orbyd's 0.25 is read against those marks, on a record that is still young — 122 resolved theses is enough to be directional, not enough to be a verdict. What makes the figure worth anything is that it can't move: the odds each tier stakes were fixed before any outcome, so the score is reproducible from the public ledger.
What a near-zero skill score tells you
The Brier skill score (−0.02) asks a sharper question than raw accuracy: does staking conviction beat simply forecasting the base rate? A positive value means the conviction tiers add information; zero means they don't yet, which is the honest read on a record this size. Publishing that plainly — rather than a self-reported headline accuracy — is the whole point.
It also draws the line between a real scoreboard and a disappearing one. Research that quietly keeps its winners and drops its losers can claim any accuracy it likes, precisely because nothing was falsifiable to begin with. A score that can sit near zero in public, attached to every resolved call, is one you can actually audit — which is what AI equity research needs to be trustworthy, and what an invalidation trigger makes possible. Whether that score reflects skill or luck is a question a sample this size can't yet settle.
The Brier score, decomposed
0.020
Reliability: the gap between the play-out rate each tier claimed and the rate it actually hit. Lower is better.
0.016
Discrimination: how far the conviction tiers pull outcomes apart from the base rate. Higher is better; near zero means the tiers don't yet separate.
0.242
The irreducible difficulty of the question itself — the variance you'd face forecasting the base rate alone, before any conviction is staked.
These three sum to the headline figure: Brier = calibration − resolution + uncertainty. The split is the honest part. A low calibration term paired with a near-zero resolution term is exactly what a skill-≈0 record looks like — confidence that tracks reality reasonably well, but conviction tiers that don't yet sort the wins from the losses. The refinement term (0.226, uncertainty minus resolution) measures the same thing from the other side: how much sharper than the base rate the forecasts have managed to be so far.
Dashed = perfect calibration. Each dot is a conviction tier; the line shows the gap between what it claimed and what happened.
Does conviction predict the outcome?
| Tier | Claimed | Actual | n |
|---|---|---|---|
| HIGH | 75% | 67% | 3 |
| MEDIUM | 60% | 63% | 27 |
| LOW | 50% | 34% | 92 |
72 of 72 invalidations had the published trigger fire first.
Play-out rate by archetype
- Compounder 50% · 12
- Cyclical recovery 27% · 30
- Theme leader 64% · 11
- Special situation 34% · 32
- Earnings inflection 56% · 16
- Retail squeeze 44% · 9
- Defensive 42% · 12
Play-out rate by regime at resolution
- Risk-on 42% · 109
- Neutral 31% · 13
Resolved · 122
Two ways a thesis resolves — played out or invalidated, both kept on the record:
- FWRD LOW Played out
- EXTR LOW Invalidated trigger fired
- MARA LOW Invalidated trigger fired
- PDFS LOW Invalidated trigger fired
- STM LOW Invalidated trigger fired
- TMC LOW Invalidated trigger fired
- KMX MEDIUM Played out
- LEGN LOW Invalidated trigger fired
- BBY LOW Played out
- CLPT LOW Invalidated trigger fired
- NXT MEDIUM Invalidated trigger fired
- PD MEDIUM Played out
- JOBY LOW Invalidated trigger fired
- STUB LOW Invalidated trigger fired
- WING LOW Invalidated trigger fired
- AMD LOW Played out
- OXM MEDIUM Played out
- PNRG MEDIUM Played out
- NEXA LOW Played out
- PENG MEDIUM Played out
- GCO LOW Invalidated trigger fired
- KALU LOW Invalidated trigger fired
- MTH LOW Invalidated trigger fired
- SHOO LOW Invalidated trigger fired
- FORM MEDIUM Invalidated trigger fired
- HIVE LOW Invalidated trigger fired
- KFRC LOW Played out
- KOS LOW Played out
- LGN LOW Invalidated trigger fired
- RMBS LOW Invalidated trigger fired
- VSH LOW Invalidated trigger fired
- IMOS MEDIUM Played out
- LPG MEDIUM Played out
- BTDR LOW Invalidated trigger fired
- CIEN LOW Invalidated trigger fired
- FLEX LOW Invalidated trigger fired
- FPS LOW Invalidated trigger fired
- KEEL LOW Invalidated trigger fired
- MRVL LOW Invalidated trigger fired
- MTSI LOW Invalidated trigger fired
- PI LOW Played out
- SAIL LOW Played out
- VIAV LOW Invalidated trigger fired
- ATI MEDIUM Invalidated trigger fired
- LRCX MEDIUM Invalidated trigger fired
- KYMR LOW Played out
- ONTO LOW Invalidated trigger fired
- TNGX LOW Invalidated trigger fired
- TSM LOW Invalidated trigger fired
- MYRG MEDIUM Invalidated trigger fired
- AESI LOW Invalidated trigger fired
- AMN LOW Played out
- CCJ LOW Invalidated trigger fired
- HUT LOW Invalidated trigger fired
- ALTO LOW Played out
- SXC LOW Invalidated trigger fired
- SYRE LOW Invalidated trigger fired
- FAC LOW Played out
- AMSC LOW Invalidated trigger fired
- CLSK MEDIUM Invalidated trigger fired
- BTBT LOW Invalidated trigger fired
- SHLS LOW Played out
- PSNL HIGH Played out
- CEVA MEDIUM Invalidated trigger fired
- TLRY LOW Invalidated trigger fired
- RGTI LOW Invalidated trigger fired
- ADPT MEDIUM Played out
- GENI LOW Invalidated trigger fired
- NBR LOW Invalidated trigger fired
- QUIK LOW Invalidated trigger fired
- TSHA LOW Played out
- VRT MEDIUM Invalidated trigger fired
- NBIS LOW Invalidated trigger fired
- OKTA MEDIUM Invalidated trigger fired
- SNOW HIGH Invalidated trigger fired
- BFLY LOW Played out
- DXCM LOW Invalidated trigger fired
- MGTX MEDIUM Played out
- ZETA MEDIUM Invalidated trigger fired
- CDNL LOW Played out
- BLZE MEDIUM Played out
- IPX LOW Invalidated trigger fired
- OPEN LOW Played out
- PTEN LOW Invalidated trigger fired
- HNGE LOW Played out
- RVMD HIGH Played out
- CECO MEDIUM Played out
- PRM MEDIUM Played out
- TBLA MEDIUM Played out
- ALGM LOW Played out
- AUGO LOW Played out
- DIOD LOW Played out
- FIVN LOW Invalidated trigger fired
- JBLU LOW Played out
- MCHP LOW Played out
- MRAM LOW Played out
- PCT LOW Invalidated trigger fired
- POWI LOW Played out
- STRO LOW Played out
- TER LOW Played out
- TRVI MEDIUM Played out
- DRTS MEDIUM Played out
- TWST MEDIUM Played out
- AZZ LOW Played out
- INDI LOW Invalidated trigger fired
- RIOT LOW Invalidated trigger fired
- CLFD LOW Invalidated trigger fired
- INOD LOW Invalidated trigger fired
- QCOM LOW Invalidated trigger fired
- WEST LOW Invalidated trigger fired
- PURR LOW Invalidated trigger fired
- WTI LOW Invalidated trigger fired
- WULF MEDIUM Played out
- GLW LOW Invalidated trigger fired
- NVDA LOW Played out
- RKLB LOW Invalidated trigger fired
- MP LOW Played out
- COHR LOW Invalidated trigger fired
- GLXY LOW Invalidated trigger fired
- IONQ LOW Played out
- SGMT LOW Invalidated trigger fired
- APLD LOW Invalidated trigger fired
Open commitments · 52 · soonest to resolve first
Every active thesis that ships a published, falsifiable invalidation trigger — fewer than the full book, since names without a stated kill criterion (or already resolved) aren't scoreable here. Distinct from the names held on the live book and the forward theme-bets scored on their own board.
-
Invalidates if Weekly close losing the rising 20-EMA and the ~$840 post-reclaim shelf (forfeits the 2026-06-01 $866.97 signal-bar reclaim); OR a peer optical name (MRVL/COHR/AAOI) losing its 50-day while LITE diverges lower either flips momentum to broken.
-
Invalidates if A weekly close below $80 says the guidance gap is being distributed and puts price back inside the pre-print base; secondarily, any push-out or reduction of FY2027 deliveries from the lead hyperscale customer, or a walk-back off the $130-150M guide at the Q1 FY27 print (~2026-10, est.).
-
Invalidates if A weekly close below $780 (loses the June breakout base and the low end of raised sell-side targets); a Q4 FY26 guide (~2026-07-22) that walks back nearline pricing or exabyte growth would confirm the theme flipping to saturated.
-
Invalidates if A weekly close below $52 forfeits the June post-launch base and the 2026-07-16 dilution-flush low ($55.01 close), confirming the delay is being repriced rather than absorbed; secondary break: the first-half-August Falcon 9 window for BlueBirds 11-13 slipping again, or the space complex rolling to saturated as RKLB/Firefly/SPCX keep underperforming.
-
Invalidates if A weekly close below $296 (loses the mid-July shelf) opens the gap toward the rising 200-day near $270 and confirms the AI-memory leg the fund now expresses is in a full unwind; an Intel miss on 2026-07-23 or a soft Lam Research guide on 2026-07-29 that extends the semi-cap derating seals it.
-
Invalidates if A weekly close below $360 gives up the early-July recovery base and confirms the down-leg; secondary confirms are the Q2 delivery milestone missing on continued European share loss, or any FSD class-certification / HW3-retrofit ruling.
-
Invalidates if A weekly close below $690 confirms the July breakdown is a cycle top rather than a reset, putting the 6.2x forward multiple on earnings due to be revised down. Secondary: a confirmed CXMT/YMTC capacity IPO, or the proposed US ban on Chinese memory chips being publicly shelved.
-
Invalidates if A weekly close below $142 (loses the Baird downside anchor and the last published Street floor outside the $99 outlier) confirms the next beta leg lower; a secondary break is the CLARITY Act Senate window passing without a vote, which removes the only dated catalyst in the next 30 days.
-
Invalidates if A weekly close below $52 loses the May breakout shelf and the rising 20-week EMA, flipping the structure from recovery to rollover; a Q2 Medicaid HBR blowout above guide on ~2026-07-28 alongside the theme flipping SATURATED confirms the cost trend is winning.
-
Invalidates if A weekly close below $48 the spread blowing out well beneath the ~$54 takeout signals rising deal-break risk and a slide toward standalone value in the ~$40s; a hard regulatory block (FCC license transfer / national-security review) or an RKLB collapse that guts the stock consideration would confirm the break.
-
Invalidates if A weekly close below $255 turns the $355→low-$300s digestion into a top and loses the post-Q1 breakout base; secondary confirmation would be H2 2026 NAND/NOR contract pricing flattening sequentially in the TrendForce updates, or the memory-shortage theme rolling to SATURATED.
-
Invalidates if A weekly close below $138 loses the June breakout base and rising 20-EMA, turning the pullback into a trend reversal; confirmed if the July 29 Q2 print misses ~$1.14B revenue or shows LTL tonnage rolling over.
-
Invalidates if A weekly close below $19.00 takes out the June breakout shelf and confirms the $28.96 July high as the cycle top; secondarily, a Q2 print on 2026-07-29 that meets revenue but compresses gross margin, showing the mature-node price war reasserting under an AI headline.
-
Invalidates if A weekly close below $355 loses the June post-Q2 reaction-low shelf and converts the correction into a structural de-rate; secondary breaks are a flat-or-lower hyperscaler capex guide in the ~2026-07-28 to 2026-07-31 earnings cluster, or a confirmed Google/Meta next-gen custom-silicon socket loss to Marvell.
-
Invalidates if A weekly close below $98 breaks the early-July acceleration base and negates the 7/11 golden-cross reclaim; secondary break is a Q2 print on 2026-07-29 showing sequential DART or crypto take-rate deceleration, or the CLARITY Act vote window passing without floor action.
-
Invalidates if A weekly close below $575 forfeits the June-low recovery base and re-opens the downtrend from ~$790; secondary: the reported $10B Anthropic compute deal being denied or shelved, Q2 (2026-07-29) ad revenue decelerating below +15% YoY, or the EU DSA finding converting to a formal 6%-of-revenue fine.
-
Invalidates if A weekly close below $190 breaks the ascending structure off the $116.68 52-week low and opens the path to the 200-day near $168; secondary confirmation is Q2 LTL operating ratio deteriorating YoY on the 2026-07-30 print or peer tonnage (SAIA/ODFL) rolling back over.
-
Invalidates if A weekly close below $305 loses the June breakout-retest shelf and the doubled-off-the-low structure; a secondary break comes if the managed-care theme flips to saturation (UNH/CVS roll over) or the star-ratings appeal is lost, keeping 2027 quality-bonus dollars impaired.
-
Invalidates if A weekly close below $400 surrenders the post-Micron-blowout base and confirms the memory-complex downtrend; secondary confirmation is the DRAM complex extending its bear-market drawdown and setting a lower low after the ~2026-07-30 print.
-
Invalidates if A weekly close below $85 breaks the July $88–90 shelf and confirms the sub-NAV discount is still widening; secondary confirmation is the 2026-07-30 Q2 print showing another quarter of net BTC sales with basic mNAV still under 1.0x, or BTC losing $60,000.
-
Invalidates if A weekly close below $64 (the April secondary at $64.25 failing as support) retraces the entire post-raise advance and opens the low-$50s; the story also breaks if the ~2026-07-31 Q2 print lands under the $34M revenue guide or the record $100M InP backlog declines QoQ.
-
Invalidates if A weekly close below $41 (loses the 2026-07-16 capitulation low of $41.03) confirms the de-rating has no floor; secondary break if the Groves criticality target passes 2026-07-31 undelivered, or a 424B takedown lands off the $1B S-3 / $400M ATM.
-
Invalidates if A weekly close below $10.50 confirms the broken structure and opens the $6.87 52-week low with no shelf in between; secondary breaks: an equity raise above $100M on top of the VAC issuance, the VAC agreement repriced or terminated, the OSC $725M commitment withdrawn, or the ASM scheme failing to implement.
-
Invalidates if A daily close below $118 fails the mid-July bounce structure and re-opens the $106.37 June low; secondary: a Q2 print on 2026-08-03 with US commercial revenue growth decelerating below 60% YoY (from +133% in Q1), or a second NATO member formally selecting Arcadia over Maven.
-
Invalidates if A weekly close below $24 loses the May–June breakout shelf and consolidation base; secondary breaks are Q2 adjusted EBITDA (~2026-08-05) printing below the $27–37M guide floor, or the RXO Curve spot index rolling over.
-
Invalidates if A weekly close below $130 loses the June breakout shelf and the rising intermediate trend; confirmation if the consumer-fintech theme flips SATURATED or the ~2026-08-05 Q2 print misses the $135.5M revenue / $1.11B GMV baseline.
-
Invalidates if A weekly close below $228 (loses the 2026-06-09 swing low that defended the May earnings-gap base); secondary break if the AI-networking-optics theme flips saturated Marvell/Coherent/Lumentum optical demand rolling over.
-
Invalidates if A weekly close below $25 loses the June breakout base and drops OSCR back into its prior $10.69–$25.58 range; secondary breaks are the managed-care theme flipping to saturated, or a Q2 MLR print above the 83.4% FY ceiling on Aug 6.
-
Invalidates if A weekly close below $79 forfeits the early-June secondary low and the reclaimed $85 offer shelf, signaling the post-offering repair has failed; secondary condition is the small-cap-ai-momentum theme flipping to SATURATED as marketplace growth decelerates below ~30% YoY.
-
Invalidates if A weekly close below $67 forfeits the post-earnings breakout above the old $67.19 52-week high; secondary break if the squeeze-momentum theme rolls to SATURATED with no new utility signings before the ~early-August Q1 FY2027 print, or the $250M buyback lapses 2026-09-21 unused.
-
Invalidates if A weekly close below $35 loses the June breakout shelf (prior 52-week high $35.65) and negates the re-rate leg; secondary: the spatial-tools theme flipping to saturated (peer downgrades, ILMN/A/BRKR stalling) or Atera H2-2026 shipments slipping into 2027.
-
Invalidates if A weekly close below $91 (loses the 2026-07-18 intraday flush low of $91.50) confirms the unwind has another leg and voids the base-building case; a secondary condition is the 2026-08-06 Q2 print cutting the FY26 >$1.1B guide or flagging a hyperscaler order push-out.
-
Invalidates if A weekly close below $228 takes out the 2026-07-01 capitulation low and confirms the second Calpine lock-up tranche is being distributed with no marginal buyer. Secondary: a second PJM state adopting New York's ≥50MW moratorium framework, or a 2026 adj-EPS guide cut below the $11 floor at the ~2026-08-06 Q2 print.
-
Invalidates if A weekly close below $12.75 breaks the 52-week floor and confirms the theme has moved from correction to full unwind, voiding any re-accumulation read. Secondary: the 2026-08-06 Q2 print showing revenue below Q1's $2.86M, or a new ATM/shelf take-down beyond the $100M Commerce issuance.
-
Invalidates if A weekly close below $8.85 confirms a breakdown to new 52-week lows and voids the base-building watch; a secondary invalidation is the KORUS $200B package finalizing with no named NuScale allocation, leaving the catalyst spent.
-
Invalidates if A weekly close below $148 forfeits the June bounce shelf and re-opens the $132.66 May low; secondarily, an Aug 7 Q2 print that reaffirms rather than raises the $6.72–7.52B FY26 EBITDA guide removes the last un-priced company catalyst.
-
Invalidates if A weekly close below $94 loses the breakout base and rising 50-day shelf that define the advance. Secondary: a Q4 FY26 print (~2026-08-10) with free cash flow not turning positive or book-to-bill under 1, or the defense-electronics theme flipping to saturated as KTOS/AVAV roll over.
-
Invalidates if A weekly close below $1,700 loses the pre-print consolidation shelf and turns the post-earnings breakout into a failed expansion; secondary confirmation would be public news that TSMC forced withdrawal of the planned equipment price increases, or an FY2026 guide cut beneath the €36B floor.
-
Invalidates if A weekly close below $58 confirms the failed OCC round-trip and opens Mizuho's $50 target and the $49.90 52-week low; secondarily, a Q2'26 print (~2026-08-11 est.) showing USDC circulation below the $77.0B Q1 mark with sequentially lower reserve income confirms the Open USD / Hyperliquid margin-compression case.
-
Invalidates if A weekly close below $95 confirms the next leg of the neocloud de-rating after the lost $104 post-Q1 shelf; the bear read only breaks on a weekly close back above $104 — that holds a higher low, or an 8/11 print raising the FY26 $12-13B guide with capex flat.
-
Invalidates if A weekly close below $293 loses the June breakout shelf and ends the momentum leg; secondary breaks are the 28-day past-due rate re-expanding above ~2.5% on the Aug 12 Q2 print (breaks the CashAI edge) or the fintech-consumer-credit theme flipping to saturated.
-
Invalidates if A weekly close below $15.78 confirms the fresh 52-week low as a breakdown rather than a washout and opens undiscovered downside; secondarily, any 424B or 8-K disclosing ATM share sales into a bounce confirms dilution is capping every rally.
-
Invalidates if A weekly close below $7.00 loses the mid-July shelf and puts the $6.18 52-week low in play, confirming the breakdown; a reclaim of $10 on expanding weekly volume, or a Q2 print (2026-08-13) that shows NHanced revenue consolidating at scale, would force a re-read.
-
Invalidates if Daily close below $18.50 (theme invalidation)
-
Invalidates if A weekly close below $265 fills the post-earnings gap and confirms the momentum leg is done, reopening the $246 gap base; a secondary break is the enterprise-AI channel theme flipping to SATURATED as the upgrade cluster exhausts and DELL/HPE/NTAP roll over together.
-
Invalidates if A weekly close below $374 loses the June recovery base and confirms the 2026-06-02 $475.40 high as a cycle top; secondary break is the ai-datacenter-infrastructure theme flipping to SATURATED as the server/memory correction extends without a replacement leg.
-
Invalidates if A weekly close below $42.00 confirms the failed blowoff and loss of the last shelf under the June range, opening the pre-earnings gap zone; secondarily, the AI-server theme flipping to SATURATED with no HPE-specific catalyst before the ~early-September Q3 print.
-
Invalidates if A weekly close below $235 breaks the freight-recovery structure and the neutral-analyst floor; a secondary trigger is a Q2 print (2026-07-15) guiding intermodal volumes flat-to-down, or a second downgrade following Morgan Stanley that flips the transport theme to saturated.
-
Invalidates if A weekly close below $142 loses the May AI-breakout shelf and rising 20-EMA, turning the round-trip into a fully failed breakout; a secondary break is the theme flipping to saturated as DELL/HPE lose leadership, or a FY27 guide cut exposing sub-8% normalized growth.
-
Invalidates if A weekly close below $30 completes the round-trip of the Microsoft-reveal re-rate and puts the spring base in the $20s in play. Secondary breaks: the Anthropic Australian tender awarded entirely to CDC/AirTrunk/NextDC/Stack, or Q4 FY26 AI cloud revenue failing to scale meaningfully above the $17.3M Q3 level.
-
Invalidates if A weekly close below $21 (loses the 7/18 flush low) confirms the breakdown is extending toward a full round-trip of the 2026 move with the $1.5B ATM as a standing lid; the constructive case only re-arms on a weekly reclaim of the $30.60 200-DMA on rising volume with a confirmed higher low behind it.
-
Invalidates if A weekly close below $135 (the IPO offer) confirms the day-one pop has fully unwound and hands the tape to lock-up/float overhang value-trap confirmed. Secondary: China's reusable-booster gap-close erodes the launch-monopoly premium, or the first public Q2 print shows Starlink margin deceleration.
How we're scored
A confident model is easy. A calibrated one is the point.
When a thesis resolves, it lands in the ledger above as played out or invalidated, dated, with a flag for whether the published trigger actually fired. As outcomes accumulate, the calibration question gets a public answer: does conviction predict the outcome — does SUPREME beat LOW, does a regime call hold? Strictly non-monetary — only whether the falsifiable claim was falsified.
The odds each tier stakes · published in advance
- SUPREME 90%
- HIGH 75%
- MEDIUM 60%
- LOW 50%
Each tier claims this chance of playing out. The Brier score above measures the gap between these stated odds and reality.
Common questions
- Do AI stock predictions actually work?
- There's no honest one-word answer — only a record. orbyd publishes one: every thesis ships a dated, falsifiable kill level, then resolves in public as played out or invalidated, and the whole set is scored with a Brier score. The point isn't a headline accuracy figure but that the calls are checkable: a model's predictions only "work" if they survive the triggers they published, and that's exactly what the scoreboard above measures.
- How is orbyd scored?
- Every thesis ships with a published invalidation trigger. When it resolves it's marked played-out or invalidated — dated, with whether the trigger fired. Across resolved theses we compute a Brier score (forecasting accuracy) and break play-out rates out by conviction tier, so you can see whether SUPREME actually beats LOW. Strictly non-monetary.
- How does a thesis resolve, exactly?
- Mechanically, against the thesis's own published kill level — no discretion. The reference is the real closing price the day the thesis was last stated. The kill is the price in the published invalidation ("close below $X"); the target is a 1:1 realisation of that same published risk (as far above the entry as the kill is below it). Walking real daily closes forward, whichever is hit first decides it: a close at/below the kill is invalidated, a close at/above the target is played-out. "Weekly close" triggers are graded on weekly closes. A thesis whose stop is within 2% (too tight to grade on price) or whose kill isn't a clean price level stays open. Graded only on whether the published claim was falsified.
- Why does the score start in June 2026?
- We rebuilt our research pipeline in June 2026. The public score measures only the calls the new system makes, so it begins fresh and fills as live theses resolve. Earlier theses were written by an approach we've since replaced — scoring them would grade the wrong model. We'd rather start honest and empty than carry a retired system's record.
- What is a Brier score?
- The mean squared error between the probability a conviction tier implies (SUPREME≈0.9 … LOW≈0.5) and the binary outcome. 0 is perfect, 0.25 is a coin flip. For a rough reference band: expert human forecasters in Tetlock's Good Judgment work are commonly cited around 0.08, and the strongest LLM forecasters in 2024–25 evaluations around 0.10. Those benchmarks score different questions, so treat them as scale markers, not a like-for-like comparison.
- What does the track record measure?
- Whether each falsifiable claim was falsified — graded against its own published kill level. It scores the reasoning, not a trade.
- What happens when a thesis is wrong?
- It's marked invalidated, dated, and kept on the public record — with whether the published trigger fired first. Being wrong in public, on the record, is the point: it's what makes the score trustworthy.