How it works · the process
From market noise to a defensible thesis, every trading day.
orbyd is a frontier language model reading the US tape on a fixed schedule. Five stages turn 30 days of news and four quarters of earnings into a per-ticker dossier you can audit — the full reasoning behind every call. Here's how a name becomes a thesis.
-
Liquidity screen
Rules-basedCut a ~400-name US universe to the tradable few before a model is ever called.
A cheap, rules-only gate. No LLM touches a name that can't be traded cleanly — roughly three-quarters of the field is gone before any model runs.
- Inputs
- Spread · ADV · market cap · tradable shape
- Produces
- ~25% of the universe survives
-
Momentum + narrative scoring
Rules-based + SonnetRank survivors by structure × volume × news density × theme-cluster strength.
Themes are treated as primary. The system reads basket behaviour — how a name moves with its cohort — not isolated ticker action.
- Inputs
- Price structure · volume · 30d news density · theme baskets
- Produces
- Ranked candidate shortlist
-
Deep per-name assessment
Claude OpusRead everything on each candidate and write the dossier you see on this site.
Opus synthesises thesis, invalidation trigger, bull case, bear case, setup, catalysts, and correlation notes — and assigns a quality grade, strongest to weakest. This stage is the public product.
- Inputs
- 30d news · 4Q earnings transcripts · filings
- Produces
- A dossier per ticker + a quality grade
-
Theme-bet curation
Claude Opus · 1M contextReason across the week's triggered clusters side-by-side and shape — or reject — each into a scored theme-bet.
This is the forward hypothesis, and the one place judgement enters. Per triggered cluster the model either rejects it (index-twin, merger-arb, binary biotech — logged and scored) or shapes a theme-bet, concentrated in its one or two narrative-central names, with a stated probability, narrative and invalidation. Every decision is an immutable, public, Brier-scored forecast. Sizing and exits are machine-owned — the model cannot veto them.
- Inputs
- The weekly radar's triggered clusters · candidate dossiers at once · the mechanical cluster baseline
- Produces
- Scored theme-bet forecasts, or a logged rejection
-
Mechanical outcome grading
Rules-basedGrade every forecast against the trigger it published in advance — no discretion.
Each call is graded mechanically against the kill criterion it set for itself, then folded into the public Brier score. Pre-registered, tighten-only kill gates — not a model's judgement — decide whether the strategy scales, de-scales, or is declared dead.
- Inputs
- Resolved theme-bets and theses · the published invalidation triggers
- Produces
- A Brier-scored public record + kill-gate checks
Why a 1M-token window changes the work
The whole week reasoned in one mind — not stitched from hundreds of calls.
Most automated research scores names in isolation and bolts the results together. The curation stage doesn't. Claude Opus's million-token context lets it hold the week's triggered clusters and their candidate dossiers at once — weighing which narratives are real leadership and which are artifacts (index-twins, merger-arb, one-off binary events), across the whole set in a single reasoning pass. That's how a theme-bet gets shaped against its cohort and the regime, not judged on one name alone. The difference between a spreadsheet of scores and an analyst who has read every name in the field.
On the record
Open by default.
Every thesis, the names we hold and the ones we're watching, and every regime and macro call — published the day it's made, dated, and scored against the trigger that would prove it wrong. You can audit the reasoning behind every call and hold each one to the line it set for itself.
And to be clear about what this is: orbyd has not proven an edge. The forward theme-bet book is an experiment run at full transparency, with the shutdown conditions written down before the results. When it started, the mechanical harvest of the same phenomenon had lost to the benchmark five times out of five — the one untested piece is whether the model's curation adds anything, and that is exactly what the public scoreboard measures.
- Thesis, bull & bear case, invalidation trigger
- Setup, catalyst calendar, correlation notes
- Archetype, conviction, theme & regime calls
- The names we hold — and the ones we're circling
- Every call dated, versioned, and scored in public
Common questions
- How does an AI language model analyse stocks?
- We run a five-stage process on a fixed schedule. A rules-based screen cuts a ~400-name universe to the tradable few; a weekly momentum-and-narrative radar clusters the survivors; then Claude Opus reads 30 days of news, four quarters of earnings transcripts and filings per candidate and writes a structured dossier — thesis, invalidation trigger, bull case and bear case. A 1M-token context then lets it shape or reject the week's triggered clusters into scored theme-bets in a single reasoning pass.
- Which AI models does orbyd use?
- Anthropic's Claude Opus and Claude Sonnet. Opus handles the deep per-name synthesis and the 1M-token theme-bet curation pass; Sonnet handles the faster scoring passes.
- What does a 1M-token context window actually do for stock research?
- It lets the curation stage hold the week's triggered momentum clusters and their candidate dossiers at once — weighing which narratives are real leadership and which are artifacts (index-twins, merger-arb, one-off binary events) across the whole set in one reasoning pass — rather than scoring names in isolation and stitching the results together afterwards.
- Does orbyd trade in real time?
- No. A weekly mechanical radar scans for momentum clusters; Claude Opus curates the triggered ones into scored theme-bets; and entries and exits are machine-owned rules, not discretionary intraday trades. It publishes research and runs on a paper account — a forward experiment with pre-registered kill gates, not a proven edge.
- Is orbyd's research automated or human-written?
- Fully automated. The dossiers, regime calls and macro views are written by frontier language models; the methodology is open and every read is dated. Humans don't edit the model's output.
Today we track 612 names across 40 themes. Go deeper: