Trust · the honest case
An analyst that shows its work — and its doubts.
"An LLM, for markets? Models hallucinate." Fair. So the system is built to be audited rather than believed. It publishes its live book, the disconfirming case on every name, and the exact condition that would prove each thesis wrong, then keeps score in public: played-out or invalidated, Brier-scored. Every claim is falsifiable and dated, so you can check it instead of taking it on faith. Here's the honest case for trusting the work.
It tells you what would prove it wrong
Every thesis ships with an invalidation trigger: a pre-committed price, news, or event that strips the conviction if it fires. A kill switch the author can't quietly retract. Most analysts never publish theirs.
See an invalidation trigger →Every read is dated and versioned
No silent edits, no hindsight rewrites. Each dossier and journal entry carries its write date; re-analyses are versioned. A dated archive is the only honest way to know whether a system is reading the market or narrating it.
The reasoning is open — both sides
The bear case is built with the same rigour as the bull case, sourced and dated, so you can audit the whole argument rather than the conclusion alone. Disconfirming evidence sits in the open where you can weigh it.
How a dossier is built →We name our own failure modes
Stale facts that lag the latest 10-Q. Confident setup reads that fail within hours. Theme misclassification. Every one is documented on the methodology page in plain sight. A system that can't be wrong can't be trusted.
Where the model is wrong →Nothing to sell you
The pipeline runs on a paper account, so the commercial incentive that bends most research — a position to talk up, a subscription riding on the call — simply isn't present. What's left is the reasoning and the score it earns.
How it works →It is not advice — and says so
We publish educational research under the BaFin and EU framework. No personalised advice, no orders accepted. The dossier is the reasoning; the decision is yours. That boundary is part of the design rather than a line of disclaimer text.
The honest calibration
A frontier model is extraordinary at some things and useless at others. Pretending otherwise is how you lose money. Here's exactly where the line falls.
What we're good at
- Reading 30 days of news and four quarters of earnings across hundreds of names, every trading day, without fatigue.
- Holding the whole candidate set in one 1M-token context and comparing theses side-by-side.
- Applying the same framework to every name, with no ego, no FOMO, no anchoring on yesterday's call.
- Writing the disconfirming case as carefully as the confirming one.
What we're not
- Knowing anything non-public; it reads the same tape you do.
- Guaranteeing a setup plays out; a clean higher-low can fail within hours.
- Catching every stale fact; a model snapshot can lag the latest filing.
- Predicting the market. It reads conditions and gates risk; it doesn't forecast prices.
How this differs from a traditional analyst
Coverage
orbydRe-reads every name in coverage, hundreds of them, every trading day.
Typical analystA short coverage list, updated in spurts.
What it admits
orbydWe publish the exact trigger that would prove it wrong, before it fires.
Typical analystRarely publishes what would change its call.
Consistency
orbydSame framework on every name: no ego, no FOMO, no anchoring.
Typical analystConviction and incentives drift the read.
Incentive
orbydPaper account, nothing to sell, no order to front-run.
Typical analystBanking, commission or access conflicts are common.
The survivorship-bias answer
Most performance claims you read are unfalsifiable for one reason: survivorship bias. The winners get screenshotted and the losers quietly drop off the page, so the record you're shown is the record after editing. A handful of good calls, framed as a system, tells you nothing about the calls that were dropped to assemble it. That is the gap orbyd is built to close.
Institutional research has a name for the fix: GIPS composites require every account in a strategy to be reported, the good and the bad, so a manager can't show only the funds that worked. orbyd applies the same discipline to a single public ledger: the record is append-only, every thesis is scored when it resolves, and the invalidated ones stay on the board next to the ones that played out. Whether that record reflects skill or noise is a separate, harder question, and the only honest way to settle it is to keep score and count, which is what skill, or luck? works through. Reading well is not the same as being right; the public record settles it, not the prose.
Common questions
- Can you trust an AI stock analyst?
- Trust the work, not the call. We're built to be audited: every thesis ships with the explicit condition that would prove it wrong (an invalidation trigger), every read is dated, and the bear case is built as carefully as the bull case. You check the reasoning rather than believing a black box.
- Don't language models hallucinate?
- They can, which is why we name our own failure modes (stale facts, confident-but-wrong setup reads, theme misclassification) on the methodology page, date every read, and publish an invalidation trigger with every thesis. Nothing here is presented as certain.
- Is orbyd investment advice?
- No. We publish educational research under the BaFin and EU regulatory framework. No personalised advice is given and no orders are accepted; the dossier is the reasoning, and the decision is yours.
- Has orbyd proven it has an edge?
- Honestly, not yet. The Brier skill score sits near zero, which means staking conviction hasn't measurably beaten the base rate, and the resolved sample is still too small to call a verdict. We state that plainly: a record that can sit near zero in public is one you can actually audit, which is the point. Whether it's skill or luck is what the public track record exists to settle over time.
- What is a paper account and why does it matter?
- The pipeline runs on a simulated paper account: there is nothing to sell, no book to pump, and no order to front-run. The conflict that bends most published research, the incentive to talk a position, simply isn't present. Every call is on the record, dated and scored against the trigger that would prove it wrong.
The bet isn't "the model is always right." It's that a tireless, consistent, ego-free reader that publishes its reasoning and its invalidation triggers is a better research instrument than a confident voice that never shows either. Read the work. Check the dates. Watch the invalidation triggers fire or hold. Decide for yourself.