ARCANE Research · The public record · Updated 25 Sep 2026

Research that says what would prove it wrong.

We test whether machines can read markets honestly. Every question is registered before the data can answer, every score comes with its margin of error, and a loss is printed in the same format as a win.

Publications
05
Lost or found nothing
02
Programmes
04
Registered, not yet answered
02

The record

Five publications. Two of them lost or found nothing.

  1. No. 05Evaluation

    Models without evidence lose to markets

    The walk-forward baseline for the ARCANE forecasting benchmark: 83 resolved questions, and the one gain that held.

    0.212 vs 0.284

    Brier score, market vs calibrated ensemble (lower is better)

    Negative result

    What would change itA prospective window, scored after publication, in which calibration no longer helps.

  2. No. 04Paper

    A learner that knows when to say nothing

    Planted structure recovered, nothing invented on noise, and an honest zero on a real credit series.

    0 of 30

    seeded runs in which a noise rule was validated

    Null result

    What would change itA rule validated on a fresh seed set built with no structure.

  3. No. 03Method card

    How a forecast becomes a record

    The rules that decide when an ARCANE Call counts, and who is allowed to say it came true.

    2

    different models must agree before a Call resolves

    Method

    What would change itAny recorded Call found to have been created or altered after its outcome was known.

  4. No. 02Pre-registration

    The graph maps our attention, not the world

    Does where an article sits in our own corpus predict whether its claim comes true? Registered before a single claim resolved.

    1 test

    fixed before any outcome existed, at p < 0.01

    Registered

    What would change itFewer than 60 resolved claims when we look, p at or above 0.01, or a correlation that reverses between the two halves of the corpus.

  5. No. 01Public conceptual note

    What a risk ceiling cannot see

    A symmetric measure of movement, and the loss path a mandate is actually meant to govern.

    2.1% vs 7.8%

    deepest fall of two paths with identical volatility and ending value

    Argument

    What would change itEvidence that, for a given mandate, a volatility limit and a direct loss-path limit bind at the same times.

Programmes

Four questions we keep asking.

  1. 01

    Forecasting & calibration

    Can a system state a probability that still holds up once the world answers?

  2. 02

    Structure & learning

    Finding a pattern is cheap. The test is whether a learner can refuse one.

  3. 03

    Evidence systems

    Every claim should carry where it came from and what would break it.

  4. 04

    Risk & mandates

    Measure the loss a mandate actually governs, not a convenient proxy for it.

How we publish

Five rules every piece keeps.

  1. 01

    Registered first

    The question, the test and what would count as a failure are written down before the data can answer.

  2. 02

    Every number has a margin of error

    A score comes with its sample size and its margin of error, or it is not reported at all.

  3. 03

    Losses are published

    A result that found nothing, or went against us, is published in the same format as a win.

  4. 04

    What would change it

    Every piece names the observation that would overturn it, the same discipline as every ARCANE article.

  5. 05

    A stated boundary

    Every piece says what it publishes and what it holds back, so the gap is visible rather than hidden.

In preparation

Registered, running, not yet answered.

Fig. 2 · The record in order
AUG 2026SEPOCTNOVDECJAN 2027FEB23 AUG · CORPUS QUESTION REGISTERED17 SEP · BACKTEST AND LEARNER RUN; HEAD-TO-HEAD REGISTERED21 SEP · RESOLUTION RULES ADOPTED22 SEP · FIVE PUBLICATIONS RELEASED25 NOV – 25 DEC · PROSPECTIVE ENROLMENT25 JAN 2027 · REVIEW, NO EARLIER
Fig. 2Every dated event in the public record, in order. Solid marks happened; the hatched window and the open mark are planned, not run.
TypeWorkStatus
EvaluationAlpha against frontier models: a pre-registered head-to-headRegistered 17 September 2026. Remaining model arms and human judging in progress. Published whatever it shows.
EvaluationThe ARCANE forecasting benchmark: first prospective windowPlanned. Held-out enrollment 25 November to 25 December 2026; review no earlier than 25 January 2027.
ProgrammeOrder is not a forecastA reading programme on order, measure and complex systems that generates testable market questions without turning history into a signal.

Every public conclusion stays open to its evidence and its limits.

Research · ARCANE