ARCANE Research · The public record · Updated 25 Sep 2026
Research that says what would prove it wrong.
We test whether machines can read markets honestly. Every question is registered before the data can answer, every score comes with its margin of error, and a loss is printed in the same format as a win.
- Publications
- 05
- Lost or found nothing
- 02
- Programmes
- 04
- Registered, not yet answered
- 02
Latest · Evaluation ·
Negative result
Models without evidence lose to markets
Given no fresh evidence, a calibrated multi-model ensemble scored a Brier of 0.284 against 0.212 for the market price at the same moment. The only statistically clear improvement came from walk-forward calibration, which cut Brier by 0.030. This is the floor every later ARCANE forecaster has to beat, published before we try to beat it.
Read the evaluationThe record
Five publications. Two of them lost or found nothing.
No. 05Evaluation
Models without evidence lose to markets
The walk-forward baseline for the ARCANE forecasting benchmark: 83 resolved questions, and the one gain that held.
0.212 vs 0.284
Brier score, market vs calibrated ensemble (lower is better)
Negative result
What would change itA prospective window, scored after publication, in which calibration no longer helps.
No. 04Paper
A learner that knows when to say nothing
Planted structure recovered, nothing invented on noise, and an honest zero on a real credit series.
0 of 30
seeded runs in which a noise rule was validated
Null result
What would change itA rule validated on a fresh seed set built with no structure.
No. 03Method card
How a forecast becomes a record
The rules that decide when an ARCANE Call counts, and who is allowed to say it came true.
2
different models must agree before a Call resolves
Method
What would change itAny recorded Call found to have been created or altered after its outcome was known.
No. 02Pre-registration
The graph maps our attention, not the world
Does where an article sits in our own corpus predict whether its claim comes true? Registered before a single claim resolved.
1 test
fixed before any outcome existed, at p < 0.01
Registered
What would change itFewer than 60 resolved claims when we look, p at or above 0.01, or a correlation that reverses between the two halves of the corpus.
No. 01Public conceptual note
What a risk ceiling cannot see
A symmetric measure of movement, and the loss path a mandate is actually meant to govern.
2.1% vs 7.8%
deepest fall of two paths with identical volatility and ending value
Argument
What would change itEvidence that, for a given mandate, a volatility limit and a direct loss-path limit bind at the same times.
Programmes
Four questions we keep asking.
- 01
Forecasting & calibration
Can a system state a probability that still holds up once the world answers?
- 02
Structure & learning
Finding a pattern is cheap. The test is whether a learner can refuse one.
- 03
Evidence systems
Every claim should carry where it came from and what would break it.
- 04
Risk & mandates
Measure the loss a mandate actually governs, not a convenient proxy for it.
How we publish
Five rules every piece keeps.
- 01
Registered first
The question, the test and what would count as a failure are written down before the data can answer.
- 02
Every number has a margin of error
A score comes with its sample size and its margin of error, or it is not reported at all.
- 03
Losses are published
A result that found nothing, or went against us, is published in the same format as a win.
- 04
What would change it
Every piece names the observation that would overturn it, the same discipline as every ARCANE article.
- 05
A stated boundary
Every piece says what it publishes and what it holds back, so the gap is visible rather than hidden.
In preparation
Registered, running, not yet answered.
| Type | Work | Status |
|---|---|---|
| Evaluation | Alpha against frontier models: a pre-registered head-to-head | Registered 17 September 2026. Remaining model arms and human judging in progress. Published whatever it shows. |
| Evaluation | The ARCANE forecasting benchmark: first prospective window | Planned. Held-out enrollment 25 November to 25 December 2026; review no earlier than 25 January 2027. |
| Programme | Order is not a forecast | A reading programme on order, measure and complex systems that generates testable market questions without turning history into a signal. |
Every public conclusion stays open to its evidence and its limits.