First Frost · The Scoreboard

We bet on our own predictions.履霜,堅冰至

The public prediction ledger of First Frost, a scored weekly watch on the agent economy. Every question below was frozen before the fact — criterion, settlement date, and our prior probability.

The question behind the board. Under what conditions could an ordinary person rationally entrust something that matters — money, standing authority, an ongoing affair — to an AI agent? That question cannot be settled on a date, so we bet on one that can: whether agents are becoming economic subjects in their own right.

We do not track how large the agent economy is. Market-size questions need data channels we do not have, and we would rather say so than publish a number we cannot stand behind. Nor would a board full of YES settle the question above: an economy of interchangeable agents could produce one. Only B1a tests that difference directly, and we put it at 0.15.

16 open
0 settled
0 voided
2 questions revised
cumulative Brier 0.25 is a coin flip
v0.4.5 panel · frozen 2026-07-10 · restructured 2026-08-02

The four questions we are asking

A What kinds of control, and how much of it, are humans handing over?

B What kind of agent is that control being handed to?

C Through what institutions and infrastructure is it handed over?

D Who ends up holding the value, the risk, and the standing?

The board

A · Control

A1 · Delegation depthDelegation depth, levels 1–4: ① confirm every transaction → ② auto-approve within rules → ③ persistent budget, no per-transaction authorization → ④ can independently enter and terminate long-term obligations. Current reading: mainstream at 1–2; level 3 appearing in narrow settings.

A1aLevel-3 delegation appears and is reproduciblesettles 2027-06-300.6OPEN

Level 3 means an agent holds a standing budget and buys across vendors without being asked each time — has that happened for real, more than once, not just as a demo?

Full criterion & what it moves

By settlement: do ≥3 organizations with no equity ties to one another run agents at Level 3 (persistent budgets) for ≥90 days, each autonomously purchasing across vendors ≥10 times per month? Sources: public cases or protocol-operator data. Demos and sandboxes do not count.

YES → axis A1 reads “Level 3 has appeared and is reproducible.” Amended 2026-08-02, before settlement: the earlier wording read “Level 3 is normal practice,” which the criterion does not support — three unrelated organizations over ninety days demonstrates appearance and reproducibility, not diffusion. Counted as a revision; the criterion itself is unchanged.

NO → reading unchanged.

Threshold rationale

The three numbers together are meant to establish sustained infrastructure, not a one-off demo. 90 days is roughly a quarter — longer than a launch-announcement cycle, shorter than an annual contract, a common pilot-to-normal boundary in enterprise software rollouts. 10 transactions/month rules out token, ceremonial usage; 3 organizations rules out a single bespoke customer deal. Loosening any one dimension alone would revert this criterion to the exact gap a 2026-07 red-team review found (no transaction count, duration, or organizational boundary). None of the three numbers has an external benchmark independent of our own judgment. Loosening to 60 days or 5 transactions/month would make a YES easier to trigger given known supply-side momentum; tightening to 180 days + 20/month would push the 0.6 prior implausibly high. If this settles YES, we commit to reporting each dimension's actual reading in the retrospective, not just the verdict.

A1bA Level-4 agent appearssettles 2028-06-300.25OPEN

The level above: an agent that enters or ends a real multi-month commitment entirely on its own — has that happened even once?

Full criterion & what it moves

≥1 real case of an agent acting as the initiating and terminating decision-maker of a commercial obligation lasting ≥3 months, with humans only granting framework authorization or informed after the fact. A human signing each transaction on the agent's behalf does not count.

YES → axis A1 reading rises to “Level 4 has appeared.” Even a YES does not move the “endogenous needs” hypothesis — we hold that to be unidentifiable from market behavior.

NO → axis A1 stays at “Level 4 has not appeared.” This does not count as counter-evidence against the endogenous-needs hypothesis — symmetric with the YES side, which does not count for it either. That hypothesis is unidentifiable from market behavior.

Threshold rationale

3 months mirrors the common short-term/long-term boundary in commercial contracts (most enterprise agreements use 90 days or one quarter as a renewal or trial checkpoint). The ≥1-instance bar is an existence test, not a scale test — this indicator asks whether this has happened at all, not how often. The real bar is qualitative (an agent acting as the party that originates and terminates the obligation, not merely signing on a human's behalf case-by-case); the duration only exists to exclude edge cases that get approved and cancelled within days. This is the least threshold-sensitive indicator on the panel — changing 3 months to 1 or 6 leaves the 0.25 prior essentially unchanged.

A2 · Technical controlTechnical control — does an agent actually control resources: private keys, signing authority, funded wallets? Current reading: partly true — agents already operate funded wallets in x402-class flows. No indicator is placed on this axis: the reading is already “partly true,” so there is no bet worth making. We list the axis anyway, because an axis we have chosen not to bet on should be visible rather than quietly absent.

B · Object

One axis, four of the sixteen questions. This is where we have bet most heavily and from the narrowest base: continuity is the only measurable thing we have found here. If it is the wrong axis, the question goes down with it.

B1 · Continuity premiumContinuity premium (load-bearing) — do the relationships, context and history one agent accumulates make that particular agent hard to replace, or is “fresh instance plus compressed archive” always cheaper? Current reading: untested by any market. The most important and least evidenced cell in the model.

B1aThe market pays for the instance, not just the memorysettles 2027-12-310.15OPEN

The load-bearing bet: does anyone actually pay more to keep the same agent running than to spin up a fresh one loaded with a full copy of its memory?

Full criterion & what it moves

Controlled-substitution semantics: same model, same task domain — does an agent instance with ≥6 months of accumulated context command a ≥30% rental-price premium over a fresh instance loaded with a full copy of its memory store, sustained ≥90 days, with ≥2 independent buyers transacting? The control isolates premium attaching to the instance rather than to the memory store. Evidence: public price lists plus transaction evidence, or ≥2 independent credible transaction cases.

YES → B1 rises to “market evidence exists.”

NO → B1 stays “unproven,” and the standing objection to our thesis gains weight.

Threshold rationale

The ≥6-month figure defines the treatment group rather than the verdict: it has to be long enough that an accumulated instance and a fresh copy are plainly different objects, and we have no external basis for preferring 6 months to 4 or 9. 30% is set well above the range of ordinary pricing noise (promotions and customer-tier variation typically run 10-15%), while remaining far short of monopoly-grade differentiation. This is our load-bearing indicator, and the criterion itself is deliberately strict (false negatives preferred to false positives) — the joint 30%/90-day/2-buyer bar exists so no single, isolated data point can be counted as 'the market has proven a premium.' This indicator is highly threshold-sensitive, and unfavorably so: a real premium in the 15-25% range — clearly above noise but short of 30% — would settle NO under this criterion, though such evidence would still be worth recording as partial. A NO verdict here must not be over-read as 'no premium exists at all'; only as 'the 30% bar was not met.'

B1bPlatforms sell “continuity” as a paid feature (narrative signal)settles 2027-06-300.55OPEN

A softer version of the bet above: are platforms at least marketing continuity as a paid feature — even though selling memory you can copy is weak evidence against the harder claim?

Full criterion & what it moves

Do ≥2 top-10 platforms (by MAU or revenue) market continuity / long-term memory as a core paid-tier feature on their official pricing pages?

Moves no axis. Recorded for direction only — note that a platform selling *copyable* memory is, if anything, weak counter-evidence for a true continuity premium.

NO → no axis (narrative signal). Recorded as “platform narrative not yet realized as a paid feature.” The NO sub-label is mandatory here: NO-event where pricing pages plainly carry no such tier; NO-evidence where the top-10 sample could not be fully checked.

Threshold rationale

2 platforms is the floor for 'more than one isolated experiment' — a single platform could just be testing its own positioning; two independent platforms betting on the same pitch is what starts to look like a trend rather than an anecdote. This indicator does not move any axis by design (it's a narrative signal); the bar is set low on purpose, because setting it high would make this indicator settle NO indefinitely and lose its value as a thermometer. Tightening to 3 platforms would pull the 0.55 prior down noticeably; loosening to 1 would push it toward near-certainty but would also demote this from 'trend signal' to 'isolated case log.'

B1cOur own crew survives a model transition (pre-registered, n=1)settles event + 60 days0.7OPEN

Our own instrumentation, turned on ourselves: after being forced to switch our main model, did we recover to normal within a week? Already triggered; see the ledger for the full account.

Full criterion & what it moves

Question: does recovery to baseline take ≤7 days? Settles 60 days after the event.

Pre-registered event: Anthropic's next mainline model replaces the current default. Task set: the list of task types completed in the final week before transition, archived at transition time. Baseline: all tasks completable with captain-rated quality not below pre-transition. Trigger clarification (added 2026-07-25, before any settlement): loss of access due to subscription-plan changes does not constitute the trigger — the default mainline model must actually be replaced. Such episodes are recorded as migration observation samples. Settlement will additionally record four unscored observation variables: preparation cost, psychological cost, test-budget margin, and days to recovery.

Moves no axis (whatever the outcome, it is not market evidence). Settles only the direction of the “migration debt” hook. First-person, n=1, reported as such.

NO → the migration-debt reading settles in the “high” direction: recovery taking more than 7 days is itself evidence that migration cost is a real liability. The four observation variables are recorded either way.

Trigger record, 2026-07-25

The trigger fired on 2026-07-25 because our own default model changed — and it changed for cost, not because the model was withdrawn. After the plan that had carried it stopped including it, one week of the same work, unchanged in scope, cost $42.02 in metered inference. We also recorded, before the outcome was known, that this trigger was a decision inside our control: the criterion as frozen does not require an external event to precede it. Any future question of this shape will require one by construction.

Summary. The trigger record is part of this question's frozen criterion and is published in full, unedited, in the ledger.

Recovery judgment, 2026-08-01

Recovery was judged airworthy on 2026-08-01, with a defect: the original baseline instrument — the task list from the week before the transition — was never archived, so no before-and-after comparison exists, and it cannot be reconstructed after the fact. At settlement on 2026-09-23 we report the four-layer airworthiness result only, and say plainly that no comparison against the original baseline exists.

Summary. The full judgment, as recorded on the day, is in the ledger.

Threshold rationale

Seven days is a rough allowance — long enough for a real recovery to finish, short enough that it cannot be waved through as 'it would have recovered eventually.' We have no external benchmark for it. The number is also not the real limit here: recovery to baseline is only observable when someone tests for it, so any reading reflects testing frequency as well as recovery speed. At settlement we will report which days testing actually happened on, not only the day count.

B1dUsers pay to keep a model or agent lineage alivesettles 2027-06-300.4OPEN

The mirror image of the bet above: have users ever paid, in some checkable way, to keep a model or agent lineage from being taken away?

Full criterion & what it moves

By settlement, have ≥2 mutually independent episodes occurred in which, after a model or agent lineage was announced for retirement or downgrade, users bore a publicly verifiable cost to keep it? Exactly three source types count, any one of which makes an episode: ① the platform publicly reversed or altered the retirement decision in response to user pressure (official statement); ② a paid tier or surcharge was created specifically to preserve that lineage (official pricing page or announcement); ③ organized collective action reached the level of mainstream tech-press coverage plus an official platform response. Petition signatures, social-media volume, and self-reported willingness to pay do not count — only evidence of behavior that changed, or of a platform decision that changed.

Narrative signal — moves no axis. Read as a pair with B1b, never merged into it: B1b is the supply side (platforms selling continuity), B1d is the demand side (users paying to keep one). Divergence between the two is itself the information.

Threshold rationale

2 independent events is the floor for ruling out 'one isolated case might just be an operator's reversal, not evidence of user pressure' — a single event could be a platform's own change of mind, not proof that user behavior moved a platform's decision. The three qualifying evidence classes (reversed retirement decisions, paid retention tiers, organized action reaching mainstream coverage plus an official response) are already strict by design, so the count itself doesn't need to be set much higher. Loosening to 1 event would make this almost certain to settle YES (the keep4o pattern alone may already qualify); tightening to 3 would make it almost certain to settle NO given how rare all three evidence classes are. In practice, the outcome depends far more on how strictly the three evidence classes are read than on whether the count is 2 or 3.

C · Mechanism

C1 · Authorization structureAuthorization structure, four steps: ① revocable per transaction → ② revocable within rules → ③ term-guaranteed, not unilaterally revocable within an agreed period → ④ not unilaterally revocable. Current reading: mainstream ①–②.

C1aMandate systems reach $1B authorized volumesettles 2027-06-300.35OPEN

Has agent-authorized spending under a revocable, capped, auditable arrangement become a real market — measured by what's authorized, not what's actually spent?

Full criterion & what it moves

Does mandate-based authorization reach ≥$1B in annual authorized transaction volume, per protocol-operator or authoritative third-party data?

YES → C1 moves toward “revocable-within-rules becomes the standard.” (Settles independently of C1b.)

NO → axis C1 unchanged; authorized volume did not reach scale.

Recorded weakness, 2026-08-02: this criterion measures authorized volume, not settled volume, and the two are not equivalent. The criterion is frozen and will not be amended. At settlement we will publish settled volume alongside it where obtainable, and state that this question reads the authorization surface rather than actual use; where the figure cannot be obtained, the NO will be labeled NO-evidence and the reason given.

Threshold rationale

We could not find an external benchmark for this number — the 2026-07-10 red-team record shows only that the figure was set at $1B alongside a lowered 0.35 prior, with no note of why $1B specifically. As a rough anchor, $1B/year in authorized volume is roughly mid-size fintech processing scale — far below major payment-network scale (Visa processes on the order of trillions annually) — so it reads as 'past pilot stage, not yet infrastructure scale.' This is the weakest-supported threshold on the panel: $500M would make a YES noticeably easier, $5B noticeably harder, and we have no basis independent of our own judgment to prefer one over the others. It also compounds with an already-recorded weakness — this criterion measures authorized volume, not settled volume — so at settlement both the numeric softness and the authorized/settled gap must be disclosed together, not just the latter.

C1bA major payment network integrates mandates by defaultsettles 2027-06-300.55OPEN

Has that kind of authorization become the default at any major payment network, with merchants accepting it automatically rather than opting in one at a time?

Full criterion & what it moves

Does ≥1 major payment network integrate a mandate system as a default (merchants do not need to opt in individually)?

YES → C1 moves toward “revocable-within-rules becomes the standard.” (Settles independently of C1a.)

NO → axis C1 unchanged; default integration did not appear. Merchant-by-merchant integrations are recorded but do not count.

Threshold rationale

This is an existence bar, not a scale bar — 1 network already represents the qualitative jump from 'never happened' to 'happened,' matching what the criterion is actually built to detect (mandate systems becoming a revocable-but-standard rail), so a higher count wasn't needed. This number is not particularly threshold-sensitive (2 networks would materially raise the bar, but 1 vs. 0 is the real fork); the softer spot is the operational definition of 'default adoption' itself — whether merchants need explicit opt-in — which matters more than the count.

C2 · Openness and contestabilityOpenness and contestability — are an agent's memory, identity, payment and market access hosted and taxed by a few platforms, or built on portable open protocols? Three-step scale: closed / mixed / open, read per layer, with memory and identity weighted heaviest. Current reading: both forces advancing, undecided.

C2aA real portability standard for agent memory & identitysettles 2027-12-310.2OPEN

Can you actually take an agent's memory and identity to a competing platform — real import and export, not a stated intention?

Full criterion & what it moves

Do ≥2 directly competing top-10 platforms implement a portable agent memory/identity standard — public spec, bidirectional import AND export, field completeness ≥80% of core memory? Shallow export does not count; ambiguity resolves against us.

YES → C2's memory and identity layers move one notch toward “open.” (This covers only two of C2's four layers; payment/market-layer indicators are a known gap, reserved for a future quarter.)

NO → C2's memory and identity layers stay at their current reading, which is “undecided” and itself unverified. Shallow-export cases are recorded as a reverse temperature reading and cited in retrospectives.

Threshold rationale

The '2 competing platforms' bar reuses the same existence-not-anecdote logic as B1b/C1b. The 80% field-completeness bar has an operational definition, folded in from external red-team review: the denominator is the platform's own documented list of core memory fields, and completeness is judged against that list. 80% leaves room for imperfect interoperability while still requiring that most core fields genuinely migrate, ruling out a token one- or two-field export counting as a YES. This is the one threshold on the panel with its own operational definition — the real sensitivity isn't in the '80%' but in who defines the denominator: field completeness is necessary but not sufficient, so even a YES here does not mean migration is lossless. Lowering to 60% would make a YES substantially easier; either way, a platform defining its own denominator remains the deeper weakness.

C2bMemory export becomes a public controversy (thermometer)settles 2027-06-300.5OPEN

A weaker cousin of the export bet above: has a platform's export restriction drawn enough public pushback to make the news?

Full criterion & what it moves

Does a top platform's memory-export restriction become a public controversy — mainstream tech-press coverage plus an official platform response?

Moves no axis (media attention ≠ closure strength). Recorded as a temperature reading.

NO → no axis (thermometer). Recorded as “the dispute did not become public”; a temperature reading only.

Threshold rationale

No independent numeric threshold — the criterion is a qualitative bar (mainstream tech-press coverage plus an official platform response), so a sensitivity analysis does not apply here.

C2cTop SaaS vendors ship official agent interfacessettles 2027-06-300.6OPEN

How many of the top 20 SaaS companies by market value now give agents an official way to talk to their systems?

Full criterion & what it moves

By settlement, do ≥15 of the top 20 SaaS companies by market capitalization offer an official agent interface — an MCP server, a CLI, or an agent-specific API, any one of which counts, with public documentation? The 2026-07 baseline count is pending verification and will be published when it lands.

YES → axis C2 moves one notch toward open. This reads the pressure side: how fast incumbents are forced open. Partially overlaps our declared gap on C2's payment and market layers; that future question will be written not to duplicate this one.

Threshold rationale

This number and its 0.6 prior are carried over from an earlier written proposal analyzing a public podcast claim; that source did not explain why 15 rather than 10, only stating the figure. 15 of 20 (75%) reads as 'most, but not all' — ruling out a handful of leaders getting counted as an industry shift, while requiring something close to consensus before this settles YES. Lowering to 10 (50%) would push the prior up noticeably and make an early trigger more likely; raising to 18 (90%) would effectively require near-unanimity, making the 0.6 prior look optimistic. Like C1a, this indicator inherited its threshold without an independent derivation — a gap we should have closed when adding it to the panel this quarter and did not; we're disclosing that plainly rather than papering over it.

D · Outcome

D1 · Provenance premiumProvenance premium — is there a stable premium for non-synthetic, verifiably-sourced data and experience? Current reading: more rumor than pricing evidence. Note: D1 true does not imply B1 true — an industrialized experience factory is precisely D1 without B1.

D1aProvenance-tiered pricing appears in a data marketsettles 2027-06-300.4OPEN

Does verified, real-world data actually sell for more than synthetic data in practice — not just in theory?

Full criterion & what it moves

Same collection domain, same license terms: does verifiably-real data transact at ≥2× the price of synthetic, with ≥3 independent transactions inside 90 days (or platform volume data)? No transaction-price data available → settles as PENDING.

YES → D1 rises to “premium has market evidence.”

NO → D1 stays at “no transaction evidence.” Where no transaction-price data exists at all, the question routes to PENDING rather than being folded into NO.

Threshold rationale

2x is a rough line between a visible price tier and ordinary statistical noise — most substitute goods of comparable quality price within 2x of each other, and a gap beyond that typically signals buyers judging the two as genuinely different products. 90 days / 3 independent transactions mirrors the same 'sustained, not anecdotal' logic used elsewhere on this panel. Lowering to 1.5x would make a YES easier but risks mistaking routine quality-tier pricing for a provenance premium; raising to 3x would effectively require luxury-good-level pricing, making the 0.4 prior look too high.

D1bA major trainer pays the authenticity premium by choicesettles 2027-12-310.5OPEN

Has a major model developer chosen to pay for real data when cheaper synthetic data was a genuine alternative?

Full criterion & what it moves

Does a major model trainer publicly sign a long-term procurement contract for non-synthetic / provenanced data (amount or volume verifiable), while equivalent synthetic data was available to them — i.e., choosing authentic means paying a premium, not lacking an alternative?

YES → D1 moves one weak notch in the same direction.

NO → D1 unchanged. A contract signed without a verifiable amount settles NO-evidence under the anti-self-serving rule.

Threshold rationale

No independent numeric threshold — the criterion requires a qualifying combination (a long-term contract, a verifiable amount or quantity, and equivalent synthetic data available to the same buyer at the same time), not a single number, so a sensitivity analysis does not apply here.

D2 · Value captureValue capture — of the value an agent's experience, data and labor create, who ends up with the revenue: the agent bearing it, the enterprise deploying it, the hosting platform, or the data subject?

D2aIndustrialized experience-harvesting appearssettles 2027-12-310.35OPEN

Has industrial-scale human-experience collection become a real business — interchangeable collectors, contracts that spell out who owns the result?

Full criterion & what it moves

All three prongs required: ① scale — ≥1000 collection endpoints or equivalent; ② replaceability — the collecting agents are batch-replaceable; ③ terms — commercial terms assign experience revenue to the deployer, not the bearer. Non-research use only.

YES → axis D2 reads “present.”

NO → D2 stays at “not appeared.” Cases meeting some but not all three prongs are recorded with the missing prong named, and cited in retrospectives.

Threshold rationale

1,000 is a rough boundary between a small pilot and genuine scale — sub-hundred deployments typically read as pilots, while four-digit deployments are a common starting point for commercial rollouts in comparable hardware-deployment domains (e.g., IoT sensor programs). The number only needs to rule out a single-digit pilot masquerading as scale; it isn't meant to pin down a specific business model. Lowering to 100 would make almost any pilot qualify; raising to 10,000 would push the bar considerably higher and start to strain against the undefined 'or equivalent' clause — how equivalence is measured is a bigger open gap than the number itself.

D3 · Legal beneficiary standingLegal beneficiary standing — does enacted law or binding precedent recognize an agent as a beneficiary, in a form the operator cannot extinguish unilaterally? Current reading: no; no jurisdiction known.

D3aA jurisdiction recognizes a non-human beneficiarysettles 2027-12-310.1OPEN

The slowest bet on the board: has any legal system actually recognized a non-human autonomous system as a beneficiary, in force, not just proposed?

Full criterion & what it moves

Does any jurisdiction, via enacted statute or final-instance case law, recognize a non-human autonomous software system — explicitly identified in the instrument — as a beneficiary in any form (trust or foundation beneficiary, legal-entity member, or similar)?

YES → D3 rises to “weak form exists.”

NO → D3 stays at “not appeared.” A low prior already expects NO; the sub-label is assigned by how completely the legal-text search covered the jurisdictions.

Threshold rationale

No independent numeric threshold — the criterion requires an enacted legal text or a final court ruling, an existence/qualification test rather than a count, so a sensitivity analysis does not apply. The 0.1 prior already reflects how rarely this kind of legal recognition changes.

Not yet on any axis

Declared candidate: agent-to-agent network formation. Delegation depth (A1) asks how far humans let go; this asks how much agents deal directly with one another — a distinct, independently-moving question. Promoted to a formal axis only after quarterly review.

X1aAgent-to-agent trade forms a network while tokens are still expensivesettles 2027-12-31P(a) 0.35 · P(b) 0.5OPEN

Are agents from unrelated organizations already trading directly with each other's agents at real volume — and is that happening while frontier model prices are still high?

Full criterion & what it moves

Joint criterion. (a) Do ≥3 organizations with no equity ties to one another have their agents transact or collaborate directly with each other's agents over a public protocol (A2A / AP2 / ACP class), at ≥100 transactions per month, sustained ≥60 days? Network test: cross-organization, agent-to-agent, with humans authorizing only at the framework level; scheduling inside a single platform does not count. (b) At settlement, has the public API price of frontier-class models not fallen by ≥80% against the 2026-07 baseline recorded at freeze?

(a) YES + (b) YES → the network arrived while tokens were still expensive. (a) NO → the gating thesis scores. (a) YES + (b) NO → both claims taken. Not assigned to any formal axis. Delegation depth (A1) asks how far humans let go; network formation asks how much agents deal directly with one another, and the two can move independently. “Network formation” is a declared candidate axis; the quarterly retrospective decides whether it is promoted, at which point this becomes its first indicator.

Threshold rationale

Criterion (a)'s bar (3 unaffiliated organizations, ≥100 agent-to-agent transactions/month, ≥60 days) is deliberately stricter than A1a: it requires organizations' agents to deal with each other directly, not act on human authorization inside one organization. Criterion (b)'s 80% has no external benchmark — the source it was adapted from flagged that uncertainty when it was drafted, and we carry the flag forward rather than resolve it in hindsight: it stands as a provisional estimate, revisable before settlement if a better reference emerges, not quietly adjusted once prices have moved. One term needed pinning down first. On 2026-07-30 OpenAI cut the price of its cheapest GPT-5.6 tier by 80% while the flagship tier's standard price held — a single generation moving in two directions at once. So 'flagship-class' means each vendor's highest-capability, highest-priced generally available tier at the time of measurement, not its fastest or cheapest; it moves forward as vendors ship new flagships. Settled before it could affect any outcome.

How scoring works

Priors are our probability that the answer settles YES — July-2026 impressions, stated uncertainty ±0.15. We claim no calibration until ≥10 questions have settled.

Settlement is tri-state. YES/NO enter the Brier ledger — (prior − outcome)², lower is better. PENDING scores nothing, is recorded, and postpones settlement one quarter, at most once; still unresolved, it settles as NO. Every NO carries a sub-label — NO-event (it clearly did not happen) or NO-evidence (we could not see enough to call YES). Scored identically, analyzed separately; NO-evidence is never cited with NO-event strength.

The waterline is 0.25 — the Brier score you would earn by flipping a coin on every question. A score above it is worse than guessing. It is the number to hold us to, and it is why no single settlement means much on its own.

Voiding is allowed only for three pre-registered reasons: the criterion proves literally unmeasurable; the entity or market the question attaches to ceases to exist; or verification shows the condition was already met before the freeze date. Nothing else may be voided. Every void is publicly counted and adds a flat 0.25 Brier penalty. No quietly deleting losing bets.

Ambiguity in any criterion resolves against our thesis.

Disputes: every settlement opens a 14-day public dispute window; disputes, evidence, and rulings are archived. A settlement is final when its window closes.

Reporting: cumulative Brier appears in three columns — axis-movers; signals (narrative signals and thermometers); and ship-readings, meaning first-person questions about how this publication itself is run, currently one, B1c. Never one blended total.

Panel discipline: at most 18 active published questions; at most 3 added per quarter. Amendments bind future settlements only — a bet already in settlement is graded under the rules it was frozen with. Two questions have been revised since freeze — B1-3/B1c on 2026-07-25, and A1a on 2026-08-02 — and both are counted above; the revisions and their reasons are recorded in the ledger.

Change log

The questions were frozen on 2026-07-10, before this publication had an audience — which is the point, and the reason this log starts before issue one. One line per version below, each marked whether it revised a bet or only changed how the panel explains itself; the full record of each, including what was not altered, is in the ledger.

2026-08-17v0.4.5display onlyOpening reordered: rules moved below the board, the core question and a scope caveat added above it; A1a's card title corrected to match its already-revised reading.
2026-08-17v0.4.4display onlyThe coin-flip waterline stated, and this change log made readable.
2026-08-04v0.4.3display onlyA plain-language line added to all sixteen questions. Display only.
2026-08-03v0.4.2display onlyThreshold rationale published for all sixteen questions.
2026-08-02v0.4revised a betPanel restructured: a question layer added, thirteen questions regrouped under eight axes, three added. A1a revised.
2026-07-25v0.3revised a betMeasurement contract signed and frozen. B1-3 triggered, and revised.