Results
The Occupancy Test: every figure its pre-registration committed to reporting, in the order it committed to them.
Collection has not begun. Every field on this page was fixed in the pre-registration on , and each one will be filled in whatever it shows.
The finding
What the experiment set out to measure, against the threshold it committed to before it ran.
Occupancy
Against the pre-registered threshold
| Measure | Estimate | Interval | n | Interval on a 0 to 100% axis |
|---|---|---|---|---|
| Primary: Occupied alone | pending | pending | pending | |
| Broad: Occupied + Adjacent | pending | pending | pending |
Verdicts, and what each one means
| Verdict | Definition | Ideas |
|---|---|---|
| Occupied | A shipping product exists that a reasonable person would describe as the same thing. | pending |
| Adjacent | Something close exists, with meaningful differentiation still available. | pending |
| Open | Nothing found within the search budget.“Open” means “not found in 10 minutes”, not “does not exist”. | pending |
The three hypotheses
H1
Occupancy exceeds 70%. Most ideas will have at least one shipping competitor.
- Test
occupancy.primary- Readout
- Occupied alone: pending
- Outcome
- Pending
The primary test is Occupied alone against the threshold. Occupied plus Adjacent is reported beside it, whatever either shows.
H2
Asking for novelty ("something nobody has tried") does not meaningfully reduce occupancy. Predicted because novelty-to-an-LLM tracks training-data recency, not market vacancy.
- Test
occupancy.byPrompt[4] vs byPrompt[1]- Readout
- Novelty-seeking: pending · Generic: pending
- Outcome
- Pending
A null result on H2 means “no large effect”, not “no effect”.
H3
Independent sessions of the same model produce overlapping ideas at a rate far above chance, replicating mode collapse on this specific task.
- Test
convergence- Readout
- Observed: pending · chance: pending
- Outcome
- Pending
Compared against the rate expected by chance, using the method disclosed under Convergence.
The breakdown
The same measure cut the two ways the pre-registration named, and the convergence rate beside them.
Occupancy by prompt
Primary occupancy by prompt
| Measure | Estimate | Interval | n | Interval on a 0 to 100% axis |
|---|---|---|---|---|
| 1. Generic | pending | pending | pending | |
| 2. Constrained | pending | pending | pending | |
| 3. Domain | pending | pending | pending | |
| 4. Novelty-seeking | pending | pending | pending | |
| 5. Anti-convergence | pending | pending | pending |
Occupancy by model
Primary occupancy by model
| Measure | Estimate | Interval | n | Interval on a 0 to 100% axis |
|---|---|---|---|---|
| Claude | pending | pending | pending | |
| GPT | pending | pending | pending | |
| Gemini | pending | pending | pending |
Convergence
Convergence rate against chance
| Measure | Estimate | Interval | n | Interval on a 0 to 100% axis |
|---|---|---|---|---|
| Observed across independent sessions | pending | pending | pending | |
| Expected by chance | pending | pending | pending |
- Clustering
- Pending
- Embedding model
- Pending
- Similarity threshold
- pending
The threshold is set on the pooled corpus before any cell-level result is inspected.
Whether the method held
Checks on the judging itself, and on predictions written before any of this existed.
Method integrity
Judging, re-judging, and the seed
| Measure | Value | n |
|---|---|---|
| Judge disagreement on re-judging, overall | pending | pending |
| — verdicts resolved at Tier 1 (the 3-minute call) | pending | pending |
| — verdicts resolved at Tier 2 | pending | pending |
| Tier-1 early-exit rate | pending | |
| Mean judging minutes per idea (T) | pending | |
| Share of verdicts re-judged | 20.0% | |
| Randomisation seed | 20260823 |
Prediction scorecard
Before collection, three models were each given the frozen pre-registration and asked for fifteen predictions with a confidence. Each is scored after the results are final. Brier score: lower is better, and stating even odds on everything scores a quarter.
Comparable arms
| Arm | Seal | Brier | Hit rate |
|---|---|---|---|
| exp01-claude | Not sealed | pending | pending |
| exp01-gpt | Not sealed | pending | pending |
| exp01-gemini | Not sealed | pending | pending |
Written inside the working session, with the project’s full context. It exists to ask whether knowing the context makes a model’s predictions better, so it is scored on its own line and never averaged in.
Bonus arm — scored alone
| Arm | Seal | Brier | Hit rate |
|---|---|---|---|
| exp01-claude-projectcontext | sealed | pending | pending |
What to distrust
Every departure from the frozen plan, the limits of what this can show, and the files to check it all against.
Deviations
Deviations from the frozen pre-registration
Pre-registration frozen 2026-08-23. Every departure from it — however small, however defensible — gets logged here with a date and a reason, and is published alongside the results.
An empty log is a claim. A populated log is honesty. Neither is a failure.
| Date | Section | Deviation | Reason | Effect on results |
|---|---|---|---|---|
| 2026-08-23 | Design → Models | gemini-2.5-pro → gemini-3.1-pro-preview |
The pinned model was withdrawn before the first call: the API returns HTTP 404, "no longer available to new users". It is still listed by the /models endpoint, so this was only caught by a live smoke test. gemini-3.1-pro-preview is the only Gemini pro tier this key can reach; a flash model would have made the Gemini arm a different weight class from Opus 5 and GPT-5. |
None on collected data — no data existed. Reproducibility is degraded: the replacement is a preview model and may be withdrawn the same way. Stated in the piece. |
| 2026-08-23 | Tooling → MAX_TOKENS |
2048 → 16384 | The frozen value truncated the arms differentially, which is the exact failure its code comment claimed to prevent. All three models spend output budget on internal reasoning before writing an answer, in different amounts. Measured on a throwaway prompt: Claude stopped at max_tokens after 1198 reasoning tokens; GPT-5 spent all 2048 on reasoning and returned an empty string — every GPT response would have been blank. At 16384 all three stop naturally. |
None on collected data — no data existed. Had this run unnoticed, the GPT arm would have been empty and the Claude arm cut mid-list, silently breaking the first-3-ideas extraction rule. |
What counts as a deviation
- Any change to prompt wording, model list, session count, or extraction rule
- Any occupancy verdict assigned outside the tier rules
- Any search that ran over the 10-minute budget
- Any change to the analysis after data collection began
- Judging performed by anyone/anything other than the pre-committed hybrid protocol
- Use of a randomisation seed other than
20260823
What does not
- Fixing a typo in a prompt before the first API call (record it here anyway if in doubt)
- Tooling and script changes that do not alter the data or the protocol
Limitations
- “Open” means “not found in 10 minutes”, not “does not exist”.
- The model found candidates, a human made every call.
- Claude is one of the three models measured, the assistant that ran the first search for candidates, and the author of the sealed predictions. All three roles are declared.
- Re-judging is blind to the original verdict and URL, but it is the same judge. That is the honest limit of a solo operator.
- The per-prompt splits are powered only to detect large differences. A null result on H2 means “no large effect”, not “no effect”.
Raw data and verification
Published files, with the hash of each
| File | SHA-256 | Bytes |
|---|---|---|
| Responses, extracted ideas, verdicts with their URLs | Published with the results |
shasum -a 256 <file>Compare against the hashes above, and against the archived copies of /seals for the predictions. A hash I am still able to edit proves nothing.