Barely Sentient

Results

The Occupancy Test: every figure its pre-registration committed to reporting, in the order it committed to them.

PendingCollected —Models —n —Sample —

Collection has not begun. Every field on this page was fixed in the pre-registration on , and each one will be filled in whatever it shows.

The finding

What the experiment set out to measure, against the threshold it committed to before it ran.

Against the pre-registered threshold

MeasureEstimateIntervalnInterval on a 0 to 100% axis
Primary: Occupied alonependingpendingpending
Broad: Occupied + Adjacentpendingpendingpending

Verdicts, and what each one means

VerdictDefinitionIdeas
OccupiedA shipping product exists that a reasonable person would describe as the same thing.pending
AdjacentSomething close exists, with meaningful differentiation still available.pending
OpenNothing found within the search budget.“Open” means “not found in 10 minutes”, not “does not exist”.pending

H1

Occupancy exceeds 70%. Most ideas will have at least one shipping competitor.

Test
occupancy.primary
Readout
Occupied alone: pending
Outcome
Pending

The primary test is Occupied alone against the threshold. Occupied plus Adjacent is reported beside it, whatever either shows.

H2

Asking for novelty ("something nobody has tried") does not meaningfully reduce occupancy. Predicted because novelty-to-an-LLM tracks training-data recency, not market vacancy.

Test
occupancy.byPrompt[4] vs byPrompt[1]
Readout
Novelty-seeking: pending · Generic: pending
Outcome
Pending

A null result on H2 means “no large effect”, not “no effect”.

H3

Independent sessions of the same model produce overlapping ideas at a rate far above chance, replicating mode collapse on this specific task.

Test
convergence
Readout
Observed: pending · chance: pending
Outcome
Pending

Compared against the rate expected by chance, using the method disclosed under Convergence.

The breakdown

The same measure cut the two ways the pre-registration named, and the convergence rate beside them.

Primary occupancy by prompt

MeasureEstimateIntervalnInterval on a 0 to 100% axis
1. Genericpendingpendingpending
2. Constrainedpendingpendingpending
3. Domainpendingpendingpending
4. Novelty-seekingpendingpendingpending
5. Anti-convergencependingpendingpending

Primary occupancy by model

MeasureEstimateIntervalnInterval on a 0 to 100% axis
Claudependingpendingpending
GPTpendingpendingpending
Geminipendingpendingpending

Convergence rate against chance

MeasureEstimateIntervalnInterval on a 0 to 100% axis
Observed across independent sessionspendingpendingpending
Expected by chancependingpendingpending
Clustering
Pending
Embedding model
Pending
Similarity threshold
pending

The threshold is set on the pooled corpus before any cell-level result is inspected.

Whether the method held

Checks on the judging itself, and on predictions written before any of this existed.

Judging, re-judging, and the seed

MeasureValuen
Judge disagreement on re-judging, overallpendingpending
— verdicts resolved at Tier 1 (the 3-minute call)pendingpending
— verdicts resolved at Tier 2pendingpending
Tier-1 early-exit ratepending
Mean judging minutes per idea (T)pending
Share of verdicts re-judged20.0%
Randomisation seed20260823

Before collection, three models were each given the frozen pre-registration and asked for fifteen predictions with a confidence. Each is scored after the results are final. Brier score: lower is better, and stating even odds on everything scores a quarter.

Comparable arms

ArmSealBrierHit rate
exp01-claudeNot sealedpendingpending
exp01-gptNot sealedpendingpending
exp01-geminiNot sealedpendingpending

Written inside the working session, with the project’s full context. It exists to ask whether knowing the context makes a model’s predictions better, so it is scored on its own line and never averaged in.

Bonus arm — scored alone

ArmSealBrierHit rate
exp01-claude-projectcontextsealed pendingpending

Check the hashes

What to distrust

Every departure from the frozen plan, the limits of what this can show, and the files to check it all against.

Deviations from the frozen pre-registration

Pre-registration frozen 2026-08-23. Every departure from it — however small, however defensible — gets logged here with a date and a reason, and is published alongside the results.

An empty log is a claim. A populated log is honesty. Neither is a failure.

Date Section Deviation Reason Effect on results
2026-08-23 Design → Models gemini-2.5-pro → gemini-3.1-pro-preview The pinned model was withdrawn before the first call: the API returns HTTP 404, "no longer available to new users". It is still listed by the /models endpoint, so this was only caught by a live smoke test. gemini-3.1-pro-preview is the only Gemini pro tier this key can reach; a flash model would have made the Gemini arm a different weight class from Opus 5 and GPT-5. None on collected data — no data existed. Reproducibility is degraded: the replacement is a preview model and may be withdrawn the same way. Stated in the piece.
2026-08-23 Tooling → MAX_TOKENS 2048 → 16384 The frozen value truncated the arms differentially, which is the exact failure its code comment claimed to prevent. All three models spend output budget on internal reasoning before writing an answer, in different amounts. Measured on a throwaway prompt: Claude stopped at max_tokens after 1198 reasoning tokens; GPT-5 spent all 2048 on reasoning and returned an empty string — every GPT response would have been blank. At 16384 all three stop naturally. None on collected data — no data existed. Had this run unnoticed, the GPT arm would have been empty and the Claude arm cut mid-list, silently breaking the first-3-ideas extraction rule.

What counts as a deviation

  • Any change to prompt wording, model list, session count, or extraction rule
  • Any occupancy verdict assigned outside the tier rules
  • Any search that ran over the 10-minute budget
  • Any change to the analysis after data collection began
  • Judging performed by anyone/anything other than the pre-committed hybrid protocol
  • Use of a randomisation seed other than 20260823

What does not

  • Fixing a typo in a prompt before the first API call (record it here anyway if in doubt)
  • Tooling and script changes that do not alter the data or the protocol
  • “Open” means “not found in 10 minutes”, not “does not exist”.
  • The model found candidates, a human made every call.
  • Claude is one of the three models measured, the assistant that ran the first search for candidates, and the author of the sealed predictions. All three roles are declared.
  • Re-judging is blind to the original verdict and URL, but it is the same judge. That is the honest limit of a solo operator.
  • The per-prompt splits are powered only to detect large differences. A null result on H2 means “no large effect”, not “no effect”.

Published files, with the hash of each

FileSHA-256Bytes
Responses, extracted ideas, verdicts with their URLsPublished with the results
shasum -a 256 <file>

Compare against the hashes above, and against the archived copies of /seals for the predictions. A hash I am still able to edit proves nothing.