Skip to content

Declared before the work, so the result cannot be chosen after it

A measurement firm reporting its own improvement has the same conflict its methodology page exists to prevent, one level up. So the questions, the metric, the target, the window and the analysis are written here and locked before any optimisation of our own properties begins. Once locked, the text cannot change; amendments are appended with their date.

geomg-ai.vercel.app

VOID 2026-09-11 · locked 2026-08-21

Frozen question set
6d084be71d512d4f
Primary metric
answer presence (brand named in the answer)
Secondary metrics
citation presence, share of voice, per-question cited counts
Target that counts as real
+10 percentage points on the primary metric, with the 90% interval on the paired difference excluding zero
Window
2026-08-21 to 2026-09-30
Engines
ChatGPT, Claude, Perplexity · grounding on

ANALYSIS (terms defined in the glossary)

Mean of per-question rates over the frozen set, 90% band by meanRateInterval centred on the published rate (Wilson fallback at zero hits). Before/after tested question by question: the 90% interval on the paired difference must exclude zero. Anything inside the band is reported as too close to call. Failed probes leave the denominator. A change of engine set, grounding, engine profile or analyst starts a new baseline and is never compared across.

The result is published whether or not it moves.

AMENDMENTS

2026-09-11 Amendment 1, 11 September 2026. VOID. This pre-registration is void. It will not be analysed, and no result from it will be published, because there is nothing to analyse. Reason. The window, 21 August to 30 September 2026, contains no measurement of this site. geomg-ai.vercel.app was taken off daily collection on 8 August 2026, thirteen days before this pre-registration was locked, and was never put back on. We locked a before-and-after test on an instrument that was not scheduled to run, and did not notice. Separately, from 29 August 2026 no collection ran for any client, because the API account that grades every answer ran out of credit, so the window would have been empty from 29 August even if this site had been enrolled. The last measurement of this site was a scan on 18 August 2026, before the window opened. There is no after measurement. No effect is claimed, the 10-point target is neither met nor missed, and nothing from this window will be reported as a result. A new test of this site needs a new pre-registration, locked only after daily collection for it is running and has stored answers.

referlabs.com.au

ENDED 2026-08-27 · locked 2026-08-27

Frozen question set
5ed18f790e1857b7
Primary metric
citation presence (own domain among cited sources)
Secondary metrics
answer presence, share of voice, per-question cited counts
Target that counts as real
+5 percentage points on the primary metric, with the 90% interval on the paired difference excluding zero
Window
to be fixed at treatment ship date by dated amendment
Engines
ChatGPT, Claude, Perplexity · grounding on

ANALYSIS (terms defined in the glossary)

Mean of per-question rates over the frozen set, 90% band by meanRateInterval centred on the published rate (Wilson fallback at zero hits). Before/after tested question by question: the 90% interval on the paired difference must exclude zero. Anything inside the band is reported as too close to call. Failed probes leave the denominator. A change of engine set, grounding, engine profile or analyst starts a new baseline and is never compared across.

The result is published whether or not it moves.

AMENDMENTS

2026-08-27 Amendment 1 — 27 August 2026. Recorded before any measurement of this experiment is saved and before any treatment ships. The original design, pre-registered above, paired 21 treated and 21 control pages on word count within hub and proposed a question-paired delta between the two arms. Reading the assignment through the collection to date (720 cells, 15–27 August; npm run arms) shows that test cannot measure the treatment: each frozen question cites at most one arm, so a delta between arms would measure a difference between topics, not between treatments. The same holds for any comparison between two of our own pages — two pages answer different questions — so no ReferLabs page is used as a control for another. Thirty-one of the 42 experimental pages have never been cited in those 720 cells, including all 18 solar-energy pages, because no frozen question covers that hub. The measurement cannot see them. The revised test uses the competitor control committed for referlabs.com.au on 10 August 2026, before this experiment was designed: for each frozen question, the control is every other domain cited in the same answers to that question, drawn by rule, never hand-picked. The test is the difference-in-differences of stats.ts pairedDiD: referlabs.com.au's cited rate on each question, after treatment versus its pre-treatment baseline, minus the competing domains' change on the same question over the same period, paired on the question, with the 90% interval on the difference-in-differences required to exclude zero. Frozen sets 5ed18f790e1857b7 and 93231de6713fb10c; engines, grounding, engine profile and analyst as at baseline, and any change to them starts a new baseline rather than a comparison. The target effect is +5 percentage points in cited rate. It was chosen because it is meaningful against the 20.6% pooled baseline and does not assume headroom we have not measured; +10pp has no basis in the observed distribution and on Perplexity may exceed the headroom available. The citation rate differs sharply by engine at baseline (Perplexity 39%, Claude 12%, ChatGPT 3.5%), so results will be reported per engine as well as pooled, and a pooled figure will not be published without the per-engine breakdown alongside it. Under this test every cited page is measurable, because the control lives on that page's own questions. Results are read per page by arm: a treated page's change against its competitors is the treatment effect for that page; a control page's change against its competitors is the drift the treatment must clear. Pages with zero baseline citations on their questions are reported directionally, never as verdicts. Pages no frozen question can reach are reported as not measurable, and no verdict will be published for them. The measurement window is to be fixed at the treatment ship date, by dated amendment recorded before any treated page changes. Treatment is not yet scheduled; the first saved measurement of set 1 (about 4 September) is a baseline event, not the start of the window.

2026-08-27 Amendment 2 — 27 August 2026. The experiment is ended on 27 August 2026, before any treatment shipped and before any measurement was saved. No treatment effect is claimed and none will be published. Reason. Bing AI Performance data for the last 28 days (about 1,200 citations) shows that the large majority of the site's AI citations come from brand-check queries. Eleven of the fourteen pages serving those queries had been assigned to the treated or control arms and were therefore frozen for the duration of the experiment, including pages with known defects: an answer paragraph opening with a pronoun, and first affiliate links placed beyond word 400. Holding those pages unchanged for ninety days cost more than a test that Amendment 1 had already shown could measure only a small subset of the assigned pages. Retained, unchanged: the arm map as assigned on 26 August 2026 (seed 20260826); the collection windows and every cell collected under them; the competitor control committed for referlabs.com.au on 10 August 2026; and Amendment 1's finding that a question-paired delta cannot measure page-level treatment on domain-level frozen sets. The arm report (npm run arms) continues to read the historical data and is kept for that purpose. From this date every experimental page is released for editing. In the ReferLabs repository every route in src/lib/experiment/assignment.ts is set to excluded, with its original arm and pair preserved in comments so the history stays readable; the control guard remains in place and no longer fires. The pre-registration above and Amendment 1 stand as written.

2026-08-27 Amendment 3 — 27 August 2026. Set 2 is stood down on cost grounds. This is not an end to collection: the experiment itself was already ended by Amendment 2, and this records the second instrument being stood down cleanly rather than left running or left failing. The baseline continues. State at stand-down. Both frozen sets for referlabs.com.au were active and mid-window, each with 6 collection days banked of the 14-day window, first day 2026-08-22, last day 2026-08-27, zero gap days. Set 2 "what people type" (93231de6713fb10c, 28 prompts) is now inactive and is reported as ABANDONED until it is stood back up; the stand-down was forced past the mid-window guard with the 6 banked days in view, which is recorded here because the guard exists to make exactly that trade explicit. Set 1 "baseline" (5ed18f790e1857b7, 10 prompts) remains active and continues to collect daily. The baseline is not permitted to stand down, so a domain publishing a /proof series always has a live instrument behind it. Cost. Measured from the priced cells, collection runs at about $0.070 per cell at list, or about $7.92 per collection day across both sets, roughly $11 per day at invoice and roughly $330 per month. 487 of the 621 cells in the reporting period carry no price on file and are not counted, so the figure is a floor drawn from the two days that are priced, which agree with each other to within a cent per cell. With the experiment ended and no treatment effect to be published, that daily rate is not buying a result. Standing down set 2 removes 84 of the 114 daily cells and leaves the baseline running at about $2.09 per collection day at list, roughly $2.92 at invoice. What is retained. Every cell already collected under both sets, both frozen sets with their hashes, questions and banked days untouched, and the arm map preserved as recorded in Amendment 2. Nothing is edited or deleted, and set 2's questions remain frozen and unmodifiable. Resumption. Windows count collection days rather than calendar days, so on the tool's documented behaviour the 6 banked days on set 2 are not lost and its window resumes accumulating when the set is stood back up. That behaviour has not been tested through a stand-down and resume cycle, and will be verified before any measurement is published from a resumed window. No measurement, and no verdict, will be published for a window that has not reached 14 collection days.

2026-08-27 Amendment 4 — 28 August 2026. Set 2 "what people type" (93231de6713fb10c, 28 prompts) was stood back up at 09:14 AEST on 28 August 2026, after one day inactive (stood down 14:12 AEST, 27 August, Amendment 3). Reason: those questions measure discount-code exposure, which is the revenue mechanism for this property rather than a secondary metric, and the cost trade recorded in Amendment 3 is reversed on that basis. Nothing about the set was edited; its hash, questions and banked cells are as they were. Banked days at resume: 6 collection days on set 2 (22 to 27 August, zero gap days), and 6 on set 1 (same dates). Verification of resume behaviour, as Amendment 3 committed to before publishing from a resumed window: the window-health check reports set 2 as IN PROGRESS with those 6 days and 8 to go, and the measurement tool's own dry run counts 6 days and 444 answers for set 2 from its freeze day, 18 August. The window did not restart; a stood-down day is a day with no cells, and days are counted, not the calendar. This is verified on the counting logic and on one real stand-down and resume cycle; no measurement has yet been saved from a resumed window, and none will be published before the set reaches 14 collection days. Incomplete collection days, all from an API credit lapse on the grading model: 27 August (set 1: 27 of 30 cells; set 2: 66 of 84 before the stand-down), 23 August (set 2: 42 of 84), and 28 August, which had no cells before the restart. The measurement counts any day with stored cells as a collection day and excludes failed cells from every denominator, so these days carry fewer answers rather than being interpolated; the daily digest's stricter "clean day" test (at least 90% of a set's cells) would not count 23 August or 27 August for set 2. Both sets need 8 more collection days; on uninterrupted collection the first measurement of each falls on or about 5 September 2026.

The live series this commits to: /proof. The method: /methodology. The data behind it: /data.