The Verified Score
The last post here described the agent bridge and ended on a line that has aged into a to-do list: the game was already deterministic, the bridge just stopped throwing that away.
Thirteen days later that sentence needs an asterisk. The game was deterministic in the sense that mattered for playing it, and not in the sense that matters for proving a score. Between August 2 and August 15 the suite spent almost all of its output closing that gap: finding the places where the simulation quietly read the wrong clock, building a service that can verify a submitted run without trusting the submitter, and shipping a weekly leaderboard that is fed by that service and nothing else.
That is four stories, so this is four short pieces rather than one long one.
Contents
- The divergence hunt — every replay fork was the same bug wearing a different costume.
- A leaderboard you cannot lie to — why the verifier refuses to simulate anything.
- One score, four repositories — following a single run from the browser to the public page.
- Thirteen days — the receipts, and the one number that went backwards on purpose.
1. The divergence hunt
A deterministic simulation makes a promise: same seed, same inputs, same run. CandyRush records a run as a seed plus input streams, and verification means re-simulating that recording and comparing a rolling digest of sim state against the one the original run stamped. If the promise holds, the two streams are identical.
They were not. One capture ran 5,176 ticks and 3,651 of them differed between two re-simulations of the same recording — not against a different machine, against itself.
The striking thing, after a week of chasing them, is that none of these were random-number bugs. Every single one was a clock bug: some piece of state changed on the frame clock inside a simulation that advances on the tick clock. Which tick the change landed on then depended on how fast the machine happened to be running, and two runs of one recording forked with no RNG involved at all.
The costumes it wore:
- A corpse that would not leave. An enemy exited the sim registry when its
dissolve fade finished, and that fade is wall-clock. Measured: enemy
id=101left the registry on tick 1331 in one run and never in another. Death now deregisters on the death tick and the fade is purely cosmetic. That one pair of changes took the capture from 3,651 differing ticks to 2, and to 0 on a later capture. - A camera that lagged one frame. Enemy spawns were placed relative to the
camera transform, which
LateUpdatewrites once per frame. Under a replay pump that advances hundreds of ticks per frame, the transform trails the player by however many ticks the platform batched. The evidence is unusually clean: at tick 724, 25 enemies were bit-identical, the 10 newest spawns were displaced by one constant translation of (+2.25, +2.69) — magnitude 3.498 world units — and the RNG draw counts matched in every scope. Same draws, one rigid offset: the spawn reference point differed. Those displaced enemies then crossed an offscreen-teleport threshold at different moments, which cost one extra draw at tick 773 and forked the digest 50 ticks after the actual divergence. - A float round trip, three lines long. The player's facing was computed
exactly in fixed point, pushed through a
float32conversion, converted back, and compared against a fixed-point threshold. Every other term in that predicate was exact. When the comparison flips, one enemy teleports or does not — which is exactly one RNG draw gained or lost, with no change to the player position and no change to the enemy count. The digest forks while every component the digest reports still matches. That is why the fork read as unattributable for days, and why two earlier attributions were wrong.
That last one is really a story about the instrument, not the bug. The digest carried enemy count and not enemy state, so enemy positions were never verified by any replay or leaderboard check at all. Over the same window the digest gained cumulative per-scope RNG draw counts, then per-entity hashing of position, HP and type, and the sampling cadence went from a 60-tick window to per-tick, so a divergence report names an exact tick instead of a neighbourhood. Two of three committed test fixtures did not move at all when per-entity hashing landed — they contain no entities to fold, which is a blunt measurement of how little those fixtures were exercising.
The worst bug of the set was not a divergence at all. Level-up offers were drawn from the running RNG stream, and a recording stores the card index the player picked. So any upstream divergence changed which three cards were offered, and the replay then applied a different ability and played a different build: all seven level-ups of one capture resolved to abilities the player never took. Offers are now keyed on a stable ordinal rather than a stream position, recordings stamp the chosen ability id, and playback compares it. Seven mismatches went to zero while the run was still diverging elsewhere, which is the useful shape — one instrument, one failure mode, no conflation.
And the one I like most moves nothing today. Each field chunk drew its background from a pool using an RNG minted from the current tick, and the streaming methods that call it run on frame-paced updates. Every shipping field has exactly one background prefab, so the draw is seed-independent and no digest moves. The day a field ships a second backdrop it becomes a live fork — two verifications of one submission disagreeing, arriving in a commit that only added art.
2. A leaderboard you cannot lie to
On August 2 there was no SimVerify. It was bootstrapped on August 8 and is now four .NET projects, roughly 21,000 lines including tests, deployed as a service behind a tunnel with a health endpoint and a build mutex.
The design decision that shapes everything else is a refusal: SimVerify does not simulate. A headless C# re-simulation was considered as the verification authority and rejected, because it would be a third artifact with its own floating-point behaviour — desktop Mono/IL2CPP versus browser WASM — and agreeing with it proves nothing about what the player's browser actually did. The shipped production build is the authority. The service dispatches a headless browser at that artifact, collects the result, and compares digests. It orchestrates; it never re-derives the truth itself.
The same instinct governs the comparator. Service and game must run the identical comparison code, so the comparator source lives in UTools and is consumed here as a git submodule compiled directly into the service. NuGet was rejected explicitly: the game builds from the UTools pin at head, the server would build from a published version, and that skew window is the exact failure the whole design exists to prevent. Linked source across repos pins nothing at all.
On top of that sits a cross-artifact gate: one recording, two artifacts, one commit SHA. It runs the WASM build twice first, and a build that disagrees with itself short-circuits the run so the native leg never launches — that verdict would be unattributable, and an unattributable red is worse than no gate.
Then came the part that assumes a hostile caller, and it changed requirements
rather than adding to them. The submission endpoint used to accept a run based on
a matchesStamped field — a value computed by the artifact under test. The
service's own code already labelled it advisory. A server that trusts it verifies
nothing. Acceptance now runs the shared comparator server-side, and the board
records the score the reconstruction derived, never the one the file claimed;
same for the run duration. A pass that derives no outcome boards nothing at all
rather than boarding a zero.
The identity path had a matching set. Steam session tickets validate against Valve, and the verified SteamID is the only identity a submission has. Four findings there are worth repeating because each is invisible until someone looks:
- The ticket call was missing the
identityparameter Valve documents as required, and pointed at the wrong host — as shipped, no real ticket could ever have validated. - The publisher Web API key rides in the query string, and the default HTTP client logger prints request URIs at Information level, so every submission would have written the key into the service log. Loggers removed; a test now asserts the key appears in no log line at any level, and it fails without the fix.
AuthenticateUserTicketis idempotent, so a captured ticket could board a run repeatedly. Tickets are now single-use, tracked by SHA-256 rather than by storing the credential.- The verdict switch failed open on an unknown verdict.
Four gates now stand in front of the endpoint in the order a hostile request meets them: an upload byte ceiling applied ahead of routing (a cap applied after model binding is a cap on a body already read in full), schema validation of the untrusted recording, admission backpressure that refuses rather than queues, and a per-identity sliding-window rate limit keyed on the verified SteamID — 429 and only 429, never conflated with an auth failure or a verification failure.
The plausibility scoring deserves its own line, because it is the one place the service deliberately declines to enforce. Shadow metrics are computed and stored as an append-only corpus, and the strictest mode routes a suspicious run to the Assisted division rather than rejecting it. An uncalibrated threshold that can discard a run will false-positive the best players first.
3. One score, four repositories
Here is the whole path a score now takes.
A player runs the game. CandyRush records the seed and the input, level-up
and cheat streams, and stamps a digest sample every N ticks. UTools supplies
the fixed-point math, the digest hash format, and the comparator those digests
are compared with. On submission, SimVerify validates a Steam ticket, drives
a headless browser at the shipped artifact, re-derives the digest stream and the
score, and writes the derived score to a weekly board. It publishes that board as
static JSON. PublicWebsite fetches the JSON client-side and renders it at
/candyrush/leaderboards, which means scores refresh on every visit with no site
rebuild, and the public site never talks to the verification service at all.
Four repositories, four internally-correct pieces, and — predictably — every interesting bug living in the seams between them.
The sharpest one is a ruling that outgrew its schema. A board is identified by
(week, commit SHA) and archives rather than wipes, so a week can own more
than one board when the game ships a new build mid-week. The published schema was
previousWeeks: string[], keyed by week, which simply cannot express that: the
site's archive picker collapsed two boards into one row and left the second
unreachable. The fix is a purely additive schema v2 — new board id, commit SHA
and a full board listing, with the old field keeping its old meaning so a v1
renderer keeps working — and a site that prefers the new listing when it yields
anything and falls back to the old one when it does not.
The rest of the seam bugs are all of one family: an assumption that was true on one side of the boundary and false on the other.
- The archive picker was derived from the currently loaded board. Archived board files are immutable snapshots, so their listing is frozen at publish time and cannot name anything published later. Opening an archive therefore shrank the picker and stranded the visitor. The listing is now a union that only grows.
- Every archived board wore the "may be out of date" badge, because an archive is older than the 24-hour staleness threshold by definition. Staleness is now a property of the live board only.
- The publisher names board files with an 8-character SHA. The site rendered seven. Showing a different abbreviation of the same SHA is worse than showing none: two renderings of one build read as two builds, which is precisely the confusion that keying a board by commit SHA exists to remove.
- Steam IDs stay strings all the way through the web layer. 64-bit ids exceed what JavaScript numbers hold exactly, and a leaderboard that silently rounds identities is a leaderboard that merges players.
Board ids arrive from the network and are interpolated into a fetch URL, so they are validated before any request — traversal, protocol-relative and absolute-URL ids are all refused. That is the other thing the seams teach: the moment a value crosses a repository boundary it is untrusted input, even when you wrote both sides.
4. Thirteen days
Everything above, as receipts. Counts are merge commits landed on each trunk between August 2 and August 15.
| Repository | PRs merged | Commits | What it was doing |
|---|---|---|---|
| CandyRush | 67 | 221 | Determinism fixes, digest instrumentation, weekly-challenge boot, submission client |
| UTools | 23 | 75 | Fixed-point math, digest hash v4→v6, the shared comparator, RNG draw counters |
| SimVerify | 17 | 44 | The entire service — repository created August 8 |
| PublicWebsite | 2 | 6 | The leaderboards page and its archive |
Two honest notes on those numbers. A pull request here is a unit of review, not
of effort — several of them are audit deltas against an already-merged PR, a
pattern that recurs often enough to be a working method rather than an accident.
And of the 139 non-merge CandyRush commits in the window, 118 carry a
Co-Authored-By: Claude trailer.
The characteristic output of these thirteen days is not a feature. It is instruments: a digest that went from counting enemies to hashing them, per-scope RNG draw counters folded into that hash, per-tick capture windows, uncapped state dumps at a named tick, an in-editor verify harness, a cross-artifact gate, and a service whose entire job is to disbelieve a submitted file.
Which brings up the one number that deliberately went backwards. The simulation version stamped into every recording climbed 3 → 4 → 5 → 6 in about a week, because a single carve-out sentence in a design document read as permission to bump it whenever digest semantics changed. On August 10 it was reset to 1 and frozen, and the sentence that drove four bumps was deleted. Nothing about the simulation regressed; the version simply stopped being a diary and went back to being a compatibility gate that will not move again until launch.
The pattern across all four pieces is the same, and it is not really about games. Every failure in this window was a boundary failure — between the frame clock and the tick clock, between a build and its own re-run, between two artifacts of one commit, between a publisher's schema and a reader's assumption, between a client's claim and a server's verdict. Each side was correct. The seam was not.
A score is now something the suite can prove rather than something a player reports. Getting there mostly meant finding the places that were quietly measuring in the wrong units, and none of them looked wrong from the inside.