The Fan-Out That Lost
Our Unity build pipeline reached 3,744 lines in a single ci.yml. One file ran
the test suites, four platform builds, artifact publishing, cross-artifact
determinism verification, a performance-regression gate, Discord announcements,
and three Steam depot lanes. This is the story of taking it apart — including
the parallelization that made things slower, and why we ended up with four
separate workflow linters.
What a monolith actually costs
The file grew the way these always do. 1,039 lines in early August. 2,082 by the 19th. 3,547 on the 22nd, then 3,744 at its peak the same day.
Size was not the real problem. Duplication was:
- rclone provisioning — ten verbatim copies
- the Unity mutex + launch + result-check block — seven copies
- the R2 environment block — fourteen copies
And they had drifted. Some rclone copies fell back to the tool cache silently, others logged the install, and the macOS ones carried an architecture switch the Windows ones did not. Of the seven Unity build blocks, only some checked that a player was actually produced — which matters, because Unity exits zero having built nothing more often than anyone would like.
The forcing function was not tidiness. It was the second game. The goal, written
down: stand up the next project by copying .github/ and setting a handful of
repository variables, with no YAML edits. Every hardcoded Steam app ID, depot
number, builder namespace and bucket prefix was an obstacle to that.
Wave one: the fan-out that lost
The suite ran as one job — EditMode, then PlayMode, then builds — 72 minutes wall clock. The obvious move was to fan out into three parallel jobs.
The arithmetic looked sound. The per-job preamble (workspace scrub, vendor pack seed, package manifest, editor resolution) measured 21 seconds. Paying it three times costs about 42 extra seconds to take roughly ten minutes off the critical path.
Then we measured it:
| Design | EditMode | PlayMode | Builds | Wall clock |
|---|---|---|---|---|
| Monolithic (one job) | 10m14s | 15m31s | ~46m | 72m09s |
| Three-way split | 24m25s | 18m27s | 69m23s | 69m |
| Two-way (suites ∥ builds) | 46m03s combined | ~55m |
The three-way split bought about 4%. Dropping back to two slots bought 25%.
EditMode — the CPU-bound suite — took the worst of it, running 2.4× slower under the three-way split than it had inside the monolith.
The reason is embarrassing in hindsight and worth stating plainly: all three runner slots are instances on one physical box. The third concurrent Unity editor added no capacity. It divided the same machine further.
The preamble arithmetic that justified the three-way split was correct and irrelevant. The question was never the cost of splitting. It was whether the machine had headroom to split into.
Collapsing back deleted a surprising amount of machinery that existed only to
serve the split: a whole report job (both result XMLs are back in one
workspace, so combining them is a step again), the artifact hand-off of those
XMLs, and two per-suite summarize steps. The measurement now lives in the
workflow's own header comment, as a standing prohibition for the next person who
has the idea.
The one thing that survived the collapse was the composite action extracted to serve it — the duplicated preamble stayed gone.
Wave two: decomposition
Three days later came the real restructure:
.github/actionlint.yaml | 27 +
.github/workflows/_tests.yml | 419 +++++
.github/workflows/cd.yml | 2532 ++++++++++++++++++++++++++++++
.github/workflows/ci.yml | 3530 ++++--------------------------------------
4 files changed, 3319 insertions(+), 3189 deletions(-)
ci.yml went from 3,273 lines to 425.
Split by trigger, not by a chain
The instinct is to chain CI into CD with workflow_run. Do not. That trigger
runs the default branch's copy of the workflow and loses the triggering ref,
which breaks every dispatch-from-a-feature-branch.
Instead the split is by trigger. ci.yml runs on pull requests; cd.yml runs
on pushes to dev/qa/main and on manual dispatch. Exactly one fires per
event. Both call the same reusable _tests.yml, so a push still runs the suite
exactly once, first, and the deploy lanes gate on its outputs.
Three lanes, one implementation
git dev -> Steam "development" automatic
git qa -> Steam "qualityassurance" automatic - the merge into qa IS the gate
git main -> default branch uploads only; promotion is a human
The Steam branch names cannot match the git names: Steamworks enforces 4–32
characters, so dev and qa are both rejected at creation. We discovered that
after wiring a map that could never have worked.
Why the lanes could not simply be one job: needs: in Actions is static. One
job cannot wait on "whichever pair ran." A single deploy job would have needed
all five upstreams plus branching logic to work out which lane it was in — the
shape that looks fine and then fails unreadably. So the 452-line deploy job
became a 694-line reusable workflow with four thin callers that differ only in
parameters.
The policy differences are inputs, not code branches. The dev lane is
deliberately not gated on the test suite — a red suite must not cost you the
artifacts. The prod lane is gated, because a bad dev build costs a re-deploy
while a bad Steam build on qa or main is permanent.
Never always()
One sweep removed always() from the job graph in five places. The reason is
specific: always() returns true even when the run is cancelled, and the CD
concurrency group cancels in progress. So: dev push A goes green through both
builds, push B supersedes and cancels run A, and run A's deploy fires anyway —
shipping a superseded commit to Steam and setting it live. Steam builds cannot
be deleted. The replacement is !cancelled().
Five composite actions
Each exists because duplication had already drifted:
unity-workspace(272 lines) — everything between checkout and invoking the editor. Deliberately does not check out the repo, because a local action is loaded from the workspace.unity/build-player(144 lines) — one Unity batch build under the host-global lock, reporting an honest result: a nonzero exit is red, and a zero exit that produced no player is equally red. Always exits 0 itself and reports via an output, so a broken Linux build does not cost the Windows zip.rclone/setup(134 lines) — resolves rclone into the tool cache once per runner. Deliberately does not take R2 credentials as inputs: that would widen them from one step to the whole job to save six lines.git-bash(76 lines) — on the self-hosted Windows runner,bashon PATH is the WSL launcher, which mangles the Windows script path.unsparse-workspace(51 lines) — see below.
Two constraints shaped all of them: composite actions run inside the caller's
job and workspace (reusable workflows are whole jobs on their own runners), and
composite actions cannot read the secrets context. That second one bit us
twice.
Four linters, four shipped defects
A doc-string took down every Unity job
First run after extracting unity-workspace: all three Unity jobs dead inside a
minute with Unrecognized named-value: 'secrets'. Line 19 of action.yml — the
input description, reading "the caller must pass ${{ secrets.UTOOLS_PAT }}".
GitHub evaluates template expressions in description fields, and an action has
no secrets context.
Then the same class recurred, in a shell comment
Three days later an explanatory comment inside a run: block read "the caller
passes ${{ vars.VENDOR_PACK_HOST_ROOT }}". A shell comment is not a comment to
the runner — the expression validates when the action loads. It took dev's CD
down on its first run after merge.
Why nothing caught it: actionlint does not lint action.yml files at all; our
syntax checker substitutes every ${{ }} for a bare token before parsing, so
the expression was already gone; and our workflow linter only walked
doc['jobs']. Three gates, three different blind spots, same defect.
The exit code that printed PASSED and then failed
GitHub appends exit $LASTEXITCODE to every shell: powershell step.
PowerShell cmdlets do not reset $LASTEXITCODE. So a step that runs a native
command, decides its failure is non-fatal, and then finishes with cmdlets will
fail anyway — carrying the exit code of the very command it just forgave.
From a real run log:
tar entry for the player: -rwxr-xr-x 0 root root 4672 linux/CandyRush
Linux executable bit PASSED (0755 in the ustar stream).
##[error]Process completed with exit code 1.
Root cause: tar.exe ... | Where-Object {...} | Select-Object -First 1
terminates the pipeline the moment it has its match, breaking the pipe into
tar.exe and leaving a nonzero $LASTEXITCODE. The step inherited the exit code
of a native command it had deliberately cut off — after its own logic had
already succeeded and said so.
Worth being precise, because it is counterintuitive: this is a false red with
a "PASSED" line in the log, not a false green. The mirror-image hazard also
shipped — a warning-only upload that would have killed the whole builds job, one
of them after writing ok=true, so the verdict and the job result disagreed.
This class had shipped four times. Each was found by a human reading a diff, which is not a process. It is now a linter rule.
Parsing every run: block as real code
A mechanical sweep replaced a hardcoded list with a variable-driven loop and left
the old block's trailing } behind. Valid YAML. actionlint: zero. Our workflow
linter: zero. A guaranteed parse error on the runner, minutes into a queued job.
So check-step-syntax.py extracts every run: block, substitutes expressions
for bare tokens (what the shell is actually handed), and feeds it to the real
parser for its language — PowerShell's own AST parser, and bash -n.
Two details worth stealing. First, it infers the shell rather than trusting
annotation, because shell: bash is not a no-op: the implicit default is
bash -e while an explicit shell: bash is bash --noprofile --norc -eo pipefail. Writing it out would newly enable pipefail on existing piped steps
and could turn a passing job red. A syntax gate must not change runtime
behaviour to install itself.
Second, the first version spawned one interpreter per step — 77 PowerShell startups, about 9 seconds of fork overhead to do perhaps 200ms of parsing. Batching into a single invocation took it from 9.7s to 0.95s.
Tracing the job graph
The fourth gate evaluates every job's if: against eleven real trigger shapes
and fails if any job is unreachable under all of them. Twenty jobs times eleven
shapes is 220 answers; a condition change silently moves some of them and none of
it is visible in a diff.
It immediately found a builds job that ran, executed the entire preamble, and
then skipped every build step — on the one dispatch that ships a permanent Steam
build. Cause: four stacked negations, each added when a new dispatch lever
appeared, which nobody could evaluate by reading.
The gates that passed while verifying nothing
The best lesson came from a failure in the linters themselves.
A run failed with could not read ".github/workflows/*.yml" — the literal glob
reached actionlint because bash had nothing to expand it to. The workspace had
been pruned: actions/checkout sets sparse-checkout and never turns it back
off, and our utility runner's workspace is shared and persistent. One job had
been doing sparse checkouts for years — correctly, while it ran on an ephemeral
hosted runner. Moving it to a self-hosted box changed the premise without
changing the checkout.
But the second half matters more. In that same pruned run, two of our three home-grown gates would have found zero files, looped zero times, and reported success having verified nothing. Only the third had an input guard, and that is the sole reason the failure surfaced as a red step rather than three silently-green ones.
All of them now fail when their inputs are absent, negative-tested from an empty directory. And because the fix could not heal a workspace already poisoned, there is now a small action that turns sparse mode off and then asserts the tree is whole.
What generalizes
Measure the machine, not the arithmetic. The three-way fan-out was justified with correct math about a preamble and wrong assumptions about hardware. Three runner slots on one box are not three machines.
A gate that depends on remembering is not a gate. Four separate defects were each caught by a human reading a diff. That works until it doesn't. Every one is now a check that exits nonzero.
A check that finds nothing must prove it looked. A linter that silently verifies an empty set is worse than no linter, because it produces a green tick that means nothing.
Duplication's cost is drift, not lines. Ten copies of rclone provisioning were not a problem because they were long. They were a problem because only some of them were still correct — and in the Unity build block, only some checked whether a build had actually produced anything.
The pipeline is now 5,588 lines of workflow across six files instead of one 3,744-line file, plus 677 lines of composite actions and 2,741 lines of scripts. That is not smaller. It is separable, parameterized, and gated — and the next game should be able to use it by setting repository variables rather than editing YAML. That was the point.