Skip to content

fix(build): restore @swc/helpers in the standalone bundle, health-check the deploy - #243

Merged
Musiker15 merged 1 commit into
mainfrom
fix/standalone-swc-helpers
Aug 17, 2026
Merged

fix(build): restore @swc/helpers in the standalone bundle, health-check the deploy#243
Musiker15 merged 1 commit into
mainfrom
fix/standalone-swc-helpers

Conversation

@Musiker15

Copy link
Copy Markdown
Member

Fixes the production outage that followed the next 16.2.12 -> 16.3.1 bump in #240.

What broke

next 16.3.1 traces only the cjs/ half of @swc/helpers into output: standalone, but its own require-hook loads the ESM entry point:

Cannot find module .../@swc/helpers/esm/_interop_require_default.js

It surfaces as an unhandledRejection during startup. The process stays alive and never binds the port, so PM2 reported online with a climbing restart counter while curl 127.0.0.1:3008 got no connection at all and every public request returned 503.

Reproduced locally on a clean build of the same commit, so this is the bundle, not the server:

source node_modules standalone bundle
@swc/helpers/cjs present present
@swc/helpers/esm present missing

The fix

copy-standalone-assets.mjs already exists for exactly this, it puts sharp and the @img closure back after Next declines to trace them. @swc/helpers joins that list.

It copies the whole package per pnpm store entry rather than just the missing directory, so a future change to which half gets traced cannot reintroduce this quietly. Version-matched per entry: during recovery we grafted the esm/ of 0.5.15 onto 0.5.23 and it worked, which is a good reason not to rely on that.

Verified on a clean rebuild:

restored @swc+helpers@0.5.23 -> .../standalone/node_modules/.pnpm/@swc+helpers@0.5.23/node_modules/@swc/helpers
$ find apps/web/.next/standalone -path "*@swc/helpers/esm/_interop_require_default.js"
apps/web/.next/standalone/node_modules/.pnpm/@swc+helpers@0.5.23/node_modules/@swc/helpers/esm/_interop_require_default.js

Why nothing noticed

The deploy ran green four times through the outage. It asked PM2, and PM2 answers about the process, not about whether anything is listening. pm2 reload returning ✓ is compatible with a bundle that cannot serve a single request.

So the workflow now curls 127.0.0.1:3008 after the reload, retries for a minute, and on failure dumps the last 40 error-log lines before exiting non-zero. That check alone would have caught this at the first deploy instead of 40 minutes later.

Not changed: concurrency with cancel-in-progress: false and the rm -rf apps/web/.next before the build were already in place, deploys never overlapped and the tree was not stale.

Follow-up worth considering

The health check reports the failure but does not undo it, pm2 reload has already swapped the process by then. An automatic rollback to the previous build would turn this from "loud failure" into "no outage", at the cost of keeping a previous build around on the server. Happy to do that separately if you want it.

…ck the deploy

next 16.3.1 traces only the cjs half of @swc/helpers into output: standalone,
while its own require-hook loads the ESM entry point. The bundle starts and
then dies on the first request needing a helper:

  Cannot find module .../@swc/helpers/esm/_interop_require_default.js

raised as an unhandledRejection. The process stays alive without ever binding
the port, so PM2 reports it online while every request 503s. That took
forms.msk-scripts.de down after the 16.2.12 -> 16.3.1 bump in #240.

copy-standalone-assets.mjs already exists for this exact class of problem, it
puts sharp and the @img closure back after Next declines to trace them. Add
@swc/helpers to it, copying the whole package per store entry rather than the
one missing directory, so a change to which half gets traced cannot bring this
back. Version-matched: grafting the esm of one version onto another happens to
work but is not something to ship.

The deploy also gets a health check. It ran green through this outage because
it only asked PM2, which answers about the process, not about whether anything
is listening. Curl the port instead, retry for a minute, and dump the error log
before failing. A deploy that cannot serve should not report success.

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 2 advisory finding(s) below merit a look before merge.


Graphify review — findings

Adds a copySwcHelpers() step to copy-standalone-assets.mjs that restores the full @swc/helpers package (per version-matched pnpm store entry) into the standalone bundle, since Next 16.3.1 only traces the cjs/ half while its require-hook loads the ESM entry. Adds a port-level health check to the deploy workflow that polls http://127.0.0.1:3008/ after pm2 save, dumps error logs and fails the job if it never returns 200.

Worth a look

  • Broken PM2 process is saved before the new health check can reject it.github/workflows/deploy.yml:60 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review Execution auto-disposal is off for this run; enable it (with sandbox isolation) to have Graphify try to confirm or refute this automatically.
  • Deploy success now requires root URL to return exactly HTTP 200.github/workflows/deploy.yml:70 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review Execution auto-disposal is off for this run; enable it (with sandbox isolation) to have Graphify try to confirm or refute this automatically.
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 9 functions depend on the 9 functions this change touches.

Health — grade A; no new coupling hotspots.

Verification — 9 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 9 function(s) in the blast radius were not formally verified this run

@Musiker15
Musiker15 merged commit ed6cf9e into main Aug 17, 2026
6 checks passed
@Musiker15
Musiker15 deleted the fix/standalone-swc-helpers branch August 17, 2026 21:46
Musiker15 added a commit that referenced this pull request Aug 18, 2026
…t serve (#246)

The health check added in #243 turned a silent outage into a loud one, but it
still leaves the site down: by the time it runs, pm2 reload has already swapped
the process. On 2026-08-17 that would have meant a red workflow and a 503 until
someone read it.

So keep the previous bundle instead of deleting it. The build moves
apps/web/.next aside rather than removing it, and if the port never answers,
the deploy puts the old release back: reset to the previous SHA, reinstall
against that lockfile, regenerate the client, restore the bundle, reload. The
job still exits non-zero, and the error says plainly which commit is live.

What this deliberately does not undo is the migration. `prisma migrate deploy`
has already run and migrations are forward-only. That is fine for additive
changes, since older code ignores columns it does not know about, but a
migration that drops or renames something still needs a human. The comment in
the workflow says so.

pm2 save moved behind the health check as well. Persisting the process list
before knowing whether the release works only makes a bad state stickier.

Verified by running the script against stubbed git/pnpm/pm2/curl in a sandbox,
since the failure path is the one that must not be discovered in production:
a healthy deploy keeps the new bundle and exits 0; an unhealthy one with a
recoverable predecessor restores the old bundle, reports that the pushed commit
is not live and exits 1; and when the rollback cannot come up either, it says
that instead of claiming success.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant