Skip to content

feat: durable message lifecycle from admission to execution - #3721

Open
Astro-Han wants to merge 18 commits into
apache:mainfrom
Astro-Han:fix/durable-message-lifecycle
Open

feat: durable message lifecycle from admission to execution#3721
Astro-Han wants to merge 18 commits into
apache:mainfrom
Astro-Han:fix/durable-message-lifecycle

Conversation

@Astro-Han

@Astro-Han Astro-Han commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Summary

This is the first end-to-end durable message lifecycle PR built from the latest main, replacing Draft PR #3633. It does not cherry-pick or continue the #3633 patch series.

The problem is simple: once turn.message.submit accepts a message, a Host crash must not make that message disappear, revive it into a queue, or execute it twice.

The solution keeps one durable authority for the message admission and one lifecycle/settlement owner. queued, leased, and in_flight are Host-memory projections rebuilt from durable facts.

The four durable values are facts, not a persisted provider state machine:

  • Accepted: SQLite admission and the canonical transcript row committed atomically.
  • HandedOff: a durable root admission/source proof, or an immutable steering event, proves that Root execution owns the message.
  • Executed: a durable provider-request proof exists downstream of that handoff proof. Steering uses its event timestamp as the lower bound.
  • Cancelled: a retract proof or terminal Stop proof closes the message without downstream execution.
flowchart LR
  S[turn.message.submit] -->|one SQLite admission + transcript transaction| A[Accepted]
  A -->|root admission/source proof| H[HandedOff]
  A -->|retract proof| C[Cancelled]
  H -->|provider request proof downstream of handoff| E[Executed]
  A -->|terminal Stop proof| C
  H -->|terminal Stop proof| C
Loading

First-principles ownership

  • SqliteSessionMetadataStore owns the durable message admission, transcript identity, ordering, edits, reorders, promotion, retract, lifecycle rows, and size boundaries.
  • RootAdmissionOwner owns the durable Root execution contract and source-message proof.
  • HostMessageCoordinator owns only the reconstructible queue projection and the shared proof classifier/settlement owner used by normal terminal cleanup and restart recovery.
  • Root admission is handed off before Runtime activation. Follow-up transcript rows are rebound to the successor turn before activation, so one submitted message keeps one canonical transcript identity.
  • Durable Session capability binding is derived from the durable Root execution contract; it does not depend on an in-memory provider marker.
  • RuntimeKernel no longer owns queue state, leasing, folding, retract, or fallback authority. The old method surface remains only as a no-state compatibility shell while CLI/Desktop projection work stays out of scope.

Crash-cut behavior

  1. Crash before SQLite admission: no accepted message exists, so there is nothing to recover.
  2. Crash after admission but before Root admission: recovery rebuilds an Accepted message into Host memory; it is not executed.
  3. Crash after Root admission but before activation: the durable Root contract is replayed; the root-source proof is handed off before activation.
  4. Crash after provider proof: recovery sees the downstream proof and settles Executed; it does not replay the message.
  5. Crash during terminal cleanup: recovery and normal terminal cleanup call the same settlement owner, so Accepted/HandedOff facts converge to Executed or Cancelled from durable proofs.

Scope

Included: atomic admission and transcript, edit/reorder/promotion/retract, successor handoff, proof-driven settlement, restart recovery, size boundaries, canonical transcript identity, durable capability derivation, and removal of the old RuntimeKernel queue authority.

Excluded: CLI/Desktop projection refactors and Side Conversation UI/hooks.

Verification

Only affected builds and tests were run; the full repository test suite was not run.

  • Exact pushed head: 2ea21a689.
  • npm --workspace @maka/runtime run build passed.
  • npm --workspace @maka/storage run build passed; SQLite metadata suite: 51 passed.
  • npm --workspace @maka/runtime-host run build passed; production SQLite + Runtime Host/UDS message, queue, and coordinator suites: 60 passed.
  • Root/recovery/canonical projection/Goal authority composition suite: 71 passed.
  • Runtime interaction regression suite: 11 passed.
  • npm run astryx:surface-inventory passed locally.
  • npm run format:check passed.
  • CI run 32758898698 passed for exact head 2ea21a689.
  • Fault coverage includes atomic transcript admission, oversized admission rejection before transcript mutation, edit/reorder/promotion/retract, ordered successor handoff, root admission without a Run, killed/graceful Host recovery, steering-event recovery, provider-proof settlement, and whole-session transcript de-duplication.

AI use

This PR was implemented with Codex assistance. The design, repository decisions, code changes, affected-test selection, review of Draft PR #3633 evidence, adversarial review, simplification audit, and final verification were directed and checked against the repository's durable authorities and production composition.

Preserve live Client capability bindings, make cancellation retries idempotent, and keep admission-backed transcripts out of compatibility Run synthesis until their root contract owns them.

Generated-by: Codex
@Astro-Han
Astro-Han marked this pull request as ready for review August 24, 2026 18:21

@M4n5ter M4n5ter left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English

I found two correctness issues on exact head 2ea21a68950502615f5e109ce914a7af94a97140: one makes durable follow-up reordering fail for every real permutation, and the other rejects otherwise valid 32–49 KiB messages as an internal failure. The exact-head hosted test check is green, but its queue test uses an in-memory lifecycle stub for reorder and does not exercise either storage boundary.

简体中文

我在 exact head 2ea21a68950502615f5e109ce914a7af94a97140 上确认了两个正确性问题:持久化 follow-up 队列的任何实际换序都会失败;另一个问题会把本来合法的 32–49 KiB 消息错误地变成内部错误。当前 head 的托管 test 检查是绿色,但队列测试对重排使用了内存 lifecycle stub,没有覆盖这两个真实存储边界。

const current = rows.map((row) => row.message_id);
if (
current.length !== unique.length ||
current.some((messageId, index) => messageId !== unique[index])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English

[P1] Compare identity membership here instead of requiring the requested order to equal the current order. The caller invokes this method only when the order changed, but this predicate rejects every such permutation before the updates run. With two accepted follow-ups persisted as message-1, message-2, calling reorderMessageAdmissions(session, [message-2, message-1]) deterministically throws SessionMetadataConflictError and leaves the old order intact. The coordinator test does not catch this because its lifecycle fake implements reorderMessageAdmissions as a no-op. Please validate equal cardinality plus the same identity set, then write the requested order in this transaction, and add a real SQLite reorder-and-restart test.

简体中文

[P1] 这里应校验 identity 集合,而不是要求请求顺序与当前顺序完全相同。调用方只会在顺序发生变化时进入这个方法,但当前条件会在更新执行前拒绝所有实际 permutation。把两个 accepted follow-up 按 message-1, message-2 持久化后,调用 reorderMessageAdmissions(session, [message-2, message-1]) 会稳定抛出 SessionMetadataConflictError,持久顺序仍保持不变。现有 coordinator 测试没有发现它,因为 lifecycle fake 把 reorderMessageAdmissions 实现成了 no-op。请改为校验数量相同且 identity 集合一致,再在同一事务中写入请求顺序,并增加真实 SQLite 的换序与重启测试。

disposition,
admittedAt: Date.now(),
};
await this.#lifecycle.commitMessageAdmission(messageAdmission);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English

[P2] Preflight the exact pending-admission envelope before this write. The earlier guards check the public queue projection (submitted content) and future root payload (prepared content) separately, while the storage record contains both and has its own 64 KiB cap. Using the real decoders on this head, a plain 32 KiB message passes both earlier guards but normalizePendingMessageAdmission throws Pending message admission exceeds size limit; the mismatch continues through 49 KiB and escapes this handler as internal_failure. Please either align the durable format/limit with the supported envelopes or return a declared capacity error before mutation, and cover the 31/32/49/50 KiB boundaries with the real lifecycle store.

简体中文

[P2] 在这里写入前需要预检完整的 pending-admission envelope。前面的守卫分别检查公共队列 projection(submitted content)和未来 root payload(prepared content),但存储记录同时保存两份内容,并另有 64 KiB 上限。在当前 head 上使用真实 decoder 验证时,普通 32 KiB 消息能通过前两个守卫,却会在 normalizePendingMessageAdmission 抛出 Pending message admission exceeds size limit;这个错配一直持续到 49 KiB,并最终被转换成 internal_failure。请让持久化格式/上限与已支持的 envelope 对齐,或在任何 mutation 前返回明确的容量错误,并使用真实 lifecycle store 覆盖 31/32/49/50 KiB 边界。

@Astro-Han Astro-Han left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this head and found blocking issues that need to be addressed before merge.

[P1] Recovery can replay raw /skill text instead of the prepared skill invocation

Idle submit persists raw modelContent first (message-coordinator.ts:906-918) and only later expands the skill in the Host (root-turn-coordinator.ts:1028-1077). If the Host exits between those steps, recovery replays the persisted raw content without re-running skill preparation, producing a root that the normal path would have wrapped.

[P1] Rejected idle submit leaves a phantom user message in the transcript

The admission and user transcript are written together, but later start/admission can still fail (skill blocked, oversized, binding failure). The cancellation only flips lifecycle state, not the transcript — a failed send remains visible and retries create duplicates.

[P2] Message reorder with identical content is rejected as a conflict

The metadata store compares target order byte-for-byte, so any non-trivial reorder is treated as a conflict. Existing tests pass only because reorders in those fixtures are no-ops.

[P2] Admission envelope can exceed storage limits undetected

Submitted and prepared payloads are checked separately, but the combined admission envelope (64 KiB limit) is not pre-validated. Inputs in the 32–49 KiB range pass early gates yet fail at admission, surfacing as an internal failure.

CI on 2ea21a689505 is test: success. These issues are independent of CI and require fixes before approval. Heads verified at time of review.

简体中文存在恢复路径与 transcript 残留等阻断问题,需修复后重审。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants