Skip to content

fix: close the kernel of a tab whose close beacon beats its websocket - #1247

Draft
maartenbreddels wants to merge 9 commits into
masterfrom
fix/page-close-unknown-page
Draft

maartenbreddels wants to merge 9 commits into
masterfrom
fix/page-close-unknown-page

Conversation

@maartenbreddels

@maartenbreddels maartenbreddels commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

A close beacon for a page that the kernel does not know no longer fails, and a tab that closes before its websocket connects no longer keeps its kernel alive until the cull.

Problem

When a browser tab closes, Solara sends a close beacon to /_solara/api/close/<kernel id>?session_id=<page id>. KernelContext.page_close looked up the page with self.page_status[page_id]. When the kernel had not registered that page, the lookup raised a KeyError, and the route returned a 500. Error trackers such as Sentry then logged an exception for each such beacon.

A kernel can miss the page in two ways:

  • An app that runs several server processes can deliver the beacon to a process that holds a stale copy of the kernel. That copy never saw the page.
  • The tab closes while it loads. The browser's connectKernel returns before the websocket opens, so the beacon can arrive after the server created the kernel but before it called page_connect. Before this change, the late page_connect then marked the closed tab as connected. The kernel lived until the cull, which is 24 hours by default.

Change

  • page_close marks an unknown page as closed and returns. It does not close the kernel or bump the cull, because the kernel may still be initializing.
  • page_connect refuses a closed page with the new PageClosedError, a subclass of RuntimeError.
  • app_loop catches PageClosedError once the kernel has initialized. It removes the refused websocket from the kernel. It then calls the new close_if_no_live_pages, which closes the kernel only when no page is connected or disconnected. The close reason is closed-before-connect, not page-close, so the persisted state stays until its TTL.
  • A kernel is visible to other websockets while its state takeover waits on the backend. A refused reconnect of the same tab can therefore close it. _restore_on_connect now attaches persistence and starts the flush worker under the context's teardown lock, and skips both once close() has begun. Initialization then skips the stream wiring, and app_loop drops the websocket, so the client reconnects to a new kernel.

Validation

I ran tests/unit/lifecycle_test.py and tests/unit/state_server_test.py after each commit. The last run gave 47 passed. Each new test failed on the commit before its fix:

  • test_kernel_lifecycle_close_beacon_for_unknown_page (two cases): failed on master with the production KeyError. It checks that the page is marked closed, the cull stays the same, the kernel stays open, and a late page_connect is refused.
  • test_app_loop_close_beacon_before_connect: failed with an unhandled PageClosedError before app_loop caught it. It checks that the kernel closes with reason closed-before-connect.
  • test_app_loop_close_beacon_before_connect_with_live_page: failed because the refused websocket stayed on a kernel that another page uses.
  • test_close_during_takeover_does_not_attach: failed with an AttributeError in _wire_kernel_streams. It checks that no persistence manager or flush worker attaches to a kernel that closed during its takeover.

Gaps

  • A beacon that arrives before initialize_virtual_kernel creates the kernel still finds nothing. That kernel then lives until the cull, as on master.
  • page_status keeps one closed entry per unknown page id until the kernel closes. Without persistence the close route needs no cookie, so a client that knows a kernel id can add entries.
  • Some windows need two pages on one kernel id, which only a shared ?kernelid= URL gives. master has the same windows for known pages:
    • close_if_no_live_pages checks for live pages, releases the lock, and then closes. A page that connects in between loses its kernel and reconnects to a new one.
    • A replacement context after a supersede does not keep the closed mark.
    • A refused websocket on the reuse branch still moves the kernel's event loop and resets the flush worker's epoch.
  • No test covers the order where close() waits for the attach under the teardown lock and then stops the new worker. Reviewers checked that order by reading the code.
  • I did not test any of this in a deployment with several server processes.

Align results

Caution

/align was not run on this change: the maintainer chose the design directly when the review found the race.

Crossreview results

The crossreview ran 7 rounds with astra, opus, and glm. All three approved commit 72d49f5. No CRITICAL or HIGH finding stands after the last round.

  • Rounds 1 to 3 approved the first design, which ignored an unknown page. The maintainer then found the race that this design left open: a beacon that beats page_connect let the late connect revive the closed tab. The defect got in because my brief marked "an unknown page changes no kernel state" as settled, although I had made that decision alone. So the reviewers did not question it. Opus saw the window in round 1, but called it "not a regression", and I accepted that.
  • Round 4 rejected the second design, which closed the kernel inside page_close. All three reviewers found that a beacon during the state takeover closed a kernel that was still initializing, and a flush worker leaked. The defect got in because I did not check that a new kernel is visible to other code before it finishes initializing.
  • Rounds 5 to 7 reviewed the design the maintainer chose. Astra showed that the takeover leak was still reachable with one tab, through a reconnect. Commit 72d49f5 fixed it, and all three approved.
  • Small fix 849c7a6: app_loop now drops a websocket whose kernel closed while it initialized, so no traceback is logged. All three reviewers raised this as LOW. I confirmed the fix as the driver, because it does only what the finding asked and adds no name.
  • Small fix: page_close now reads the page status once with .get(), instead of a membership check and a second lookup. The maintainer flagged that pattern. It was safe under the lock, but the single read makes that obvious. I confirmed the fix as the driver, because it changes no behavior and adds no name.

Constitution change that could have prevented the first defect: a crossreview brief may mark as settled only the decisions that the user made. It must give decisions that the driver made alone as claims for the reviewers to attack.

🤖 Generated with Claude Code

A close beacon can name a page that the kernel never registered. This
happens when an app runs several server processes and the beacon
reaches a process with a stale copy of the kernel, or when a tab closes
before its websocket connects. page_close then raised a KeyError, so the
close route returned a 500 and error trackers logged an exception for
something harmless.

page_close now logs the unknown page and returns, as it already does
for a closed kernel or a page that is already closed. The kernel and
its known pages stay as they are.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review found that the test only had a connected page, where page_close
never bumps the cull, so it could not catch a bump. The test now uses
a disconnected page with a scheduled cull. The log line redacts the ids
like the other persistence-era log lines do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the 0.2 s cull timeout, a stalled CI runner could let the cull
close the kernel before the assertions ran. The test now uses the
default cull timeout and closes the kernel through page_close, so no
timing is involved.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The server creates a kernel first and connects the page after. When the
tab closes in between, the close beacon arrives first. Ignoring that
beacon let the late websocket mark the closed tab as connected, and the
kernel then lived until the cull timeout.

page_close now marks an unknown page as closed, so page_connect refuses
it. A kernel with no live pages closes, as for a known page. A page that
never connected never cancelled the cull, so its close does not bump it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closing the kernel inside page_close (previous commit) could hit a kernel
that was still initializing: a beacon during the persistence takeover
closed it, and the takeover then attached a flush worker to the closed
kernel, which leaked. The refused late connect also raised in the
websocket handler, so every tab closed during load logged an error.

page_close now only marks an unknown page closed and leaves the kernel
and its cull alone. page_connect refuses that page with PageClosedError.
The websocket handler catches it, which happens after initialization,
and closes the kernel if no page is live. The reason is not page-close,
because no page used the kernel, so its persisted state is kept.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When two pages share a kernel and one closes before it connects, the
kernel stays alive for the other page. The refused websocket stayed in
the kernel's websocket set, so the next broadcast tried to send to a
closed socket and logged errors. The handler now removes it before it
returns. The docstring of close_if_no_live_pages now gives the real
reason the persisted state is kept.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…over

A new context is in `contexts` while its state takeover waits on the
backend. A reconnect of the same tab can reuse it, and after the tab's
close beacon that reconnect is refused and closes the kernel. The first
handler then attached persistence and started a flush worker on the
closed kernel, which nothing stopped, and crashed while wiring streams
to the closed kernel.

The takeover now attaches under the teardown lock and skips the attach
once close() has begun, and initialization skips the stream wiring for
a closed kernel. The refused-websocket cleanup reads the kernel session
only once, because a concurrent close can clear it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ized

When a kernel closes during its state takeover, initialization returns
the closed kernel, and page_connect then raised an unhandled error. That
logged a traceback for an expected close. In a narrow case the websocket
also waited for a client message before it closed, which delayed the
reconnect by a few seconds. app_loop now returns at once, so the
websocket closes and the client reconnects to a new kernel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@maartenbreddels maartenbreddels changed the title fix: ignore a close beacon for a page the kernel never saw fix: close the kernel of a tab whose close beacon beats its websocket Oct 7, 2026
A membership check followed by a second lookup looks like a race to a
reader. It is safe under the lock, but one read makes that obvious.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch had an error being deployed

1 failed deployment
fix/page-close-unknown-page - solara-stable PR #1247 — d0c591fe Deployed Oct 7, 2026 by maartenbreddels
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant