Skip to content

Stop propeller before admin on single-binary shutdown - #21

Merged
pfernandes21 merged 1 commit into
masterfrom
devin/flyte-single-ordered-shutdown
Aug 6, 2026
Merged

pfernandes21 merged 1 commit into
masterfrom
devin/flyte-single-ordered-shutdown

Conversation

@Vervious

@Vervious Vervious commented Aug 6, 2026

Copy link
Copy Markdown

Tracking issue

Related to #16 (HA mode for flyte-binary)

Why are the changes needed?

In single-binary mode, propeller publishes workflow/node/task events to admin over localhost:8089 — i.e. to the admin running inside its own pod. cmd/single/start.go starts admin, propeller and datacatalog as unordered siblings of one errgroup, and admin installs its own process-level SIGTERM handler. So on pod termination admin closes its listeners immediately while propeller keeps reconciling and keeps renewing the leader lease for the rest of the grace period.

That window is not cosmetic. A propeller round that emits an event during it fails to publish, and the failure escalates:

RuntimeExecutionError: max number of system retry attempts [27/10] exhausted.
Last known status message: ErrorRecordingError: failed to publish event, caused by:
EventSinkError: Error sending event, caused by [rpc error: code = Unavailable
desc = "transport: Error while dialing: dial tcp [::1]:8089: connect: connection refused"]

In-flight executions fail during an ordinary rolling restart. This only became reachable with #16: a Recreate singleton never had a terminating pod overlapping a live one.

Holding the lease to expiry is the second half of the same bug — it is why leader handover takes ~46s rather than the ~25s the lease config implies.

What changes were proposed in this pull request?

Ordered shutdown, scoped to single-binary mode:

  • cmd/single/start.go — the root command owns SIGTERM/SIGINT via signal.NotifyContext and drives shutdown explicitly: cancel propeller first, wait for it to exit (bounded at 5s so a wedged propeller cannot eat the pod's grace period), then cancel admin and datacatalog and wait for them to return. The errgroup is replaced by per-service contexts plus a result channel; a service failing outside shutdown still tears the process down as before.
  • flyteadmin/pkg/server/service.go — the gateway waits on its context when given a cancellable one, and falls back to its own signal handler otherwise. Standalone flyteadmin passes context.Background(), so its behaviour is unchanged; only the single binary takes the new path. HTTP shutdown now uses a live timeout context rather than the cancelled one, so graceful shutdown isn't skipped.
  • flytepropeller — new leader-election.release-on-cancel config, plumbed to client-go's ReleaseOnCancel. The single binary enables it, so a terminating leader releases the lease instead of renewing until it expires. Losing leadership because our own context was cancelled is now an ordinary shutdown, not logger.Fatal.

Defaults for standalone deployments are untouched.

How was this patch tested?

Unit tests (./cmd/single, ./pkg/server, ./pkg/controller/..., ./pkg/leaderelection/...) plus a control-first A/B on a local 3-node kind cluster. Control is clean master 6bac6f6f3, fix is this branch; identical Helm values apart from the image tag. 2 replicas, 6–8 concurrent 15-task executions in flight, then kubectl delete pod on the current lease holder (only the leader reconciles, so killing a non-leader proves nothing).

control fix
Executions, 120s grace 8/8 FAILED 8/8 SUCCEEDED
Executions, 30s grace 1/6 FAILED (intermittent) 6/6 SUCCEEDED
connection refused to [::1]:8089 2880 0
EventSinkError 1292 0
retry-exhaustion 257 0
Lease handover, 30s grace 45.8s / 47.0s 1.19s
Lease renewals after termination 16 / 61 0
Admin gRPC probe failures 0 0

Ordering is shown positively rather than only by absent errors — the terminating pod logs

16:57:31.428 start.go:272      Shutdown requested. Stopping Propeller before Admin.
16:57:31.431 service.go:473    Servers gracefully stopped
16:57:31.433 controller.go:376 Stopped leading during shutdown.

with zero propeller reconcile lines after admin stopped. The handover improvement is mechanically visible too: the terminating leader emits no new renewTime after termination under the fix, versus 16 (30s grace) and 61 (120s grace) on master.

Regression check on the default path: single replica, Recreate, no PDB, pod deleted mid-execution — clean shutdown, no panic, new pod Ready with 0 restarts, admin reachable ~1s later, in-flight and post-restart executions both succeeded.

Caveats worth a reviewer's attention:

  1. At the default 30s grace the bug is intermittent — it reproduced in 1 of 2 control runs. The clean 8/8-vs-8/8 comparison required widening the window to terminationGracePeriodSeconds=120, applied identically to control and fix. Both sets of numbers are above.
  2. leader-election.release-on-cancel was only exercised via the programmatic default the single binary sets; the config/chart path itself is untested.
  3. The 5s bounded wait never fired locally — propeller stopped in ~5ms every time.
  4. Grepping logs for localhost:8089 finds nothing even while the bug fires: localhost resolves to IPv6, so the logs read [::1]:8089.

Labels

fixed

Check all the applicable boxes

  • I updated the documentation accordingly.
  • All new and existing tests passed.
  • All commits are signed-off.

Related PRs

#16

Link to Devin session: https://app.devin.ai/sessions/137e980b425d43668660198d80b6e21d
Requested by: @Vervious

Stop Propeller and release its leader lease before shutting down Admin in single-binary mode, while preserving standalone FlyteAdmin signal handling.

Assisted-by: Devin:claude-sonnet-4.5

Co-Authored-By: benchan <ben@vervious.com>
@Vervious Vervious self-assigned this Aug 6, 2026
@devin-ai-integration

Copy link
Copy Markdown

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@pfernandes21
pfernandes21 merged commit 809f402 into master Aug 6, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants