Skip to content

fix(kfoperators): don't fail gang-scheduled jobs still waiting for PodGroup admission as "operator hasn't updated the CR" - #24

Merged
pfernandes21 merged 1 commit into
masterfrom
devin/1788410854-kfoperator-stale-check-gang-wait
Sep 3, 2026
Merged

pfernandes21 merged 1 commit into
masterfrom
devin/1788410854-kfoperator-stale-check-gang-wait

Conversation

@devin-ai-integration

Copy link
Copy Markdown

Tracking issue

Related to exa-labs/monorepo#138180 (kf-operator timeout 1m → 10m on delphi).

Why are the changes needed?

The pytorch/tensorflow/mpi plugins fail a task with kubeflow operator hasn't updated the <kind> custom resource since creation time when status.startTime is still nil kf-operator.timeout after the CR was created. The check was written to detect an absent operator, and a nil startTime is the wrong proxy for that under gang scheduling.

With Volcano gang scheduling the training operator (v1.9.0, pkg/controller.v1/common/job.go ReconcileJobs) deliberately does not create pods — and therefore never reaches the code that stamps startTime — until the PodGroup leaves Pending. While it waits it writes only status.lastReconcileTime:

if jc.PodGroupControl.DelayPodCreationDueToPodGroup(pg) {
    log.Warnf("PodGroup %v unschedulable", jobKey)
    syncReplicas = false
}
if !syncReplicas {
    now := metav1.Now()
    jobStatus.LastReconcileTime = &now
    return jc.Controller.UpdateJobStatusInApiServer(job, &jobStatus)
}

So a healthy, queued gang looks "stale" to Flyte and is deleted as soon as scheduler admission takes longer than the timeout. The timeout can bound how long we wait for the operator; it should not bound how long the scheduler takes to admit a gang.

RCA on delphi (all read-only, Loki flyte/kubeflow/volcano namespaces), 6 × e3a-i4lc-bs256gc-hn2-4x (4×8 GPU) executions killed 2026-09-02 23:16Z–2026-09-03 00:27Z, e.g. am4k45t4rg8qpcmflv9c:

t actor event
23:16:31 flyte PyTorchJob + PodGroup am4k…-fvcmtxhi-0 created
23:16:31 operator reconcile → PodGroup … unschedulable → status write (lastReconcileTime), no pods
23:16:31 – 23:17:44 volcano scheduler never touches the PodGroup (mid-session; sessions were ~130 s)
23:17:36 flyte kubeflow operator hasn't updated the pytorch custom resource since creation time 23:16:31 → CR deleted
23:17:44 volcano Job … was deleted

No worker pod ever existed for these six; the "SIGTERM to healthy workers" seen on ak5ppmf4lzktk59hbqtb attempts 0/1/2 is a different mechanism — Volcano preempt evicting the prio -6 gang for a prio 0 task (a4bf…), which the operator then records as failed=N — not this check. Attempts 1/2 of that execution sat queued for 71 min and 64 min with startTime set only because their PodGroups happened to be admitted within ~10 s, i.e. whether a gang survives the check today depends on where the scheduler is in its session loop when the CR lands. The delphi bump to 10m (exa-labs/monorepo#138180, live since 00:37Z, zero kills since) makes that race rare, but any gang whose PodGroup stays Pending past the configured timeout — saturated queue, long sessions — is still killed for no fault of the operator.

What changes were proposed in this pull request?

common.OperatorNeverReconciled(status) = status.StartTime == nil && status.LastReconcileTime == nil, used by all three plugins in place of the bare StartTime == nil test. Either timestamp proves the operator reconciled the CR; the check still fires when neither is set (operator down / not watching the namespace), which is the case it was written for. No config or behaviour change for jobs that reach pod creation.

How was this patch tested?

  • TestOperatorNeverReconciled (common): empty status and Created-condition-only status are "never reconciled"; StartTime, LastReconcileTime, or both are not.
  • TestGetTaskPhase in pytorch/tensorflow/mpi: new case — CR created 1h ago, StartTime nil, LastReconcileTime set → PhaseQueued, no error. Verified it fails on the pre-patch predicate (kubeflow operator hasn't updated …) and passes with the patch; the existing "operator did not modify the job" and "suspended" cases are unchanged and still pass.
  • go test ./go/tasks/plugins/k8s/kfoperators/..., go build ./... (flyteplugins), gofmt -l clean.

Roll-out note: delphi pins the flyte-binary image by digest in infra/core/delphi/training-stack.ts (FLYTE_IMAGE_TAG), so this needs a fork image build + a monorepo bump after merge; not done here.

Labels

fixed

Check all the applicable boxes

  • I updated the documentation accordingly.
  • All new and existing tests passed.
  • All commits are signed-off.

Link to Devin session: https://app.devin.ai/sessions/84936c9760074a4793d713c917952f29
Open in Devin Desktop: https://app.devin.ai/desktop/session/84936c9760074a4793d713c917952f29?variant=devin
Requested by: @jld-adriano

… stale-CR check

The pytorch/tensorflow/mpi plugins fail a job when status.startTime is still
nil kf-operator.timeout after creation, on the assumption that a missing
startTime means the training operator never saw the CR. With Volcano gang
scheduling the operator deliberately withholds pod creation (and therefore
startTime) until the PodGroup is admitted, writing only
status.lastReconcileTime while it waits. A healthy queued gang therefore
looks 'stale' and is deleted as soon as scheduler admission takes longer
than the timeout.

Count either timestamp as proof the operator reconciled the job; the check
still fires when the operator is genuinely absent.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Author

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR. Add '(aside)' to your comment to have me ignore it.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@pfernandes21
pfernandes21 merged commit da4f23b into master Sep 3, 2026
43 of 46 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants