Skip to content

Fix Kubernetes locker ownership with client-go Lease election - #983

Open
pood1e wants to merge 3 commits into
openconfig:mainfrom
pood1e:fix/k8s-lease-lifecycle
Open

pood1e wants to merge 3 commits into
openconfig:mainfrom
pood1e:fix/k8s-lease-lifecycle

Conversation

@pood1e

@pood1e pood1e commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #982.

The Kubernetes locker can reject valid runtime target identities and can continue collection after Lease renewal is no longer reliable. This replaces the custom acquisition and renewal loop with client-go leader election and LeaseLock, gives every acquisition a unique holder identity, and reports renewal-deadline failures through the collector's stop-and-retry path.

Release verifies the session owner, uses UID and resource-version preconditions, and gives the elector shutdown, Lease read, and Lease delete independent timeout budgets. Shutdown releases run with bounded concurrency. The collector cancels its subscription before release I/O, outside the operational mutex.

A shared Lease informer serves ownership queries. Its API list/watch is restricted to app=gnmic, and an original-key prefix index avoids scanning every cached Lease for each List call. qps and burst make the client-go request budget explicit.

Existing keys supported by the previous locker keep their slash-to-hyphen Lease names and legacy labels, while annotations retain the exact key and value. Previously unsupported keys use gnmic-<sha256> names. This lets old and new replicas coordinate during a rolling update. The deprecated renew-period and retry-timer settings remain aliases during that rollout; the documented sequence removes them after every replica is upgraded.

Kubernetes setup, lifecycle constraints, RBAC, limitations, and the rolling-upgrade sequence are documented in docs/user_guide/ha_kubernetes.md, alongside the EndpointSlice discovery documentation merged in #980.

Local validation at 359cd3e1a5ff37834279a218e2f675791e576a3f:

  • ./tests/run_tests.sh: passed for all three Go modules.
  • ./tests/run_tests.sh --race: passed for all three Go modules.
  • go vet ./pkg/lockers/k8s_locker ./pkg/app ./pkg/collector/managers/targets ./pkg/collector/managers/cluster: passed.
  • Staticcheck on the changed packages has no new findings; the unsuppressed run reports the repository's existing pkg/app/app.go SA1019 deprecation.
  • Regression tests cover cancellation channel semantics, renewal failure, independent release budgets, bounded shutdown concurrency, rolling Lease compatibility, prefix-index filtering, ownership-safe deletion, and configuration aliases.

The earlier comments contain combined Kubernetes deployment evidence from this change and #980 before the review-fix rebase. The review fixes above are covered by the new local unit and race runs; GitHub checks on the updated head remain authoritative.

@pood1e

pood1e commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Local runtime checks passed for 2ba66e42ce82979aab28f3266add4122fb1687a4, combined with the EndpointSlice implementation from #980 at 92a5d658fa2cdb4b418e5e2ab0bdbaf7e67bdd0b:

  • Target identities are preserved when mapped to Lease names.
  • Pod replacement recovers active target ownership and fresh Subscribe responses.
  • Failed Lease writes cause collection to stop; collection recovers after write access is restored.
  • Sampled active owners agree with Lease annotations at converged checkpoints.

The stop check uses removal of active targets from the runtime API and unchanged receive counters, since UP gauges update asynchronously. Blocked API requests and release I/O are covered by the regression tests in this PR. These observations describe the exercised cases, not an exactly-once guarantee.

@pood1e

pood1e commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Additional local integration checks passed for target ownership reconciliation, subscription recovery and Remote Write with this PR combined with #980. The PR's local unit and race checks also passed.

GitHub Actions requiring maintainer approval remain pending; these results describe local validation.

@pood1e

pood1e commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor Author

Final combined deployment validation is now complete for this Lease lifecycle change together with #980:

  • Kubernetes v1.36.2+k3s1, 3 nodes, 6 gNMIc replicas distributed 2/2/2.
  • 641 target Leases and 4,487 subscriptions remained healthy through a 7,204-second steady-state run.
  • Active owners remained unique; Remote Write failures, dropped messages, subscription failures, and container restarts had zero increase.
  • Dedicated member/leader replacement recovery passed (11.458s and 25.639s business-sample recovery, both below the existing 60s limit).
  • Lease write denial still stopped active collection and recovery restored all owners in the earlier fault test; the final run completed with no cleanup errors.

The PR is mergeable and all GitHub checks are green.

@karimra, could you review this together with #980 when available?

Comment on lines +104 to +111
case <-ctx.Done():
session.cancel(ctx.Err())
case <-session.stopped:
}
if err := ctx.Err(); err != nil {
errs <- err
close(done)
return

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

on ctx.Done() this sends err on errs chan. errs should only carry renewal failures.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 359cd3e1. Caller cancellation now cancels the Lease session and closes done without sending on errs; errs is used only when the lock cannot be maintained. The blocking app-side error-channel drain was removed as well. TestKeepLockCancellationCompletesWithoutErrorReceiver covers this path.

Comment on lines +128 to +135
ctx, cancel := context.WithTimeout(ctx, k.Cfg.RetryPeriod)
defer cancel()
session.cancel(context.Canceled)
select {
case <-session.stopped:
case <-ctx.Done():
return ctx.Err()
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this context can be mostly consumed by session.stopped. That leaves almost nothing to client.Get and client.Delete coming afterwards.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 359cd3e1. release now creates a fresh RetryPeriod timeout for each stage: waiting for the elector to stop, reading the Lease, and deleting it. TestReleaseUsesIndependentTimeouts verifies that the stop wait cannot consume the Get or Delete budget.

Comment on lines +76 to +79
leases, err := k.leases.List(labels.Everything())
if err != nil {
return nil, err
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would be good to filter during listing rather than getting all leases and filtering them afterwards.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 359cd3e1. The informer list/watch remains server-side filtered to app=gnmic, and now maintains an original-key-prefix index. List uses ByIndex(prefix), so each call receives only matching cached Leases instead of enumerating the whole informer cache. TestLeaseCacheFiltersAndTracksOwnership checks the index cardinality and API-free reads.

}
}
return errors.Join(errs...)
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This functions uses a shared 5s timeout context and passes it to k.release, which applies a RetryPeriod (2s), for each session.
Either use a bounded worker pool to release the sessions and size the context timeout accordingly, or size the timeout based on len(sessions).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 359cd3e1. Stop uses errgroup.SetLimit(16) to bound concurrent releases. Every release has independent bounded stop/Get/Delete stages, so one shared deadline no longer starves later sessions. TestStopBoundsConcurrentReleases verifies the worker limit and that all sessions complete.

func leaseName(key string) string {
digest := sha256.Sum256([]byte(key))
return "gnmic-" + hex.EncodeToString(digest[:])
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This change breaks rolling upgrades.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 359cd3e1. Keys representable by the previous locker keep the same slash-to-hyphen Lease name and legacy label, while annotations carry the exact key/value. Digest names are used only for keys the previous implementation could not represent. The compatibility test starts from an old-format active Lease, verifies that a new replica cannot acquire a second object, and verifies dual metadata on newly created compatible Leases.

Comment thread docs/user_guide/ha_kubernetes_locker.md Outdated
Comment on lines +64 to +68
All members sharing a cluster must stop before this upgrade. The previous and
new encodings refer to different Lease objects, so a mixed-version rolling
upgrade would create independent ownership domains. After all old members have
stopped, apply the new configuration and RBAC, then start the upgraded members.
Obsolete Lease objects can be removed after confirming their holders have stopped.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is not great... we should provide an upgrade path

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reworked in 359cd3e1. The stop-all instruction is gone. The existing ha_kubernetes.md now gives an ordered rolling path: grant watch RBAC first, retain supported keys and legacy timing fields while rolling the binary, then rename the deprecated fields after every replica is upgraded. It also documents the boundary for previously unsupported keys and the retained legacy collision behavior.

@pood1e
pood1e force-pushed the fix/k8s-lease-lifecycle branch from 2ba66e4 to 359cd3e Compare September 10, 2026 05:36
@pood1e

pood1e commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Review feedback is addressed in 359cd3e1a5ff37834279a218e2f675791e576a3f, rebased onto current main (including #980).

Validation on the updated head:

  • ./tests/run_tests.sh — passed for ., pkg/api, and pkg/cache.
  • ./tests/run_tests.sh --race — passed for ., pkg/api, and pkg/cache.
  • go vet ./pkg/lockers/k8s_locker ./pkg/app ./pkg/collector/managers/targets ./pkg/collector/managers/cluster — passed.
  • go mod tidy -diff and git diff --check — clean.
  • Staticcheck on ./pkg/lockers/k8s_locker ./pkg/app passes after excluding the existing pkg/app/app.go:29 SA1019 deprecation; no finding is introduced by this PR.

The new regression coverage exercises caller cancellation without an errs send, renewal failure reporting, independent release timeout budgets, the 16-worker shutdown bound, legacy Lease interoperability during a rolling update, indexed prefix lookup, and deprecated configuration aliases.

The Test workflow is waiting for maintainer approval because this is a forked PR. I have left the review threads open for maintainer verification.

@karimra, all six review points have corresponding fixes and replies; this is ready for re-review when available.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Kubernetes locker cannot safely maintain target ownership

2 participants