I'm planning to upgrade a production Kubernetes cluster from 1.34.8 to 1.35, and Elasticsearch is currently running in the cluster. I'd like advice on checking compatibility and reducing risk before starting. In particular, what should I verify for Elasticsearch, ECK or another operator, StatefulSets, storage classes and CSI drivers, ingress, CRDs, deprecated APIs, autoscaling, and other dependencies? Should the operator or Elasticsearch be upgraded before Kubernetes, and should the control plane and worker nodes be handled separately? I'm also looking for a practical pre-upgrade checklist, backup and rollback plan, recommended tools or commands for finding deprecated APIs, and a safe production rollout strategy that minimizes downtime. Experiences with similar upgrades would be especially helpful.
5 Answers
If the operator needs an update, use a version that supports both sides of the transition and follow its documented upgrade order. In many cases the operator is upgraded first, followed by Kubernetes, but the exact sequence is product- and version-dependent, so do not guess. Before production, inventory all workloads and CRDs, scan manifests for removed APIs with tools such as kubent or Pluto, run server-side dry runs against a 1.35 test cluster, and check events, logs, storage provisioning, ingress, autoscaling, and Elasticsearch health after each stage.
The safest approach is to test the upgrade on a separate cluster that closely matches production. Verify the ECK or Elasticsearch operator version supports both your current Kubernetes version and 1.35, then review the Kubernetes 1.35 API and release-change documentation for removals or behavior changes. StatefulSet and storage APIs are generally stable, but the operator, CSI driver, ingress controller, autoscaler, and any custom resources still need individual checks.
Treat Elasticsearch snapshots as your recovery option rather than assuming a Kubernetes rollback will restore everything. Take and verify snapshots in durable object storage, test restoring them into a separate cluster, and confirm that the snapshot repository, credentials, storage, and retention policy all work. Keep the existing Elasticsearch version available so you can rebuild the cluster consistently if necessary.
Upgrade the control plane and worker nodes separately. Start with a staging environment, then upgrade the control plane, validate cluster health, and roll worker node pools gradually. Cordon and drain nodes using the normal disruption controls so Elasticsearch pods can relocate safely. Make sure you have enough hot, warm, and master capacity for one node or pod to be unavailable at a time, and wait for the Elasticsearch cluster to return to a healthy state between disruptions. ECK normally manages appropriate disruption behavior, but you should still verify the PodDisruptionBudgets and scheduling capacity.
Several production users report that managed-cluster upgrades work smoothly when Elasticsearch has adequate capacity and the operator is compatible. A good rollout is: clone or recreate production in a test cluster, restore a recent Elasticsearch snapshot, upgrade and observe there, confirm backups and monitoring, upgrade the production control plane, then rotate worker pools gradually. Define clear stop conditions—unhealthy Elasticsearch, pending pods, failed volume attachment, operator errors, or rising application error rates—and pause the rollout if any appear.

A practical snapshot strategy is to run regular repository snapshots, monitor them for failures, and periodically perform a restore test. A snapshot that has never been restored is not a complete rollback plan.