What Kyverno Policy Migration Actually Means on EKS

Kyverno policy migration is the structured process of moving an existing set of namespace-scoped or cluster-wide Kyverno policies from one EKS cluster, one Kubernetes minor version, or one Kyverno release to another without losing enforcement coverage or breaking workloads. On Amazon Elastic Kubernetes Service specifically, migration rarely happens in isolation. Teams typically run into the question when EKS is moving from a version like 1.30 to 1.32 or 1.33, when Kyverno itself jumps from a 1.11.x line to 1.12.x or 1.13.x, or when the organization is consolidating multiple dev clusters into a single governed production platform. The Nasscom EKS 1.33 implementation guide notes that AWS has steadily expanded hybrid and policy-as-code use cases in the 1.32 and 1.33 release windows, which means more teams now treat policy as a first-class deployment artifact rather than a post-install add-on.

Also worth reading: What Should Enterprises Include in a ModelOps Evaluation Checklist in 2026? · What should an enterprise LLM red teaming checklist cover before a model goes live in 2026? · How Should Organizations Structure an Enterprise AI Evaluation Checklist for 2026 Deployments?

The phrase "policy migration checklist" is slightly misleading because Kyverno policies are stored as Kubernetes Custom Resources (ClusterPolicy and Policy) inside the cluster. They are declarative YAML files that can be versioned in Git, exported with kubectl get, and re-applied with kubectl apply -f. The real checklist is therefore a sequence of verification, conversion, and validation steps rather than a literal file copy. A mature migration plan touches five layers: the Kyverno controller, the CRDs, the policies themselves, the policy exception (PolicyException) objects, and the generated ConfigMap-backed reports that downstream auditors may rely on.

The Pre-Migration Audit Layer

Before touching a single manifest, the migration owner should run a full inventory of every ClusterPolicy, Policy, and PolicyException in the source cluster. The standard kubectl command for this is kubectl get clusterpolicies,polices,policyexceptions -A -o yaml > policies-backup.yaml, but enterprise teams usually split the dump into per-namespace files so that Git diffs stay readable. At this stage, count how many policies use background scanning versus the ones that gate admission only. Background-scanning policies (the ones with a spec.backgroundScan: true and a generated PolicyReport) tend to dominate migration risk because they produce cron-driven reports that have to keep functioning after the upgrade.

It is also worth checking the Kyverno version currently running. The project ships breaking changes between minor lines: the 1.10 release deprecated the old match.block syntax, 1.12 removed the v1beta1 Kyverno CRD definitions, and 1.13 reorganized the engine reporter under a new ConfigMap. A cluster still running Kyverno 1.9 cannot simply roll forward to 1.13 without an intermediate 1.11 stop because the CRD conversion strategy changed. Treat the version delta as a hard dependency on the checklist and not as a footnote.

Finally, catalog the Generate and Mutate rules separately. Generate rules create child resources like NetworkPolicies or default resource quotas, and they re-run on every cluster event. Mutate rules patch incoming objects, often with webhook order dependencies. A checklist that treats all rules the same will quietly break the most subtle ones.

Building the Actual Migration Checklist

A practical Kyverno policy migration checklist for EKS can be written as a twelve-step sequence. The team should freeze new policy authoring in a feature branch, capture the inventory described above, and pin the source Kyverno Helm chart to the exact chart version that matches the source cluster. From there, the team should diff the upstream Kyverno release notes between the source and target versions and flag every rule that mentions a deprecation, a renamed annotation, or a changed default for failurePolicy. The next step is to install the new Kyverno version in a non-production EKS cluster that mirrors production topology, and apply the existing policies under a dryRun audit mode.

The middle of the checklist focuses on the policies themselves. Validate rules should be tested against a controlled workload that is expected to fail; mutate rules should be tested against a workload that exposes the merge order; generate rules should be tested by deleting the child resource and confirming the controller recreates it within the documented reconciliation window. PolicyException objects should be rewritten with the new labelSelector fields that 1.12 introduced, and any cluster-scoped exceptions should be re-anchored to the new cluster-scoped ClusterPolicy CRD. The final acceptance step is a production dry run, where the new Kyverno version is installed in audit mode and the PolicyReport is reviewed for a full 24-hour window before enforce is switched back on.

Comparing Migration Strategies

There are three realistic strategies for moving policies between EKS clusters, and the right one depends on cluster topology and the team's tolerance for downtime. The table below compares them on the dimensions that matter most for a regulated enterprise.

StrategyBest forDowntime windowRollback speedGitOps fitMain risk
Blue/green cluster with Argo CD syncLarge regulated enterprises running multi-tenant EKS0 minutes for workloads, 5-15 minutes for policy cutoverUnder 5 minutes via Git revertExcellentRequires duplicate node groups and a load balancer shift
In-place Helm upgrade with audit-first rolloutSingle-cluster pilots and small production estates1-3 minutes of admission webhook 503s10-20 minutes via Helm rollbackGood if Helm is the source of truthMismatch between old and new CRD versions can leave policies in Error state
Reapply manifest set after fresh Kyverno installDisaster recovery, region failover, and lab-to-prod promotion10-30 minutes totalManual re-apply requiredFairRisk of dropping the PolicyException CRDs if not exported first
Blue/green with Argo CD is the safest option for any cluster serving regulated workloads, because the source cluster continues to enforce policy while the destination cluster is being primed. In-place Helm upgrade is the cheapest option and the one most teams use, but it assumes that Helm values, not raw manifests, are the source of truth. The reapply strategy is fine for short-lived environments but is the most common source of orphaned PolicyReport objects that confuse auditors months later.

Common Mistakes and How to Avoid Them

The first mistake is treating Kyverno policies as "just YAML" and applying them with a single kubectl apply -f policies/. This skips the schema validation step and the engine dependency check. A small typo in a CEL expression or a JMESPath query can cause the policy controller to crash-loop and silently drop admission enforcement. The fix is to run kubectl apply --server-side --validate=true and to keep kubectl get clusterpolicy -o yaml in the rollback branch.

The second mistake is forgetting PolicyException objects. These are first-class resources in Kyverno 1.11 and later, and they are not migrated by the standard Helm chart. A team that only migrates ClusterPolicy will see a flood of legitimate workloads suddenly fail admission because their previously exempted Pods are no longer covered.

The third mistake is ignoring the background scan window. When a policy is reapplied with generateExisting: true, Kyverno will run a full cluster scan that can take hours on a 5,000-node estate. During that window, PolicyReport entries are noisy and incomplete, and many on-call engineers will misread them as a bug.

A fourth, less obvious mistake is failing to coordinate with the AWS Load Balancer Controller and the EKS add-ons lifecycle. EKS 1.32 and 1.33 changed the default ordering of admission webhooks, and a Kyverno install that worked perfectly on 1.30 may experience timeout races on 1.33 if the AWS LB Controller is not on a matching version. The Nasscom implementation guide calls this out explicitly under its hybrid networking section.

When the Migration Should Actually Happen

Timing is a separate problem from the checklist itself. Most teams should run Kyverno migrations inside a wider EKS upgrade window, not as a standalone event. The reason is that admission webhook configurations are versioned against the Kubernetes API server, and EKS API server deprecations are announced at least two minor versions ahead. A team planning an EKS 1.30 to 1.33 jump should schedule the Kyverno migration no later than the 1.31 intermediate, because that is the last release where the deprecated v1beta1 webhook conversions are still honored.

For enterprises governed under SOC 2, ISO 27001, or the EU AI Act, the migration window is also the right place to refresh the policy evidence pack. PolicyReports older than 90 days should be archived, and a fresh baseline report should be generated so that auditors see a clean before/after picture. The cost of doing this as a side-effect of the migration is near zero; the cost of doing it as a separate audit project is typically two to three engineering days per quarter.

Cost, Tooling, and Pricing Reality

Kyverno itself is open source under Apache 2.0, so the software cost is zero. The real cost is the engineering time to execute the checklist. For a mid-sized EKS estate of 30 clusters and 200 policies, an experienced platform team can complete the full migration in 4 to 6 working days, broken down roughly as 1.5 days for inventory, 2 days for the staged rollout, and 1.5 days for validation and report archival. At a fully loaded platform engineer rate of $150 to $220 per hour in 2026 US market conditions, that translates to a one-time project cost between $5,000 and $10,500.

Tooling choices matter at the edges. Enterprises running Argo CD can lean on the Kyverno plugin to detect policy drift between Git and the cluster. Teams that prefer a SaaS layer often evaluate tools such as Nirmata, Magalix (now part of Cisco), or Styra DAS, which charge between $8 and $25 per cluster per month in 2026 for managed policy authoring. The enterprise AI labs audience is a special case, because model evaluation pipelines typically need policy hooks on inference namespaces, and a SaaS layer that understands both OPA-style and Kyverno-style rules can save roughly one full-time engineer compared with maintaining the policy controller in-house.

A Concrete Walkthrough on EKS 1.33

A representative migration on EKS 1.33 starts by tagging the source cluster with the current Kyverno Helm release name and namespace, then running helm get values kyverno -n kyverno > old-values.yaml. The next command is helm repo update followed by helm diff upgrade kyverno nirmata/kyverno --version 3.2.0 -n kyverno -f old-values.yaml, which surfaces CRD and webhook changes before they hit the cluster. After the diff is reviewed, the team sets failurePolicy: Audit on every policy in a temporary overlay, applies the Helm upgrade, and watches kubectl get events -n kyverno for 30 minutes. Once the engine reports healthy, the overlay is removed and the team waits one full background scan cycle, which on EKS 1.33 with the default 15-minute cron is roughly 4 to 6 hours for a 2,000-namespace estate.

If the scan cycle is clean, the team promotes the change to production by repeating the overlay dance in the real cluster, this time keeping the new engine in Audit for 24 hours before flipping back to Enforce. PolicyException objects are reapplied in the same window. A 24-hour audit soak is the single most valuable insurance policy in the whole checklist, because it surfaces webhook ordering races, slow CEL expressions, and missing Generate rules that a 30-minute review cannot catch.

Final Verification Before Closing the Ticket

The checklist is not finished until four artifacts are archived: the source cluster's policy dump, the new cluster's policy dump, the PolicyReport from the 24-hour audit soak, and the Helm values diff. These four files together form a complete evidence chain that an auditor can replay to prove that no policy was lost, dropped, or silently weakened during the migration. A migration ticket that closes without these artifacts should be reopened, because the next EKS upgrade will rely on them and they are cheap to produce now and expensive to reconstruct later.

For a team running on enterprise AI labs infrastructure, the same checklist also doubles as a guardrail for model evaluation sandboxes, where policy exceptions are routinely granted to allow non-standard images, GPU node selectors, and pull secrets. Treating the migration as a recurring quarterly practice, rather than a one-time event, keeps those exceptions auditable and prevents the slow accumulation of shadow policies that every regulated AI platform eventually has to deal with.

Key Takeaways

A Kyverno policy migration on EKS is a 12-step process anchored on inventory, version pinning, audit-first rollout, exception preservation, and report archival. The biggest risks are forgotten PolicyException objects, webhook ordering changes introduced by EKS 1.32 and 1.33, and incomplete background scan windows. Blue/green cluster promotion is the safest strategy, in-place Helm upgrade is the cheapest, and a bare reapply is acceptable only for non-regulated environments. Budget four to six engineering days for a mid-sized estate, and always close the work with a 24-hour audit soak and a four-file evidence pack.