GKE Agentic Migration Stops at the Pull Request—Production Proof Starts There
Google Cloud’s GKE Agentic Migration changes how AI-generated infrastructure enters an EKS-to-GKE migration. It does not give a model direct authority to rebuild or deploy a cluster. Instead, it reads infrastructure-as-code and Kubernetes configuration, generates target-state artifacts, validates defined properties, and submits the result for human review as a pull request.
That is a useful control boundary. It is not evidence that the migrated workload will preserve permissions, routing, capacity, latency, data integrity, or recovery behavior at runtime.
Google announced the open-source release on September 24, 2026. Its September 25 product roundup identified the project as Public Preview. Those descriptions are compatible: “open source” concerns distribution and licensing, while “Public Preview” describes product maturity. As of September 28, 2026, evaluations should treat the framework as preview software—not as a generally available migration service.
What Google released
The Google-maintained repository describes GKE Agentic Migration as a client-side, serverless, spec-driven framework built on the Model Context Protocol. “Serverless” here does not mean that the workflow is a hosted Google control plane: orchestration runs through agent skills and a local MCP server, with no centralized backend operated by the user.
The documented operating model includes:
- Agent harnesses: Antigravity and Claude Code.
- Identity: The framework uses the operator’s Application Default Credentials rather than a service-account key file stored on disk.
- Durable state: Migration progress is persisted in a Cloud Storage bucket called the ledger.
- License: The repository is published under the Apache License 2.0.
- Output: Generated Terraform, Kubernetes manifests, documentation, runbooks, and pull requests in a Git repository controlled by the user.
The repository divides work between probabilistic and deterministic components. LLM workers read infrastructure-as-code, choose target shapes, and author HCL, YAML, and documentation. Compiled MCP tools handle operations that the project defines as deterministic and auditable, including state transitions, selected cloud mutations, Git writes, transforms, and validation.
The documented preflight checks include:
terraform validate- Kubernetes manifest-structure checks
- Output contracts
Human approval remains part of the workflow. The repository and onboarding guide explicitly state that the optional live EKS scan is read-only, the source AWS environment is not modified, and the framework does not run terraform apply. Merging and applying the generated pull request remain human-controlled actions.
The important boundary: validation is not equivalence
The framework can reject malformed or contract-breaking artifacts. It cannot prove properties that its validators do not observe.
A successful terraform validate run establishes that Terraform accepted the configuration’s syntax and internal structure. Manifest checks can establish that generated Kubernetes objects conform to implemented rules. Neither result demonstrates that the deployed system will behave like the EKS source.
Google documents several cloud-specific translations in its announcement and repository:
| Translation | Confirmed generated target | Runtime evidence still required |
|---|---|---|
| AWS IRSA | Workload Identity Federation for GKE configuration | Effective allow and deny behavior, namespace isolation, token exchange, and access to external resources |
| AWS ALB ingress | GKE Gateway API resources, including HTTPRoute |
Host and path routing, TLS, redirects, timeouts, health checks, session behavior, and dependency reachability |
| Karpenter node claims | GKE Node Auto Provisioning or Custom Compute Classes | Scheduling under scarcity, scale-out latency, quotas, disruption handling, zone failure, and Spot interruption behavior |
| EBS CSI storage classes | GKE Persistent Disk CSI or Hyperdisk mappings | Data transfer, access-mode compatibility, performance, restore integrity, and recovery objectives |
The first column is documented automation. The final column is an assurance requirement derived from the behavior that production systems must preserve; it is not a claim that Tech Trend Insight performed those tests.
Google’s broader EKS-to-GKE architecture guidance reinforces the distinction. It tells teams to inventory cluster topology, networking, quotas, Kubernetes versions, RBAC, network policies, autoscalers, storage, secrets, dependencies, taints, affinities, authentication, custom resources, backup processes, and operational workflows. Those dimensions extend well beyond artifact syntax.
The practical consequence is straightforward: a validated pull request is a stronger proposal, not a verified replacement environment.
Configuration translation is deliberately narrower than migration
The repository defines GKE Agentic Migration as a Kubernetes-domain translator. Its automated scope includes common Kubernetes objects, Helm charts, Kustomize overlays, plain manifests, Terraform, and selected mappings for ingress, identity, and storage.
It also specifies clear handoff boundaries:
- Stateful-data transport: Not performed by the plugin.
- Cloud-managed services: RDS, ElastiCache, S3, EFS, Secrets Manager, and similar dependencies produce manual runbooks and can block workload progress until the underlying data or service has moved.
- Application-source refactoring: Java, C#, Python, and other application-code changes are outside automated scope.
- Enterprise identity federation: External IdP and Active Directory integration are outside scope.
- Cross-cloud networking: SD-WAN, Direct Connect, and related network architecture are handed off.
- Non-containerized workloads: Virtual machines and bare-metal workloads are excluded.
Google’s announcement says stateful migration should be handed to purpose-built services such as Database Migration Service or Storage Transfer Service. A generated StatefulSet, PersistentVolumeClaim, or storage-class mapping describes a desired target. It does not copy bytes, control concurrent writes, reconcile divergence, validate restore integrity, or prove recovery objectives.
The ledger enables resumability—but requires its own access review
The repository assigns work across three personas:
- Migration Admin: Initializes the Cloud Storage ledger and manages membership and permissions.
- Platform Engineer: Performs discovery and assessment, resolves blockers, designs the GKE landing zone, and produces the platform infrastructure pull request.
- App Developer: Translates a selected application component and produces a workload pull request.
Because progress is stored in the ledger rather than only in a conversation, another session or operator can rejoin the recorded workflow. The repository uses Cloud Storage managed folders and IAM to separate persona-specific paths.
That separation should be evaluated using effective IAM, not assumed from folder names. Current Cloud Storage managed-folder documentation states that:
- Permissions on nested managed folders are additive.
- Managed folders require uniform bucket-level access.
- IAM deny policies cannot be attached directly to buckets or managed folders.
- Managed folders cannot be targeted as resources in deny-policy conditions.
Inherited project, folder, bucket, and managed-folder grants therefore need to be reviewed together.
The current onboarding guide also documents a residency constraint: bootstrap creates the ledger in the US multi-region, with no region selector. The guide says the bucket contains migration state rather than workload data. Organizations should still classify the contents before use because migration state can include inventories, repository metadata, design decisions, generated artifacts, blockers, and handoff records.
A safe cutover requires evidence the plugin does not generate
Google’s official migration architecture guide recommends validating each migration-plan step and maintaining a rollback strategy. After workloads are deployed to GKE—but before exposing them—it recommends integration, load, compliance, reliability, and other workload-specific testing, supported by logs, metrics, and error reports.
The same guide recommends gradually shifting traffic to GKE, monitoring behavior as load increases, backing up source data, and decommissioning the source only after the target serves requests correctly.
An official live-traffic migration tutorial demonstrates one implementation for an HTTP service:
- Deploy the application to both EKS and GKE.
- Place a temporary NGINX proxy in GKE.
- Route traffic through a GKE Gateway and
HTTPRoute. - Keep EKS serving while weights shift incrementally toward GKE.
- Monitor a continuous Locust workload.
- Remove the proxy after the route reaches the GKE backend.
The tutorial reports results from its own example run, under its stated regions, infrastructure, and network conditions. Those numbers are not a benchmark for GKE Agentic Migration and should not be generalized to production workloads. The pattern also does not automatically cover stateful services, asynchronous consumers, long-lived connections, transactional cutovers, or different network topologies.
Editorial recommendation: Treat the generated pull request as the first artifact in an assurance sequence:
- Apply it in an isolated target environment.
- Compare effective identity and authorization behavior.
- Validate routes, TLS, dependencies, secrets, and external services.
- Exercise scheduling, autoscaling, disruption, and capacity limits.
- Prove backup, restore, rollback, and recovery behavior.
- Shift traffic through predeclared stages and stop thresholds.
- Retain the EKS path until acceptance criteria are met.
- Decommission only after technical and business approval.
Governance signals are still early
Repository governance should be evaluated separately from the framework’s architecture.
The current onboarding guide says it applies to commit c012c0f, identifies the plugin as version 1.0.0, and carries an update date of September 22, 2026. As checked on September 28, 2026:
- The repository’s GitHub Releases page showed no published releases.
- The GitHub Security page reported no
SECURITY.md. - The same page showed no published security advisories.
These are not evidence of a vulnerability. They leave operational questions that adopters should resolve before relying on the project: how supported versions will be identified, how breaking validator changes will be communicated, where security reports should go, and how upgrade guidance will be published.
For an auditable trial, record at least:
- Repository commit
- Plugin and harness versions
- Model endpoint and configuration
- Validator versions and output
- Approval decisions
- Generated diffs and runbooks
- Final merged changes
- Post-deployment test results
Open source does not mean zero migration cost
The software license does not cover the infrastructure and engineering required to use it.
The documented workflow uses a Cloud Storage ledger and background model calls. Cloud Storage pricing includes storage, operation, data-processing, replication, and network components. Agent Platform generative-AI pricing meters model usage according to the selected model, endpoint, token volume, and applicable pricing conditions.
The live-traffic tutorial also identifies billable EKS, load-balancing, data-transfer, GKE, Cloud NAT, Cloud Router, external-IP, and network-egress components. A real migration can add image replication, data movement, validation environments, monitoring, engineering review, remediation, and parallel EKS/GKE operation.
No reviewed primary source provides an end-to-end migration estimate for GKE Agentic Migration. Nor do the sources establish a measured speedup, lower failure rate, migration-cost reduction, or production reliability improvement attributable to the plugin.
A defensible economic evaluation should therefore measure:
- Model input and output usage
- Cloud Storage capacity and operations
- Image-copy and data-transfer volume
- AWS and Google Cloud network charges
- Temporary infrastructure
- Parallel operation duration
- Review and remediation hours
- Validation and rollback work
Proposed evaluation protocol
The following tests are proposed. Tech Trend Insight did not run them and has no benchmark results.
Reproducibility
Run a pinned EKS fixture through multiple clean sessions using one repository commit, harness version, model configuration, and validator set. Compare target-shape decisions, generated diffs, runbooks, parked work, validator output, and pull-request contents.
Variation is not automatically a defect. Unexplained variation should, however, block claims of deterministic migration output.
Behavioral parity
Deploy into an isolated GKE project and test cold-start identity, least-privilege access, ingress paths, TLS, secret injection, dependency reachability, node scarcity, autoscaling, disruption budgets, and rollback. Define thresholds before observing results.
Ledger isolation and recovery
Interrupt the workflow between discovery, platform design, and workload translation. Rejoin from clean sessions, revoke and restore persona access, and inspect effective Cloud Storage permissions. Confirm that recorded state, parked work, and generated artifacts remain coherent.
Cutover safety
Keep EKS available while shifting controlled traffic percentages. Monitor workload-specific indicators such as errors, latency, saturation, queue depth, failed dependencies, and rollback time. Stop the shift when predeclared thresholds are crossed.
Cost
Capture model usage, ledger activity, generated-artifact volume, image transfer, cross-cloud egress, infrastructure charges, review effort, remediation effort, and coexistence duration. Compare costs only against a declared source workload and migration baseline.
Decision rule
GKE Agentic Migration is a reasonable trial candidate for teams that want AI-generated infrastructure to enter existing Git review and CI/CD controls instead of a live cluster.
Its strongest design choice is the separation of LLM authoring from deterministic operations, human approval, and deployment authority. That separation narrows one category of risk: uncontrolled model-to-cluster mutation.
Proceed toward production only when all four conditions are true:
- The source EKS environment and rollback path remain recoverable.
- Workload-specific behavioral tests and acceptance thresholds exist.
- Migration, validation, data-transfer, and coexistence costs are measured.
- The pull request is treated as a proposal—not as proof of a safe cutover.
If any condition is missing, the appropriate action is to delay deployment, not to place more trust in the generated configuration.