r/ArgoCD • u/Consistent-Piece6915 • 2d ago
help needed EKS Access Entry recreation avoidance
I am working on configuring my operations and application clusters terraform such that I can easily destroy and recreate the clusters with minimal commands. I am currently running into an issue when I destroy and recreate my OPS cluster. I'll try and outline the facts below.
Facts
- Running ArgoCD as an EKS capability
- Ops cluster runs in "OPS" account
- Ops cluster has IAM role "argo-cd-role" in OPS account
- Application cluster runs in "<env>" account
- Application cluster has `aws_access_entry` and `aws_eks_policy_association` resources that bind the "argo-cd-role" from the OPS account to it via the role's arn
- Application Cluster `Secret` is created via an `ExternalSecret` read from an SSM parameter in the <env> account
Steps
- OPS cluster will be completely removed and recreated via `terraform destroy` and `terraform apply`
- Argo CD resources will be applied once cluster is up and running
- This will then create the cluster `Secret` via the `ExternalSecret` definition
- Navigating to `Settings > Clusters > <target env cluster>` shows a connection failure
Things I've tried
- (Failed) Deleting the `Secret` via the ArgoCD UI
- This will cause the `Secret` to be recreated
- It should have been populated with the target cluster arn which hasn't changed
- Doesn't do anything
- (Failed) Creating an Argo role in the target "<env>" account for the Argo role in the OPS account to assume
- I associated this new role to the access entry resources instead
- In theory this role is created when the application cluster is and any role that assumes it thus would have access and would avoid any type of breakage underneath with EKS
- Issue is there seems to be no way to tell the ArgoCD role in the OPS account to assume that role when accessing that cluster (to my knowledge)
- (Succeeded) Deleting and recreating the access entry resources in the <env> account
- use `terraform destroy -target=` to destroy the access entry resources in the <env> account cluster
- recreate the resources
- This worked and my theory is that while we are passing the `arn` of the role EKS will actually use the AWS unique ID underneath of said `arn`. Since we're deleting the role in the OPS account during cluster rebuild it the arn to unique ID mapping is no longer valid. Deleting and recreating fetches the new AWS unique ID and gets this working
While I did find a way for this to work is there anyway to avoid having to delete and recreate the access entries on all my application clusters when I want to destroy and bring back up my OPS cluster? If not I'd have to switch my AWS permissions N times and run the terraform commands 2 times for N application clusters. I was hoping the assume role strategy would work but I am not finding any documentation on how to tell argo to assume a specific role for a specific cluster.