Terraform apply still fails with 409 RoleAssignmentExists after upgrading past 3.8.6¶
One-sentence summary: the resource is marked tainted in state from a prior failed apply, so Terraform still plans a replace regardless of the ignore_changes fix landed in 3.8.6, and terraform untaint (not a config change) clears it.
🚨 Symptom¶
terraform apply (or a Terraform Cloud plan/apply run) fails with:
Error: unexpected status 409 (409 Conflict) with error: RoleAssignmentExists: The role assignment already exists. The ID of the existing role assignment is <guid>.
with module.saif-appservices.module.external_identity.azurerm_role_assignment.okta_secret_reader["External"],
on ../saif-external-identity/main.tf line <n>, in resource "azurerm_role_assignment" "okta_secret_reader":
This appears even after the module has already been upgraded to saif-appservices >= 3.8.6, which is expected to have fixed this exact error via a lifecycle { ignore_changes = [scope] } block.
📌 Applies to¶
| Aspect | Value |
|---|---|
| Component | saif-external-identity module, azurerm_role_assignment.okta_secret_reader / azurerm_role_assignment.oidc_secret_reader |
| Forge versions | Workspaces that hit the original casing-drift bug (pre-3.8.6) and have not been untainted since |
| Related versions | saif-appservices module >= 3.8.6 |
🧠 Cause¶
Before 3.8.6, azurerm_role_assignment.okta_secret_reader/oidc_secret_reader had no lifecycle block. AzureRM normalizes resource group names to PascalCase, but Azure's RBAC API returns the casing used when the assignment was first created, so a casing-only diff forced a replace on every apply. If that replace's create step ever failed partway through (for example, a run that errored on the 409 itself), Terraform can leave the resource marked tainted in state.
Upgrading to 3.8.6 adds ignore_changes = [scope], which stops Terraform from planning a new replace from config. It does not clear a taint flag that a prior run already wrote to state — a taint forces a destroy/recreate on the next apply independent of what the current config says. So the workspace keeps failing with the same 409 even though the module fix is in place, because the state, not the config, is now the source of the replace.
See 3.8.6 release notes for the underlying config fix; this article covers the state cleanup some workspaces still need after upgrading.
✅ Fix¶
-
Confirm the resource is tainted rather than genuinely drifted. Check state directly for the
taintedflag on this resource address:function Get-AllStateResources($module) { $module.resources foreach ($child in $module.child_modules) { Get-AllStateResources $child } } $state = terraform show -json | ConvertFrom-Json Get-AllStateResources $state.values.root_module | Where-Object { $_.address -like '*okta_secret_reader*' } | Select-Object address, tainted(The resource lives inside nested modules, so it only shows up under
root_module.child_modules[].resources, notroot_module.resourcesdirectly — walk the tree instead of readingroot_module.resourcesalone.)Or check
terraform planoutput: a tainted resource is called out with# <resource address> is tainted, so must be replacedright above the resource block, with no attribute diff. A genuine config-driven replacement instead shows# forces replacementnext to the specific attribute line causing it (this is what the casing-drift bug looked like before 3.8.6). Don't use# forces replacementalone as a taint test — it flags a real config diff, not a taint. -
Untaint the resource directly; do not run a full destroy/import or download the entire module tree just to do this. A minimal working directory containing only the
terraform { cloud { ... } }block (matching the workspace's org/name) and pinnedrequired_providersversions matching what wrote the state is enough for state-only operations (init,state list,untaint,show): -
Pin
required_providersto the versions that last wrote the state (check the workspace's last successful apply, or the module'sversions.tf) before runninginitin a bare directory. An unpinned config resolves the latest provider, which can fail to decode older state — for example,azurerm5.x removed attributes present in 4.x state, surfacing as an unrelated-looking "unsupported attribute" error duringinit/untaint. That is a provider/state version mismatch, not state corruption.
🔬 Verify¶
No output. The next plan against the upgraded module shows no replace for okta_secret_reader/oidc_secret_reader: