Why SSM Agent Does Not Immediately Pick Up a Newly Attached IAM Role
Why SSM Agent Does Not Immediately Pick Up a Newly Attached IAM Role
After attaching an instance profile with AmazonSSMManagedInstanceCore to a running EC2 instance, the instance may still be absent from Systems Manager. It might recover after a service restart—or on its own almost half an hour later.
This is often the combined result of control-plane propagation, credential lifetime scheduling, and authentication backoff rather than an incorrect IAM policy.
Four sources of delay
Instance profile propagation to IMDS
After AssociateIamInstanceProfile, the role must become visible through the instance metadata service. Until IMDS returns the role name, restarting the agent cannot produce usable credentials.
TOKEN=$(curl -sS -X PUT \
-H 'X-aws-ec2-metadata-token-ttl-seconds: 21600' \
http://169.254.169.254/latest/api/token)
curl -sS -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/iam/security-credentials/Refresh at the credential half-life
The SSM Agent credential refresher schedules renewal around 50% of the current credential lifetime. Core treats instance-role credentials as valid for at least one hour, so the normal refresh point can be about 30 minutes later. A console-side role change does not automatically wake a process holding valid old credentials.
Long backoff after AccessDenied
If both the instance-profile path and Default Host Management credentials fail, authentication/authorization errors use a long sleep. In the agent source, getLongSleepDuration() waits about 25 minutes plus up to five minutes of jitter.
Even if IAM becomes correct seconds later, a sleeping process does not retry early.
Runtime and process caches
The agent stores credential runtime state in identity_config.json, and some identity information is cached in memory. Removing the runtime credential cache can trigger a refresh; restarting the service rebuilds the complete process state.
Read the log before changing configuration
sudo grep -iE \
'CredentialRefresher|credential|rotation|Sleeping|RemoteRetrieve|AccessDenied' \
/var/log/amazon/ssm/amazon-ssm-agent.log | tail -50| Log text | Meaning |
|---|---|
Next credential rotation will be in ... | Waiting for the regular refresh point |
Sleeping for ... before retrying | Failed retrieval backoff |
Credential config file missing | Cache removal was detected |
Successfully connected with instance profile role credentials | Instance-profile path succeeded |
Run the built-in diagnostics as well:
sudo ssm-cli get-diagnosticsRecommended recovery order
- Wait until IMDS returns the expected role.
- Flush the cached credentials:
sudo ssm-cli flush-cached-credentials- If the command is unavailable, back up and remove the runtime file:
sudo cp -a /var/lib/amazon/ssm/runtimeconfig/identity_config.json \
/var/tmp/identity_config.json.backup 2>/dev/null || true
sudo rm -f /var/lib/amazon/ssm/runtimeconfig/identity_config.json- If the instance is still offline, restart the agent:
sudo systemctl restart amazon-ssm-agent
sudo systemctl --no-pager --full status amazon-ssm-agent- Verify from the control plane:
aws ssm describe-instance-information \
--filters 'Key=InstanceIds,Values=i-xxxxxxxxxxxxxxxxx'If it is still absent, continue with VPC endpoints, DNS, time, firewall, and service-endpoint connectivity. Credential caching explains the specific post-attachment delay; it does not explain every SSM outage.
Better automation
Attach the instance profile through the Launch Template, Auto Scaling Group, or RunInstances so it is already present when the agent first starts.
For a running-instance role change, automation should associate the profile, poll IMDS for the expected role, flush cached credentials, restart only after a timeout, and finally verify PingStatus from SSM.
Source references
- Credential refresher
- Backoff policy
- EC2 role provider
- Credential flush command
- AWS SSM Agent troubleshooting
Summary
A common sequence is: the role is not yet visible, the agent's first validation fails, and the process enters a 25–30-minute backoff. Check IMDS first, then the agent log, then flush the cache, and restart last. This sequence recovers quickly without destroying the evidence that identifies the actual layer at fault.
