Safely Upgrading PV, ENA, and NVMe Drivers on EC2 Windows Domain Controllers
Safely Upgrading PV, ENA, and NVMe Drivers on EC2 Windows Domain Controllers
Moving an older Windows Server domain controller from an early EC2 virtualization stack to Nitro may require PV, ENA, and NVMe driver work. Unlike an ordinary member server, a DC also carries AD DS, SYSVOL, FSMO, DSRM, and System State recovery requirements.
The following process assumes two domain controllers. Hostnames, volumes, and regions are examples; validate the exact Windows version, instance generation, and current AWS packages before execution.
Change principles
- Change one DC at a time and keep the other healthy.
- Move FSMO roles away from the DC under maintenance.
- Validate AD, DNS, SYSVOL, and replication before touching drivers.
- Preinstall ENA/NVMe support before moving to Nitro.
- Keep AMI/EBS snapshots and System State, but do not treat a snapshot as the only AD recovery plan.
- Confirm the DSRM password and an offline/serial recovery path.
Health checks
Run on both DCs:
dcdiag /e /c /v /f:C:\Temp\dcdiag-before.txt
repadmin /replsummary
repadmin /showrepl * /csv > C:\Temp\repl-before.csv
net shareConfirm SYSVOL and NETLOGON shares and resolve existing AD DS, DNS, DFSR, storage, or event-log errors before the maintenance window.
Record drivers and platform details:
Get-CimInstance Win32_ComputerSystem | Select-Object Manufacturer,Model
Get-CimInstance Win32_PnPSignedDriver |
Where-Object {$_.DeviceName -match 'Amazon|ENA|NVMe|Xen'} |
Select-Object DeviceName,DriverVersion,DriverDateTransfer FSMO and verify DSRM
Move-ADDirectoryServerOperationMasterRole `
-Identity 'DC2' -OperationMasterRole 0,1,2,3,4
netdom query fsmoConfirm or reset the local DSRM password through ntdsutil. DSRM uses a local recovery account, not necessarily the domain-administrator password. Store it in the approved secrets process, never in scripts or plaintext change notes.
Backups
Prepare at least three layers:
- An EC2 AMI or snapshots of all related EBS volumes.
- A Windows Server Backup System State copy.
- A verified healthy, writable peer domain controller.
wbadmin start systemstatebackup -backuptarget:E: -quiet
wbadmin get versions -backuptarget:E:AD restore still has VM-Generation ID and supported-recovery constraints. Do not clone a snapshot and bring two identical domain-controller identities online.
Prepare target drivers
Use official AWS distribution locations, verify signatures and hashes, and test the same packages on a clone of the production OS. ENA and NVMe can be staged before the platform move so Windows can discover the network and boot volume on Nitro.
pnputil /enum-drivers | Select-String -Pattern 'Amazon|ENA|NVMe' -Context 0,5
Get-PnpDevice -PresentOnly | Where-Object Status -ne 'OK'PV component upgrade/removal may require DSRM when legacy Xen components or filter drivers are involved. Never remove the active storage driver without a confirmed recovery route.
Staged sequence
For the first DC:
- Verify FSMO has moved.
- Verify replication, DNS, and SYSVOL.
- Create System State and EC2 backups.
- Preinstall ENA/NVMe.
- Upgrade PV components according to the supported procedure.
- Stop the instance and change the instance family/platform.
- Boot and validate storage, network, time, and AD.
- Observe stability before touching the second DC.
Only repeat the process for the second DC after the first has fully rejoined and replication is healthy.
Post-boot acceptance
Get-Disk
Get-NetAdapter
Get-NetIPConfiguration
Get-PnpDevice -PresentOnly | Where-Object Status -ne 'OK'
dcdiag /test:dns /v
repadmin /replsummary
repadmin /syncall /AdeP
nltest /dsregdns
net shareAlso validate SSM/CloudWatch, Windows time, business ports, event logs, and stability after another reboot. “The installer succeeded” is not acceptance for a domain controller.
Recovery paths
If DSRM works, roll back the recent driver/component, inspect SetupAPI/System/Directory Service logs, and verify boot/storage/network drivers before returning one DC to normal mode.
If Windows does not boot, stop repeated attempts, attach a copy of the root volume to a rescue instance in the same Availability Zone, preserve evidence, and repair drivers or the registry offline. Restore from the pre-change image when appropriate while following supported AD virtualization recovery rules.
If Windows boots but ENA does not, use Serial Console or offline repair and verify that the target driver is in Driver Store. Loss of RDP does not by itself mean AD data is corrupt.
Summary
The essential control is reducing the failure domain: transfer FSMO, change one DC at a time, pre-stage target drivers, retain layered backups, and validate every layer. Once PV/ENA/NVMe work is embedded in the AD recovery design, a failed driver or platform change has a defined and testable return path.
