The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Microsoft Configuration Manager (the current name for SCCM) has no single high-availability switch. A resilient deployment combines a manually promoted passive site server, SQL Server high availability, redundant client-facing roles, protected content storage, and independently tested backups. Each layer solves a different failure: a passive server does not protect a failed database, and SQL replication does not restore lost content or certificates.
For a typical production primary site, use one passive site server, a supported SQL Server Always On availability group or failover cluster instance, multiple management points/distribution points/software update points/SMS Providers, a remote highly available content library, and off-site immutable backups. Validate the design against your required recovery time objective (RTO) and recovery point objective (RPO).
HA, redundancy and disaster recovery are different
High availability (HA) keeps a service running through a component failure, usually with a short RTO. Redundancy means more than one instance of a role or dependency; it may be automatic, manual, active/active or active/passive. Disaster recovery (DR) restores service after a datacenter loss, ransomware, unrecoverable corruption or destroyed infrastructure. RTO is the maximum acceptable outage; RPO is the maximum acceptable data loss.
- A failed management point is a site-system-role outage.
- A failed active site-server computer is a site-server HA event.
- A lost SQL database is a database HA or recovery event.
- A lost datacenter is a full DR event involving identity, storage, networking and every ConfigMgr dependency.
Which ConfigMgr components can be protected?
| Layer | What it protects | What it does not provide |
|---|---|---|
| Passive site server | Replacement for a failed CAS or primary-site server | Automatic failover, database protection, or content backup |
| SQL Always On AG | Replicated site database and planned/automatic SQL failover when configured | Site-server binaries, content, package sources, certificates or WSUS files |
| SQL failover cluster instance (FCI) | Instance-level SQL failover between cluster nodes | Independent database copies or protection from shared-storage failure |
| Redundant site-system roles | Alternate management points, distribution points, software update points and SMS Providers | Recovery of a completely destroyed site |
| Remote content storage | Availability of the content library to active and passive servers | Protection unless the storage itself is replicated and backed up |
| ConfigMgr Backup Site Server task | Site database and site-server recovery information | Content library, package sources and all supporting infrastructure |
| Infrastructure DR or rebuild | Recovery after site, database or datacenter loss | Instant service continuity |
Microsoft documents these capabilities separately in its high-availability options guidance.
#1 Best Overall
Passive site-server high availability
ConfigMgr supports one additional site server in passive mode for each central administration site (CAS) or primary site. Secondary sites cannot use this feature. The passive server shares the site database and content library but does not process site-server work while passive. See Microsoft’s site-server high-availability documentation.
What it protects
- Failure of the active site’s hardware or operating system.
- Replacement or migration of the site-server computer.
- Some recovery cases where the database and shared content remain usable.
Prerequisites and restrictions
- Both computers must be in the same Active Directory domain and run supported operating systems and matching ConfigMgr source versions.
- The passive computer needs appropriate SQL permissions and access to the same remote content library.
- Move the content library to a remote network location and grant both site-server computer accounts the required share and NTFS permissions, as described in Microsoft’s HA guidance.
- Install the server as passive before assigning site-system roles. It cannot be a distribution point and cannot be a node in the Always On availability group that hosts the ConfigMgr database.
- Keep the service connection point and other roles on servers that will remain available during a site-server change.
Promotion is not automatic
Only one passive site server is supported, and promotion is an administrative operation. During an unplanned outage, the passive server attempts to contact the active server for up to 30 minutes before forcibly updating the site configuration and becoming active. Clients can continue using available management points, software update points and distribution points during this interval, but site-server processing is unavailable.
- Verify that the passive computer can reach the SQL instance or availability-group listener and the remote content library.
- Connect the console to an available SMS Provider. The passive server installs an SMS Provider, which is essential if providers on the failed server are unreachable.
- Use the site-server promotion action. For an unplanned failure, allow the documented contact/wait behavior to complete.
- Check SMS Executive, site components, database connectivity, replication, management points, software update points, distribution points and console access.
- Repair or rebuild the former active server, then return it to passive mode only after it meets the prerequisites again.
SQL Server protection for the site database
Always On availability groups
ConfigMgr supports SQL Server Always On availability groups for CAS and primary-site databases on premises or in Azure; secondary sites cannot use an availability group. SQL Server instances in an AG require Windows Server Failover Clustering and must be members of the same WSFC cluster. The ConfigMgr database uses the FULL recovery model in an AG; non-AG configurations use SIMPLE. Review the supported-version matrix at Supported SQL Server versions, the AG preparation requirements, and Microsoft’s configuration procedure.
Synchronous commit is generally suited to low-latency local replicas and can achieve near-zero data loss, at the cost of transaction latency. Asynchronous commit suits a remote datacenter or Azure region but can leave a non-zero RPO. Failover behavior depends on replica state, quorum, listener/DNS, firewall rules and the selected SQL failover mode; it is not a ConfigMgr guarantee by itself.
- Confirm the ConfigMgr release, SQL version/edition and AG feature support.
- Place SQL instances in separate fault domains and configure WSFC, quorum and networking.
- Enable Always On, create or restore the site database on the intended primary replica, and create the AG.
- Add synchronous local and, where needed, asynchronous remote replicas.
- Configure the site for the AG listener and define where SQL backups and transaction-log backups run.
- Test planned and unplanned failover, then verify site components, SMS Provider access, replication, inventory, policy requests, updates and reporting.
Failover cluster instances
An Always On FCI moves the SQL Server instance between cluster nodes. It can be a good fit when the organization already operates WSFC and reliable shared storage. The trade-off is that shared storage remains a correlated failure domain, and an FCI does not provide independent database copies or geographic DR. Microsoft describes this model in Failover cluster instance for the site database.
Rank #2
Redundant site-system roles
Management points
Deploy management points across hosts, racks or datacenters, then configure boundary groups and fallback relationships so clients can discover an alternate. Replicate certificates, authentication settings and capacity; two servers on one failed host are not datacenter redundancy.
Distribution points
Place distribution points according to content volume, concurrent downloads, branch bandwidth, pull-distribution needs and WAN-outage behavior. A rebuilt DP may require a large redistribution and therefore a long RTO. Protect the underlying operating-system and content-library data where rapid recovery matters.
Software update points
Multiple SUPs provide alternatives, but WSUS databases, synchronization partners and upstream/downstream relationships still need protection and testing. A second SUP alone does not guarantee continuous update service.
Recommended Free Tools
SMS Providers
Install multiple SMS Providers so administrators can connect when the active site server is unavailable. If every provider is offline, the console cannot manage or promote the site.
Content storage is part of the service
The site database contains metadata about packages, applications, deployments and content locations; it does not contain the actual content files. ConfigMgr’s built-in backup also omits the content library and package/application source files. The distinction is critical:
Rank #3
- Content library: ConfigMgr-managed files used by the site and distribution points.
- Package and application sources: Original files needed to create or update content.
- Distribution-point copies: Replicas that can be rebuilt, but potentially with substantial WAN traffic.
- Database metadata: Records that do not restore missing files.
Use a highly available SMB share, clustered file server, enterprise NAS, replicated storage or another storage service only after validating SMB semantics, latency, permissions and support for the exact ConfigMgr version. Replicate the storage to another failure domain and keep immutable or offline backups. Restoring the database without restoring the content library and sources leaves deployments present but unusable or unupdatable. See Backup sites.
Backups and complete site recovery
What to back up
- ConfigMgr site database and, where applicable, SQL transaction logs.
- ConfigMgr site-server backup set.
- Content library and all package/application source shares.
- WSUS database and content, reporting databases and custom reports.
- Certificates and private keys, boot images, drivers and task-sequence dependencies.
- Custom configuration files, service-account information and documentation of site code, database name, ports, boundaries and dependencies.
Keep copies outside the production server and storage system. Immutable, offline or logically isolated copies are necessary for ransomware resilience because replication can copy encryption, corruption or malicious deletion.
Recovery runbook
When a failure exceeds normal HA, preserve evidence and identify whether it is hardware loss, corruption, ransomware or datacenter loss. Establish the last known-good backup and its RPO, then:
- Provision a replacement server with the required identity, operating system, networking and permissions.
- Obtain matching
CD.Latestsource files. - Run ConfigMgr Setup and select Recover a site, choosing the documented site-server and database recovery options.
- Restore or reconnect the site database, then restore the content library and package sources.
- Recover WSUS, reporting and other supporting databases as required.
- Reenter passwords for accounts reset during recovery and complete the post-recovery action list.
- Validate SMS Provider access, site components, database replication, certificates, client communication, content distribution, software updates, reporting and deployments.
For unattended recovery, Microsoft documents setup.exe /script C:TempConfigMgrUnattend.ini. A full site-server replacement generally requires the failed server’s hostname and fully qualified domain name. Restoring to a newer SQL Server version may be supported, but changing SQL edition during recovery is not supported by the cited guidance; verify the current matrix before changing either. See Recover a Configuration Manager site.
Secondary sites require a different plan
Secondary sites cannot use SQL Always On and their databases are not covered by the normal primary/CAS site-backup model. Recovery is performed by reinstalling or recovering the secondary from its parent primary site. The replacement needs the same installation path, FQDN and SQL configuration. Do not apply primary-site passive-server or AG designs to a secondary site.
Rank #4
- Mastering Active Directory: Design, deploy, and protect Active Directory Domain Services for Windows Server 2022, 3rd Edition
- ABIS BOOK
- Packt Publishing
Reference architectures
Cost-conscious primary site
Use one primary site, multiple client-facing roles where needed, the ConfigMgr backup task, independent SQL and file backups, and an off-site immutable copy. This suits environments with a several-hour RTO and limited budget; recovery depends on restore speed and operator skill.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesBalanced production HA
Use one active and one passive site server, a SQL AG, remote highly available content, redundant management points, DPs, SUPs and SMS Providers, plus tested off-site backups. This is appropriate for many medium and large environments, but passive promotion remains manual.
Local HA plus remote DR
Combine a passive site server, synchronous local SQL replica, asynchronous replica in another datacenter or Azure region, replicated content storage and independent backups. This can deliver low RPO, but DNS, Active Directory, certificates, WSUS, storage and network latency often determine the real regional recovery time.
Azure-assisted deployment
ConfigMgr can run on Azure virtual machines. The site database must use SQL Server in a VM, not Azure SQL Database; Microsoft recommends SQL Server AGs for database HA. Azure VM availability features can support redundant roles when the complete topology is validated. See the Configuration Manager on Azure FAQ. Azure Site Recovery can replicate and orchestrate VM infrastructure, but it does not automatically restore ConfigMgr content, certificates or application state; integrate it with the supported ConfigMgr recovery process. See the Azure Site Recovery FAQ.
Choose by RTO, RPO and operating capability
| Requirement | Practical starting point |
|---|---|
| Several-hour RTO and limited budget | Native site backup, independent SQL/file backups and a tested rebuild |
| Fast recovery from site-server hardware failure | Passive site server |
| Fast database failover | Supported AG or FCI |
| Minimal local database loss | Synchronous AG replica |
| Datacenter-level database DR | Asynchronous AG replica or supported infrastructure DR |
| Client service during site-server outage | Multiple management points, DPs and SUPs |
| Rapid content recovery | Remote HA content storage plus protected source shares |
| Cyber recovery | Immutable/offline backups, not replication alone |
| Lowest operational complexity | Fewer HA layers with stronger recovery testing |
For every dependency, document its failure domain, RTO, RPO, failover mode, operational burden, test method, security controls and total cost. Two VMs on the same host or storage array are not independent redundancy.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Test the design before an outage
- Perform a planned SQL failover and confirm listener resolution, Service Broker, permissions and ConfigMgr processing.
- Promote the passive site server in a maintenance window and verify SMS Providers, site components and console access.
- Restore the content library and package sources into isolated storage.
- Run a complete isolated site recovery using the documented backups and
CD.Latest. - Exercise ransomware recovery with immutable copies.
- Measure time to restore the database, content, certificates and credentials, and time until clients receive policy and deployments work.
- Document failback, patching, ownership, escalation contacts and every dependency.
Commercial and platform considerations
ConfigMgr’s HA capabilities are part of supported deployments; licensing is normally handled through Microsoft agreements. SQL AGs and FCIs add SQL Server, Windows Server, WSFC, hosts, storage, monitoring and specialist operating costs. Azure VM pricing varies by size, disks, region, reservations, licensing and data transfer; use the Azure pricing calculator rather than an undated estimate.
Azure Backup supports documented SQL Server availability-group scenarios, subject to restrictions; see Back up SQL Server Always On availability groups. Azure Site Recovery pricing is listed at Azure Site Recovery pricing. Rubrik and Commvault can provide broader immutable backup and cyber-recovery coverage, but pricing is generally quote-based: Rubrik Microsoft protection and Commvault backup and recovery for Microsoft Azure.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




