October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Always On

ConfigMgr (SCCM) High Availability, Redundancy and Disaster-Recovery Options

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Microsoft Configuration Manager (the current name for SCCM) has no single high-availability switch. A resilient deployment combines a manually promoted passive site server, SQL Server high availability, redundant client-facing roles, protected content storage, and independently tested backups. Each layer solves a different failure: a passive server does not protect a failed database, and SQL replication does not restore lost content or certificates.

For a typical production primary site, use one passive site server, a supported SQL Server Always On availability group or failover cluster instance, multiple management points/distribution points/software update points/SMS Providers, a remote highly available content library, and off-site immutable backups. Validate the design against your required recovery time objective (RTO) and recovery point objective (RPO).

HA, redundancy and disaster recovery are different

High availability (HA) keeps a service running through a component failure, usually with a short RTO. Redundancy means more than one instance of a role or dependency; it may be automatic, manual, active/active or active/passive. Disaster recovery (DR) restores service after a datacenter loss, ransomware, unrecoverable corruption or destroyed infrastructure. RTO is the maximum acceptable outage; RPO is the maximum acceptable data loss.

  • A failed management point is a site-system-role outage.
  • A failed active site-server computer is a site-server HA event.
  • A lost SQL database is a database HA or recovery event.
  • A lost datacenter is a full DR event involving identity, storage, networking and every ConfigMgr dependency.

Which ConfigMgr components can be protected?

Layer What it protects What it does not provide
Passive site server Replacement for a failed CAS or primary-site server Automatic failover, database protection, or content backup
SQL Always On AG Replicated site database and planned/automatic SQL failover when configured Site-server binaries, content, package sources, certificates or WSUS files
SQL failover cluster instance (FCI) Instance-level SQL failover between cluster nodes Independent database copies or protection from shared-storage failure
Redundant site-system roles Alternate management points, distribution points, software update points and SMS Providers Recovery of a completely destroyed site
Remote content storage Availability of the content library to active and passive servers Protection unless the storage itself is replicated and backed up
ConfigMgr Backup Site Server task Site database and site-server recovery information Content library, package sources and all supporting infrastructure
Infrastructure DR or rebuild Recovery after site, database or datacenter loss Instant service continuity

Microsoft documents these capabilities separately in its high-availability options guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Passive site-server high availability

ConfigMgr supports one additional site server in passive mode for each central administration site (CAS) or primary site. Secondary sites cannot use this feature. The passive server shares the site database and content library but does not process site-server work while passive. See Microsoft’s site-server high-availability documentation.

What it protects

  • Failure of the active site’s hardware or operating system.
  • Replacement or migration of the site-server computer.
  • Some recovery cases where the database and shared content remain usable.

Prerequisites and restrictions

  • Both computers must be in the same Active Directory domain and run supported operating systems and matching ConfigMgr source versions.
  • The passive computer needs appropriate SQL permissions and access to the same remote content library.
  • Move the content library to a remote network location and grant both site-server computer accounts the required share and NTFS permissions, as described in Microsoft’s HA guidance.
  • Install the server as passive before assigning site-system roles. It cannot be a distribution point and cannot be a node in the Always On availability group that hosts the ConfigMgr database.
  • Keep the service connection point and other roles on servers that will remain available during a site-server change.

Promotion is not automatic

Only one passive site server is supported, and promotion is an administrative operation. During an unplanned outage, the passive server attempts to contact the active server for up to 30 minutes before forcibly updating the site configuration and becoming active. Clients can continue using available management points, software update points and distribution points during this interval, but site-server processing is unavailable.

  1. Verify that the passive computer can reach the SQL instance or availability-group listener and the remote content library.
  2. Connect the console to an available SMS Provider. The passive server installs an SMS Provider, which is essential if providers on the failed server are unreachable.
  3. Use the site-server promotion action. For an unplanned failure, allow the documented contact/wait behavior to complete.
  4. Check SMS Executive, site components, database connectivity, replication, management points, software update points, distribution points and console access.
  5. Repair or rebuild the former active server, then return it to passive mode only after it meets the prerequisites again.

SQL Server protection for the site database

Always On availability groups

ConfigMgr supports SQL Server Always On availability groups for CAS and primary-site databases on premises or in Azure; secondary sites cannot use an availability group. SQL Server instances in an AG require Windows Server Failover Clustering and must be members of the same WSFC cluster. The ConfigMgr database uses the FULL recovery model in an AG; non-AG configurations use SIMPLE. Review the supported-version matrix at Supported SQL Server versions, the AG preparation requirements, and Microsoft’s configuration procedure.

Synchronous commit is generally suited to low-latency local replicas and can achieve near-zero data loss, at the cost of transaction latency. Asynchronous commit suits a remote datacenter or Azure region but can leave a non-zero RPO. Failover behavior depends on replica state, quorum, listener/DNS, firewall rules and the selected SQL failover mode; it is not a ConfigMgr guarantee by itself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Confirm the ConfigMgr release, SQL version/edition and AG feature support.
  2. Place SQL instances in separate fault domains and configure WSFC, quorum and networking.
  3. Enable Always On, create or restore the site database on the intended primary replica, and create the AG.
  4. Add synchronous local and, where needed, asynchronous remote replicas.
  5. Configure the site for the AG listener and define where SQL backups and transaction-log backups run.
  6. Test planned and unplanned failover, then verify site components, SMS Provider access, replication, inventory, policy requests, updates and reporting.

Failover cluster instances

An Always On FCI moves the SQL Server instance between cluster nodes. It can be a good fit when the organization already operates WSFC and reliable shared storage. The trade-off is that shared storage remains a correlated failure domain, and an FCI does not provide independent database copies or geographic DR. Microsoft describes this model in Failover cluster instance for the site database.

Redundant site-system roles

Management points

Deploy management points across hosts, racks or datacenters, then configure boundary groups and fallback relationships so clients can discover an alternate. Replicate certificates, authentication settings and capacity; two servers on one failed host are not datacenter redundancy.

Distribution points

Place distribution points according to content volume, concurrent downloads, branch bandwidth, pull-distribution needs and WAN-outage behavior. A rebuilt DP may require a large redistribution and therefore a long RTO. Protect the underlying operating-system and content-library data where rapid recovery matters.

Software update points

Multiple SUPs provide alternatives, but WSUS databases, synchronization partners and upstream/downstream relationships still need protection and testing. A second SUP alone does not guarantee continuous update service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

SMS Providers

Install multiple SMS Providers so administrators can connect when the active site server is unavailable. If every provider is offline, the console cannot manage or promote the site.

Content storage is part of the service

The site database contains metadata about packages, applications, deployments and content locations; it does not contain the actual content files. ConfigMgr’s built-in backup also omits the content library and package/application source files. The distinction is critical:

  • Content library: ConfigMgr-managed files used by the site and distribution points.
  • Package and application sources: Original files needed to create or update content.
  • Distribution-point copies: Replicas that can be rebuilt, but potentially with substantial WAN traffic.
  • Database metadata: Records that do not restore missing files.

Use a highly available SMB share, clustered file server, enterprise NAS, replicated storage or another storage service only after validating SMB semantics, latency, permissions and support for the exact ConfigMgr version. Replicate the storage to another failure domain and keep immutable or offline backups. Restoring the database without restoring the content library and sources leaves deployments present but unusable or unupdatable. See Backup sites.

Backups and complete site recovery

What to back up

  • ConfigMgr site database and, where applicable, SQL transaction logs.
  • ConfigMgr site-server backup set.
  • Content library and all package/application source shares.
  • WSUS database and content, reporting databases and custom reports.
  • Certificates and private keys, boot images, drivers and task-sequence dependencies.
  • Custom configuration files, service-account information and documentation of site code, database name, ports, boundaries and dependencies.

Keep copies outside the production server and storage system. Immutable, offline or logically isolated copies are necessary for ransomware resilience because replication can copy encryption, corruption or malicious deletion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recovery runbook

When a failure exceeds normal HA, preserve evidence and identify whether it is hardware loss, corruption, ransomware or datacenter loss. Establish the last known-good backup and its RPO, then:

  1. Provision a replacement server with the required identity, operating system, networking and permissions.
  2. Obtain matching CD.Latest source files.
  3. Run ConfigMgr Setup and select Recover a site, choosing the documented site-server and database recovery options.
  4. Restore or reconnect the site database, then restore the content library and package sources.
  5. Recover WSUS, reporting and other supporting databases as required.
  6. Reenter passwords for accounts reset during recovery and complete the post-recovery action list.
  7. Validate SMS Provider access, site components, database replication, certificates, client communication, content distribution, software updates, reporting and deployments.

For unattended recovery, Microsoft documents setup.exe /script C:TempConfigMgrUnattend.ini. A full site-server replacement generally requires the failed server’s hostname and fully qualified domain name. Restoring to a newer SQL Server version may be supported, but changing SQL edition during recovery is not supported by the cited guidance; verify the current matrix before changing either. See Recover a Configuration Manager site.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Secondary sites require a different plan

Secondary sites cannot use SQL Always On and their databases are not covered by the normal primary/CAS site-backup model. Recovery is performed by reinstalling or recovering the secondary from its parent primary site. The replacement needs the same installation path, FQDN and SQL configuration. Do not apply primary-site passive-server or AG designs to a secondary site.

Rank #4
Sale
Mastering Active Directory: Design, deploy, and protect Active Directory Domain Services for Windows Server 2022
  • Mastering Active Directory: Design, deploy, and protect Active Directory Domain Services for Windows Server 2022, 3rd Edition
  • ABIS BOOK
  • Packt Publishing

Reference architectures

Cost-conscious primary site

Use one primary site, multiple client-facing roles where needed, the ConfigMgr backup task, independent SQL and file backups, and an off-site immutable copy. This suits environments with a several-hour RTO and limited budget; recovery depends on restore speed and operator skill.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Balanced production HA

Use one active and one passive site server, a SQL AG, remote highly available content, redundant management points, DPs, SUPs and SMS Providers, plus tested off-site backups. This is appropriate for many medium and large environments, but passive promotion remains manual.

Local HA plus remote DR

Combine a passive site server, synchronous local SQL replica, asynchronous replica in another datacenter or Azure region, replicated content storage and independent backups. This can deliver low RPO, but DNS, Active Directory, certificates, WSUS, storage and network latency often determine the real regional recovery time.

Azure-assisted deployment

ConfigMgr can run on Azure virtual machines. The site database must use SQL Server in a VM, not Azure SQL Database; Microsoft recommends SQL Server AGs for database HA. Azure VM availability features can support redundant roles when the complete topology is validated. See the Configuration Manager on Azure FAQ. Azure Site Recovery can replicate and orchestrate VM infrastructure, but it does not automatically restore ConfigMgr content, certificates or application state; integrate it with the supported ConfigMgr recovery process. See the Azure Site Recovery FAQ.

Choose by RTO, RPO and operating capability

Requirement Practical starting point
Several-hour RTO and limited budget Native site backup, independent SQL/file backups and a tested rebuild
Fast recovery from site-server hardware failure Passive site server
Fast database failover Supported AG or FCI
Minimal local database loss Synchronous AG replica
Datacenter-level database DR Asynchronous AG replica or supported infrastructure DR
Client service during site-server outage Multiple management points, DPs and SUPs
Rapid content recovery Remote HA content storage plus protected source shares
Cyber recovery Immutable/offline backups, not replication alone
Lowest operational complexity Fewer HA layers with stronger recovery testing

For every dependency, document its failure domain, RTO, RPO, failover mode, operational burden, test method, security controls and total cost. Two VMs on the same host or storage array are not independent redundancy.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test the design before an outage

  • Perform a planned SQL failover and confirm listener resolution, Service Broker, permissions and ConfigMgr processing.
  • Promote the passive site server in a maintenance window and verify SMS Providers, site components and console access.
  • Restore the content library and package sources into isolated storage.
  • Run a complete isolated site recovery using the documented backups and CD.Latest.
  • Exercise ransomware recovery with immutable copies.
  • Measure time to restore the database, content, certificates and credentials, and time until clients receive policy and deployments work.
  • Document failback, patching, ownership, escalation contacts and every dependency.

Commercial and platform considerations

ConfigMgr’s HA capabilities are part of supported deployments; licensing is normally handled through Microsoft agreements. SQL AGs and FCIs add SQL Server, Windows Server, WSFC, hosts, storage, monitoring and specialist operating costs. Azure VM pricing varies by size, disks, region, reservations, licensing and data transfer; use the Azure pricing calculator rather than an undated estimate.

Azure Backup supports documented SQL Server availability-group scenarios, subject to restrictions; see Back up SQL Server Always On availability groups. Azure Site Recovery pricing is listed at Azure Site Recovery pricing. Rubrik and Commvault can provide broader immutable backup and cyber-recovery coverage, but pricing is generally quote-based: Rubrik Microsoft protection and Commvault backup and recovery for Microsoft Azure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.