Writing

Building a Safe Lifecycle for Hybrid Windows Devices

How I automated the identification, quarantine, and removal of inactive devices across Active Directory, Microsoft Entra ID, and Intune

11 min read
Security Automation Published
Active DirectoryMicrosoft Entra IDMicrosoft IntunePowerShellFastAPIIdentity Security

Hybrid device management has a lifecycle problem.

A computer can be replaced, reinstalled, abandoned, or disconnected from the corporate environment while its records remain across Active Directory, Microsoft Entra ID, and Microsoft Intune. Over time, stale objects reduce inventory accuracy, complicate investigations, and increase the number of identities and endpoints that administrators must review.

I built DeviceLifecycle to address this problem without treating deletion as a simple cleanup operation.

The project is a PowerShell-based automation that evaluates device activity, correlates identities across Microsoft platforms when those records exist, and moves eligible devices through a controlled lifecycle. I also built DeviceLifecycle-API, a separate read-only extension that exposes reports and logs to authorized internal systems.

Source code:

The engineering problem

In a hybrid Microsoft environment, one physical computer may be represented by several records:

  • an Active Directory computer object;
  • a Microsoft Entra ID device object;
  • a Microsoft Intune managed-device record;
  • other enrollment or provisioning records;
  • locally stored lifecycle state, reports, and logs.

These records do not necessarily appear, update, or disappear at the same time.

An Intune record may be removed after a Retire action while the Active Directory object remains in quarantine. An Entra ID object may still exist after its on-premises computer has stopped communicating. A computer may also exist only in Active Directory because it was never enrolled in Intune or because a cloud record has already been removed.

This means that a safe lifecycle process cannot rely on a single timestamp, on the device name alone, or on the assumption that every source must always contain a record.

The automation must answer a more precise question:

What evidence actually exists for this device, and is the available evidence reliable enough to justify changing its identity?

That question shaped the entire design.

Design principle: fail closed without confusing absence with inconsistency

The central rule of DeviceLifecycle is to automate routine cases while stopping on evidence that is ambiguous, inconsistent, or unusable.

The important distinction is that a missing cloud record is not automatically an inconsistency.

If a query to Entra ID or Intune succeeds and returns zero matching records, that source is treated as unavailable evidence. The lifecycle continues using the sources that actually exist. If a source does exist, however, its identity and activity data must pass validation before an administrative action can be allowed.

That is different from a Graph query failure, a duplicate match, an identifier mismatch, or a missing activity timestamp on an existing record. Those conditions remain fail-closed.

This distinction became important in production: requiring a match in every platform caused old Active Directory objects with legitimately absent cloud records to remain in manual review indefinitely, even when the available evidence clearly supported quarantine.

Architecture

I separated the solution into two repositories with different trust boundaries.

DeviceLifecycle

The main project owns the privileged workflow. It:

  • inventories Active Directory devices and available Entra ID/Intune records;
  • correlates identities across the platforms when cloud records exist;
  • evaluates inactivity thresholds using validated available signals;
  • produces CSV reports and execution logs;
  • stores lifecycle state in JSON;
  • quarantines or removes eligible devices according to the configured mode;
  • supports recovery of quarantined devices.

DeviceLifecycle-API

The API is an optional extension. It:

  • reads the latest CSV report and execution log;
  • authenticates consumers with an API key;
  • exposes CSV, JSON, metadata, and log endpoints;
  • performs no lifecycle action;
  • does not contact Active Directory, Entra ID, Intune, or Microsoft Graph.

The separation keeps reporting consumers outside the control plane.

Active Directory ----\
Microsoft Entra ID ----> DeviceLifecycle ---> CSV reports
Microsoft Intune ----/                       Execution logs
                                             Persistent state
                                                    |
                                                    v
                                          DeviceLifecycle-API
                                                    |
                                  Dashboards, monitoring, audits,
                                    inventory and internal tools

If the API is unavailable, DeviceLifecycle continues operating normally. The lifecycle engine remains the authoritative producer of state and reports.

Correlating identities safely

Computer names are useful labels, but they are weak identity keys.

A device may be renamed, reinstalled, or replaced by another computer that receives the same standardized name. For this reason, DeviceLifecycle uses stable identifiers whenever a cloud record is available:

  1. the Active Directory computer SID is matched to Entra ID onPremisesSecurityIdentifier;
  2. when exactly one valid Entra record exists, its deviceId is matched to Intune azureADDeviceId;
  3. activity timestamps from every existing validated source are evaluated against the configured thresholds.

The current activity policy is intentionally asymmetric:

  • Active Directory is always required because it is the authoritative local inventory source for this workflow;
  • a missing Entra ID record is treated as unavailable evidence;
  • a missing Intune record is treated as unavailable evidence;
  • an existing Entra or Intune record must be unique, identity-consistent, and provide the required activity timestamp;
  • recent activity in any existing source prevents a stale AD timestamp from producing a false quarantine candidate.

Records are sent to manual review when, for example:

  • multiple records match the same identifier;
  • a required timestamp is unavailable on a source that exists;
  • identifiers are inconsistent;
  • the object is already in an unexpected administrative state;
  • a protection or exclusion rule blocks automated action.

A successful source query returning zero records and a source query that fails are not treated as the same event. The former means “no evidence from this source”; the latter means the system could not establish whether that evidence exists and therefore remains fail-closed.

A staged lifecycle instead of immediate deletion

The workflow uses multiple stages.

1. Attention

After the configured attention threshold — 75 inactive days by default — a device can be classified as approaching the quarantine threshold.

No change is made. This stage gives administrators time to investigate devices before they become eligible for quarantine.

2. Quarantine candidate

After 90 inactive days across all validated signals that actually exist, a device may become a quarantine candidate.

When Quarantine or Enforce is enabled, the automation can:

  • send a Retire command to Intune when a managed-device record exists;
  • disable the computer account in Active Directory;
  • move the object to a dedicated quarantine organizational unit;
  • record identifiers and lifecycle timestamps in persistent state.

The JSON state file is essential because Intune may remove the managedDevice record after processing Retire. The lifecycle must remain traceable even after a source record disappears.

3. Permanent removal

After the configured quarantine retention period — 30 additional days by default — the device may become eligible for final removal.

In Enforce mode, the workflow can:

  • delete the Active Directory computer object;
  • remove any remaining Intune record when applicable;
  • trigger a Microsoft Entra Connect delta synchronization;
  • verify whether a residual Entra ID object remains;
  • remove the residual Entra object after an additional grace period.

This sequence creates several opportunities to detect mistakes before an irreversible action occurs.

A critical Entra Connect detail

The quarantine organizational unit must remain inside the Microsoft Entra Connect synchronization scope.

If the OU is excluded, moving a computer into quarantine may cause the corresponding Entra ID object to disappear immediately. That would bypass the intended retention period and change the lifecycle behavior.

This was an important design lesson: an operation that appears safe inside Active Directory may have an unintended cloud-side effect because of directory synchronization.

Automation must account for the complete system, not only the command being executed.

Progressive operating modes

DeviceLifecycle provides three operating modes.

ReportOnly

This is the default and safest deployment mode.

The automation inventories and evaluates devices but performs no administrative change. Reports include outcomes such as:

  • active devices;
  • devices approaching the quarantine threshold;
  • quarantine candidates;
  • ambiguous correlations;
  • missing activity timestamps on existing sources;
  • protected or excluded records;
  • records requiring manual review.

A cloud source with no matching record is represented by the absence of that source’s identifiers and timestamp, not by an artificial MissingEntraMatch or MissingIntuneMatch review reason.

Quarantine

This mode enables reversible containment actions. It can retire an existing Intune record, disable the Active Directory account, and move the object to quarantine, but it does not perform final deletion.

Enforce

This mode enables the complete lifecycle, including permanent removal after the configured retention and grace periods.

I promote environments to Enforce only after reviewing ReportOnly, testing the exact path with Enforce -WhatIf, and validating a controlled first real execution.

Separate lifecycle and snapshot tasks

The installed service model uses two scheduled tasks with different responsibilities.

Primary lifecycle task

{OrganizationName} - Device Lifecycle runs daily at the configured TaskTime and uses the configured operational mode. Once production validation is complete, this task may run persistently in Enforce.

Snapshot task

{OrganizationName} - Device Lifecycle Snapshot runs periodically — every 30 minutes by default — and always forces -ModeOverride ReportOnly.

The snapshot task exists so that DeviceLifecycle-Latest.csv remains fresh for inventory consumers, dashboards, monitoring, and the read-only API without causing administrative changes every time the report is refreshed.

Both tasks use a locking wrapper to prevent concurrent lifecycle executions.

Safety controls

The safety mechanisms are part of the core design, not additional features.

The project includes:

  • ReportOnly as the default deployment mode;
  • SID/GUID-based identity correlation when cloud records exist;
  • manual review for ambiguous or inconsistent records;
  • manual review for missing activity timestamps on sources that exist;
  • exclusion and protection rules for administrative edge cases;
  • a configurable maximum number of actions per run;
  • persistent JSON state;
  • CSV reports and execution logs;
  • delayed cleanup of residual Entra ID objects;
  • PowerShell -WhatIf support;
  • explicit validation, recovery, and uninstall procedures.

The primary scheduled task runs as NT AUTHORITY\SYSTEM. In Active Directory, it therefore operates through the computer account of the server hosting the automation.

Instead of assigning broad administrative privileges, permissions can be delegated only on the managed and quarantine organizational units. This follows the principle of least privilege and reduces the impact of a compromised automation host.

Recovery is deliberately manual

A quarantined device is not automatically restored merely because it communicates again.

A later Intune synchronization may only indicate that a device received an earlier management action. It does not prove that the computer should return to production.

Recovery therefore requires an explicit administrative decision.

The recovery script can:

  • re-enable the Active Directory computer account;
  • move it back to the production OU;
  • remove lifecycle markers and quarantine state;
  • trigger directory synchronization when configured;
  • support revalidation of join and enrollment state.

It also supports -WhatIf, allowing the recovery path to be reviewed before execution.

Automated quarantine and deliberate recovery provide a safer balance than silently reversing a security decision.

Reporting and operational state

Each execution produces structured evidence:

  • a current DeviceLifecycle-Latest.csv report;
  • per-execution reports;
  • execution logs;
  • a persistent JSON state file.

The state file preserves identifiers and lifecycle timestamps between runs. This is necessary because the workflow itself may remove records from source systems.

The reports can be reviewed directly on the server or consumed through the optional API.

The read-only API extension

I created DeviceLifecycle-API with FastAPI and Uvicorn to make operational data available to internal dashboards and services without giving them filesystem access or lifecycle privileges.

The main endpoints are:

EndpointPurpose
/api/v1/healthService and source-file availability
/api/v1/metadataReport and log metadata
/api/v1/report.csvOriginal CSV report
/api/v1/reportCSV converted to JSON
/api/v1/log?lines=500Latest log lines
/api/v1/log/fileComplete current log

The API deliberately has no endpoints for quarantine, deletion, restoration, or device modification.

This prevents a reporting integration from becoming a privileged management interface.

API security model

The API is intended for internal use, but the internal network is not treated as inherently trusted.

Its controls include:

  • API-key authentication through X-API-Key;
  • randomly generated API keys;
  • constant-time key comparison;
  • runtime secrets stored outside the repository;
  • network allowlisting;
  • disabled interactive API documentation in production runtime;
  • rejection of client-controlled filesystem paths;
  • stable file reads to avoid partially updated responses;
  • Cache-Control: no-store responses;
  • configurable log-tail limits;
  • request logging without exposing authentication secrets.

For routed or untrusted networks, the service should be placed behind an HTTPS reverse proxy.

Production validation

On August 28, 2026, I validated the updated correlation policy in a real environment.

The validation sequence was deliberately conservative:

  1. generate a fresh ReportOnly snapshot and compare the classifications with the previous behavior;
  2. confirm that old AD objects without Entra ID or Intune records became normal lifecycle candidates instead of permanent manual-review items;
  3. confirm that valid recent cloud activity still prevented quarantine;
  4. execute Enforce -WhatIf and review every proposed action;
  5. execute a controlled real Enforce run under the configured MaximumActionsPerRun limit;
  6. verify the resulting report and quarantine state before allowing the daily task to remain in Enforce.

The controlled batch completed as expected. The primary lifecycle task could therefore be promoted to persistent Enforce, while the frequent snapshot task continued running only in ReportOnly.

I intentionally do not publish the real environment’s device names, identifiers, or counts in the public repository.

What this project taught me

The hardest part of automation is not replacing commands with code. It is defining what kind of missing information matters and when software has enough evidence to make a decision safely.

A Windows device lifecycle must account for:

  • eventual consistency;
  • legitimately absent records;
  • duplicated or inconsistent records;
  • synchronization delays;
  • reused computer names;
  • asynchronous retirement operations;
  • dependencies between on-premises and cloud systems.

One of the most useful lessons from the production rollout was that a conservative system should not simply classify every missing record as an error. A better model asks whether the source was queried successfully, whether a record exists, whether that record is unique and trustworthy, and what the remaining evidence says.

The project also reinforced the value of separating capabilities. DeviceLifecycle owns the privileged workflow. DeviceLifecycle-API provides operational visibility. The snapshot task provides fresh inventory without receiving lifecycle authority of its own.

The principle I intend to carry into future security automation work is:

Reliable automation should distinguish unavailable evidence from contradictory evidence, preserve operational context, and stop when the evidence that does exist cannot be trusted.

Source code

The complete projects and their installation documentation are available on GitHub: