Ready to read this content aloud.
Executive summary
This case study documents a private device and asset integration platform built to reconcile information that lives in systems with fundamentally different responsibilities.
The environment already had technical device evidence, Microsoft Entra ID and Intune, Snipe-IT for patrimonial control, and an internal operations portal. The difficult part was not connecting APIs. It was deciding which system is allowed to be authoritative for each fact, how to represent disagreement, and how to execute external mutations without turning transient failures into duplicate or contradictory operations.
The resulting architecture uses a local Device Registry as the control state for correlation, temporal links, possession, reviews, operations, and audit history. The internal portal acts as the human control plane. External systems remain authoritative only for the domains they actually own.
The production source, organization identity, real device names, users, asset tags, IDs, API paths, runtime locations, feature-flag names, and infrastructure topology remain private. Examples in this article are intentionally generic.
My role
I designed and implemented the integration architecture end to end: the domain model, Device Registry, FastAPI contracts, scheduled reconciliation, Snipe-IT and Microsoft Graph adapters, operation state machines, rollout gates, recovery procedures, and the Portal workflows used for human review and confirmation.
A large part of the work was operational rather than purely application-level. Each feature needed a safe pilot path, rollback procedure, explicit write boundary, observable state, and a way to stop when production evidence did not match the assumptions made in development.
The project builds on two broader workstreams documented elsewhere on this site: endpoint-management modernization and the enterprise operations portal.
The problem was authority, not connectivity
At first glance, the requirement sounds simple: integrate device inventory with Intune and an asset-management platform.
The systems, however, answer different questions.
- A lifecycle inventory source can tell me which technical identities were observed.
- Microsoft Entra ID can tell me about directory device and user identities.
- Intune can tell me which managed-device enrollment exists and which Primary User is currently observed.
- Snipe-IT can tell me which physical asset exists and who currently has it checked out.
- An internal portal can capture an operator’s explicit decision when evidence is ambiguous.
Those facts overlap, but they are not interchangeable.
If I had implemented a generic two-way synchronization layer, several unsafe behaviors would have been possible:
- an old Intune Primary User could accidentally become the new patrimonial owner;
- a hostname change after reimage could create a second asset or steal an existing one;
- a timeout after an external POST could cause the same mutation to be replayed;
- absence from a technical report could be interpreted as physical disposal;
- an empty or placeholder hardware serial could make a real new device invisible to the process;
- a duplicate serial could be resolved automatically even though the evidence is ambiguous.
The design therefore starts with a matrix of authority.
Authority matrix
| Information | Authority | How it is used |
|---|---|---|
| Technical device identity | Device Registry, populated from strong source identifiers | Correlation and lifecycle history |
| Directory identifiers | Microsoft Entra ID / source directory evidence | Strong identity evidence, never possession |
| Managed-device enrollment | Microsoft Intune | Technical management target |
| Physical asset, catalog and checkout state | Snipe-IT | Patrimonial authority |
| Logical device ↔ physical asset relation | Device Registry AssetLink | Temporal, human-validated relation |
| Confirmed possession | Active Registry assignment backed by the patrimonial workflow | Desired state for user projection |
| Intune Primary User | Microsoft Graph observed state | Projection to reconcile, not authority for possession |
| Ambiguous condition | Human review queue | Explicit decision instead of heuristic mutation |
This distinction is what makes the rest of the architecture predictable.
Sanitized architecture
The public diagram intentionally removes real hosts, database names, internal URLs, credentials, and exact route names.
Mermaid source
flowchart LR
Portal[Internal portal\nHuman control plane] -->|preview / confirm| Service[Integration API + worker]
Lifecycle[Lifecycle reporting\nread-only evidence] -->|GET snapshots| Service
Graph[Microsoft Graph / Intune] -->|GET identity + observed state| Service
Snipe[Snipe-IT] -->|GET assets + checkout| Service
Service -->|local state + audit| Registry[(Device Registry)]
Service -->|gated projection write| Graph
Service -->|gated patrimonial write| Snipe
note[Hostname is presentation only.\nAmbiguous evidence stops at review.] -.-> Service
The integration service has two execution surfaces.
The API handles portal queries, previews, and explicitly confirmed commands. The worker continuously ingests evidence, enriches the Registry, produces review items, reconciles desired and observed state, retries safe operations, and records job health.
Both use the same domain state. The worker is not a second source of business truth; it is an automated executor of deterministic reconciliation rules.
Logical device is not the physical computer
This was the most important modeling decision in the project.
A Windows reinstallation can produce a new technical identity while the physical chassis, patrimonial asset tag, and ownership history remain the same.
I therefore separated two concepts:
LogicalDevice
technical identity and lifecycle
PhysicalAsset
patrimonial object in Snipe-IT
The relation between them is a temporal AssetLink.
Physical asset A
↓ valid during period 1
Logical device X
Physical asset A
↓ valid during period 2
Logical device Y
A reimage does not require merging X and Y. The old link is closed and a new human-validated link is opened to the same physical asset.
This preserves both histories: the technical lifecycle remains truthful and the patrimonial lifecycle remains continuous.
Why hostname is deliberately weak evidence
Hostnames are convenient for people and unreliable as master keys.
They can be reused, renamed, recreated during imaging, generated automatically, or copied from old processes. The architecture therefore allows hostname to appear in search results, tables, alerts, and confirmation screens, but it does not use hostname to establish physical identity automatically.
Correlation is based on stronger identifiers from the source systems. When those identifiers disagree, the process stops rather than silently falling back to a visually plausible machine name.
That tradeoff creates more review items, but dramatically reduces the risk of a confident wrong association.
Reconciliation instead of synchronization
The integration is designed around a simple loop:
authoritative desired state
↓
read observed state
↓
compare
↓
no difference → no write
safe difference → gated operation
ambiguity/error → wait or manual review
This matters especially for external systems because “the API call succeeded” and “the business state is correct” are not the same event.
The worker repeatedly evaluates state until the system is converged or a condition becomes unsafe to automate.
The mutation contract
Mutations use a layered contract instead of executing directly from a form submission.
Mermaid source
flowchart TD
Human[Human action] --> Auth{Authentication + authorization}
Auth -->|fail| Stop[Stop / manual review]
Auth --> Preview[Read-only preview]
Preview --> Structural{Structural state valid?}
Structural -->|no| Stop
Structural --> Live[Live GET from target system]
Live --> Observed{Observed state compatible?}
Observed -->|no| Stop
Observed --> Confirm[Explicit human confirmation]
Confirm --> Flag{Write gate enabled?}
Flag -->|no| Stop
Flag --> Intent[Persist operation + idempotency + write intent]
Intent --> Write[External mutation]
Write --> Verify[Read / confirm resulting state]
Verify --> Audit[Persist terminal audit state]
The key point is that write intent is persisted before the external request for operations where an uncertain response could otherwise create a duplicate.
If the process crashes or times out after the target accepted the mutation, the next execution can recognize that a write may already have happened and move into verification instead of replaying the request blindly.
Operation and step state machines
Each business change is represented by an idempotent operation. Operations contain smaller steps targeted at local state, Snipe-IT, or Microsoft Graph.
The persisted model records concepts such as:
- operation type;
- actor;
- idempotency key;
- original request;
- target entity;
- before/after evidence for each step;
- attempt count;
- retry scheduling;
- error code and message;
- whether an external write was attempted;
- completion state.
The system therefore has enough evidence to answer not only “did this fail?” but also what was intended, what had been observed, whether an external mutation was already attempted, and what may safely happen next.
New asset discovery without automatic asset creation
The worker continuously compares managed devices and the Registry with the current Snipe-IT inventory.
When a trustworthy serial identifies a device that does not yet have a patrimonial match, the service can create a local candidate and an explicit human review.
It still does not create the Snipe-IT asset automatically.
The operator receives a preview that checks the candidate, target logical device, current Snipe-IT state, catalog selection, serial, and possible duplicates. Only after confirmation and the appropriate write gate can the integration create or link the asset.
This turns automatic discovery into automatic preparation for a human decision, which is a much safer boundary for patrimonial data.
What happens when the hardware serial is unusable
Some managed devices report no useful hardware serial at all. Others expose firmware placeholders that are technically non-empty but useless as identity.
Ignoring them would create a blind spot: a legitimate new computer could be fully managed but never enter the asset-review workflow.
The solution was a controlled fallback:
- recognize null, blank, and known placeholder serials as untrusted;
- confirm that the logical device does not already have a trustworthy serial from another active enrollment;
- confirm that no active physical-asset link already exists;
- allocate a short generated corporate serial from a reserved namespace;
- check collisions against existing local candidates and current Snipe-IT assets;
- persist the generated value once;
- mark the candidate as requiring human validation.
A public example could look like:
CORP42KPX
The exact production prefix is intentionally omitted.
Most importantly, the generated patrimonial serial is not written back to Intune and is not treated as proof that two future reimages represent the same chassis. It solves the asset-registration workflow, not technical identity.
Possession is a separate domain from identity
Physical asset ownership at a given moment is represented by an active assignment in the Registry, created from an explicit patrimonial checkout workflow.
A user mapping relates the person known by Microsoft Entra ID to the corresponding Snipe-IT user. That mapping allows one confirmed human action to produce consistent identifiers for both systems without making either platform’s incidental user field authoritative for the other.
The resulting rule is:
Snipe / Portal confirmed possession
↓
Registry Assignment
↓
Desired Microsoft Entra user
↓
Intune Primary User reconciliation
This order matters. The existing Primary User in Intune never automatically creates a patrimonial assignment.
Primary User as a projection
The Primary User reconciler reads the current assignment, requires a unique active Intune enrollment, then reads the current managed-device users from Microsoft Graph before considering a write.
Mermaid source
flowchart TD
Assignment[Active Registry assignment\nconfirmed possession] --> Enrollment{Exactly one active\nmanaged enrollment?}
Enrollment -->|none| Wait[Wait for enrollment]
Enrollment -->|multiple| Review[Manual review]
Enrollment -->|yes| Graph[GET managed-device users]
Graph --> Ambiguous{More than one observed user?}
Ambiguous -->|yes| Review
Ambiguous -->|no| Compare{Desired vs observed}
Compare -->|same| Done[Already converged\nzero external write]
Compare -->|desired user differs| Set[SET Primary User\ngated write]
Compare -->|desired none, observed user| Remove[REMOVE Primary User\ngated write]
Set --> Verify[Verify later]
Remove --> Verify
Verify -->|converged| Done
Verify -->|write outcome uncertain| VerifyOnly[Verify-only path\nno blind replay]
VerifyOnly -->|still unknown| Review
Typical outcomes are deliberately explicit:
desired user U + observed U → already converged, zero write
desired user U + observed ∅ → set Primary User
desired user U + observed V → set Primary User to U
desired ∅ + observed V → remove Primary User
multiple observed users → manual review
no active enrollment → wait
The worker also has a hard per-run budget for new Graph mutations. This limits the blast radius even when a larger number of intents become eligible at the same time.
The most important production incident: an unknown write result
One of the best design tests came from a real Primary User mutation whose HTTP request ended with a transport timeout.
A timeout is not evidence that the target rejected the request.
There are two possible realities:
A. Graph never accepted the request
B. Graph accepted it, but the response never reached the client
Repeating the write immediately assumes A and ignores B.
The operation therefore moved into an uncertain verification state. The integration recorded that a Graph write had already been attempted and refused to replay it automatically. Subsequent checks were read-only: inspect the observed state, mark the step successful if it eventually converged, retry the read later if propagation was still incomplete, or escalate to manual review after the bounded verification window.
That incident reinforced a general distributed-systems rule:
A request failure and a business-operation failure are different things.
Reimage and reprovisioning
Reimage handling is another case where a convenient automatic rule would be dangerous.
The safe workflow validates that:
- source and target logical devices are distinct;
- the source has exactly one active asset link;
- the physical asset is not actively linked somewhere else;
- the target has no competing asset link;
- there is no unresolved active possession that should first be checked in;
- no conflicting operation is already in progress;
- trustworthy serial evidence agrees between the physical asset and the target managed enrollment;
- a live Snipe-IT read confirms the expected asset;
- the asset is not currently assigned to another user.
The reassociation itself is Registry-only: close the old temporal link and create the new one. The physical Snipe-IT asset does not need to be recreated or mutated just because Windows produced a new logical identity.
If trustworthy physical evidence is unavailable, the workflow requires explicit human identification rather than weakening the correlation rule to hostname.
Offboarding is a workflow, not a single check-in
User departure crosses several state domains: directory status, active assignments, physical returns, Primary User projection, retention periods, and unresolved assets.
The offboarding model therefore freezes the relevant identifiers and assets into a workflow record instead of querying mutable current state at every step.
The process can then distinguish:
- user disabled but assets still awaiting return;
- returned asset with local assignment closed;
- Intune user projection still waiting to converge;
- retention period still active;
- workflow completed.
The first full operational acceptance must use a legitimate offboarding event. The system was intentionally not tested by disabling a real user just to manufacture a passing scenario.
Absence is not disposal
A lifecycle report can show that a logical device disappeared. That does not prove the physical computer was discarded.
The same symptom can be produced by reimage, rejoin, reprovisioning, a reporting gap, or a temporarily unavailable source.
The disposal-observation logic therefore remains review-only. A physical asset enters an observation episode only when its linked technical identity is absent by strong identifiers. Eligibility requires persistence over time and multiple snapshots, plus a set of negative checks against current asset, assignment, operation, Entra, and Intune state.
Examples of blockers include:
- an active assignee;
- an active Registry assignment;
- a pending operation;
- current directory or Intune evidence;
- evidence that another logical device may represent the same physical machine after reprovisioning;
- insufficient absence duration or snapshot count;
- unavailable external evidence.
An unavailable source is a blocker, never permission to dispose.
The current production design intentionally does not perform automatic patrimonial disposal.
The internal portal as control plane
The portal is not a second integration engine. It is the human interface for the domain state already maintained by the service.
It provides surfaces for:
- inventory and current correlation state;
- active patrimonial possession;
- new-asset review;
- assignment, transfer and check-in;
- reimage reassociation;
- offboarding;
- lifecycle observations and blockers;
- operation and review history.
Potentially destructive or externally mutating actions are shown as preview-first workflows. The confirmation modal reuses the server-rendered payload and anti-forgery controls rather than inventing a parallel client-side contract.
The UI also makes uncertainty visible. A missing backend response should appear as unavailable data, not as a false “no assignment” or “no issue” state.
Why the system uses feature gates even after the feature works
Each class of external mutation was introduced behind a separate rollout gate.
During pilots, the gates allowed one write surface to be enabled while others remained closed. After operational acceptance, the same mechanisms become incident kill switches rather than something that must be toggled for every routine operation.
The safety model therefore has multiple independent controls:
human permission
+ preview
+ structural preflight
+ live external preflight
+ explicit confirmation
+ feature gate
+ idempotency
+ bounded retries / write budget
+ post-write verification
No single flag is expected to carry the entire safety responsibility.
Worker design and observability
The worker executes jobs with independent cadences for tasks such as:
- lifecycle ingestion;
- lifecycle-disposal observation;
- user mapping;
- Intune enrichment;
- new-asset discovery;
- assignment reconciliation;
- retrying pending operations;
- aging review items;
- notification preparation;
- retention cleanup when explicitly approved.
Each job records attempt, success, error, and a bounded result summary in local settings. External adapters also use retry and circuit-breaker behavior for transient failures.
The operational question is not merely “is the service process running?” It is also:
- when did each job last succeed?
- what did it plan or change?
- how many local and external writes occurred?
- which review or operation is waiting?
- did a source become unavailable?
- did an uncertain external write require verification?
That evidence is what makes the automation supportable in production.
Why there are no live Portal screenshots here
This project would be easy to illustrate with screenshots, but an operational inventory screen can expose more than it appears to: usernames, hostnames, asset tags, serials, device IDs, pending reviews, timestamps, and organizational workflow state.
For that reason, this public version uses sanitized reconstructed diagrams instead of production captures. The diagrams preserve the architecture and decision logic without requiring increasingly fragile redaction of real operational data.
The Mermaid source for each diagram is included with the article, while the website serves static SVG renderings. That keeps the page reproducible without introducing a third-party client-side diagram runtime or weakening the site’s Content Security Policy.
Security and trust boundaries
The security model is intentionally broader than secret storage.
Important boundaries include:
- the browser never receives credentials for Snipe-IT or Microsoft Graph;
- the portal calls a narrow authenticated internal API;
- Graph uses application authentication rather than a logged-in operator token;
- Snipe-IT is accessed through its REST API, not by writing its database directly;
- secrets and certificates remain outside version control;
- external integration errors are redacted before persistence or display;
- ambiguous domain state fails closed;
- local audit state is preserved before risky external writes;
- source systems retain authority for the domains they own.
These controls allow the portal to be convenient without turning it into a privileged browser-side integration client.
Results
The project produced a cohesive operating model rather than another point-to-point synchronization script.
One technical identity model
Directory, lifecycle, and Intune evidence can converge into a durable logical-device history without treating the display name as identity.
One patrimonial relation model
Physical assets remain in Snipe-IT while the Registry records which logical identity represented each asset over time.
One possession model
Confirmed checkout becomes the desired possession state and can be projected to Intune without allowing Intune to redefine patrimonial ownership.
Safe new-device handling
Devices with a real serial and devices with unusable firmware serials can both reach a human-reviewed asset workflow without automatic asset creation.
Safer failure recovery
Idempotency and persisted write intent make it possible to distinguish retryable reads from potentially already-applied external writes.
Explicit uncertainty
Ambiguous evidence becomes a review item or waiting state instead of an implicit heuristic.
No productivity percentage, financial saving, inventory-accuracy percentage, or incident-reduction metric is claimed here because those measurements have not been prepared for public release.
Lessons learned
Define authority before writing integration code
Most dangerous synchronization bugs are ownership bugs disguised as API bugs. If two systems are both allowed to be “truth” for the same field, the conflict will eventually reach production.
Identity and possession must stay separate
The person currently associated with a managed endpoint and the person who is patrimonially responsible for a physical asset are related concepts, not identical data fields.
A timeout is a state, not a reason to repeat the request
Distributed systems require an explicit unknown-outcome path. Retrying a mutation is safe only when idempotency or external state proves that replay cannot duplicate the effect.
Human review is part of the architecture
A manual review queue is not a failure of automation. It is the correct destination for evidence that the system cannot prove deterministically.
Reimage proves why temporal models matter
If asset links were permanent one-to-one rows, reimage would either destroy history or force logical identities to be merged incorrectly. Temporal validity keeps both stories intact.
Safety gates should survive the pilot
A rollout flag is useful during testing, but after acceptance it becomes even more valuable as an emergency kill switch for one write surface without disabling the entire service.
Read-only automation can still create significant value
Continuous discovery, reconciliation, review generation, observability, and preflight can remove a large amount of manual investigation even when the final patrimonial mutation remains human-confirmed.
Limitations and deliberately unfinished edges
The architecture intentionally leaves some boundaries conservative.
- Automatic physical-asset disposal is not enabled.
- Some operational acceptance depends on legitimate future events such as a real offboarding or reimage rather than synthetic production disruption.
- A device with no trustworthy physical identifier still requires human judgment when proving that it is the same chassis after reprovisioning.
- External APIs can return incomplete or eventually consistent state; the system must sometimes wait instead of resolving immediately.
- The internal portal remains coupled to an existing HESK/PHP platform even though the integration domain itself is separated behind FastAPI.
These are explicit constraints, not hidden gaps.
Why I am not publishing the production repository
A source tree can contain sensitive operational information even after obvious credentials are removed.
Potential exposure includes:
- exact feature-flag and workflow names;
- internal API topology and filesystem paths;
- schemas that reveal organizational processes;
- real identifiers in fixtures, migrations, logs, comments, or commit history;
- recovery procedures optimized for the real environment;
- assumptions about network and service boundaries;
- old secrets that may survive in Git history even when absent from the current branch.
For that reason, I consider a copied-and-redacted production repository weaker than a deliberately public reference implementation.
If I publish code for this architecture later, the safer approach is a separate educational repository built from scratch with fake adapters, generated test data, a reduced Registry schema, and no shared history with the production system.
That repository can demonstrate the engineering patterns without pretending that “sanitized production code” is automatically safe.
What this case study demonstrates
The interesting part of enterprise integration is rarely the HTTP client.
The harder work is defining ownership, modeling time, detecting ambiguity, making mutations idempotent, separating desired from observed state, and creating an operational path for the moment when an external system gives an answer that is incomplete, delayed, or unknown.
The architecture ultimately treats automation as a controlled participant in the business process:
observe
→ correlate
→ reconcile
→ explain
→ ask for a human decision when necessary
→ persist intent
→ mutate narrowly
→ verify
→ preserve history
That is the difference between an integration script that works in the happy path and a control plane that can remain trustworthy after production starts behaving unexpectedly.
Confidentiality statement
This is a sanitized description of a private production implementation. Organization names, internal source code, real endpoints, hostnames, addresses, user information, asset identifiers, device identifiers, credentials, certificate material, exact runtime paths, exact rollout settings, and sensitive operational metrics are intentionally excluded or generalized.