Summaries now show elapsed wall-clock time since the run started, and every summary (interim and final) overwrites a results.csv (Object ID, UPN, persona/status) next to the audit log, so an operator has a plain export without parsing NDJSON. ConfigPath also now defaults to ./config/persona-engine.json instead of requiring -ConfigPath every run. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7.9 KiB
Operations runbook
What to do when the engine is running, and what to do when it should not be.
Before any run
-
Validate the configuration. It costs seconds and needs no tenant.
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -ValidateOnly -NonInteractive -
Run the rules against synthetic fixtures. This shows what the rule set does before it sees a real account.
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -TestDataPath <fixtures-dir> -ValidateOnly -NonInteractive -
Preview a single user before previewing the tenant.
pwsh ./Invoke-PersonaEngine.ps1 -ConfigPath ./config/persona-engine.json -UserObjectId <ACCOUNT-OBJECT-ID> -WhatIf -Verbose
Never skip step 3. A rule set that behaves correctly against fixtures can still request a property your tenant does not populate, and finding that out on one account is cheaper than on fifty thousand.
Reading a run
Per-user lines appear immediately, one per account, colour-coded by action:
| Action | Meaning |
|---|---|
Unchanged |
Calculated value already matches the stored value. Nothing to do. |
WouldUpdate |
A change is proposed. Preview mode, or the per-user gate refused. |
Updated |
The attribute was written. |
UpdateFailed |
The write was attempted and rejected. Stored value is untouched. |
Skipped |
No write is possible: EvaluationError, or a blank or unapproved target. |
Skipped on every account almost always means the target attribute is blank or unapproved — check
the header line and PE-SAF-001.
A summary appears every summaryInterval accounts and once at the end, listing every rule
including disabled and zero-match ones. A rule that never fired and a rule that is not in the
configuration look identical if zero-match rules are omitted, and that distinction is usually what
you are looking for. Its header shows elapsed wall-clock time since the run started.
Each summary also (re)writes logging.resultsFileName (default results.csv, next to the audit log)
with one row per account processed so far — Object ID, UPN, and assigned persona/status — for an
operator who wants the current population breakdown without parsing NDJSON.
Exit codes
| Code | Meaning | First thing to check |
|---|---|---|
| 0 | Success | — |
| 1 | Configuration validation failed | The findings printed above it; no connection was attempted |
| 2 | Authentication or authorization failed | Scopes, consent, and whether the account can sign in |
| 3 | User enumeration failed | Graph availability. No accounts were processed — a partial population is never used |
| 4 | evaluationErrorThreshold exceeded |
Group or role endpoint health. Nothing was changed |
| 5 | Reconciliation failed | An engine defect. Open an issue with the EngineDefect record |
| 6 | Unexpected fatal error | The message and -Verbose stack trace |
Exit 5 is never a data condition. Outcomes are assigned by the engine, exactly one per account, so if
Processed does not equal Matched + Unclassified + EvaluationError the engine lost a user or
double-counted one. Report it rather than re-running.
Exit 4 means the population was classified from data that could not be trusted. Stored values were preserved, so nothing is damaged — but do not draw conclusions from the run.
Kill switch
In increasing order of severity. Pick the lowest one that addresses the problem.
1. Stop writing — immediate, no deployment
Add -WhatIf to the invocation. Reads, evaluation, output, and audit records continue unchanged;
zero write requests are constructed.
2. Stop classifying — one configuration change
Set enabled: false on every rule and deploy.
Every account becomes
Unclassified. In an enforcing run that proposes clearing every stored persona. Combine with-WhatIf, or use option 1 instead, unless clearing is what you want.
3. Stop running — Stage B only
Disable the Automation schedule. Nothing is in flight; the next run simply does not start.
4. Remove the capability — the one that holds if the code is the problem
Revoke User.ReadWrite.All from the execution identity. The engine keeps running and every write
becomes UpdateFailed, which is loud, logged, and harmless.
5. Remove the path — for a suspected defect in the write path
Remove the write deployment stage from the release pipeline so no build can restore write capability by accident.
Options 1 through 3 rely on the engine behaving correctly. Options 4 and 5 do not, which is why they exist.
Rollback (OTD-010)
Every Updated audit record carries previousValue, captured before the PATCH. That is what
makes rollback possible; reading the value back afterwards would return the new one.
To roll back a run:
- Find the run's records by
runId. - Select records where
recordTypeisUserEventandactionisUpdated. - For each, write
previousValueback toaccountObjectId.
# Reads the NDJSON audit file and lists what a rollback would restore.
# Review this output before writing anything back.
Get-Content <LOG-OUTPUT-PATH> |
ForEach-Object { $_ | ConvertFrom-Json } |
Where-Object { $_.runId -eq '<RUN-ID>' -and $_.recordType -eq 'UserEvent' -and $_.action -eq 'Updated' } |
Select-Object accountObjectId, userPrincipalName, previousValue, calculatedPersona
A rollback is itself a directory write and is subject to the same V-4 gate and the same test-account restriction as any enforcement run. A rollback tool is not implemented in v1 — the data needed to build one is captured, deliberately, because it cannot be reconstructed retroactively.
If a rollback is genuinely intended at the configuration level, publish it as a new higher
version rather than reusing the old number. Two different rule sets sharing one configVersion
makes the audit trail unable to tell them apart (PE-SAF-004).
Common situations
Every account is EvaluationError. A required data source is disabled or unreachable. Check
dataSources.groups.enabled and dataSources.roles.enabled against what the rules need — PE-SAF-003
catches this at validation time, so a run reaching this state usually means validation was bypassed.
Every account is Unclassified. Either every rule is disabled, or no rule matches. The summary
table distinguishes these: disabled rules are dimmed, zero-match enabled rules show 0.
The run is slow. Check the cache hit ratio via -Verbose. Membership lookups dominate: one
account needing both direct and transitive facets plus roles is three requests. Narrowing rules to a
single membership mode roughly halves that.
A write failed with 403. Non-retryable by design — retrying would hide a configuration or
authorization defect behind a timeout. Check that the identity holds User.ReadWrite.All and that
the target attribute exists on the application registration.
Audit file sink warnings. The sink warns once per run and processing continues. A locked or full log file is an operational problem with the sink, not a reason to abandon a run mid-population and leave the directory half-reconciled.
Concurrency
Two concurrent enforcing runs against the same tenant would race on the same attributes. Until the OTD-009 run-start concurrency check is implemented (T120, Stage B), the schedule is the lock: do not start a manual enforcing run while a scheduled one may be in flight.
Preview runs are read-only and safe to run concurrently.
What to attach to a bug report
- The exit code
- The
RunCompleterecord for the run — it carries the counters, timing, and exit code even when the run died early - Any
EngineDefectrecord - The
configVersionandconfigurationHashfrom any record - The rule set, sanitized
Do not attach raw audit records containing real UPNs or Object IDs to anything that leaves the organization. They are approved for internal logs, not for public issue trackers.