Completes 109 of 121 tasks. Every remaining task needs a tenant connection
(T055, T056, T101-T103) or an Azure Automation account (T115-T121).
354 offline Pester tests PASS
Engine purity (Principle IV) PASS
Sanitization (SC-013) PASS (156 files)
Graph module loaded in tests none (SC-008 holds)
What landed
- Four-layer configuration validation with stable finding codes, covering
every VR-002 and VR-003 condition, plus a 23-fixture invalid-config corpus
- Run loop, audit records (NDJSON through a single sink), summaries,
reconciliation, and exit codes 0-6
- Persistence behind a single write-body builder whose result always has
exactly one key
- Invoke-PersonaEngine.ps1 and Edit-PersonaEngineConfig.ps1
- Six docs, two pipelines, traceability matrix, V-5a and sanitization records
Three deviations from tasks.md, each recorded in its status block
T033 is not in Resolve-UserPersona. evaluationErrorThreshold is run-level
state and the rule engine is pure; a counter there would break Principle IV.
It lives in New-PersonaRunCounter and is applied in the run loop.
A new src/Engine/ layer holds Invoke-PersonaEngineRun. The entry script
imports the manifest, which requires Microsoft.Graph.Authentication, so a
loop living only inside it could not run on a machine without the Graph SDK
and SC-004 could not be proven at all. The entry script is now a thin
wrapper and what ships is what is tested.
The invalid-config corpus is generated by a committed script, with the
generated fixtures committed too, so a reviewer sees the fixture in the diff.
Defects found by running the code, not by reading it
Group and role ID lists were double-wrapped: @(Get-PersonaGroupIdPage ...)
around a comma-returned array collapsed every membership list into one
bogus space-joined entry. That is a silent false non-match, exactly what
FR-013 exists to prevent.
A 403 whose status appears only in the exception message parsed as $null,
which the retry policy treats as a transport error - five requests per
account against a tenant already refusing. Status extraction now falls back
to the message text, bounded to 400-599.
The sanitization scan walked tracked files only, so it covered 34 of 156
files and none of this phase's code. It now scans untracked non-ignored
files too, and a negative control confirms it catches a planted leak.
Test-Json reports one error per violating location, not first-failure-only
as the V-5a draft claimed. Record and pin corrected.
Enforcement remains blocked on the V-4 security sign-off.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7.6 KiB
Operations runbook
What to do when the engine is running, and what to do when it should not be.
Before any run
-
Validate the configuration. It costs seconds and needs no tenant.
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -ValidateOnly -NonInteractive -
Run the rules against synthetic fixtures. This shows what the rule set does before it sees a real account.
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -TestDataPath ./tests/TestData -ValidateOnly -NonInteractive -
Preview a single user before previewing the tenant.
pwsh ./Invoke-PersonaEngine.ps1 -ConfigPath ./config/persona-engine.json -UserObjectId <ACCOUNT-OBJECT-ID> -WhatIf -Verbose
Never skip step 3. A rule set that behaves correctly against fixtures can still request a property your tenant does not populate, and finding that out on one account is cheaper than on fifty thousand.
Reading a run
Per-user lines appear immediately, one per account, colour-coded by action:
| Action | Meaning |
|---|---|
Unchanged |
Calculated value already matches the stored value. Nothing to do. |
WouldUpdate |
A change is proposed. Preview mode, or the per-user gate refused. |
Updated |
The attribute was written. |
UpdateFailed |
The write was attempted and rejected. Stored value is untouched. |
Skipped |
No write is possible: EvaluationError, or a blank or unapproved target. |
Skipped on every account almost always means the target attribute is blank or unapproved — check
the header line and PE-SAF-001.
A summary appears every summaryInterval accounts and once at the end, listing every rule
including disabled and zero-match ones. A rule that never fired and a rule that is not in the
configuration look identical if zero-match rules are omitted, and that distinction is usually what
you are looking for.
Exit codes
| Code | Meaning | First thing to check |
|---|---|---|
| 0 | Success | — |
| 1 | Configuration validation failed | The findings printed above it; no connection was attempted |
| 2 | Authentication or authorization failed | Scopes, consent, and whether the account can sign in |
| 3 | User enumeration failed | Graph availability. No accounts were processed — a partial population is never used |
| 4 | evaluationErrorThreshold exceeded |
Group or role endpoint health. Nothing was changed |
| 5 | Reconciliation failed | An engine defect. Open an issue with the EngineDefect record |
| 6 | Unexpected fatal error | The message and -Verbose stack trace |
Exit 5 is never a data condition. Outcomes are assigned by the engine, exactly one per account, so if
Processed does not equal Matched + Unclassified + EvaluationError the engine lost a user or
double-counted one. Report it rather than re-running.
Exit 4 means the population was classified from data that could not be trusted. Stored values were preserved, so nothing is damaged — but do not draw conclusions from the run.
Kill switch
In increasing order of severity. Pick the lowest one that addresses the problem.
1. Stop writing — immediate, no deployment
Add -WhatIf to the invocation. Reads, evaluation, output, and audit records continue unchanged;
zero write requests are constructed.
2. Stop classifying — one configuration change
Set enabled: false on every rule and deploy.
Every account becomes
Unclassified. In an enforcing run that proposes clearing every stored persona. Combine with-WhatIf, or use option 1 instead, unless clearing is what you want.
3. Stop running — Stage B only
Disable the Automation schedule. Nothing is in flight; the next run simply does not start.
4. Remove the capability — the one that holds if the code is the problem
Revoke User.ReadWrite.All from the execution identity. The engine keeps running and every write
becomes UpdateFailed, which is loud, logged, and harmless.
5. Remove the path — for a suspected defect in the write path
Remove the write deployment stage from the release pipeline so no build can restore write capability by accident.
Options 1 through 3 rely on the engine behaving correctly. Options 4 and 5 do not, which is why they exist.
Rollback (OTD-010)
Every Updated audit record carries previousValue, captured before the PATCH. That is what
makes rollback possible; reading the value back afterwards would return the new one.
To roll back a run:
- Find the run's records by
runId. - Select records where
recordTypeisUserEventandactionisUpdated. - For each, write
previousValueback toaccountObjectId.
# Reads the NDJSON audit file and lists what a rollback would restore.
# Review this output before writing anything back.
Get-Content <LOG-OUTPUT-PATH> |
ForEach-Object { $_ | ConvertFrom-Json } |
Where-Object { $_.runId -eq '<RUN-ID>' -and $_.recordType -eq 'UserEvent' -and $_.action -eq 'Updated' } |
Select-Object accountObjectId, userPrincipalName, previousValue, calculatedPersona
A rollback is itself a directory write and is subject to the same V-4 gate and the same test-account restriction as any enforcement run. A rollback tool is not implemented in v1 — the data needed to build one is captured, deliberately, because it cannot be reconstructed retroactively.
If a rollback is genuinely intended at the configuration level, publish it as a new higher
version rather than reusing the old number. Two different rule sets sharing one configVersion
makes the audit trail unable to tell them apart (PE-SAF-004).
Common situations
Every account is EvaluationError. A required data source is disabled or unreachable. Check
dataSources.groups.enabled and dataSources.roles.enabled against what the rules need — PE-SAF-003
catches this at validation time, so a run reaching this state usually means validation was bypassed.
Every account is Unclassified. Either every rule is disabled, or no rule matches. The summary
table distinguishes these: disabled rules are dimmed, zero-match enabled rules show 0.
The run is slow. Check the cache hit ratio via -Verbose. Membership lookups dominate: one
account needing both direct and transitive facets plus roles is three requests. Narrowing rules to a
single membership mode roughly halves that.
A write failed with 403. Non-retryable by design — retrying would hide a configuration or
authorization defect behind a timeout. Check that the identity holds User.ReadWrite.All and that
the target attribute exists on the application registration.
Audit file sink warnings. The sink warns once per run and processing continues. A locked or full log file is an operational problem with the sink, not a reason to abandon a run mid-population and leave the directory half-reconciled.
Concurrency
Two concurrent enforcing runs against the same tenant would race on the same attributes. Until the OTD-009 run-start concurrency check is implemented (T120, Stage B), the schedule is the lock: do not start a manual enforcing run while a scheduled one may be in flight.
Preview runs are read-only and safe to run concurrently.
What to attach to a bug report
- The exit code
- The
RunCompleterecord for the run — it carries the counters, timing, and exit code even when the run died early - Any
EngineDefectrecord - The
configVersionandconfigurationHashfrom any record - The rule set, sanitized
Do not attach raw audit records containing real UPNs or Object IDs to anything that leaves the organization. They are approved for internal logs, not for public issue trackers.