169 lines
7.6 KiB
Markdown
169 lines
7.6 KiB
Markdown
|
|
# Operations runbook
|
||
|
|
|
||
|
|
What to do when the engine is running, and what to do when it should not be.
|
||
|
|
|
||
|
|
## Before any run
|
||
|
|
|
||
|
|
1. Validate the configuration. It costs seconds and needs no tenant.
|
||
|
|
|
||
|
|
```bash
|
||
|
|
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -ValidateOnly -NonInteractive
|
||
|
|
```
|
||
|
|
|
||
|
|
2. Run the rules against synthetic fixtures. This shows what the rule set *does* before it sees a
|
||
|
|
real account.
|
||
|
|
|
||
|
|
```bash
|
||
|
|
pwsh ./Edit-PersonaEngineConfig.ps1 -ConfigPath ./config/persona-engine.json -TestDataPath ./tests/TestData -ValidateOnly -NonInteractive
|
||
|
|
```
|
||
|
|
|
||
|
|
3. Preview a single user before previewing the tenant.
|
||
|
|
|
||
|
|
```bash
|
||
|
|
pwsh ./Invoke-PersonaEngine.ps1 -ConfigPath ./config/persona-engine.json -UserObjectId <ACCOUNT-OBJECT-ID> -WhatIf -Verbose
|
||
|
|
```
|
||
|
|
|
||
|
|
Never skip step 3. A rule set that behaves correctly against fixtures can still request a property
|
||
|
|
your tenant does not populate, and finding that out on one account is cheaper than on fifty thousand.
|
||
|
|
|
||
|
|
## Reading a run
|
||
|
|
|
||
|
|
Per-user lines appear immediately, one per account, colour-coded by action:
|
||
|
|
|
||
|
|
| Action | Meaning |
|
||
|
|
| --- | --- |
|
||
|
|
| `Unchanged` | Calculated value already matches the stored value. Nothing to do. |
|
||
|
|
| `WouldUpdate` | A change is proposed. Preview mode, or the per-user gate refused. |
|
||
|
|
| `Updated` | The attribute was written. |
|
||
|
|
| `UpdateFailed` | The write was attempted and rejected. Stored value is untouched. |
|
||
|
|
| `Skipped` | No write is possible: `EvaluationError`, or a blank or unapproved target. |
|
||
|
|
|
||
|
|
`Skipped` on every account almost always means the target attribute is blank or unapproved — check
|
||
|
|
the header line and `PE-SAF-001`.
|
||
|
|
|
||
|
|
A summary appears every `summaryInterval` accounts and once at the end, listing **every** rule
|
||
|
|
including disabled and zero-match ones. A rule that never fired and a rule that is not in the
|
||
|
|
configuration look identical if zero-match rules are omitted, and that distinction is usually what
|
||
|
|
you are looking for.
|
||
|
|
|
||
|
|
## Exit codes
|
||
|
|
|
||
|
|
| Code | Meaning | First thing to check |
|
||
|
|
| --- | --- | --- |
|
||
|
|
| 0 | Success | — |
|
||
|
|
| 1 | Configuration validation failed | The findings printed above it; no connection was attempted |
|
||
|
|
| 2 | Authentication or authorization failed | Scopes, consent, and whether the account can sign in |
|
||
|
|
| 3 | User enumeration failed | Graph availability. **No accounts were processed** — a partial population is never used |
|
||
|
|
| 4 | `evaluationErrorThreshold` exceeded | Group or role endpoint health. Nothing was changed |
|
||
|
|
| 5 | Reconciliation failed | **An engine defect.** Open an issue with the `EngineDefect` record |
|
||
|
|
| 6 | Unexpected fatal error | The message and `-Verbose` stack trace |
|
||
|
|
|
||
|
|
Exit 5 is never a data condition. Outcomes are assigned by the engine, exactly one per account, so if
|
||
|
|
`Processed` does not equal `Matched + Unclassified + EvaluationError` the engine lost a user or
|
||
|
|
double-counted one. Report it rather than re-running.
|
||
|
|
|
||
|
|
Exit 4 means the population was classified from data that could not be trusted. Stored values were
|
||
|
|
preserved, so nothing is damaged — but do not draw conclusions from the run.
|
||
|
|
|
||
|
|
## Kill switch
|
||
|
|
|
||
|
|
In increasing order of severity. Pick the lowest one that addresses the problem.
|
||
|
|
|
||
|
|
### 1. Stop writing — immediate, no deployment
|
||
|
|
|
||
|
|
Add `-WhatIf` to the invocation. Reads, evaluation, output, and audit records continue unchanged;
|
||
|
|
zero write requests are constructed.
|
||
|
|
|
||
|
|
### 2. Stop classifying — one configuration change
|
||
|
|
|
||
|
|
Set `enabled: false` on every rule and deploy.
|
||
|
|
|
||
|
|
> Every account becomes `Unclassified`. **In an enforcing run that proposes clearing every stored
|
||
|
|
> persona.** Combine with `-WhatIf`, or use option 1 instead, unless clearing is what you want.
|
||
|
|
|
||
|
|
### 3. Stop running — Stage B only
|
||
|
|
|
||
|
|
Disable the Automation schedule. Nothing is in flight; the next run simply does not start.
|
||
|
|
|
||
|
|
### 4. Remove the capability — the one that holds if the code is the problem
|
||
|
|
|
||
|
|
Revoke `User.ReadWrite.All` from the execution identity. The engine keeps running and every write
|
||
|
|
becomes `UpdateFailed`, which is loud, logged, and harmless.
|
||
|
|
|
||
|
|
### 5. Remove the path — for a suspected defect in the write path
|
||
|
|
|
||
|
|
Remove the write deployment stage from the release pipeline so no build can restore write capability
|
||
|
|
by accident.
|
||
|
|
|
||
|
|
Options 1 through 3 rely on the engine behaving correctly. Options 4 and 5 do not, which is why they
|
||
|
|
exist.
|
||
|
|
|
||
|
|
## Rollback (OTD-010)
|
||
|
|
|
||
|
|
Every `Updated` audit record carries `previousValue`, captured **before** the PATCH. That is what
|
||
|
|
makes rollback possible; reading the value back afterwards would return the new one.
|
||
|
|
|
||
|
|
To roll back a run:
|
||
|
|
|
||
|
|
1. Find the run's records by `runId`.
|
||
|
|
2. Select records where `recordType` is `UserEvent` and `action` is `Updated`.
|
||
|
|
3. For each, write `previousValue` back to `accountObjectId`.
|
||
|
|
|
||
|
|
```powershell
|
||
|
|
# Reads the NDJSON audit file and lists what a rollback would restore.
|
||
|
|
# Review this output before writing anything back.
|
||
|
|
Get-Content <LOG-OUTPUT-PATH> |
|
||
|
|
ForEach-Object { $_ | ConvertFrom-Json } |
|
||
|
|
Where-Object { $_.runId -eq '<RUN-ID>' -and $_.recordType -eq 'UserEvent' -and $_.action -eq 'Updated' } |
|
||
|
|
Select-Object accountObjectId, userPrincipalName, previousValue, calculatedPersona
|
||
|
|
```
|
||
|
|
|
||
|
|
> A rollback is itself a directory write and is subject to the same V-4 gate and the same test-account
|
||
|
|
> restriction as any enforcement run. A rollback tool is **not implemented in v1** — the data needed
|
||
|
|
> to build one is captured, deliberately, because it cannot be reconstructed retroactively.
|
||
|
|
|
||
|
|
If a rollback is genuinely intended at the configuration level, publish it as a **new higher
|
||
|
|
version** rather than reusing the old number. Two different rule sets sharing one `configVersion`
|
||
|
|
makes the audit trail unable to tell them apart (`PE-SAF-004`).
|
||
|
|
|
||
|
|
## Common situations
|
||
|
|
|
||
|
|
**Every account is `EvaluationError`.** A required data source is disabled or unreachable. Check
|
||
|
|
`dataSources.groups.enabled` and `dataSources.roles.enabled` against what the rules need — `PE-SAF-003`
|
||
|
|
catches this at validation time, so a run reaching this state usually means validation was bypassed.
|
||
|
|
|
||
|
|
**Every account is `Unclassified`.** Either every rule is disabled, or no rule matches. The summary
|
||
|
|
table distinguishes these: disabled rules are dimmed, zero-match enabled rules show `0`.
|
||
|
|
|
||
|
|
**The run is slow.** Check the cache hit ratio via `-Verbose`. Membership lookups dominate: one
|
||
|
|
account needing both direct and transitive facets plus roles is three requests. Narrowing rules to a
|
||
|
|
single membership mode roughly halves that.
|
||
|
|
|
||
|
|
**A write failed with 403.** Non-retryable by design — retrying would hide a configuration or
|
||
|
|
authorization defect behind a timeout. Check that the identity holds `User.ReadWrite.All` and that
|
||
|
|
the target attribute exists on the application registration.
|
||
|
|
|
||
|
|
**Audit file sink warnings.** The sink warns once per run and processing continues. A locked or full
|
||
|
|
log file is an operational problem with the sink, not a reason to abandon a run mid-population and
|
||
|
|
leave the directory half-reconciled.
|
||
|
|
|
||
|
|
## Concurrency
|
||
|
|
|
||
|
|
Two concurrent enforcing runs against the same tenant would race on the same attributes. Until the
|
||
|
|
OTD-009 run-start concurrency check is implemented (T120, Stage B), **the schedule is the lock**: do
|
||
|
|
not start a manual enforcing run while a scheduled one may be in flight.
|
||
|
|
|
||
|
|
Preview runs are read-only and safe to run concurrently.
|
||
|
|
|
||
|
|
## What to attach to a bug report
|
||
|
|
|
||
|
|
- The exit code
|
||
|
|
- The `RunComplete` record for the run — it carries the counters, timing, and exit code even when the
|
||
|
|
run died early
|
||
|
|
- Any `EngineDefect` record
|
||
|
|
- The `configVersion` and `configurationHash` from any record
|
||
|
|
- The rule set, sanitized
|
||
|
|
|
||
|
|
Do **not** attach raw audit records containing real UPNs or Object IDs to anything that leaves the
|
||
|
|
organization. They are approved for internal logs, not for public issue trackers.
|