Implement Stage A: rule engine, validation, audit, and safety gates

Completes 109 of 121 tasks. Every remaining task needs a tenant connection
(T055, T056, T101-T103) or an Azure Automation account (T115-T121).

  354 offline Pester tests      PASS
  Engine purity (Principle IV)  PASS
  Sanitization (SC-013)         PASS  (156 files)
  Graph module loaded in tests  none  (SC-008 holds)

What landed
  - Four-layer configuration validation with stable finding codes, covering
    every VR-002 and VR-003 condition, plus a 23-fixture invalid-config corpus
  - Run loop, audit records (NDJSON through a single sink), summaries,
    reconciliation, and exit codes 0-6
  - Persistence behind a single write-body builder whose result always has
    exactly one key
  - Invoke-PersonaEngine.ps1 and Edit-PersonaEngineConfig.ps1
  - Six docs, two pipelines, traceability matrix, V-5a and sanitization records

Three deviations from tasks.md, each recorded in its status block

  T033 is not in Resolve-UserPersona. evaluationErrorThreshold is run-level
  state and the rule engine is pure; a counter there would break Principle IV.
  It lives in New-PersonaRunCounter and is applied in the run loop.

  A new src/Engine/ layer holds Invoke-PersonaEngineRun. The entry script
  imports the manifest, which requires Microsoft.Graph.Authentication, so a
  loop living only inside it could not run on a machine without the Graph SDK
  and SC-004 could not be proven at all. The entry script is now a thin
  wrapper and what ships is what is tested.

  The invalid-config corpus is generated by a committed script, with the
  generated fixtures committed too, so a reviewer sees the fixture in the diff.

Defects found by running the code, not by reading it

  Group and role ID lists were double-wrapped: @(Get-PersonaGroupIdPage ...)
  around a comma-returned array collapsed every membership list into one
  bogus space-joined entry. That is a silent false non-match, exactly what
  FR-013 exists to prevent.

  A 403 whose status appears only in the exception message parsed as $null,
  which the retry policy treats as a transport error - five requests per
  account against a tenant already refusing. Status extraction now falls back
  to the message text, bounded to 400-599.

  The sanitization scan walked tracked files only, so it covered 34 of 156
  files and none of this phase's code. It now scans untracked non-ignored
  files too, and a negative control confirms it catches a planted leak.

  Test-Json reports one error per violating location, not first-failure-only
  as the V-5a draft claimed. Record and pin corrected.

Enforcement remains blocked on the V-4 security sign-off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-20 21:48:19 -04:00
parent c59c85dd55
commit cdc6bb33d3
124 changed files with 16638 additions and 199 deletions
@@ -0,0 +1,76 @@
# Sanitization scan result (SC-013)
**Status**: PASS
**Date**: 2026-08-20
**Task**: T113
**Scanner**: [tests/Test-Sanitization.ps1](../../../tests/Test-Sanitization.ps1)
**Files scanned**: 156
## What is scanned
`git ls-files --cached --others --exclude-standard` — tracked files **and** untracked files that are
not gitignored.
The original scanner walked `git ls-files` alone, which covered only tracked files. That made the
gate useless where it matters most: a leaked identifier in a file that has not been committed yet is
precisely the one worth catching, and scanning only what is already in history means the scan passes
right up until the commit that makes it too late. At the time this was found, the scan was covering
34 of the repository's 156 files and none of the implementation written in this phase.
`--exclude-standard` keeps gitignored build output and local scratch files out, so the scan covers
exactly what a commit would add.
## Patterns
| Pattern | Exemptions |
| --- | --- |
| GUIDs | Placeholder-shaped GUIDs (`00000000-0000-0000-0000-0000000000a0`); the module manifest's own `GUID =` identity line |
| Email addresses and UPNs | RFC 2606 / RFC 6761 reserved domains: `example.com/net/org`, `.invalid`, `.test`, `.localhost` |
| `onmicrosoft.com` domains | none |
| JWT and bearer-token shapes | none |
| Assigned secret, password, or key literals | none |
| PEM private key blocks | none |
Two exemptions were added during this scan, both narrow and both for things that cannot be replaced
with a placeholder:
**Reserved domains.** `alex.employee@example.invalid` is guaranteed by RFC to be unresolvable.
Rejecting reserved domains would push fixtures toward addresses that merely *look* fake, which is
worse — the difference between "obviously synthetic" and "probably nobody's" is the entire reason the
reserved list exists.
**The module manifest GUID.** A PowerShell module manifest must carry a genuine unique GUID as its
identity; it is what distinguishes this module from another of the same name. It identifies the
module, not a tenant. The exemption is **line-level** (`^\s*GUID\s*=`), not file-level: exempting the
whole manifest would let a real identifier land anywhere in it.
## Verification of the scanner itself
A negative control was run: a scratch file containing an email address on a real-world commercial
domain and a randomly generated real-shaped GUID was added to the working tree **without** committing
it. The scan failed with two findings and named both, by file and line. The file was then removed and
the scan returned to PASS.
The offending values are described here rather than quoted, because quoting them would make this
record itself a finding — which the scan promptly demonstrated when an earlier draft did exactly
that. That is the control working.
Without a negative control, a scanner that had silently stopped matching would report the same green
result as one that is working.
## Result
```
Sanitization scan passed: no tenant data, credentials, or real identifiers found.
```
## Standing obligations
This is a point-in-time result, not a property of the repository. The scan is **gate 1** of
[pipelines/validate.yml](../../../pipelines/validate.yml) and runs before every other gate on every
pull request — deliberately first, because a leaked identifier is a problem whether or not the code
compiles, and every later gate prints file contents into build logs.
Runtime audit records legitimately contain real UPNs and Object IDs, which are approved for logs. No
such value may ever be committed. When attaching evidence to a verification record (V-1, V-2, V-3),
redact identifiers to placeholders first.