Implement Stage A: rule engine, validation, audit, and safety gates

Completes 109 of 121 tasks. Every remaining task needs a tenant connection
(T055, T056, T101-T103) or an Azure Automation account (T115-T121).

  354 offline Pester tests      PASS
  Engine purity (Principle IV)  PASS
  Sanitization (SC-013)         PASS  (156 files)
  Graph module loaded in tests  none  (SC-008 holds)

What landed
  - Four-layer configuration validation with stable finding codes, covering
    every VR-002 and VR-003 condition, plus a 23-fixture invalid-config corpus
  - Run loop, audit records (NDJSON through a single sink), summaries,
    reconciliation, and exit codes 0-6
  - Persistence behind a single write-body builder whose result always has
    exactly one key
  - Invoke-PersonaEngine.ps1 and Edit-PersonaEngineConfig.ps1
  - Six docs, two pipelines, traceability matrix, V-5a and sanitization records

Three deviations from tasks.md, each recorded in its status block

  T033 is not in Resolve-UserPersona. evaluationErrorThreshold is run-level
  state and the rule engine is pure; a counter there would break Principle IV.
  It lives in New-PersonaRunCounter and is applied in the run loop.

  A new src/Engine/ layer holds Invoke-PersonaEngineRun. The entry script
  imports the manifest, which requires Microsoft.Graph.Authentication, so a
  loop living only inside it could not run on a machine without the Graph SDK
  and SC-004 could not be proven at all. The entry script is now a thin
  wrapper and what ships is what is tested.

  The invalid-config corpus is generated by a committed script, with the
  generated fixtures committed too, so a reviewer sees the fixture in the diff.

Defects found by running the code, not by reading it

  Group and role ID lists were double-wrapped: @(Get-PersonaGroupIdPage ...)
  around a comma-returned array collapsed every membership list into one
  bogus space-joined entry. That is a silent false non-match, exactly what
  FR-013 exists to prevent.

  A 403 whose status appears only in the exception message parsed as $null,
  which the retry policy treats as a transport error - five requests per
  account against a tenant already refusing. Status extraction now falls back
  to the message text, bounded to 400-599.

  The sanitization scan walked tracked files only, so it covered 34 of 156
  files and none of this phase's code. It now scans untracked non-ignored
  files too, and a negative control confirms it catches a planted leak.

  Test-Json reports one error per violating location, not first-failure-only
  as the V-5a draft claimed. Record and pin corrected.

Enforcement remains blocked on the V-4 security sign-off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-20 21:48:19 -04:00
parent c59c85dd55
commit cdc6bb33d3
124 changed files with 16638 additions and 199 deletions
@@ -0,0 +1,57 @@
# V-5a — `Test-Json -SchemaFile` failure behaviour (local PowerShell)
**Status**: CLOSED
**Date**: 2026-08-20
**Environment**: PowerShell 7.6.5, Windows 11, local workstation (Stage A1 — offline, no tenant)
**Task**: T061
**Pinned by**: [tests/Configuration/TestJsonBehaviour.Tests.ps1](../../../tests/Configuration/TestJsonBehaviour.Tests.ps1)
## Question
OTD-005 selected `Test-Json -SchemaFile` for layer 2 validation. The documented risk was that
`Test-Json` reports schema failure inconsistently across PowerShell versions — returning `$false`,
writing a non-terminating error, or throwing. Layer 2 cannot be written until the actual behaviour
on the target build is observed rather than assumed.
## Observed behaviour
| Scenario | Return value | Error stream | Terminating? |
| --- | --- | --- | --- |
| Valid document | `$true` | empty | no |
| Type mismatch (`"a": 123` against `"type": "string"`) | `$false` | 1 error: `The JSON is not valid with the schema: Value is "integer" but should be "string" at '/a'` | no |
| Missing required property | `$false` | 1 error: `The JSON is not valid with the schema: Required properties ["a"] are not present at ''` | no |
| **Schema file itself unparseable** | **`$true`** | 1 error: `Cannot parse the JSON schema.` | no |
Exception type on the error record is `System.Exception` in every failing case — there is no
distinct exception type to branch on, so the wrapper must branch on the message text or, better, on
error presence alone.
## Findings that shape the implementation
1. **Non-terminating, not throwing.** With the default `$ErrorActionPreference = 'Continue'` the
cmdlet writes to the error stream and execution continues, returning `$false`. It does not throw.
`-ErrorAction SilentlyContinue -ErrorVariable` is therefore sufficient to capture failures, as
the editor contract requires.
2. **An unparseable schema returns `$true`.** This is the load-bearing observation. A wrapper that
trusted the return value alone would report a configuration as schema-valid when the schema never
ran. Layer 2 MUST treat "error variable is non-empty" as failure regardless of the return value,
and MUST distinguish the `Cannot parse the JSON schema.` message so it can surface exit code 4
(schema file not found or itself invalid) rather than exit code 1 (configuration invalid).
3. **One error per violating location, but not exhaustive.** Two independent property violations
yield two error records, each with its own JSON pointer. Deeper or nested subschema failures may
still be reported as a single error at the outermost failing location. The wrapper therefore
emits one finding per collected error rather than assuming a single one, and the author may still
need more than one validation pass to see everything. That residual limitation is documented
rather than worked around — full violation reporting would require replacing `Test-Json` with a
third-party validator, which OTD-005 rejected.
## Consequences recorded elsewhere
- Layer 2 implementation: [src/Configuration/Test-PersonaConfiguration.ps1](../../../src/Configuration/Test-PersonaConfiguration.ps1)
- Regression pin: `tests/Configuration/TestJsonBehaviour.Tests.ps1` fails if a future PowerShell
build changes any row of the table above.
- **V-5b remains open**: this observation is for PowerShell 7.6.5 only. The Azure Automation runtime
version is unverified, and finding 2 in particular is version-sensitive. Re-run this probe there
before Stage B (T116).
@@ -0,0 +1,76 @@
# Sanitization scan result (SC-013)
**Status**: PASS
**Date**: 2026-08-20
**Task**: T113
**Scanner**: [tests/Test-Sanitization.ps1](../../../tests/Test-Sanitization.ps1)
**Files scanned**: 156
## What is scanned
`git ls-files --cached --others --exclude-standard` — tracked files **and** untracked files that are
not gitignored.
The original scanner walked `git ls-files` alone, which covered only tracked files. That made the
gate useless where it matters most: a leaked identifier in a file that has not been committed yet is
precisely the one worth catching, and scanning only what is already in history means the scan passes
right up until the commit that makes it too late. At the time this was found, the scan was covering
34 of the repository's 156 files and none of the implementation written in this phase.
`--exclude-standard` keeps gitignored build output and local scratch files out, so the scan covers
exactly what a commit would add.
## Patterns
| Pattern | Exemptions |
| --- | --- |
| GUIDs | Placeholder-shaped GUIDs (`00000000-0000-0000-0000-0000000000a0`); the module manifest's own `GUID =` identity line |
| Email addresses and UPNs | RFC 2606 / RFC 6761 reserved domains: `example.com/net/org`, `.invalid`, `.test`, `.localhost` |
| `onmicrosoft.com` domains | none |
| JWT and bearer-token shapes | none |
| Assigned secret, password, or key literals | none |
| PEM private key blocks | none |
Two exemptions were added during this scan, both narrow and both for things that cannot be replaced
with a placeholder:
**Reserved domains.** `alex.employee@example.invalid` is guaranteed by RFC to be unresolvable.
Rejecting reserved domains would push fixtures toward addresses that merely *look* fake, which is
worse — the difference between "obviously synthetic" and "probably nobody's" is the entire reason the
reserved list exists.
**The module manifest GUID.** A PowerShell module manifest must carry a genuine unique GUID as its
identity; it is what distinguishes this module from another of the same name. It identifies the
module, not a tenant. The exemption is **line-level** (`^\s*GUID\s*=`), not file-level: exempting the
whole manifest would let a real identifier land anywhere in it.
## Verification of the scanner itself
A negative control was run: a scratch file containing an email address on a real-world commercial
domain and a randomly generated real-shaped GUID was added to the working tree **without** committing
it. The scan failed with two findings and named both, by file and line. The file was then removed and
the scan returned to PASS.
The offending values are described here rather than quoted, because quoting them would make this
record itself a finding — which the scan promptly demonstrated when an earlier draft did exactly
that. That is the control working.
Without a negative control, a scanner that had silently stopped matching would report the same green
result as one that is working.
## Result
```
Sanitization scan passed: no tenant data, credentials, or real identifiers found.
```
## Standing obligations
This is a point-in-time result, not a property of the repository. The scan is **gate 1** of
[pipelines/validate.yml](../../../pipelines/validate.yml) and runs before every other gate on every
pull request — deliberately first, because a leaked identifier is a problem whether or not the code
compiles, and every later gate prints file contents into build logs.
Runtime audit records legitimately contain real UPNs and Object IDs, which are approved for logs. No
such value may ever be committed. When attaching evidence to a verification record (V-1, V-2, V-3),
redact identifiers to placeholders first.