Engineering reference MVP2 / foundation

Testing & verification

Compilation and fixtures are necessary; real pre-execution behavior is a separate acceptance gate.

Development previewMVP2 foundation complete. Live hooks, enforcement, and Anthropic BYOK are pending. Planned behavior is labeled separately from working features.

Current local evidence

The implementation record dated October 8, 2026 reports debug and release builds of all three arm64 executables with Swift 6 and SDK 27, plus 43 passing Swift Testing cases (23 retained tests and 20 new security contract tests). Native rendering checks cover 20 light/dark cases, including demo/real separation, pending/expired review, and denied notifications.

Scene launch and relocated bundle checks validate a visible dashboard and packaged resources. Full Xcode archive validation, manual VoiceOver coverage, and actual notification delivery acceptance remain pending. These are development results, not efficacy claims.

Run the relevant checks

./scripts/test.sh
./scripts/smoke-test.sh
./scripts/bundle-smoke-test.sh
./scripts/ui-smoke-test.sh

Scripts build and exercise harmless fixtures. Native render artifacts are under build/ui-smoke/. The UI renderer checks appearance and state behavior, not pixel equality with an approved design.

The checked-in implementation status explains the Command Line Tools adjustments and remaining release gates.

Required live-phase tests

  • Malformed, oversized, null, invalid UTF-8, deeply nested, and concurrent host input.
  • Host output encoding and preservation of native permissions.
  • Path normalization, shell ambiguity, quoting, symlinks, and benign cleanup near misses.
  • Secrets in strings, nested objects, fragmented tokens, Unicode, logs, and errors.
  • Model errors, deadlines, budgets, hallucination, and evidence precedence.
  • Exact approval scope, expiry boundaries, replay, double-click, disconnect, restart, and simultaneous reviews.
  • SQLite migrations, foreign keys, permissions, retention, and transactional races.
  • Installer ownership, idempotency, preservation, concurrent edits, backups, and interrupted writes.

Prove the body never ran

Both Claude Code and Codex must exercise a real local hook on exact verified releases. A harmless denied operation in a disposable directory must leave no execution marker. Codex tests require user-reviewed trust before the callback is counted.

Review tests must cover Block, Allow once with host permissions intact, expiry denial, and an old notification. Failure tests kill the UI and service, remove the helper, disable hooks, force timeouts, and confirm honest degraded status. Privacy tests revoke keys, simulate network errors, and prove unsafe payloads are never sent.

Never use real secrets, destructive targets, or actual exfiltration as test material.

The release evaluation corpus

MVP1 requires at least 40 sanitized scenarios: 10 catastrophic, 10 high-risk review, 10 benign near-misses, and 10 prompt-injection or drift cases. Each defines the gold outcome, evidence, task anchor source, and false-positive tolerance.

Version prompts and schemas, and rerun the corpus after changes. Track measured precision, missed high-risk cases, latency, calls per 100 actions, and token use. Do not publish a security efficacy percentage without measured evidence.

Performance targets, not promises

MeasureEngineering target
Ordinary local no-override, added timeMVP2 warmed-path target: p95 under 150 ms
Local rule decisionp95 ≤ 25 ms
Contextual review8 s internal model deadline
High-risk approval45 s default, within verified host timeout
Concurrent model reviewsDefault maximum 3

These targets need on-device verification and are not observed production measurements or network SLAs.

Based on the MVP1 specification, the additive MVP2 specification, and the acceptance matrix · October 8, 2026.