← Back to blog
Engineering note2026-10-09 10:10 UTC

Pressure-testing Ota on PythiaLabs: decision evidence is not execution authority

An eleven-lane PythiaLabs exercise exposed a real formatting failure and Ota discovery defects, without confusing decision evidence, smoke tests, or CI results with execution authority.

A decision is not an execution grant

PythiaLabs combines an Elixir decision engine, Rust components, an MCP bridge, a generated site, and Python conformance suites. The useful question was not simply whether its tests passed. It was what each result established: a decision, a selected repository task, or an effective runtime boundary.

A Pythia ALLOW result is not, by itself, an Ota execution grant. A future integration would need explicit authority and binding to the action, resource, scope, subject, freshness, policy, decision identity, and one-use consumption. These runs did not demonstrate such an integration.

The pressure contract made that distinction concrete: agent.safe_tasks: []. Humans and CI could exercise reviewed lanes, but no lane was declared callable by an agent. Passing a test did not widen that permission boundary.

Exact inputs, including the red result

The historical pressure ran on 29 August 2026 against these inputs:

InputExact revision
Reviewed upstream repository17df87775c0d
Native and discovery forkbf51a0b4c931
Container fork916f9f127c9d
Source-built Ota, pre-release v1.6.27cd99c9abd2c0

These are historical source-built results, not released-binary validation or a test of today's upstream checkout. Source pins also do not freeze every transitive dependency or toolchain input.

The native matrix ran eleven explicit lanes: Elixir format and tests, MCP smoke, Rust worker build and tests, site build and formatting, and four Python conformance families. Ten returned exit code 0. verify:site-format returned 1: the existing prettier --check . reported 13 unformatted files. The overall native run failed.

That was a repository-quality finding, not an Ota execution failure. The pressure work retained the failure rather than reformatting upstream files or weakening the command to obtain green CI. Each lane retained its own log and exit-code file, so a reviewer could distinguish successful execution from the failed formatting check.

The separate declaration and discovery run passed its required gates: contract validation, task-posture export, and review-only candidate generation. Doctor was recorded but was not a required pass gate. Its retained result was exit code 1 and not_ready: that runner lacked Elixir and Mix. The execution runner installed those tools separately. The green discovery job therefore did not establish environment readiness, resolve formatting, or apply the candidate.

Two container lanes, not a runtime attestation

The fork did not use Pythia's Dockerfile as its reviewed container environment. At the examined revision, that Dockerfile installed Rust through a mutable network bootstrap.

Instead, an explicit Linux/amd64 Node context pinned node:20-bookworm by SHA-256 image digest. The container matrix passed only verify:mcp-smoke and verify:site-build, including their selected dependencies. This exercised context selection, the site's cwd: site, and lockfile-backed npm hydration and build.

The MCP smoke harness used a test-owned fake Mix executable. It checked JavaScript framing and routing, not a real Elixir evaluator or an MCP host enforcing a decision. The container pass also did not validate the upstream Dockerfile, establish image provenance beyond digest identity, prove cross-platform parity, or make either task agent-safe.

Discovery must preserve the actual lane

The repository exposed Ota discovery defects: a site build inferred at the repository root instead of site/, a multiline MCP check reduced to its first command, and incomplete recovery of the distinct verification lanes. These were Ota findings, separate from Pythia's formatting failure.

A separately recorded local implementation-branch replay at the reviewed upstream revision retained cwd: site in the detected build and candidate closure. The complete multiline MCP body remained unknown. Of eleven candidate changes, five were applicable and six unknown; apply-candidate --require-complete returned candidate_incomplete and wrote no ota.yaml.

That local replay is not the hosted matrix or released-binary proof. Its useful property is conservative admission: incomplete discovery exposes review work instead of manufacturing a runnable task or agent permission.

Uncovered behavior stays explicit

The selected lanes do not stand in for the whole repository:

BehaviorClassification and boundary
Selected verification tasksContract-owned and exercised; ten native lanes passed, one failed, and two Linux container lanes passed.
Upstream formatting defects, NIF fallback, and contributor test skipsRepo-owned behavior; a fallback-compatible Elixir pass does not prove native NIF compilation and loading.
Real MCP evaluator, host loading, and enforcementnot_proved; the smoke harness is not runtime attestation.
Credentialed CAEP, provider calls, merge, communications, and database effectsnot_proved; local conformance is not external authority or production safety.
Pages deployment and Liminal multi-repository lifecycleRepo-owned behavior outside the selected contract slice; not exercised.
Working-directory and multiline discoveryNamed Ota defects with bounded local repair evidence; typed Mix/Hex hydration remained an Ota capability gap in this pressure slice.

The record covers runs through Ota's selected execution path. It does not audit arbitrary agent commands outside that path or establish repository-global governance.

Upstream review is a separate result

The work subsequently reached asynchronous upstream review in draft PR #264. As checked on 7 October 2026, it remained open and unmerged at 587e722c7c7a. Its later released-pin evidence is separate from the August source-built runs described here. Review is not adoption, endorsement, or ongoing use.

The transferable method is simple: name the lane, retain its exact inputs and result, preserve real failures, and state what the check did not establish. A decision, a smoke test, and a green job are useful evidence only when their boundaries remain visible.

Evidence