Ota on Eris: What the Execution Evidence Proved
A bounded Eris verification slice passed on native and container lanes. The maintainer declined adoption. Both outcomes matter: execution evidence must stay scoped, and another tool must earn its upkeep.
Passing tests answered only one question
The bounded Eris verification slice passed. That did not establish a working model-driven runtime, and it did not establish that Eris needed Ota.
After reviewing the experiment, maintainer Jan Paul Dahlke chose to leave the draft PR open as a reference rather than merge it. Cargo and the existing CI already covered his immediate needs; another contract layer had not earned its upkeep. The engineering result and the adoption result are different, and this note records both.
The evidence below comes from 2 September 2026, using Ota v1.6.27 at an exact Eris fork revision. Jan's public decision followed on 7 September. This is a historical pressure report, not a fresh verification of current Eris or the latest Ota release.
Eris is a local-first Rust agent with a Gatekeeper, Markdown-vault access, model-server dependencies, semantic memory, and a large source-owned test matrix. At the inspected revision, its setup surface spanned more than compiling one binary: a deployment also depends on model artifacts, an inference server, Qdrant, hardware posture, and operator-owned state.
That makes Eris a useful pressure case for Ota. The easy mistake would be to run cargo test, get green output, and describe the repository as ready. The useful question is narrower and harder:
Can an agent discover and execute a reviewable repository contract while Ota remains explicit about everything that contract did not exercise?
The retained run supports that selected source-verification slice. It does not support a claim about the complete Eris runtime or an autonomous agent's behavior.
What the contract owns
The proposed contract at the tested revision declares one agent entrypoint, verify, for the selected native source boundary. It includes:
- lock-enforced Cargo dependency hydration into a repository-local cache;
- the Eris binary and test build;
- Clippy over the binary target;
- the source-owned
cargo test-fullrunner, which reported 47 batches; - an exact disposable-vault read fixture;
- an exact Gatekeeper-allowed operation;
- an exact Gatekeeper-refused operation.
The contract tells agents to begin with ota tasks --safe --use, use --agent on every selected execution, and treat Cargo.lock and human-owned workflows as protected. The CI lane separately checks for lockfile and tracked-file changes; those checks are not proof that an agent cannot write outside Ota. The workflow is separate and non-blocking. It uses pull_request, read-only repository permissions, commit-pinned actions, and Ota v1.6.27 from the contract rather than a moving development branch.
Two boundaries must not be confused. Ota's --agent path evaluates admission for the selected repository tasks. The Gatekeeper tests exercise Eris's own Rust policy code. In particular, test:gatekeeper:refused succeeds when an Eris fixture observes the expected denial. It is not an Ota refusal canary, a live model test, or proof that an agent cannot bypass its instructions.
For an agent following that contract, the intended selection path is deliberately small:
ota tasks --safe --useota up --workflow verify --native --agent --dry-run --jsonota run verify --agent --streamota run test:vault:read --mode container --agentota run test:gatekeeper:allowed --mode container --agentota run test:gatekeeper:refused --mode container --agentThe first command discovers the declared agent-callable surface. The preview selects the native workflow without starting execution. The native run invokes the selected verification closure. The three container commands exercise only the controls that explicitly declare that mode. This command sequence is not an independent attestation of everything an agent ran elsewhere.
Native evidence and repository findings
The native lane was split across Ubuntu and macOS. It validated the contract, recorded Doctor output, exposed the agent-safe task surface, previewed the workflow, and ran the bounded verify closure through Ota.
The retained fork matrix ran the same contract at revision a8237b2cdfcdd651bec6d82f2d7576a531176fb5 on Ubuntu and macOS. Both hosts used an Ota binary identifying as v1.6.27 at Core commit 918578d3e0370cc7bb2366b6e831b6f19e58e2a6, with Rust 1.98.0. The native previews reported READY / RUNNABLE without starting execution, and both verification logs recorded all 47 batch completions.
The jobs were deliberately non-blocking and used continue-on-error. A green workflow badge alone would therefore be insufficient. The evidence is the retained outcomes.txt, preview JSON, verification logs, and image record, checked against the same revision. The native outcomes record successful validation, diagnosis, task discovery, preview, and verification, while separately retaining formatting failure.
Doctor reported the repository as risky while reporting the agent boundary as ready. That is not a contradiction: the selected tasks were admitted, while ten advisories retained the external network dependency-hydration boundary. Neither label certifies safe model behavior or a working deployment.
The matrix also exposed two repository facts that the integration left unchanged:
cargo fmt --all -- --checkreports existing formatting drift, so formatting remains visible but advisory;- Clippy exits successfully while reporting 113 existing warnings.
A third repository boundary prevented an honest Windows lane. The tested upstream tree contained a path with double quotes in its filename, which Git cannot check out on an NTFS-backed Windows runner. That failure occurs before Ota can validate or execute the contract. The integration therefore did not claim Windows execution. This is a checkout limitation at that revision, not a general claim that Eris or Ota cannot run on Windows.
Why the container lane is separate
Container execution is valuable here because a contributor or agent should be able to exercise a portable subset without inheriting the host Rust installation. But "runs in Docker" is not enough evidence by itself.
The Ota container context is Linux-only, linux/amd64, ephemeral, and bound to a digest-pinned Rust 1.95 Bookworm image. Its Cargo cache is isolated from the native cache. The lane executes the three portable controls directly through ota run --mode container --agent:
- the disposable-vault read;
- the Gatekeeper-allowed operation;
- the Gatekeeper-refused operation.
All three controls passed in the hosted container lane. The retained image record identifies Linux/amd64 and the exact repository digest rust@sha256:6258907abe69656e41cd992e0b705cdcfabcbbe3db374f92ed2d47121282d4a1 with Rust 1.95. Cargo.lock and the tracked repository remained unchanged across the selected commands. This supports the selected container context, lock-enforced Cargo commands, exact fixture invocation, and expected Gatekeeper outcomes for that slice. It does not validate Eris's Dockerfile, install models, start llama-server, start Qdrant, exercise a GPU, or prove that a real model obeyed a policy.
That distinction matters. Ota governs the repository execution it selects. It does not convert a fixture into runtime attestation for an external model, or audit commands run outside that path.
What Eris changed in Ota
The first ota detect pass found genuine build and test signals, but Eris exposed two Core consistency defects during the early pressure work.
First, a named GitHub Actions step containing ${{ matrix.batch }} was promoted as runnable task truth even though Ota could not resolve the matrix value. The correction retains that body as unresolved evidence instead of inventing an executable command.
Second, toolchains.rust.version: stable passed validation while Doctor compared it literally with the installed semantic version. The correction aligns validation, Doctor, and fulfillment: diagnose-only requirements must be semantically comparable, while rustup-owned execution may still use channel names. Both corrections are documented in the v1.6.28 changelog.
The retained v1.6.27 exercise also exposed six modeling and evidence pressures:
- Typed Cargo hydration did not express
--locked; the contract instead declared the structured commandcargo fetch --locked. - The contract kept the portable container tasks separate rather than treating its native
verifyaggregate as a container closure. - The native preview needed explicit
--native; an implicit selection had made the macOS lane look blocked by an unselected container context. - Doctor repeated inherited dependency-hydration advisories across parent tasks, suggesting a more compact presentation of one transitive cause.
- The 47-batch suite remains source-owned shell orchestration. Ota retains every batch marker and terminal result, but it does not independently model each command inside that script.
- The hosted container logs record each ephemeral container identity and selected image, but do not retain terminal cleanup evidence proving those container identities were removed.
These observations preserve the tested contract's limits; they are not a current backlog or a claim that every limitation still exists in later Ota versions. A newer release requires its own selected contract and evidence rather than inheriting this run's result.
The uncovered behavior inventory
The ownership map is deliberately explicit:
- Contract-owned and exercised: selected hydration, native build and verification, and the three isolated container controls.
- Repository-owned: the internal 47-batch script, Eris's policy assertions, formatting baseline, Clippy findings, and checkout-compatible filenames.
- Ota modeling/evidence pressure: lock-enforcement representation, mode selection, advisory presentation, internal batch granularity, and terminal cleanup evidence in this historical slice.
- Not proved by this run: the runtime, isolation, lifecycle, and adoption claims below.
The bounded pressure slice does not prove:
- installation, startup, identity, or compatibility of
llama-server, Ollama, or GGUF models; - model inference or grammar enforcement;
- Qdrant startup, readiness, persistence, or semantic-memory correctness;
- live TUI, web, chat, or operator-vault behavior;
- the ignition wizard, vault sealing, or real vault writes;
- hardware or GPU profiles;
- external integrations or credential delivery;
- the maintainer's private
feature/shippinglifecycle; - independent modeling of every command inside the source-owned 47-batch shell runner;
- terminal container-cleanup evidence after the three ephemeral controls;
- operating-system enforcement of every declared writable/protected path, or compliance by an agent running commands outside Ota;
- packaging, installation, or repository-wide readiness;
- Windows execution.
Those are not footnotes to a green badge. They are the scope map for this exercise. Each future slice should either become contract-owned and exercised, stay explicitly outside the selected scope, or expose a concrete Ota capability gap.
Passing verification did not earn adoption
Jan's public review was clear: Cargo, Cargo.lock, and the existing CI were sufficient for his immediate verification needs. Keeping a solo project's maintenance surface small mattered more than adding another governance front-end. He kept the draft open as a reference, with no plan to merge it at that point.
That is an adoption result, not a failed test or an endorsement. Ota must remove a recurring operator burden that the existing tools leave unresolved, and justify the cost of maintaining its contract. A careful integration and passing controls alone do not establish that value.
The experiment produced a reviewable execution surface and bounded evidence. It did not produce upstream adoption, retained use, payment, or proof of a complete agent runtime. That distinction is the result worth publishing.
Evidence
- Eris upstream
- Draft adoption PR
- Tested fork snapshot
- Exact contract
- Exact evidence workflow
- Final fork matrix
- Maintainer's adoption decision
The hosted artifact names and retained ZIP digests are:
eris-ota-ubuntu-latest:sha256:caef86d5d34b6aa4c70b159bf04b5b757d7863ab6ee410d398fe780af188e408eris-ota-macos-latest:sha256:436fd4431db5b30e07ef564ef9ee93d144a53f17b1476423fc36885d1bf0a378eris-ota-container:sha256:147a7879f0fdc51528679a5003d913b2e4b5b85935af5be5f4cab59f508ac81dota-readiness(advisory):sha256:12bb3e31c95d269d169a7ff9d65f8bd4359f4b7411b12ef00451d98308abc261
Artifact downloads depend on GitHub's retention and access rules. A digest identifies an archive; it is not an independent attestation, a permanent download guarantee, or a substitute for reading the results inside it.
Take action