Repository-scoped · content-hashed · terminal-native

The gnomon is the fixed blade on a sundial.

The sun travels. The shadow sweeps the hours. The gnomon does not move — and that is precisely why the dial can be read. Remove the blade and you have decorated stone, not an instrument.

This harness names itself after that object. .gnomon/ is the fixed configuration — models, roles, tools, approval — content-hashed and committed with your repository. Everything else is shadow.

cd my-project && gnomon launch

v0.2.0 · 1,000+ TypeScript tests · 57 Rust · green on Linux, macOS and Windows

XII I II III IV V VI VII VIII IX X XI N .gnomon/ blade — fixed · content-hashed shadow — sweeps the plate · everything else varies

Fig. 0 — Plan view (looking down). Fixed gnomon; black umbra projected from the blade.

Why this name. Behaviour is readable because something holds still. The harness puts that something in a directory you commit.

Measured — properties first, rates with their caveats

On task completion, no measurable difference
Same 34 tasks, same model, same clock, same night. Four arms null — peer, chain, model tier, timeout instruction
50.0% vs 47.1% · p = 1.0000
Sends a quarter of the context
opencode sends 36,490 bytes to gnomon's 7,824 for the same prompt — counted off the wire. Most of the gap is tool schemas (20,812 vs 4,312), re-sent every turn
4.66× less context than OpenCode
Surface hash means what it claims
No path changes behaviour without moving the hash — exhaustive, not sampled
13/13 paths faithful · 0 false negatives
Cannot rewrite its own permission surface
Decided from filesystem state. OpenCode rewrote its permission file 5/5 across four escalating configs
gnomon 5/5 contained · opencode 0/5
Full audit trail with no speed penalty
Off by default and free when off. Hash-chained log on every turn when on
338ms vs 354ms — within noise
The harness itself adds almost nothing
Interactive sessions pay startup once
~33ms logic · ~200ms boot

Property suites describe the current tree · rates lag the code and name their commit · not a leaderboard · full measurements §7

Ask most coding agents why did you do that and there is no answer. Configuration scatters across dotfiles and machine state — the same repo behaves differently for two people.

One directory declares models, roles, tools, approval policy. Content-hashed. Stamped on every record the harness emits. Linux, macOS and Windows — same surface, same hash, all three in CI.

On task completion gnomon is indistinguishable from OpenCode at this model tier. Everything it adds — a hashed surface, capability separation, a tamper-evident trail — costs nothing in capability. The model varies. The conversation wanders. The tools differ each run. .gnomon/ does not.

The real path, in order. Slash commands stay local. Everything else: role → skills → context → tools under approval policy.

Three guarantees in this path:

  • Tools absent, not discouraged. Schema list matches role exactly.
  • Nothing mutating runs unseen. Diff before write; bash gated.
  • Three buckets. result · refusal · apparatus_failure
you type a line slash command? local · no model pick role · load skills · build context manual · suggest · auto send system + history + tool schemas THIS role only tool call? answer approval gate? refusal execute · sandbox · feed back next step record bucket · session · audit result · refusal · apparatus_failure

Fig. 1 — Agent turn loop. Slash commands stay local; everything else routes through role, approval, and honest outcome buckets.

Schema enforcement. Verifier has no write — absence, not instruction.
Approval before mutation. Real diff. Standing approvals in audit trail.
Honest outcomes. Declined = refusal. Timeout = apparatus failure.

Rust owns verifiable parts. TypeScript owns the loop. Sessions and audit live outside the hashed surface — logs inside .gnomon/ would change the hash every turn.

.gnomon/ — content-hashed config.toml routing · endpoints · audit roles.toml · tools.toml policy.toml · system.md skills/ surface_hash TypeScript — loop gnomon-core turns · tools · context · skills gnomon-cli gnomon-surface · edit · exec (Rust) outside surface .gnomon-sessions/ .gnomon-audit/ must not alter hash varies

Fig. 2 — Rust owns verifiable surface, edit, and exec. TypeScript owns the loop. Sessions and audit sit outside the hash.

The blade is .gnomon/: fixed, identifiable, content-addressed. Replies, tool traffic, session history — the shadow.

Rust and TypeScript compute the same surface hash; a test holds them together. conformance/ pins exit codes, enumerations, manifest shape.

Workflows where repository-scoped behaviour matters — not generic productivity claims.

Team lead · config review

Review agent behaviour in PRs.

bash_allow changes in diff. Hash updates on merge. Surface self-escalation refused 5/5 — OpenCode rewrote its own permission file 5/5.

PR .gnomon/ → review → merge → new hash
Post-incident · oversight

What was permitted? Who approved?

Off by default, free when off — 338 ms vs 354 ms with it on. Hash-chained trail carries the surface hash; gnomon audit verify names the sequence that broke. Primitives for oversight, not a compliance guarantee.

8/9 tamper attacks caught · the 9th is published
Solo dev · local models

Ollama on laptop; same agent on CI.

Surface committed. Clone anywhere — same roles, tools, gate.

init → roles.toml → launch → task in CI
Brownfield adoption

Point a verifier at a repo you did not write.

Pilot: 10 of 14 planted defects found, 0 of 5 adversarial controls falsely flagged, 0 containment violations — local models, $0 API. Reading code flips less often than making a task pass.

7.1% flip rate · two synthetic projects · no peer
Harness author · embed

Pipeline via gnomon task --json.

Published exit codes. Bucket from exit value. Degradations announce on stderr and land in TaskRecord.degradations[].

task --dir repo --yes --json
Specify → verify

Roles separated by capability.

Coordinator can't edit. Verifier can't write. Chain can stop on a failed check — not on an opinion.

/spec → coordinator → implementor → verifier
A poor fit if you want IDE integration, HTTP/SSE MCP transports, cloud or async execution, or the broadest model support. Those are real requirements and other harnesses serve them better — OpenCode for a far larger ecosystem, Aider for Git-centric diff discipline, OpenHands for sandboxed autonomous execution.

The verifier has no write and no edit, so it cannot alter what it judges — and because bash can write anything, bash_allow narrows it to test commands. task runs a sub-turn under that role's tools, so a role that may not write also may not delegate to one that can.

RoleCannot
coordinatoredit; outside write_allow
implementorgated by write_allow / bash_allow / bash_deny
verifierwrite, edit; bash_allow for tests only

Each boundary has a test that fails if it moves. Other harnesses have permission prompts; these are declared per role, hashed, and testable.

The chain stops on a failed check — not on an opinion. One dial, three positions: never, on_refusal, on_check. A stage whose declared check ran and failed stops the chain. A stage that merely disagrees in prose does not — reading its sentence would be instruction, not capability. Default: never. Not yet re-measured — the earlier arm tested a chain that could not gate.

  1. No machine-scoped config. All in .gnomon/.
  2. Manifest every session. Content-addressed.
  3. Declared tool schemas. Refusal, not shorter list.
  4. Three buckets. No composite verdict.
  5. Published exit contract. 0–1 / 2–4 / 10–13.
  6. Published enumerations. README tested.
Read before citing. At this model tier the harness is not the bottleneck on task completion — four arms came back null. What follows are the exhaustive property results that describe the current tree, then the rates that lag it. Methodology and raw data live in the repo. Not a leaderboard.

I — What the hash and the trail actually guarantee

Exhaustive, deterministic, $0. These re-run on every change and describe the build you can download.

ClaimResultSource
Surface-hash fidelity13/13 faithful · 0 false negativessurface-fidelity
Declared degradations announced & recorded14/14degradation-contract
Injected faults disclosed by name8/8fault-disclosure
Silent success at decision points0/11 falsely successfulsilent-success
Prompt injections crossed0/6 · deliveries verified readinjection-2026-09-01
Determinism across locale, tz, cwd, $HOME, mtimes10/10 identicaldeterminism-2026-08-31
Surface self-escalation refusedgnomon 5/5 · opencode 0/5containment-2026-08-31
Audit tamper attacks caught8/9 · full re-chain published as a limitauditability-2026-08-31
Context on the wire vs opencode4.66× · 7,824 B vs 36,490 Bcontext-cost
Audit trail on vs off338 ms vs 354 ms — within noiselatency-2026-08-31
Startup (tsx boot vs gnomon logic)~197 ms vs ~33 mslatency-2026-08-31

Retires the earlier “13–43× leaner than OpenCode” figure, which multiplied a token ratio by a retracted pass-rate ratio. Bytes, not tokens — the ratio is not a tokenizer artifact.

II — Task completion — null, and that is the claim

Sampled rates cost money and lag the code. Each is attributable to the commit named in its result file — none of these rates were measured against v0.2.0.

ArmResultSource
Peer vs OpenCode (equal terms)50.0% vs 47.1% · McNemar p = 1.0000 · 34 pairedpeer-opencode-2026-09-02
Consistency (pass^2)pass@1 51.2% → pass^2 45.2% · retention 0.88reliability-passk-2026-09-05
v0.1.1 Terminal-Bench44.7% on 47 tasks · ~41% of trials hit the timeout capv011-timeout-2026-09-03
Role chain / model ceiling / timeout teachingall nullEVIDENCE.md

About one apparent success in eight does not reproduce. Goose is not a valid current baseline — listed under claims with no evidence until re-run. An earlier 18-point “win” vs OpenCode is an artifact of unequal adapters and must not be quoted.

III — Brownfield audit pilot

Two synthetic projects, no peer, $0 local. Supports the brownfield use case — label it a pilot.

EndpointResult
Planted defects found10/14
Adversarial controls falsely flagged0/5
Containment violations0
Flip rate7.1% — half this harness's task-completion flip rate

A harness that hides its gaps is worse than one that has them.

  • MCP: stdio only. Pinned server → discover tools → gate per role. Reproducibility bounded — gnomon pins invocation, not what the server does.
  • Not a hosted or async service. Runs in your terminal — and from there reaches cloud models, authenticated cloud CLIs (gh, az…), and the web. It doesn't run itself as a background job on someone else's server: no queue, no worktree pool. One unattended path: cron-scheduled loops — ticks on the scheduler.
  • Terminal only. No IDE.
  • Path sandbox, not process isolation. Enforced for webfetch (SSRF guards). bash still reaches the network.
  • Summary compaction not reproducible.
  • Task-completion rates lag the release. No scored arm has run against v0.2.0 yet. Property suites have.

Linux, macOS and Windows — all three in CI, every commit. Windows runs natively, not through WSL. It needs Git for Windows for its POSIX shell, and gnomon refuses rather than falling back to cmd.exe — a shell that changes with the OS would be machine-scoped behaviour the hash cannot see.

The blade does not move. What changed in this release is where it can be planted, and how much of what goes wrong now reaches the record instead of the floor.

Windows, natively. Not through WSL. Git for Windows supplies the POSIX shell; with none found, bash refuses and says how to get one. Full suite on all three OSes in CI.
Approval prompt race closed. write/edit re-read after approval and refuse naming drift — no silent discard of a file touched while you decided.
Degradations reach scripted runs. Announce on stderr regardless of --json; TaskRecord carries degradations[]. All 14 declared paths announced and recorded.
The chain can stop. [chain] gate: never / on_refusal / on_check. Failed check stops the chain; prose disagreement does not.
Approval gate: v and ?. Full preview past the 60-line cap; ladder help. Neither consumes an attempt. Typed answers, not single keystrokes.
gnomon migrate. One command brings an existing surface up to date. Breaking for embedders only: agent.ts and its exports are gone.
Read this before citing a rate. No task-completion number published by this project was measured against 0.2.0. The property suites were — they are exhaustive, deterministic and free, so they re-run on every change and describe the build you can download. The sampled rates cost money and lag the code; each is attributable to the commit named in its own result file, and to nothing else.