blob: a1e73707d5a5814c495f0af972ce4909929c8c94 [file] [view]
# Maka AHE Target Protocol
Maka exposes an AHE-facing target contract so AHE can run the outer
self-iteration loop without becoming part of Maka's interactive runtime.
Maka remains the source of truth for runtime events, tool execution,
permissions, artifacts, and official evaluation feedback.
The source-backed TypeScript contract lives in
`packages/headless/src/ahe-target-protocol.ts`. Consumers should use the
`@maka/headless/ahe-target-protocol` subpath after build.
## Protocol Files
AHE should treat these files as the initial file protocol for a Maka target
snapshot and candidate iteration:
- `target-snapshot.json`: a `MakaAheTargetSnapshot` with
`protocolVersion: "maka.ahe-target.v2"`, the source label, content-addressed
snapshot id, SHA-256 source manifest, optional Git metadata, and the
component map from `MAKA_AHE_CURRENT_COMPONENTS`.
- `harness-results.json`: a `MakaAheHarnessResults` object containing per-task
`maka.ahe.run_result.v1` rows. Each current row names its `taskRunId` and
content-bound execution-lineage ref. `official_pass` and `official_fail` are
valid only when `scoreAuthority` is `official_verifier` or `official_scorer`.
- `trace-index.json`: a `MakaAheTraceIndex` that maps every task id to Maka
execution lineage, payload-safe Task Events and AgentRun inspections,
optional canonical Runtime Event sources, messages, transcripts, tool
results, and artifacts.
- `change-manifest.json`: a `MakaAheChangeManifest` describing the staged
patch, source-backed components changed, failure evidence, root cause,
targeted fix, predicted fixes, risk cases, validation dataset, and rollback
criteria.
- `change-evaluation.json`: a `MakaAheChangeEvaluation` comparing baseline and
candidate official cells, including transition labels, observed and missed
predictions, regressions, infra/excluded tasks, and self-check-only tasks.
All refs are opaque file/blob/url refs. The protocol does not require AHE
Python, NexAU, E2B, ADB, tmux, or Rive code in Maka runtime.
## Evidence Export
After recording Headless task runs, export a source-backed evidence snapshot with:
```sh
maka eval ahe export <taskRunId...> \
--store <out>/runs \
--repo <maka-repo-root> \
--out <evidence-dir> \
[--run-id <id>] \
[--source-label <label>] \
[--harbor-trial-dir <dir>] \
[--include-events]
```
The command validates the component map, reads existing `TaskRunProjection` rows, and writes the target snapshot, harness results, trace index, and per-run trace exports. It does not run AHE, apply patches, or change the Desktop runtime.
Each trace has a ledger-explicit layout:
```text
traces/<taskRunId>/
├── task-run.json
├── task-events.jsonl
├── execution-lineage.json
├── messages.json
└── agent-runs/<agentRunId>/
├── inspect.json
└── runtime-events.jsonl # only with --include-events
```
`task-events.jsonl` is the payload-safe Task Event export; it is never labeled
as Runtime Events. `execution-lineage.json` binds the target snapshot, Task
Event coverage, attempts, AgentRuns, and immutable Runtime Event coverage using
the shared `ExecutionEvidenceRef` contract. Missing links, source files, or
coverage are exported as explicit gaps. AgentRun `inspect.json` is payload-safe
by default. `--include-events` opts into raw canonical Runtime Event payloads.
If requested Runtime sources are unavailable, the lineage records
`requested_with_gaps` rather than claiming that raw evidence was included.
Every materialized local file ref carries its SHA-256 digest and byte size.
The v2 snapshot identity is derived from the exact bytes of every registered
source ref plus the canonical component topology. `createdAt`, `sourceLabel`,
and optional Git metadata do not affect the id: the same managed target bytes
produce the same id, while any managed source change produces a new id. Source
refs must resolve to regular files inside the repository; symbolic-link escapes
are rejected. Readers can still validate legacy v1 snapshots, but v1 has no
content-binding claim and is never reinterpreted as v2.
`--harbor-trial-dir` imports official verifier/scorer evidence for exactly one task run. Without it, a single-run export also detects the standard Harbor trial layout relative to the task-run store. Local checks remain non-authoritative evidence.
Programmatic callers use `buildMakaAheTargetSnapshot`, `makaAheEvidenceFromTaskRunProjections`, `writeMakaAheEvidenceExport`, and `validateMakaAheSourceRefs` from `@maka/headless/ahe-evidence-export`. Snapshot readers should use `validateMakaAheTargetSnapshot` from `@maka/headless/ahe-target-protocol` before trusting the document's internal content binding.
## Current Component Map
The first target component map is intentionally current-state only:
- `maka-system-prompt`: desktop prompt and workspace instruction sources.
- `maka-heavy-task-policy`: heavy-task policy and benchmark wrapper
expectations.
- `maka-tool-contracts`: Maka tool descriptions, input schemas, gating, and
selected desktop wrappers.
- `maka-context-management`: context budget, tool-result pruning, and semantic
compaction.
- `maka-permission-policy`: permission modes, pre-tool-use policy, runtime
enforcement, and dynamic tool availability.
- `maka-runtime-evidence`: canonical runtime events, AgentRun records, and
runner output. This is evidence rather than an editable harness component.
- `maka-headless-evaluation`: headless result format and Terminal-Bench smoke
runner protocol.
AHE manifests must use component ids from this map. Future product surfaces can
be added only when they have source-backed contracts.
## Patch Gate
AHE may propose a patch only through `change-manifest.json` with
`patch.applyMode: "staged_patch"`. The manifest must include:
- at least one changed component id;
- failure evidence from baseline traces/results;
- a root-cause statement and targeted fix;
- predicted fixed cases and risk cases;
- a validation dataset with task ids;
- rollback criteria.
Maka validators reject manifests that cite unknown components or omit
falsifiable evidence. When `patch.changedFiles` is supplied, every path must be
a repo-relative POSIX path, must avoid generated/dependency/repository-control
content, and must be listed as a source ref on one of the editable changed
components. Evidence-only components such as `maka-runtime-evidence` cannot be
patched. A self-check can be recorded, but it cannot be reported as
`official_pass` or `official_fail`.
## Evaluation Feedback
Candidate evaluation must compare baseline and candidate cells with official
verifier/scorer authority wherever pass/fail is claimed. The transition matrix
is the feedback AHE should use for the next iteration:
- `fail_to_pass` confirms a fixed baseline failure.
- `pass_to_fail` is a regression and should trigger rollback criteria when it
matches the manifest.
- `infra_or_excluded` and `selfCheckOnlyTaskIds` are accounting buckets, not
official success.
The protocol is deliberately file-first. `maka eval ahe export` generates
and validates the source-backed export boundary. Import and candidate-evaluation
CLI workflows remain outside this contract.