> ## Documentation Index
> Fetch the complete documentation index at: https://docs.getmillwork.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Build a verifiable agent handoff

> Adapt a tested RFQ handoff pattern to one workflow with approved authority, verifiable evidence, and a receiver who checks before acting.

A supplier quote changes hands twice before a sourcing agent acts. In the [runnable RFQ example](/cookbook/output-checks/primitive-agent-email), the receiver checks the signed record and rejects an altered copy. Use that local pass and failure to design one handoff of your own: who may approve it, which evidence stays with it, what the receiver checks, and when the work must stop.

**Goal:** define one bounded handoff whose receiver can check the recorded decision, the exact action received, and whether the record changed.

**You are done when:** you have reproduced the RFQ pass and altered-copy rejection, documented your own actors and allowed transitions, and shown what your receiver checks before acting. A synthetic or shadow verdict does not approve a live pilot.

<div className="cookbook-action-card">
  <Card title="Run the RFQ example" href="/cookbook/output-checks/primitive-agent-email#run-the-example" arrow="true">
    Start with invented quotes and test-only keys. Verify one handoff, then change a copy and see it rejected.
  </Card>

  <Card title="See the shadow decision" href="/cookbook/output-checks/primitive-agent-email#measure-it-in-shadow" arrow="true">
    Compare the ordinary and missed-defect synthetic trials before choosing measures for your workflow.
  </Card>
</div>

## One transfer, from authority to action

Choose a consequential transfer, such as a quote another agent will use. Your system decides who may start it, what evidence it needs, and when it must stop. The output check assesses a frozen packet. The receiving agent verifies the record and the exact action it received before applying its own authority.

An **output check** applies your rules to the frozen handoff before send. This diagram shows the successful path and the two places that stop it. In the offline RFQ example, Primitive delivery and the Millwork check and receipt are simulated; a real run needs separate approval.

```mermaid theme={null}
---
config:
  flowchart:
    nodeSpacing: 20
    rankSpacing: 24
    padding: 10
    diagramPadding: 8
    curve: stepAfter
  themeVariables:
    textColor: "#211F1B"
---
flowchart TD
  A["<b>Approved case</b><br/>named parties and allowed transfer"] --> B["<b>Your service</b><br/>freezes evidence, recipient and action"]
  B --> C["<b>Millwork</b><br/>calls your output check"]
  C --> D{"<b>Pass, receipt and<br/>signed observation</b><br/>agree?"}
  D -->|Fail or no verdict| X["<b>Stop before send</b><br/>review or reject"]
  D -->|Yes| E["<b>Your service</b><br/>signs a checkpoint and sends the handoff"]
  E --> F{"<b>Receiver</b><br/>record, keys and received action match?"}
  F -->|No| Y["<b>Stop at receiver</b><br/>do not act"]
  F -->|Yes| G["<b>Receiver</b><br/>acts under its own policy"]

  classDef customer fill:#EEF4F6,stroke:#7895A0,color:#211F1B,stroke-width:1px
  classDef service fill:#F3F0F8,stroke:#9A86B8,color:#211F1B,stroke-width:1px
  classDef decision fill:#F8F6F3,stroke:#B92B2B,color:#211F1B,stroke-width:1px
  classDef stopped fill:#F8F6F3,stroke:#B92B2B,color:#211F1B,stroke-width:1px,stroke-dasharray:4 3
  class A,B,E,G customer
  class C service
  class D,F decision
  class X,Y stopped
  linkStyle default stroke:#4B5563,stroke-width:1px
```

The RFQ example models a Primitive email handoff; even its optional live Millwork mode keeps mail simulated. The record, controller and output check belong to your service. Its [handoff explanation](/cookbook/output-checks/primitive-agent-email#how-it-works) names the exact packet, receipt, signed checkpoint, and receiver checks. The diagram above describes a pattern to adapt, not a second ready-to-run integration.

## Prove the RFQ path locally

Use the [checked RFQ download and canonical commands](/cookbook/output-checks/primitive-agent-email#run-the-example). The example needs Node.js 22 and `npm ci`; it uses invented mail, simulated Primitive and Millwork responses, and test-only signing keys. It sends no email and starts no paid run.

<Steps>
  <Step title="Run and verify one retained handoff">
    Run `npm run cases`, then use the complete `verify-record` command it prints for `fixtures/rfq-042`. The receiver should report `record verified` for the retained mail and record through the printed sequence and head. Later entries remain unknown to that witness.
  </Step>

  <Step title="Change a copy and see it stop">
    Copy that fixture's `record.jsonl`. Change one hex digit of `packet_sha256` in its `transfer_prepared` entry, then run the same printed command against the copy. Expect `hash_mismatch` and exit code 1. Leave the original fixture and its test keys untouched. The [failure table](/cookbook/output-checks/primitive-agent-email#when-a-check-fails) explains the other rejection codes.
  </Step>

  <Step title="Compare two synthetic shadow verdicts">
    Run the [ordinary and missed-defect trials](/cookbook/output-checks/primitive-agent-email#measure-it-in-shadow). Read their `shadow_verdict` and rule rows: the ordinary trial passes; the variant misses a planted defect after the witnessed head and fails both S2 and S3. **Both commands exit 0.** Process success means the report ran, not that the workflow passed its decision rules.
  </Step>
</Steps>

These checks prove that the example implements its stated boundaries on invented inputs. They do not prove a supplier's quote is true or that another workflow is implemented.

## Keep the proof; decide the business policy

The RFQ page's [three-file customization](/cookbook/output-checks/primitive-agent-email#make-this-your-case) changes that RFQ example's policy, reply, and quote rules. It is not a three-file port to GTM or another workflow. Design and approve the domain choices below before an agent changes lifecycle code.

<div className="completion-evidence-overview primitive-handoff-adaptation" role="region" aria-label="Proof steps and local design choices">
  | Preserve and verify | Design and approve for your workflow |
  | - | - |
  | Append-only entries linked by hashes; verification against a known head. | What one case is, who may start it, its intake evidence, and its allowed state transitions. |
  | A signed checkpoint with stated coverage; entries after its witnessed head remain unknown. | Who holds the signing keys, how the receiver pins public keys independently, and who may revoke them. Test keys are never production keys. |
  | A frozen packet bound to the exact recipient and action the receiver gets. | Which fields, evidence sources, policy version, and output-check rules make a transfer eligible. |
  | A check that reads your frozen packet by opaque attempt ID; a qualifying receipt and signed observation before a consequential send. | What a pass authorizes, what a person must approve, and what the receiver may do after verification. |
  | Rejection or no verdict stops the dependent action; missing evidence never becomes a pass. | Who investigates mismatches, unavailable evidence, retries, and recovery. |
  | Exclusions, unsigned records, denominators, and unmeasured claims remain visible in the shadow report. | Observation window, baseline provenance, measures, thresholds, and stop owner, fixed before local shadow observation. |
</div>

A hash chain shows that entries have not changed **relative to a known head**. A signature shows that a pinned key endorsed what it signed. Neither establishes business truth, the signer's authority, or delivery of an email. The receiver must compare its retained handoff with the signed action and act only under its own policy. Verification cannot authenticate the sender of copied mail or detect a second send of the same signed head. The RFQ [evidence table](/cookbook/output-checks/primitive-agent-email#what-each-record-shows) explains what each record can and cannot show.

<span id="adaptation-checklist" />

## Adapt one bounded workflow

Use test data until your own owners approve the authority, evidence, and policy. The output of each step is something a person can inspect and an agent can test.

<Steps>
  <Step title="Name the case and its decision owners">
    **Input:** one unit of work, its starting event, authorized parties, and a person who owns exceptions. **Artifact:** a case boundary and authority table. **Completion check:** each proposed transition names who can cause it; unknown senders and out-of-scope cases stop before admission.
  </Step>

  <Step title="Define states and evidence custody">
    **Input:** intake records, their source, permitted states, and allowed transitions. **Artifact:** a state-and-evidence map naming who retains each record. **Completion check:** missing evidence and invalid transitions have an explicit reject, quarantine, pause, or no-verdict path; an agent's own claim is not treated as independent proof.
  </Step>

  <Step title="Freeze the transfer and its check">
    **Input:** exact payload, recipient, action, policy version, and output-check criteria. **Artifact:** a versioned packet schema and ruleset. **Completion check:** the check reads the frozen packet from your store by opaque attempt ID; changing the packet, recipient, or received action breaks the binding. A candidate's prose cannot supply missing evidence.
  </Step>

  <Step title="Specify the record and receiver verification">
    **Input:** entry format, checkpoint coverage, independently pinned public keys, and the handoff the receiver retains. **Artifact:** record, witness, key-trust, and receiver-check specifications. **Completion check:** the receiver detects a changed entry or action, checks the signed head through its sequence, and treats later or unsigned entries according to their stated limits before acting.
  </Step>

  <Step title="Exercise pass, failure, and recovery locally">
    **Input:** synthetic approved and malformed cases, plus unavailable-evidence and missing-receipt cases. **Artifact:** local test results and a code-to-action table. **Completion check:** a valid handoff verifies; a tampered copy fails; a failed check, unknown receipt, revoked key, or no verdict cannot trigger the dependent action. Do not reuse the RFQ's green tests as evidence that a newly designed workflow works.
  </Step>

  <Step title="Preregister and measure in shadow">
    **Input:** eligible cases, local records and review log, planted defects, observation window, units, denominators, baseline source, thresholds, and stop owner. **Artifact:** a dated preregistration and a local report that names exclusions and unsigned observations. **Completion check:** the report distinguishes measurements from missing evidence and leaves pilot-only measures unmeasured. A separately authorized person decides whether any pilot is worth requesting.
  </Step>
</Steps>

The downloaded RFQ example includes [the preregistration](/cookbook/output-checks/primitive-agent-email#preregistration) in `shadow/PREREGISTRATION.md` and [the interview kit](/cookbook/output-checks/primitive-agent-email#interview-kit) in `shadow/INTERVIEW_KIT.md`; the RFQ page shows [the local input shapes](/cookbook/output-checks/primitive-agent-email#measure-your-approved-records-locally) the kit reads. [Download and inspect the checked example](/cookbook/output-checks/primitive-agent-email#run-the-example) before adapting those files. Its four-week window, thresholds, test keys, and synthetic time baseline are example-specific choices, not defaults for your organization.

## Example adaptation: a GTM launch handoff

Consider an invented campaign inviting people to a webinar on **14 October**.
A drafting agent has the dated event brief and an approval record from the
campaign owner. A publishing agent may schedule only the exact approved copy.

<div className="completion-evidence-overview primitive-handoff-adaptation" role="region" aria-label="Hypothetical GTM adaptation">
  | Handoff step | What happens in this hypothetical campaign |
  | - | - |
  | Admit the case | Bind one campaign ID to its owner, dated event brief and permitted publishing agent. Keep a missing approval in review. |
  | Freeze and check | Bind the approved 14 October copy, audience, channel, recipient and evidence references into one packet. The check compares the date with the approved brief. |
  | Reject a changed date | Alter a copy to say 15 October. The check rejects it; no handoff is sent. An approved edit needs a new packet and check. |
  | Verify before scheduling | The receiver uses independently pinned keys and its retained handoff to verify the record and exact action. It schedules only under its own release policy. |
  | Measure in shadow | Count eligible campaign handoffs, plant changes in copies, and record review minutes against measures fixed before observation. |
</div>

This is an adaptation sketch, not a runnable GTM engine. Supply your own approved
policy, evidence, keys, state transitions and tests. The RFQ controller, packet
schema, receiver verifier and shadow interpretation contain domain choices; the
example's three-file edit does not implement this campaign workflow.

## Read a shadow result without overstating it

<div className="completion-evidence-overview primitive-handoff-adaptation" role="region" aria-label="Evidence stages and limits">
  | Evidence stage | What it supports | What it does not authorize or prove |
  | - | - | - |
  | **Synthetic RFQ trial** | The bundled kit produces its stated pass and failure verdicts from invented cases. | Customer savings, real-world reliability, supplier truth, or a live run. |
  | **Partner-owned local shadow** | Approved case records and review logs can be checked locally against measures fixed before observation. Excluded, unsigned, missing, and no-verdict evidence remains visible. | A Millwork run, publication to another party, or automatic permission for a pilot. |
  | **Separately approved pilot** | May test the pilot-only questions after its own authorization and observation window. | A result already measured by the offline trial or shadow report. |
</div>

The shadow kit starts no Millwork run and sends no customer data. A partner's approved records stay on its own machine while the kit processes them locally; the report contains counts, rates, minutes, reason codes, dates, and digests, not mail content. Cases with no signed checkpoint can still be counted, but their hash-chained records are local observations, not independently witnessed evidence. A checkpoint covers entries only through the signed head; a planted change after that head can be missed. Read the [ordinary and missed-defect results](/cookbook/output-checks/primitive-agent-email#measure-it-in-shadow) beside the preregistered rules, not as a forecast or a launch decision.

## Give this to your coding agent

Use **Copy page** on this guide and the [Primitive RFQ example](/cookbook/output-checks/primitive-agent-email), and attach or paste both Markdown copies with the brief. Fill the bracketed inputs or ask the agent to list what is missing. Keep authority, key trust, thresholds, and live actions subject to your approval; the brief does not authorize contacting suppliers or starting paid runs.

<div className="agent-prompt">
  ```text theme={null}
  Read the current Markdown attached for this guide and the Primitive RFQ example
  first. If either is missing, you may read the page URLs below; use them
  only if both are available and include the described shadow instructions.
  Otherwise stop and ask me for the current Copy page Markdown. Do not infer steps
  from titles or search excerpts.

  Page URLs:
  https://docs.getmillwork.dev/cookbook/agent-workflows/verifiable-handoff.md
  https://docs.getmillwork.dev/cookbook/output-checks/primitive-agent-email.md

  Docs discovery:
  Index: https://docs.getmillwork.dev/llms.txt
  Full export: https://docs.getmillwork.dev/llms-full.txt
  MCP connection discovery: https://docs.getmillwork.dev/.well-known/mcp/server-card.json
  Use the MCP connection only if your environment supports remote MCP.
  These sources do not replace the required current Markdown for BOTH pages or
  the stop above when either page or its shadow instructions is missing.

  Work locally with synthetic data. You may download the checked example and install
  its pinned dependencies. Do not call Primitive or Millwork APIs, deliver email,
  start paid runs, use customer records, or change production settings.

  Workflow and one eligible case: [describe]
  Authorized sender, checker, receiver, and exception owner: [names or roles]
  Intake evidence and who retains it: [sources and custody]
  States, allowed transitions, and stop/no-verdict paths: [locally approved rules]
  Frozen transfer payload, recipient, received action, and policy version: [spec]
  Output-check criteria and independently pinned key sources: [spec]
  Receiver's permitted action after verification: [locally approved decision]
  Shadow window, measures, denominators, baseline source, thresholds, and stop owner:
  [preregister before observing real cases]

  First reproduce the RFQ example's retained-handoff pass and altered-copy failure.
  Then run both synthetic shadow variants and report their shadow_verdict and every
  failed rule; do not infer the evaluation verdict from process exit. The missed-defect variant fails S2 and S3 even though its final line names S2 first.

  For my workflow, return: an authority/state table; evidence and custody map;
  packet/check/receiver specifications; local pass, tamper, unavailable-evidence
  and no-verdict tests; a dated shadow plan; and a list of choices still needing
  human approval. Distinguish chain integrity, signer verification, authority,
  and business truth. Do not claim the RFQ's three-file edit implements my workflow.
  ```
</div>

## Present the case study

Keep this text skeleton with the local report. It works for RFQ and for a future workflow once you have evidence; leave unknowns marked as unknown.

```text theme={null}
Problem and bounded unit of work:
Actors, authority, states, and handoff point:
Evidence sources, custody, and signer/key assumptions:
What the output check decides; what the receiver verifies before acting:
Offline pass and altered-copy failure (fixture and observed code):
Observation window; eligible unit; exclusions and unsigned records:
Preregistered measures, units, denominators, baseline source, and stop rules:
Observed results (label synthetic or local shadow):
Verdict and every failed rule (read the report, not the shell exit):
What the record cannot prove and what remains unmeasured:
Locally chosen adaptation decisions and unresolved approvals:
Separately authorized next step, if any:
```

Use the completed skeleton to record the offline result and unresolved local approvals. Then ask the workflow owner to decide whether a separately authorized local shadow or pilot step is warranted. Keep missing, unsigned, excluded, and no-verdict evidence beside the result it limits.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.