Skip to content

PLAUD Capture Pipeline

Turn the PLAUD recorder into the front door of the BluTown authority stack: speech in, governed artifacts out, nothing written without a routing decision and a receipt.

0. Correction that changes the plan

github.com/orgs/Plaude is not the recorder vendor. That org is plaudeai.com — an in-app AI messenger SDK for financial services (@plaude/react-native, plaude-cli, plaude-flutter, docs). Similar name, unrelated company. Nothing in it is usable here. Discard it.

The recorder is PLAUD (plaud.ai). Its supported integration surface today:

Surface State Use
Plaud MCP Available; already installed in Cowork as mcp__plaud__*, currently unauthenticated Primary read path — list, search, transcripts, notes
Plaud CLI Available Batch export, cron-driven pulls, scripting
Public REST API Private beta / waitlist only Do not design against it
Zapier Available Fallback trigger only; avoid — no receipt trail

Design consequence: the pipeline is pull-based, not webhook-based. Poll on a schedule via CLI/MCP. Do not architect for push until PLAUD ships a public API.


1. What already exists (verified on /Volumes/AgentPlatform)

Component Path / endpoint Role in this pipeline
OpenClaw runtime /Volumes/AgentPlatform/.openclaw/ — 21 skills, plugins, nodes, state Ingress + scheduling + channel notify
OpenClaw source BluTown/openclaw/ Skill authoring target
Existing skills cron, channel_message, file_reader, docx, pdf, self-improving-agent Reuse — do not rebuild
Beads / Dolt .beads/ (dolt on :3308), gt / bd CLI Operational work ledger (ADR-0007)
Master docs BluCity-Docs/ Authority — Engineering-Standard, Research, Products, Evidence, Playbooks
Customer material Customers/ (currently DOI/) Client-scoped derivatives
Knowledge vault Knowledge/ + index-to-qdrant.mjs Vector index for retrieval
Governance Engineering-Standard/governance/, cedar_policies/ Gate on every write

You are not missing infrastructure. You are missing one skill and one routing contract.


2. The core problem to solve

A recording is not one artifact. A single 40-minute conversation typically contains four different things that belong in four different places:

  1. Commitments — "I'll send the estimate Friday" → work ledger
  2. Decisions — "we're using OpenTofu, not Terraform" → ADR candidate
  3. Client context — pricing, scope, personnel → Customers/<Client>/
  4. Raw thinking — architecture riffs, research leads → Research/

The failure mode of every recorder workflow is dumping all four into one markdown file nobody reads again. The routing contract is the product.


3. Routing contract

classify → route → gate → write → receipt

Class Destination Gate Auto-write?
Commitment / action bd bead, assigned + dated none Yes
Decision (architectural) Research/decisions-inbox/ as ADR candidate human promote No — proposal only
Customer meeting Customers/<Client>/notes/YYYY-MM-DD-<slug>.md client-scope check Yes
Research / idea Research/capture/YYYY-MM-DD-<slug>.md none Yes
Standard change Never auto-written Thomas only No
Everything else Research/capture/_unrouted/ weekly sweep Yes

Hard rules

  1. Engineering-Standard/ is never written by the pipeline. It is the authority tier. Recordings can only produce candidates in Research/decisions-inbox/ that you promote by hand. This is the single most important line in the plan — it is what keeps your standards trustworthy once voice capture is automated.
  2. Customer audio is classified. Anything routed to Customers/ carries a client: frontmatter key and never lands in the general Qdrant index without an explicit collection scope. Client conversations are the highest confidentiality material you generate.
  3. Every write emits a receipt per Execution-Receipt-Specification.md, carrying the PLAUD file_id as the provenance anchor. One recording → one receipt → N artifacts. This is your dogfooding story: the standards work applied to your own day.
  4. Nothing is created without a bead reference. Commitments get a bead; documents reference the bead that produced them.

4. Architecture

PLAUD device
    │  (device syncs to PLAUD cloud — not controllable)
    ▼
PLAUD cloud
    │  pull: plaud-cli / mcp__plaud__list_files (cursor = last_seen file_id)
    ▼
OpenClaw skill: plaud_capture          [.openclaw/skills/plaud_capture/]
    ├─ 1. poll        cron skill, every 30 min
    ├─ 2. fetch       get_note (summary/actions) + get_transcript (paged)
    ├─ 3. classify    LLM → strict JSON schema, one entry per extracted item
    ├─ 4. gate        Cedar check per destination class
    ├─ 5. preview     channel_message → iMessage/Slack, approve/reject
    └─ 6. commit      write artifacts + bd create + receipt
    ▼
┌──────────────┬─────────────────┬──────────────────┬──────────────┐
│ beads/Dolt   │ Research/       │ Customers/<C>/   │ Qdrant       │
│ (actions)    │ (ideas, ADR-c)  │ (client notes)   │ (retrieval)  │
└──────────────┴─────────────────┴──────────────────┴──────────────┘
    ▼
Evidence/receipts/RX-PLAUD-<file_id>.md

Why OpenClaw owns this and CoPaw does not

OpenClaw is already the channel gateway and already has cron and channel_message. The approve/reject loop is a messaging problem, and OpenClaw is where messaging lives. CoPaw earns a role later, only if classification needs multi-step tool use — which it will not for v1.

Classification contract

Constrain the model to a schema. Free-form summarization is what makes these pipelines untrustworthy.

{
  "file_id": "string",
  "recorded_at": "ISO-8601",
  "client": "string | null",
  "items": [{
    "class": "commitment|decision|customer|research|unrouted",
    "text": "string",
    "owner": "string | null",
    "due": "YYYY-MM-DD | null",
    "confidence": 0.0,
    "transcript_ref": "utterance index"
  }]
}

confidence < 0.7 → force human review, never auto-write. transcript_ref is mandatory — every artifact must be traceable back to the sentence that produced it.


5. Phased rollout

Phase 0 — Unblock (30 min)

  • Authenticate the Plaud MCP: mcp__plaud__login. It is installed and failing auth right now, which is why nothing works today.
  • Install plaud-cli; confirm list + transcript against one real recording.
  • Decide the polling identity: which machine runs cron — Mac or Oracle. Oracle, if you want capture to continue when the laptop is closed.

Exit: you can pull a transcript by file_id from the command line.

Phase 1 — Read-only classify + preview (1–2 days)

  • Scaffold .openclaw/skills/plaud_capture/, modeled on news and cron.
  • Implement poll → fetch → classify → channel_message preview only.
  • Zero writes. Output goes to your phone as a proposed routing table.

Exit: ten consecutive recordings classified; you agree with ≥80% of the routing. Tune the prompt until you do. Do not proceed on worse than that — a mis-routing pipeline writing autonomously is worse than no pipeline.

Phase 2 — Write the two safe classes (2–3 days)

  • Enable auto-write for research → Research/capture/ and commitment → bd create.
  • Emit Evidence/receipts/RX-PLAUD-<file_id>.md per run.
  • Decisions and customer material still preview-only.

Exit: a week of recordings produces beads you actually work from.

Phase 3 — Customer routing (2 days)

  • Add client detection (calendar attendee match beats transcript inference).
  • Write to Customers/<Client>/notes/, scoped Qdrant collection per client.
  • Cedar policy: deny cross-client reads.

Exit: DOI conversations land in Customers/DOI/ and nowhere else. Verify by grepping the general index for a DOI-specific term and getting zero hits.

Phase 4 — Decision inbox + morning brief (2 days)

  • Research/decisions-inbox/ receives ADR candidates in ADR-TEMPLATE.md shape, status candidate, decision_status UNKNOWN.
  • Morning brief over OpenClaw: yesterday's commitments vs. calendar vs. beads.

Exit: you promote one recording-derived ADR candidate into Engineering-Standard/decision-records/ by hand.


6. What not to build

  • No parallel work ledger. Actions become beads. ADR-0007 already forbids inventing a sidecar ledger in markdown or JSON.
  • No custom transcription. PLAUD already transcribes with speaker attribution. Re-transcribing costs money and loses diarization.
  • No Zapier in the path. No receipt, no Cedar gate, external dependency on client-confidential audio.
  • No writes to Engineering-Standard/ ever. Restated because it is the rule most likely to erode under convenience pressure.

7. Risks

Risk Severity Mitigation
Client audio derivatives leak into general index High Client scope assigned before any write; per-client Qdrant collection; Cedar deny cross-client; Phase 3 grep verification
PLAUD private-beta API changes / MCP breaks Medium Depend on CLI + MCP only; keep classification decoupled from fetch so the fetch adapter is swappable
Classification drift silently mis-routes Medium transcript_ref mandatory; confidence floor 0.7; weekly _unrouted/ sweep
Recordings of conversations without consent High Out of scope for tooling — one-party vs two-party consent varies by state. Establish your own disclosure practice before automating retention of client conversations
Bead spam from casual speech Medium Only commitment class with an owner creates beads; no owner → _unrouted/
Dolt fragility under automated writes Medium Batch bead creation once per poll cycle, not per item; follow the Dolt diagnostics path in BluCity-Docs/CLAUDE.md before any restart

8. Validation checklist

  • [ ] mcp__plaud__get_current_user returns an authenticated user
  • [ ] plaud-cli exports one transcript with speaker labels + timestamps
  • [ ] 10 recordings classified at ≥80% routing agreement
  • [ ] Receipt file exists for every processed file_id, no orphans
  • [ ] Every generated artifact resolves back to a transcript_ref
  • [ ] DOI-specific term returns zero hits in the general Qdrant collection
  • [ ] Engineering-Standard/ git log shows zero pipeline-authored commits
  • [ ] Cron survives a laptop-closed cycle (if Oracle-hosted)

9. Rollback

Each phase is independently reversible.

  1. Disable: .openclaw/disable-launchagent or remove the cron entry. Capture stops; nothing else is affected.
  2. Unwind writes: artifacts are file-based and git-tracked in BluCity-Docs / Customers — revert the commit range.
  3. Unwind beads: bd close the run's beads, identified by the RX-PLAUD-<file_id> receipt reference carried on each bead.
  4. Unwind vectors: drop the per-run Qdrant point IDs (namespace point IDs by file_id at write time so this is a one-liner, not a re-index).

10. Open decisions for Thomas

  1. Poll host — Mac or Oracle. Recommend Oracle for continuity.
  2. Consent practice — what you disclose before recording client calls, and retention period for client audio derivatives. Tooling cannot decide this.
  3. Client detection source — calendar attendees (accurate, needs calendar connector) vs. transcript inference (no dependency, less reliable). Recommend calendar.
  4. Whether commitments auto-create beads or stage for approval in Phase 2. Recommend auto-create — a bead is cheap and closeable; a missed commitment is not.