Skip to content

MVP, risk, evidence

MVP, correctly

MVP is not a crappy first version. MVP is the smallest product experiment capable of testing the most important business hypothesis.

Sometimes that is software. Often it is: a landing page with a call to action, a manual or concierge service, a prototype, a spreadsheet, a mockup, a video, a paid pilot with one design partner, a workshop sold at a price.

Write:

CORE HYPOTHESIS      the belief that, if false, kills the product
RISKS                ranked (see below)
SMALLEST TEST        the cheapest thing that produces a real signal on the top risk
SUCCESS CRITERIA     numbers and a date, decided before the test runs

An MVP tests uncertainty; it does not merely reduce scope.

The eight risks

MARKET RISK          Does anybody care?
PRODUCT RISK         Can we create the outcome?
TECHNOLOGY RISK      Can it actually work at the required quality and cost?
DISTRIBUTION RISK    Can we reach buyers?
ECONOMIC RISK        Can it make money?
ADOPTION RISK        Will users change behavior?
TRUST RISK           Will they allow us to do it (data, authority, agents acting)?
REGULATORY RISK      Are we allowed to?

Rank them for this product. The MVP tests the top one. For governed agents trust risk usually outranks technology risk; the demo that matters is the receipt, not the agent.

Authority is earned per class, never granted globally

For any product that takes operational responsibility inside a customer's systems, autonomy is a ladder the product climbs with evidence, not a tier list the customer picks from on day one.

OBSERVE    detect, classify, prioritize, report. No changes to customer systems.
ASSIST     detect, diagnose, prepare the change, test it, produce evidence, open it for review.
           A human approves deployment. Expect this to be the default operating mode for a long time.
OPERATE    execute an approved remediation class within explicit policy. Humans move to
           exceptions. This is the destination, not the starting offer.

Operate authority is scoped to a tuple, never to the account:

CUSTOMER + REMEDIATION CLASS + VERIFIED HISTORY + POLICY + ACCEPTANCE

A failure or regression in a class removes that class's authority automatically. Say so in the offer; the automatic downgrade is a feature the buyer is paying for, not a caveat.

Do not sell Operate as the default. Selling full autonomy before it is earned converts trust risk, which is usually the top risk, into the first thing the customer tests and the first thing that breaks.

The acceptance gate for a remediation class

A class is not billable because an agent exists or a workflow ran once:

DEFINED_SUCCESS_CRITERIA
VERIFIED_RESULT
ZERO_ACCEPTANCE_REGRESSIONS
EVIDENCE_RECORDED
EXECUTION_COST_RECORDED
HUMAN_ESCALATION_BEHAVIOR_PROVEN
ROLLBACK_AND_FAILURE_PATH_PROVEN
SECOND_ESTATE_PROOF           (before the class is treated as reusable)

The same gate answers three questions at once: can it be sold, can human approval be safely reduced, and can the method enter the reusable library.

Product Evidence Ladder

Not all evidence is equal. Lower to higher confidence:

Founder believes it
Customer says it sounds useful
Survey says people want it
User joins waiting list
User tries prototype
User repeatedly uses product
User pays
User renews
User expands usage
User recommends it without being asked

Map onto the corpus epistemic hierarchy: L1 runtime or cryptographic proof, L2 verified system state, L3 reviewed document, L4 stated intent, L5 unverified assertion, LLM output, or chat transcript. Grade every claim in the definition. A definition resting on L4 and L5 is a hypothesis document and its first line must say so.

Evidence State table (close every definition with it)

Tag every load-bearing claim with one state. The table is the "what is real now versus what remains unproven" section, and it is the part a generic product framework misses.

State Meaning Example
CANONICAL_POLICY Governed doctrine in BluCity-Docs says so Drupal remains system of record; agents cannot replace Drupal authority
CANONICAL_PRODUCT_DIRECTION / CURRENT_PRODUCT_DIRECTION The current product documents say so; not yet proven in source or market ContextControl = governed agent operations
CURRENT_SOURCE Verified in the repository at a named ref recipe_blucity -> recipe_amcs -> site_template_amcs exists
HISTORICAL_SUPERSEDED Exists in the corpus but is explicitly superseded; must not be reused Old AMCS $200/$800/$2K tiers
NOT_ESTABLISHED Nobody has proven it; the definition treats it as a hypothesis Clean install + two-run proof; customer #2 portability; paid pilot demand; market-validated pricing
NO A dependency people assume that is not required Cedar/ContractPlane required for AMCS MVP: NO; MCP required: NO, edge capability

Also list explicit non-dependencies. Half the value of the table is stating what the MVP does NOT require, because that is what stops the platform construction program from replacing the product experiment.

The demo that sells

The demo is never the chatbot and never the module list. It is a side-by-side:

BEFORE   Human performs 12 steps across 4 tools.
AFTER    Agent performs 9 bounded steps. Human reviews 3.
         All changes remain revisions. The authority boundary holds. Evidence is retained.

Or, for an operations product, a receipt: who acted, under whose authority, what was allowed, what context was used, what changed, what was denied, who approved the boundary crossing, whether the outcome was re-verified, whether recovery is proven. If the demo cannot be drawn in this shape, the activation event is not yet understood.

Validation sequence

PROBLEM DISCOVERY -> PROBLEM VALIDATION -> SOLUTION VALIDATION -> MVP -> EARLY ADOPTION -> PMF -> SCALE

Problem validation: ten conversations with people who match the ICP, where they describe the problem unprompted and name the workaround and its cost. Solution validation: they react to a concrete offer with a price. Early adoption: someone pays and uses it repeatedly. Do not plan scale before that.

Kill criteria (write them now)

Define under what evidence you stop. Example:

After 50 qualified prospects contacted by <date>:
fewer than 5% agree the problem is significant
AND zero willingness to pay at the draft price
-> stop, or reframe the product thesis and reset the clock once.

The corpus rule for exploration: time-box it and require an exit decision, one of sell, validate with prospects, defer, or stop. A strategy without kill criteria becomes permanent sunk-cost justification.