The toolkit

Reference cards

Keep this page open during simulations. Everything here is a tool for one of the five questions — the header on each card tells you which letter of SOLVE: Surface · Observe · Learn · Vet · Earn.

S–O · The three Bs

LevelTestVillage-well example
BuildDid we make/do the thing?The well is built.
BehaviorCan I watch someone behave differently?Villagers spend less time carrying water.
Bottom lineIs it a big number many factors feed?Standard of living rises.

Quick test in the wild: if a “goal” could be achieved while creating zero value (a shipped feature nobody uses), it's a build. If no single team could plausibly own it, it's a bottom line. Behaviors live in between: specific, observable, measurable. (Elsewhere you'll hear this ladder as outputs → outcomes → impact — same levels, different words.)

S · Opening checklist

O · The three behavior questions

1. What are people (customers, users, employees) doing today?
2. What would they be doing differently if we succeeded?
3. What behaviors predict the result we care about?

And the linked form every target should take: If we create [this customer/user behavior], it will deliver [this business result].

O · Leading vs. lagging

A lagging indicator tells you how you did (revenue, NPS, churn) — real, but too late to steer by. A leading indicator is a behavior that predicts the lagging one (newsletter opens predicting return visits; early client–vendor meetings predicting deal success). Target behaviors that are leading indicators; report bottom lines that are lagging ones. If you can't name the leading indicator for your project, you don't yet know how it creates value.

L · Issue trees & MECE

L · Journey map with boosters & blockers

  1. Lanes for each actor (customer, staff, system). Left to right: what they actually do, step by step — the real process, not the official one. Follow one real case end to end.
  2. Pass two, the behavior question: at each step, what behaviors predict success (boosters, green) and failure (blockers, red)?
  3. Harvest: “increase the rate of [booster]” and “decrease the rate of [blocker]” are your candidate target behaviors.

V · The hypothesis template

We believe that [this change]
will create [this behavior change — observable and measurable].
We'll know we're right when [this evidence, observable by this date].

Then list what must be true for the claim to hold, and rank by shakiness. Your first test attacks the shakiest condition — not the easiest one. Vet before you bet.

V · The experiment ladder

Cheapest first. Climb only as high as the decision requires.

RungWhat it looks likeCost
LookQuery data you already have; find the correlation or its absence.Hours
AskInterview or survey the people whose behavior must change.Days
Fake itPaper prototype, concierge test, landing page, pump-before-the-well.Days–weeks
PilotParallel run or one-site/one-category rollout with a baseline comparison.Weeks
Prove itRandomized test (A/B) on the real metric.Weeks–months

A note on the word “MVP”: it gets used two different ways — for a cheap experiment (the lower rungs of this ladder), and for a first sellable release with a deliberately chosen feature set, which is the standard meaning in product management. Both are legitimate; they are different things. This site says cheapest test for the first and first release for the second.

V · The evidence triad

KindQuestion it answersExamples
AnalyticalDoes the logic and math hold?Model backtest, sensitivity analysis, savings model
BehavioralDo people actually do what the claim requires?Pilot usage, override rates, interviews, parallel runs
OutcomeDid the metric move?A/B result, before/after on the leading indicator

Which kinds you need follows from your deliverable: a pure recommendation maxes out at analytical + behavioral before the decision; a built artifact must eventually reach outcome evidence — until it does, it's a demo.

E · Conclusion-first synthesis

Recommendation — one sentence, the answer.
Because — two or three reasons, each backed by a number.
Worth — the value, in dollars and days, with its range.
Risks — the top one or two, each with a mitigation.
Next — what happens Monday, and who owns it.

The sixty-second spoken version keeps the same order and cuts everything that isn't load-bearing. Never narrate your process chronologically; nobody's decision depends on what you did second.

E · The adoption argument

Every recommendation asks somebody to change their behavior — which means your stakeholders are customers, and adoption is a behavior change you must design for, not hope for. The heavier the behavior change, the heavier the argument: name who changes what, what's in it for them, what makes the new way easier than the old way, and how overrides/exceptions are governed. Then close with the measurement plan: the two or three leading indicators that will show, within weeks, whether the must-be-trues are holding.

Any letter · Case math habits

The failure-mode table

StepThe classic failureThe fix
S · SurfaceAccepting the requested thing as the problem“If this worked perfectly, what number changes — compared to what?”
O · ObserveTargeting a build or a bottom line and calling it the behavior“What would I watch someone do differently?”
L · LearnReciting a canned framework; or analysis without endStructure this problem; rank branches; go deep on few
V · VetConfirming instead of falsifying; testing by building everythingAttack the shakiest must-be-true with the cheapest test
E · EarnChronological narration; no adoption planAnswer first; treat stakeholders as customers