Skip to main content

Applications for the autumn mentorship cohort are open until 30 September. Apply now

Data & Evidence

KPI Design: What Can Actually Be Measured?

How to tell whether a proposed success measure can actually be observed, attributed, defended against gaming and verified after the money is spent - and how to rewrite the ones that cannot.

IntermediateReading + exercise136 minv1.0.0 · effective Aug 9, 2026
Course catalogue

Learning outcomes

  • Lay out a results chain from inputs to impacts without treating the chain as evidence that it holds
  • Apply a four-part test — observable, attributable, gaming-resistant, reasonably verifiable — to any proposed KPI
  • State attribution claims proportionately, distinguishing correlation, confounding and counterfactual reasoning
  • Specify baselines, denominators, cadence and ownership, and anticipate Goodhart-style distortion
  • Rewrite a vague metric into a source-specific, time-bound measure with its limitations stated
136 minutes · 5 lessonsVersion 1.0.0 · effective 8/9/2026
Lesson 1 of 524 min read

From inputs to impacts

Inputs, activities, outputs, outcomes and impacts - a results chain as a structure for asking questions, never as evidence that the chain holds.

What you will be able to do

Most disputes about whether a funded programme succeeded are not really disputes about numbers. They are disputes about what the numbers were supposed to show. One side points to a delivered tool; the other asks whether anyone uses it; a third asks whether usage changed anything. All three can be arguing in good faith from the same report, because the report never said which link in the chain it was measuring. This lesson gives you the vocabulary and the diagram that make that argument tractable before the money is committed rather than after.

  • Place a proposed metric on the correct link of a results chain.
  • Separate what a team controls from what it can only influence.
  • Write the assumption sitting on each arrow.
  • Treat the chain as a hypothesis to be tested, not as an argument that it works.

Definitions

Inputs
The resources committed: funds, staff time, equipment, existing infrastructure.
Activities
What the team does with the inputs: building, writing, running workshops, operating a service.
Outputs
The direct products of the activities, largely within the team control: a released tool, a published report, a number of sessions delivered.
Outcomes
Changes in behaviour, capability or condition among people outside the team, which the programme can influence but not control.
Impacts
Longer-term, wider changes to which the programme is one contributor among many.
Results chain
The ordered sequence inputs → activities → outputs → outcomes → impacts, used as a common evaluation vocabulary.
Theory of change
The results chain plus the explicit assumptions on each arrow explaining why one link should produce the next.

The five links, and why the boundary matters

The vocabulary here is not the Institute invention. It comes from the evaluation literature, where the terms in the OECD development-assistance glossary have been the common reference for decades. Adopting standard words has a practical benefit: it stops each proposal inventing its own hierarchy of goals, deliverables, milestones and objectives, in which the same word means something different from one page to the next.

The single most important line in the chain falls between outputs and outcomes. Everything to the left of it is within the team gift. If they receive the funds and do the work, the outputs appear; if the outputs do not appear, that is a delivery question of the kind the delivery-risk material addresses. Everything to the right depends on other people responding. A team can ship an excellent explanatory guide and still see no change in how anyone writes rationales, because the world contains competing demands on attention that the team does not command.

Both directions of confusion cause damage. Holding a team contractually to an outcome - treating "raise participation" as a deliverable - punishes them for the behaviour of third parties and pushes them towards whichever cheap proxy is easiest to move. Accepting an output as though it were an outcome - reporting that six workshops ran and calling that capability-building - retires the interesting question without answering it. The productive position is to fund outputs, measure outcomes honestly with attribution stated carefully, and say plainly which is which.

Impacts sit furthest right and are usually the reason anyone cares. Better-informed representation, healthier treasury stewardship, wider participation: these are the stated purposes of much governance funding, and they are also the things a single programme can least plausibly claim to have caused. That is not a reason to omit them from the chain. It is a reason to state, at the outset, that the programme contributes to them and that nobody will be able to measure that contribution cleanly.

A worked chain

Fictional composite for training

Request: 180,000 units over eight months to run a rationale-writing programme for Cardano DReps - a written guide, eight live workshops, and a public library of worked examples. Inputs: 180,000 units; two facilitators at half time; one editor at quarter time; an existing community forum for distribution. Activities: write the guide; run eight workshops; collect, anonymise and publish worked examples; answer questions in the forum. Outputs: one published guide at a stable URL; eight workshops delivered, with attendance recorded; at least twenty-four worked examples published. Each is countable, and each is squarely within the team control. Outcomes: participating representatives publish rationales more often; published rationales more frequently name the evidence relied on; delegators report that rationales are easier to compare. Each depends on choices made by people outside the team. Impacts: delegation decisions across the ecosystem are better informed. One programme among many possible contributors. Now place the proposal own metrics. It offers three: number of workshop attendees, guide page views, and "improved rationale quality across the ecosystem". The first two are output metrics. The third is an impact claim with no metric attached at all - no definition of quality, no source, no population. The outcome link, which is the one the funder actually cares about, has nothing on it. That gap is the finding, and it is visible in about five minutes with the diagram in front of you.

The gap is also fixable, and cheaply. An outcome metric might be: the count of published rationales, from the public on-chain vote records and their linked metadata, among the named cohort of participating representatives, in the three months after their workshop, compared with the three months before. That is not a perfect measure - later lessons will show what it cannot support - but it sits on the right link, it comes from a source nobody has to trust the team about, and it can be computed by a stranger.

The arrows carry the risk

Practitioner guidance on theories of change makes a point worth repeating: the value is in the arrows, not the boxes. Between each pair of links sits an assumption, and the programme fails at whichever assumption turns out to be false. Writing them down converts a persuasive narrative into a list of testable statements.

  • Inputs → activities: this holds only if the named facilitators are actually available for the committed hours.
  • Activities → outputs: this holds only if workshops can be scheduled at times a distributed community can attend.
  • Outputs → outcomes: this holds only if the reason representatives publish few rationales is that they lack a method - rather than lacking time, or seeing no audience for the effort.
  • Outcomes → impacts: this holds only if delegators read rationales and act on what they read.

Look at the third assumption. The whole programme rests on it, and it is a claim about why people currently behave as they do. If the real constraint is time rather than method, an excellent guide changes nothing, and no amount of delivery competence will rescue it. That is exactly the kind of question a representative can ask before funding and cannot answer afterwards from a completion report. The proportionate response is usually not rejection: it is a small piece of discovery work, or an early outcome check with an agreed decision point.

Label each assumption evidenced, plausible or untested, and be honest that most will be plausible. Plausible is an acceptable basis for funding. Plausible presented as evidenced is not, and the label is the difference.

Counterexample: the chain as rhetoric

A proposal includes a full-page diagram with colour-coded links running from a modest tooling budget all the way to "a more decentralised and resilient Cardano". Every arrow is drawn; no assumption is written on any of them. The diagram is doing persuasive work - it makes the leap from a tool to an ecosystem property feel like a series of small steps - while establishing nothing. The correct response is not to dismiss the proposal, but to ask which arrow the team considers most likely to break and what they would observe if it did. A team that has thought seriously about its own work answers that question readily and often improves the proposal in the process.

Common mistakes

  • Reporting outputs and describing them as outcomes.
  • Contracting a team to outcomes as though they were deliverables within its control.
  • Leaving the outcome link with no metric while over-instrumenting the output link.
  • Drawing arrows without writing the assumption each one carries.
  • Marking an assumption evidenced when the evidence is the proposal own confidence.
  • Treating a completed diagram as an argument that the programme works.
  • Extending the chain to grand impacts and then quietly measuring only the first link.

What this establishes

You can lay out a programme as inputs, activities, outputs, outcomes and impacts using standard vocabulary; place every metric a proposal offers onto the link it truly measures and see immediately which links are unmeasured; separate what the team controls from what it can only influence; and write each arrow assumption as a testable sentence labelled evidenced, plausible or untested.

What remains unknown

The chain does not tell you whether any assumption is true, how large an effect to expect, or whether the outcome is worth the input. It cannot rank two coherent programmes against each other. Nothing in this lesson quantifies anything, and no dataset, study or effect size is asserted anywhere in it. Whether a plausible-but-untested chain is a sufficient basis for funding is a judgement each representative makes independently.

Takeaways

Inputs and activities are effort, outputs are products, outcomes and impacts are changes in the world.

Fund outputs, measure outcomes, and never let one be reported as the other.

The assumptions on the arrows are where programmes fail; write them down and label them.

A results chain organises the questions. It never answers them.

Takeaways

  • Inputs and activities are effort; outputs are what the team produces; outcomes and impacts are what changes in the world.
  • Teams control outputs and only influence outcomes. Holding them to outcomes as if they were outputs is unfair; accepting outputs as if they were outcomes is uncritical.
  • Every arrow in the chain hides an assumption. The assumptions, not the boxes, are where a programme usually fails.
  • A tidy chain proves nothing. It organises the questions; it does not answer them.

Applied activity

Draw the chain and name the assumptions

Take any funding request and draw its results chain on one page: inputs, activities, outputs, outcomes, impacts. Place every metric the proposal offers onto the link it actually measures, and mark the links where the proposal offers no metric at all. Then write the assumption on each arrow as a sentence beginning "this holds only if". Mark each assumption as evidenced, plausible or untested, and say what would make an untested one evidenced. Do not invent data or estimate effect sizes.

Deliverable: A one-page results chain with every proposed metric placed on a link, unmeasured links marked, and each arrow carrying a labelled assumption. · about 40 minutes

Sources

  • DRep Institute Standards of Practice

    Cardano DRep Institute

    The Institute process standards referenced by the classification lab.

Practical exercise

KPI and monitoring plan

Take three KPIs from the fictional composite request in this course, or from a real proposal of your choice, and produce a monitoring plan. For each KPI record: the original wording; which of the four tests it fails and why; a rewritten version naming the data source, the population and denominator, the measurement window and cadence, the baseline value or the fact that the baseline must be established first, and the person or role who publishes it; the attribution claim you would consider defensible, phrased proportionately; one way the metric could be gamed cheaply and the countermeasure or the accepted residual risk; and one stated limitation. Close with a three-sentence note saying what the whole set would and would not establish. Do not invent datasets, studies, sample sizes or results, do not state that a measurement proves value, and do not state how anyone should vote.

Deliverable: Three rewritten KPIs with source, denominator, window, cadence, baseline, owner, attribution claim, gaming risk and limitation, plus a three-sentence note on what the set does and does not establish.

  1. 1.What exactly is being asked for, in your own words, without the proposer's framing?
  2. 2.Which claims did you verify against a primary source, and which source was it?
  3. 3.Which claims could you not verify, and what would it take to verify them?
  4. 4.What is the strongest argument against your current reading of the evidence?
  5. 5.What would you publish so someone who disagrees with you can audit your reasoning?
Autosaves on this device0 / 20,000

Signed out: this draft stays on this device. Sign in to save it to your account and have it count towards completion.

Final assessment

KPI Design assessment

Eight questions on results chains, the four-part test, attribution, baselines, gaming and rewriting metrics. Every answer turns on process, never on whether a proposal deserves funding.

8 questions · pass mark 6%Private result · unlimited attempts · never published

Sign in to submit an attempt. Assessments are graded on the server, so they cannot be taken anonymously.

  1. 1. A proposal states: "Outcome: 5,000 people will download our governance handbook." Where does this sit on the results chain, and what follows?

    Distinguish what the team controls from what it hopes to change.

  2. 2. A metric reads: "improved coordination between stakeholders." Which part of the four-part test does it fail first?

    Ask what you would have to do to obtain the number at all.

  3. 3. A proposal sets a target but no baseline exists for the indicator, and none can be reconstructed. What is the appropriate reviewer response?

    Baselines are cheap before and impossible after.

  4. 4. A funded education cohort votes more often after the programme than before. Which claim is proportionate to that design?

    Consider who joined the cohort and why.

  5. 5. Which change most improves "cumulative registered users: 12,400"?

    What could this number ever tell you that is bad news?

  6. 6. A grant pays per training session delivered. Sessions multiply while submissions stay flat. What does this best illustrate?

    This is about incentives, not honesty.

  7. 7. You finish a monitoring plan and find that three of four links in the results chain cannot be measured at proportionate cost. What belongs in your write-up?

    Recall the boundary this course keeps.

  8. 8. A report shows "representatives publishing rationales" rising from 40 to 46. What is the first thing to establish?

    Two people should be able to compute the same figure.

Reference sheet

KPI and monitoring worksheet

A printable worksheet for testing, rewriting and monitoring a proposed success measure the same way every time.

Definitions

Inputs
The resources committed: funds, staff time, equipment, existing infrastructure.
Activities
What the team does with the inputs: building, writing, running workshops, operating a service.
Outputs
The direct products of the activities, largely within the team control: a released tool, a published report, a number of sessions delivered.
Outcomes
Changes in behaviour, capability or condition among people outside the team, which the programme can influence but not control.
Impacts
Longer-term, wider changes to which the programme is one contributor among many.
Results chain
The ordered sequence inputs → activities → outputs → outcomes → impacts, used as a common evaluation vocabulary.
Theory of change
The results chain plus the explicit assumptions on each arrow explaining why one link should produce the next.
Observable
A stranger outside the team could obtain the number from a named source without the team cooperation.
Attributable
The programme is a plausible cause of movement in the metric - a claim about mechanism, not yet about proof.
Gaming-resistant
There is no shortcut that moves the number substantially at low cost without producing the intended effect.
Verifiable
Someone else, later, could reproduce the same figure from the same definition and source.
Indicator
A quantitative or qualitative variable that provides a simple and reliable means to measure achievement or reflect change.
Proxy
A metric used because the thing you care about cannot be measured directly. Every proxy carries a gap, which must be stated.
Correlation
Two quantities moving together. It constrains explanations without selecting among them.
Causation
A claim that the intervention produced the change, such that the change would not have occurred in the same way without it.
Confounder
A third factor influencing both the intervention and the outcome, producing an association that is not causal.
Counterfactual
What would have happened to the outcome without the intervention. Never directly observed; only ever approximated.
Selection effect
Distortion arising because participants differ systematically from non-participants in ways that also affect the outcome.
Contribution claim
A statement that the intervention was one plausible contributor among several, rather than the sole cause.
Attribution
The ascription of an observed change to a specific intervention, together with an account of the reasoning and its limits.
Baseline
The value of the indicator before the intervention, against which subsequent values are compared.
Denominator
The quantity a count is divided by to make it comparable across time or populations - per user, per epoch, per unit of spend.
Population definition
The explicit inclusion and exclusion rules determining who or what is counted.
Cadence
How often the metric is computed and published, and for how long.
Vanity metric
A number that reliably rises regardless of whether anything of value happened, and which therefore cannot distinguish success from failure.
Goodhart-style effect
The degradation of a measure once it is adopted as a target, because effort shifts towards the measure rather than the thing it stood for.
Monitoring line
A single sentence containing source, population, denominator, window, cadence, baseline and owner - everything needed to compute the metric without asking the team.
Limitation sentence
One sentence stating what the metric cannot show, published wherever the metric is published.
Unmeasured link
A link in the results chain for which no acceptable metric exists, recorded explicitly rather than left blank.
Monitoring plan
The full set of monitoring lines, attribution claims and limitations for a programme, agreed before funding rather than assembled at reporting time.

Checklist / method

  • Results chain: inputs → activities → outputs → outcomes → impacts. Which link does this metric sit on?
  • The chain is a hypothesis, not evidence. Which link is assumed rather than shown?
  • Test 1 - Observable: could a stranger obtain this number without asking the team?
  • Name the data source explicitly: which ledger query, which registry, which published report?
  • Test 2 - Attributable: could this programme plausibly move this number, and what else could?
  • Test 3 - Gaming-resistant: what is the cheapest way to make this number move without the intended effect?
  • Test 4 - Verifiable: could the same number be reproduced twelve months later, by someone else?
  • Record the failure rather than deleting the metric: a failed test is a finding.
  • Baseline: what is the value today, and if unknown, is establishing it a funded first step?
  • Denominator: per what - per user, per epoch, per unit of spend, per eligible population?
  • Population: who is counted, who is excluded, and does the definition change over time?
  • Window: over what period is it measured, and is that period long enough for the effect to appear?
  • Cadence: how often is it published, and does the schedule outlive the funding period?
  • Owner: which named role publishes the number, and what happens if they stop?
  • Counterfactual: what would you expect the number to do if the programme did not happen?
  • Confounders: name at least two other plausible causes of movement in this metric.
  • Proportionate claim: consistent with / suggests / is evidence of - pick the weakest wording the data supports.
  • Vanity check: would this number look good even if the programme achieved nothing?
  • Goodhart check: once someone is paid against it, what behaviour does it reward?
  • Limitation: write one sentence stating what this metric cannot show.
  • Never invent datasets, studies, sample sizes or figures; label illustrative numbers as fictional.
  • Measurement informs judgement. A met KPI does not prove value, and a missed one does not prove waste.

Final template

  1. 1.What was asked - a plain restatement of the request.
  2. 2.What I verified - claim, source, and what the source actually says.
  3. 3.What I could not verify - the open list, with the question still outstanding.
  4. 4.Strongest counter-argument - stated in its best form.
  5. 5.Disclosures - relationships, holdings or history relevant to this action.
  6. 6.How I would publish this - the rationale a reader could audit.