Article 18 · Design the measurement system

Building a metric tree

Before building a metric tree, decide whether the relationship is mathematical, hypothesised, or diagnostic. Each structure makes a different claim.

Teams often call any hierarchy of measures a metric tree.

That is convenient, but dangerous. A branch can mean that one measure is mathematically part of another, that the team believes it may influence an outcome, or simply that it is a useful place to investigate. Those are not the same claim.

Before drawing the tree, decide which relationship you are representing.

Three different structures

Not every metric tree makes the same claim

Choose the relationship before drawing the hierarchy, and label the connection explicitly.

Mathematical decomposition

calculated from

  1. Parent measure A calculation that reconciles to defined components.
  2. Child measures Share compatible units, populations and time windows.
  3. Evidence needed Definitions and arithmetic that can be checked.

Use this to explain how a measure is calculated.

Hypothesised drivers

may influence

  1. Outcome or behaviour The result the team wants to understand.
  2. Possible drivers Beliefs about factors that might affect the result.
  3. Evidence needed Experiments, analysis, research or stronger comparison designs.

The branches are hypotheses, not proven causes.

Diagnostic workflow map

diagnosed through

  1. Headline signal The movement that triggered investigation.
  2. Workflow measures Stage, timing, failure and coverage evidence.
  3. Evidence needed Reliable workflow records and useful dimensions.

Use this to locate where observed behaviour may have changed.

Do not use the same arrow to mean calculation, influence and diagnosis.

Three structures are commonly mixed together

1. Mathematical decomposition

A decomposition shows how a measure is calculated from components that reconcile to it.

For the service-quotes workflow:

Accepted eligible requests
=
Eligible requests
× Quote coverage rate
× Acceptance rate among quoted requests

This relationship can be checked mathematically if the definitions, populations, and time windows align.

A decomposition is useful for explaining the arithmetic of a headline measure. It should not quietly include a metric such as customer confidence or provider responsiveness unless that metric is genuinely part of the formula.

2. Hypothesised driver tree

A driver tree shows factors the team believes may influence an outcome.

Request acceptance rate
├── Provider availability
├── Time to first quote
├── Quote relevance
├── Ease of comparison
└── Customer confidence

These branches are hypotheses, not components of a formula and not proven causes.

Faster responses may be associated with higher acceptance because suitable providers respond quickly. Or both may be influenced by service category, geography, request quality, or provider supply. The tree helps the team state and test its beliefs. It does not prove them.

3. Diagnostic workflow map

A diagnostic map organises measures around points where a workflow may be changing or failing.

Eligible request published
→ suitable providers notified
→ first quote received
→ additional quotes received
→ quotes viewed or compared
→ quote accepted

Supporting measures might include:

  • notification coverage;
  • quote coverage rate;
  • median time to first quote;
  • multi-quote coverage rate;
  • quote view rate;
  • request acceptance rate;
  • withdrawal or expiry rate.

This structure helps the team decide where to investigate when the headline metric moves. It does not claim that each measure is a mathematical child or causal driver of acceptance.

Label the relationship, not just the measure

The problem is rarely the box or the branch. It is the unlabeled claim between them.

A useful tree or map makes the relationship explicit:

Edge label Claim being made Evidence needed
calculated from The child measures reconcile mathematically to the parent Shared definitions, units, populations, and windows
may influence The child is a hypothesis about a driver Analysis, experiments, research, or other supporting evidence
diagnosed through The child helps locate where movement may be occurring Reliable workflow evidence and useful dimensions

Do not use the same arrow to mean all three.

Counts and rates do not automatically belong in one hierarchy

A common tree might place request count, quote count, provider count, response rate, acceptance rate, and customer satisfaction underneath one headline outcome.

Those measures may all be useful, but their units and relationships differ:

  • requests are customer opportunities;
  • quotes are provider responses;
  • providers are supply-side actors;
  • rates depend on explicitly defined denominators;
  • satisfaction may be available only for a self-selected subset;
  • completed work may occur well after quote acceptance.

Putting them into one visual hierarchy does not make them comparable or explanatory.

Start by naming the unit of every branch. Then ask what relationship the branch actually claims.

Dimensions normally sit beside the structure

Dimensions help the team interpret a measure without becoming new branches in the core tree.

For example:

Metric:
Request acceptance rate within 30 days

Useful dimensions:
service_category
geography
provider_availability_band
request_route
workflow_version

A lower rate in one geography may point towards a supply or service-coverage problem. It does not require a separate permanent metric-tree branch for every location.

Keep the definition stable and use dimensions to investigate where the pattern differs.

Build from a real decision

Use this sequence:

  1. Name the headline question. What movement is the team trying to understand?
  2. Choose the relationship. Is the structure a decomposition, a set of hypotheses, or a diagnostic map?
  3. Define the unit and window. Do not mix requests, quotes, customers, and providers without making the change visible.
  4. Add only useful branches. Each branch should change the calculation, hypothesis, or investigation.
  5. Label assumptions. Make uncertain or unvalidated relationships explicit.
  6. Test a scenario. When the headline moves, does the structure help the team calculate, investigate, or test something specific?
  7. Review after change. Product, population, instrumentation, and operating conditions can make an old tree misleading.

One outcome may need more than one structure

A team investigating request acceptance could use all three structures, provided it keeps them separate:

  • a decomposition to explain the arithmetic;
  • a driver tree to record hypotheses about supply, speed, relevance, and confidence;
  • a diagnostic workflow map to locate where observed behaviour changed.

Combining them into one large tree may look comprehensive, but it weakens the meaning of every branch.

The practical rule is simple: state what the connection means before using the connection to explain anything. A metric tree becomes trustworthy when its relationships are as carefully defined as its metrics.