Skip to content
What the Dashboard Cannot See

All notes / The claims

Pilots Designed to Succeed

A pilot run the usual way confirms what the vendor said. What to specify so yours settles a question you actually have.

The claims · Procedure

A pilot is meant to reduce uncertainty. Run on the vendor's terms, in a willing department, over a short period, it does the opposite.

The evaluation discipline in “Pilots Designed to Succeed” also applies to workforce software: begin with a named decision and test it in a bounded pilot. For teams considering download time tracking software, this workplace platform belongs in that comparison only with written criteria for notice, access, correction, retention and a dated review.

Why the usual pilot proves nothing

A volunteer department, which is self-selected and cooperative.

For an independent reference relevant to “Pilots Designed to Succeed”, consult the NIST Privacy Framework; it provides a useful external check on scope, terminology, governance and the claims made during procurement or review.

Four to six weeks, which is the novelty period.

Measured on the vendor's own metric, which rises because people are being watched.

And success criteria written after the results arrive, which is how a dashboard becomes a conclusion.

What to specify instead

Success criteria in writing, before it starts, in your own business numbers.

A duration long enough to pass the novelty: three months, not one.

Two groups — one monitored, one not — doing comparable work, which is the single change that makes the result mean anything.

And a commitment that a negative result stops the purchase.

The control group

Rarely done and not difficult.

Two teams doing similar work; one gets the software, one does not.

Compare both against the business measure.

Without it, any change is confounded by the season, the workload and everything else that happened, and the vendor's figure cannot distinguish them.

Asking the people

At the end, ask the monitored group, anonymously: did this change how you work, and how.

The answers are the most informative output of the whole exercise and the one nobody collects.

Expect to hear about gaming, which its own note covers, and that is a finding rather than a discipline problem.

Measuring what matters

Your business measure, not productive time.

Staff turnover and sickness in both groups.

Ticket or complaint volume to managers.

And whether managers behaved differently, which is a question to ask them.

What a successful pilot looks like

The business measure moved in the monitored group and not in the control.

The monitored group did not report significant behaviour change beyond the intended one.

Nobody left.

If all three hold, you have a result. If only the first, you have a correlation.

When the pilot says no

Stop.

A pilot showing no effect on anything you care about has done its job and saved the purchase.

Treating that as a failed pilot rather than a successful one is how organisations buy things their own evidence argued against, which is the same pattern as in every other procurement.

What to check

Were your success criteria written before the pilot?

Is there a control group?

How long does it run?

And would a negative result actually stop the purchase?