Design the pilot around a later decision

A successful pilot does not mean the robot looked impressive. It means the evaluation produced enough trustworthy evidence to make the next decision. Start by writing that decision: continue testing, stop, change the task, expand to another room, or consider a commercial commitment after current terms are confirmed. Then choose one or two bounded routines and the people, spaces, dates, supervision, and KOKO version covered by the result.

Record a baseline before introducing KOKO. Measure how the routine works today, including completion time, human effort, errors, interruptions, stress, and any existing device or service cost. Without a baseline, a pilot can report activity but cannot show improvement. Do not borrow performance claims from another setting or treat a roadmap description as the starting result; observe the exact task under the agreed household or study conditions.

Use the KOKO price and value hub

Prepare a bounded KOKO demo

Measure outcomes and the human work behind them

Define a successful task in observable terms: what must happen, by when, within which area, and with what quality. Track attempts, completions, partial completions, and failures rather than reporting only the best demonstration. Time can be useful, but faster is not automatically better if the result creates cleanup, confusion, or extra checking. Measure the outcome the household values instead of a convenient proxy that does not change daily life.

Count every form of human contribution: setup, commands, repeated prompts, supervision, object preparation, route clearing, recovery, charging, troubleshooting, and post-task correction. Separate expected onboarding work from recurring assistance. A task completed after ten minutes of invisible human preparation may still be useful, but its value is different from an independent completion. This labor record is also necessary for a fair cost and accessibility evaluation.

  • Attempt, completion, partial completion, and failure counts
  • Elapsed time plus setup, supervision, recovery, and cleanup time
  • Outcome quality measured against a written acceptance standard
  • Conditions that made performance better, worse, or impossible

Track reliability, interventions, and safe recovery

Reliability is consistency across ordinary variation, not one flawless run. Repeat the routine at different approved times and with realistic, controlled changes in lighting, furniture, background activity, or network conditions. Record the denominator with every rate. Nine successful attempts out of ten says more than '90 percent' alone, and neither should be generalized beyond the tested build, route, objects, people, and conditions.

Log interventions by type and consequence: clarification, remote help, physical repositioning, emergency stop, task abandonment, or recovery after an error. Define in advance which events pause the pilot and who has stop authority. Near misses, privacy surprises, inaccessible controls, distress, pet reactions, and unexpected data collection are outcomes even when no damage occurs. A safe stop can be a successful protective response while still revealing a product-fit limitation.

Apply the KOKO household safety checklist

Separate capability evidence from roadmap

Include experience, consent, privacy, and accessibility

Ask each affected person about usefulness, predictability, comfort, perceived control, and willingness to continue. Preserve individual responses rather than averaging away a serious concern. Participation should be voluntary, and guests, children, care recipients, researchers, or workers may require different consent and oversight. Adoption is not the same as compliance: a person using the system because someone else insists is not evidence that the experience fits their needs.

Test whether intended users can start, pause, understand, correct, and stop the interaction through methods they can reliably use. Record workarounds and assistance instead of calling the interface accessible by assumption. For privacy, document which spaces, inputs, accounts, transmissions, logs, and retention settings were actually used. The pilot should stay within approved data boundaries and stop when the current data flow cannot be explained.

Evaluate accessible KOKO interaction

Review KOKO privacy questions

Set thresholds before seeing the results

For each metric, define a target, a minimum, and a stop threshold before the pilot begins. Also define the number of attempts and minimum observation period. The final scorecard should show results, missing data, exceptions, participant feedback, incident notes, human hours, and evaluation costs. Avoid combining everything into one attractive score; a serious safety or consent problem should remain visible even when other measures are strong.

End with go, revise, extend, or stop, plus the evidence behind that choice. A good pilot can conclude that the task is unsuitable, that more controlled testing is needed, or that value depends on different terms. Searches for 'the price is right robot KOKO' become answerable only after the result and confirmed cost are compared. The pilot does not establish a universal KOKO value; it supports a dated decision for this setting.

  • Predefined target, minimum, and stop threshold for each metric
  • Complete result counts with exceptions and missing observations
  • Human hours and pilot costs included beside outcome measures
  • A dated next decision and the conditions required to revisit it

Frequently asked questions

How many KOKO pilot metrics should I use?

Use a small balanced set that covers the intended outcome, reliability, human effort, safe recovery, user experience, privacy, accessibility, and cost. More metrics do not help if the pilot cannot collect them consistently.

Is task completion rate enough to prove value?

No. Include the number and conditions of attempts, outcome quality, setup and intervention time, failure consequences, participant experience, and total cost. A completion can still require more work than the prior routine.

Can a KOKO demo count as a pilot?

A demo can answer narrow questions, but a pilot usually needs a defined baseline, repeated observations, approved users and spaces, consistent records, and decision thresholds over a suitable period.

Does a successful pilot guarantee the same result after purchase?

No. Record the build, environment, support, configuration, people, and dates tested. Confirm that any later offer covers the same relevant conditions and identify what changed.

Sources & further reading

  1. Homebot One: Official KOKO overview (opens in a new tab)
  2. NIST AI Risk Management Framework (opens in a new tab)
  3. GAO Technology Readiness Assessment Guide (opens in a new tab)
  4. NIST Privacy Framework (opens in a new tab)

From Homebot One, the team building KOKO in Fremont, California.