Methodology

Hands-on testing that can be repeated.

A review should distinguish product documentation, external benchmarks, editorial assessment, and actual hands-on evidence. These protocols define the last category.

Protocol version 1.0 · 2026-08-22

Publication threshold

A review may say “hands-on tested” only when the test log records the product plan, exact model/version when visible, test date, three real tasks, inputs, timing, corrections, usage consumed, failures, and retained output evidence. Until then, pages display “independent hands-on result not yet published.”

Core three-task design

  1. Typical task: the workflow the product most directly promises.
  2. Stress task: longer context, noisier input, multi-file work, or harder constraints.
  3. Failure task: a case designed to expose hallucination, consistency, permission, rights, or recovery limitations.

What to record

  • Time to first output and time to a usable result
  • Number of material corrections
  • Plan, credit, token, or task consumption
  • Instruction following and output quality on a 1–5 rubric
  • Errors, refusals, retries, and manual recovery
  • Privacy, consent, ownership, and commercial-use concerns
  • Output screenshot or evidence URL where publication rights allow

Assistant and research tools

Use the same source packet and three prompts across products. Score citation accuracy separately from writing quality. Verify every cited source manually. Record the model/version shown by the product, because the product default may change.

Coding assistants

Use one contained feature, one bug with tests, and one unfamiliar multi-file change. Start from the same repository commit. Record tests run, regressions, tool calls, token/credit use, manual interventions, and whether the generated patch was accepted.

Automation tools

Build the same trigger → transform → branch → notification workflow. Record setup time, operation/task counts, failure recovery, logging quality, credential scope, and the monthly cost at an identical run volume.

Image, video, and voice

Use matched prompts/scripts and fixed output settings. Keep modalities separate. Record failed generations, consistency across reruns, prompt adherence, editing control, latency, credit cost, watermarks, consent requirements, and commercial-use terms. Do not compare Elo scores from different arena pools.

Meeting and productivity tools

Use a consented recording with known speakers and a ground-truth action-item list. Measure speaker attribution, missed commitments, invented actions, search retrieval, export quality, and retention/privacy controls.

Download the log

Download the test-log CSV

Current status

The site has published trial checklists and external benchmark evidence where available. Independent hands-on result sets are being added only when the full protocol and evidence can be retained. Missing evidence is labeled rather than simulated.