Knowledge centerOutcome definition

Outcome Design

An Outcome Must Be Defined Before It Is Measured

A reproducible endpoint specifies the variable, participant-level metric, method of aggregation, time point, assessor, and decision rules.

Dr. Andrie Udal, PhD, MSHDSHealthcare Data Science Lead, ProWritePublished 27 July 2026Updated 27 July 2026
ProWrite editorial framework: make every critical decision explicit, reviewable, and traceable to the research record.

“Improvement in pain” sounds like an outcome, but it does not tell a research team what to collect or an analyst what to calculate. Improvement could mean a lower score at week 12, change from baseline, a 30% reduction, time until sustained relief, or a participant’s global judgment. Each definition can produce a different result.

An outcome must be operational before it is measured. The SPIRIT-Outcomes 2022 extension identifies core components that turn a broad domain into a reproducible endpoint: the specific measurement variable, participant-level analysis metric, method of aggregation, and time point. In practice, teams should add source, assessor, instrument version, and decision rules.

01Name the variable actually collected

Begin with the measurement variable: a laboratory value, validated questionnaire score, adjudicated event, device reading, clinical assessment, or record-derived field. Name the instrument, edition, scale direction, unit, and source. “Quality of life” is a domain; the exact instrument and score are the variable.

If an event requires adjudication, define the event criteria, evidence sources, committee process, and handling of uncertainty. If data come from electronic records, specify the code sets, encounter types, look-back window, and validation of the extraction logic.

02Specify the participant-level metric

The metric is the value carried from each participant into analysis. Common choices include final value, change from baseline, percentage change, time to event, maximum value, or presence of an event. These are not interchangeable.

If change from baseline is used, define baseline. Is it the last measurement before treatment, the mean of several measurements, or a fixed visit? If a responder threshold is used, justify the cutoff and state whether equality counts. If participants can have repeated events, decide whether the endpoint uses first event, total events, or event rate.

03Define aggregation and comparison

Explain how participant values become a group result: mean, median, proportion, risk, rate, survival probability, or another summary. Identify the contrast between groups or conditions. A composite outcome requires definitions for every component and a rule for how components combine.

The estimand should also address intercurrent events. What happens to outcome interpretation after treatment discontinuation, rescue therapy, switching, or death? ICH E9(R1) makes these scientific choices part of the treatment effect being estimated, not end-stage data-cleaning decisions.

04Fix the time point and allowable window

“At follow-up” is not reproducible. State the target time and acceptable window. If several observations fall in the window, define the selection rule. If a visit is late, decide whether it remains eligible and how timing enters the analysis. The WHO Trial Registration Data Set asks primary and secondary outcomes to include the name, measurement method, and time point, underscoring that timing is part of the public definition.

Time-to-event outcomes need a start time, event definition, censoring rules, and follow-up horizon. Without these, two analysts can construct different survival times from the same records.

05Align every study document

The protocol, registry, case-report form, data dictionary, statistical analysis plan, table shells, and manuscript should use the same outcome definition or explain authorized changes. A registry entry reading “mortality at 30 days” should not become “in-hospital mortality” in the paper without a documented reason; discharge practices make those endpoints different.

Create an outcome specification table with one row per primary and secondary outcome and columns for domain, variable, metric, aggregation, time point, source, assessor, missingness, intercurrent events, and analysis. Review it before the collection system is finalized.

06Test the definition with a reproduction exercise

Give the specification and a small set of mock participant records to an independent analyst. Ask that person to derive the endpoint without verbal help. Disagreement reveals ambiguity while correction is still cheap.

Test edge cases deliberately: a measurement just outside the visit window, two observations equally close to the target date, a partial date, an event followed by reversal, rescue treatment before assessment, and death before follow-up. The written rule should produce a determinate, clinically defensible result for each case or explicitly route it to adjudication. Record how adjudicators handle insufficient or conflicting evidence.

The final test is simple: could another qualified team identify the same outcome for the same participant at the same time using only the written rule? If not, the project has an outcome idea—not yet a reproducible outcome.

References

  1. SPIRIT-Outcomes 2022 Extensionoutcome variable, metric, aggregation, time point, composites, and target differences. Accessed 27 July 2026.
  2. SPIRIT 2025 Statementcurrent minimum protocol items for randomized trials. Accessed 27 July 2026.
  3. ICH E9(R1): Estimands and Sensitivity Analysis in Clinical Trialsoutcomes, intercurrent events, treatment effects, and estimands. Accessed 27 July 2026.
  4. WHO Trial Registration Data Setrequired names, measurement methods, and time points for registered outcomes. Accessed 27 July 2026.

This article is educational and intended for research purposes. It does not provide individual medical advice, diagnosis, or treatment. No patient data were used.