Outcome Numbers Depend on Who a Vendor Enrolls
Two vendors can report nearly identical improvement rates on populations that share almost nothing.
During renewal reviews, outcome numbers from competing behavioral health vendors get lined up in the same table. One reports that 68% of members improved. Another reports 71%. The table treats those as comparable figures.
They usually are not. An improvement rate is a percentage of a population, and vendors do not enroll the same populations. Eligibility criteria, intake screening, and the way a program markets itself all shape who ends up in the denominator before any care happens. Two programs can run similar protocols, report similar percentages, and be describing groups of people with little in common.
Who gets into the denominator
Most digital behavioral health programs make decisions at onboarding that filter the population. Some of those are explicit and appear in the contract. A program may exclude members with active suicidal ideation, an active substance use disorder, an eating disorder diagnosis, or a recent inpatient stay.
Others are softer and never get written down. A program that presents itself as a stress and resilience app attracts people who describe their problem as stress. A program that requires a 20-minute onboarding flow before a first session loses the members with the least activation, who are frequently the ones carrying the heaviest symptom load.
None of this is concealed, though it rarely appears in the outcomes summary a buyer sees. By the time a vendor reports an improvement rate, the members who would have been hardest to move are often already outside the count. The reported figure is accurate for the members it covers, which is a narrower group than most buyers assume.
Two directions a number can move
Baseline severity pulls a reported percentage in opposing directions, which is part of why these comparisons resist eyeballing.
Programs serving mostly mild or subclinical presentations tend to post strong figures for the share of members who reached the normal range, because their members started close to it. The absolute change on a symptom scale can be small.
Programs serving members with severe presentations tend to show larger absolute point reductions and a lower share reaching the normal range, because the distance is longer. A member who moves from extremely severe depression to moderate depression has had a substantial change that a normal-range threshold does not register.
Neither pattern establishes that one program works better than the other. Both are partly a function of where the population started. Without the baseline distribution, an improvement rate cannot be interpreted.
What to ask for during a comparison
A handful of requests will tell you which vendors can answer at this level of detail.
Baseline severity distribution. Ask for the share of the measured population that fell in each severity band at onboarding, on a validated instrument, for the specific book of business being quoted.
Written exclusion criteria. Ask what happens at onboarding to a member who screens moderate-to-severe, and whether that member is served, routed elsewhere, or declined. Then ask whether declined members appear anywhere in the outcomes denominator.
Outcomes reported by baseline band. A vendor measuring continuously can report improvement separately for members who entered elevated and members who entered in the normal range. A vendor that cannot produce that cut is quoting a blended number.
Measurement cadence and completion rates. A single onboarding-and-discharge pair will overstate results relative to repeated measurement, because the members who stop responding are rarely the ones doing well.
What Wave reports
Wave's peer-reviewed outcomes come from a controlled engagement and outcomes study published in JMIR Formative Research (Pickover and Adler, 2025, N=64). More than half of participants presented with severe or extremely severe depression, anxiety, or stress at baseline. Coaching participants showed statistically significant greater symptom reduction than app-only controls across all three DASS-21 domains. The authors ran sensitivity analyses excluding participants who were in the normal range at intake, and the findings held among participants with elevated baseline symptoms.
Separately, and not from that study, Wave's internal book-of-business measurement shows 72% of engaged members reaching clinically meaningful symptom improvement within eight weeks. Those are operational figures from ongoing measurement across Wave's member population and should be read as such.
Wave administers the DASS-21 every 30 days for active members, which produces a longitudinal series for anyone who stays engaged. Members are not screened out for severity at onboarding. Coaches are National Board Certified and work with ongoing supervision. They do not diagnose or manage a medical condition, and they route to licensed care when a presentation calls for it.
Population composition determines whether an outcome number generalizes to the group covered under your contract. A workforce with a meaningful share of moderate-to-severe presentations will not behave like the population of a program that excluded those members at the door.
When two vendors put similar percentages in front of you, ask each one for the baseline severity distribution behind the number before you rank them.
Wave is a mental health coaching platform for US employers and health plans, built on National Board Certified coaches, measurement-based care, and outcomes partners can audit.
Want to learn more? Reach out to us at partners@wavelife.io.

