OE-007 Outcome Evaluation
Evaluating the credibility of published outcomes
A practical checklist for parents, funders, and partners to judge whether a youth program's published outcomes data is credible before trusting it.
Entry checked on
Q.Why can't I just trust the numbers a program publishes?

To judge whether a youth program's published outcomes are credible, check whether the data was gathered through a defined evaluation process, whether the methods and limitations are disclosed, and whether anyone outside the program had a hand in reviewing the findings. A number on a website with no explanation of how it was measured is a claim, not evidence.
This question matters for parents choosing a program, funders deciding where to invest, and partner organizations deciding whether to refer families. Most of us cannot audit a program's data ourselves, but we can learn to spot the signs that separate a credible evaluation from a marketing number.
Why can't I just trust the numbers a program publishes?
Because a published outcome, such as "90% of youth improved their grades," tells you almost nothing on its own. It does not say how improvement was defined, how many youth were actually measured, who measured them, or what happened to the youth who did not improve. The Centers for Disease Control and Prevention's program evaluation framework lists gathering credible evidence as one of six distinct steps in a sound evaluation, separate from describing the program and from drawing conclusions (cdc.gov). A credible number is the product of that full process. A number with no visible process behind it should be treated as unverified.
What should a credible published outcome actually include?
A credible outcome claim should let you answer four questions without contacting the program: What exactly was measured, who was measured and how many, over what time period, and compared against what. The CDC framework names five standards that mark high-quality evaluation: "Relevance and utility, Rigor, Independence and objectivity, Transparency, Ethics" (cdc.gov). Of these, transparency and rigor are the two you can check yourself as an outside reader, because they should be visible in how the program describes its own results, not hidden inside an internal file you will never see.
What are the warning signs of a weak or misleading claim?
Treat a published outcome with caution when it shows any of these patterns:
| Warning sign | Why it matters |
|---|---|
| No sample size given | A 90% improvement rate means something different for 10 youth than for 400. |
| No comparison point | "Improved" compared to what: a pre-test, a prior year, a similar program? |
| Vague measurement | "Felt more confident" is not the same as a validated survey instrument. |
| Only the best year shown | Cherry-picked results hide the normal variation that real programs have. |
| No mention of who was excluded | Youth who dropped out or were not measured can quietly disappear from the count. |
| No outside review | Self-reported numbers with no external check carry less weight than reviewed ones. |
None of these signs alone prove a program is dishonest. Smaller nonprofits often lack the staff to run a full outside evaluation. But a pattern of several warning signs together should lower your confidence in the claim.
How does independence and objectivity change how much I should trust a result?
The CDC framework lists independence and objectivity as a standard separate from rigor, which means a methodologically sound evaluation can still be weakened if the same people who run the program are the only ones who measured and reported its success (cdc.gov). This is not an accusation of dishonesty; it is simply a structural risk that any self-reported result carries. A program that discloses who collected the data, and whether anyone outside the program reviewed it, is giving you the information you need to weigh that risk yourself. A program that does not disclose this is asking you to take it on faith.
What questions can I ask a program directly?
If a published outcome matters to your decision, it is reasonable to ask the program these questions before relying on it:
- How many youth were included in this result, and over what time period?
- How was the outcome measured, and was the measurement tool validated elsewhere?
- What happened to youth who left the program before it ended?
- Did anyone outside the program's staff review the data or the methods?
- Is a fuller evaluation report available beyond the summary number on the website?
A program with a genuinely credible evaluation will usually be able to answer most of these without defensiveness, because the answers already exist somewhere in its own process. A program that cannot answer any of them is likely reporting an impression, not a measured result.
What does this mean for how a program should publish its own results?
The same standards work in reverse. A program that wants its outcomes to be trusted should publish the sample size, the time period, the measurement method, and whether the data was reviewed by anyone outside its own staff, following the same transparency and rigor standards named in the CDC framework (cdc.gov). For guidance on choosing which indicators to track in the first place, see our article on choosing outcome indicators for youth programs, and for turning results into action once they are gathered, see improving programs with evaluation data.
Credibility is not about perfection. It is about whether a program shows its work. A result you cannot trace back to a method is a claim worth questioning, and a result that is transparent about its limits is usually the one worth trusting more, not less.

