Blog · Research · August 24, 2026 · 8 min read
We Scored 1,655 Videos to See If Packaging Predicts Views. It Does Not.
We build a tool that scores video packaging, so we had an obvious commercial interest in the answer to one question: does a higher score mean a video performs better? We ran the study properly, and the answer was no. Here is the whole thing, including the parts that make our own product look worse, because a result you only publish when it flatters you is not a result.
Why we could not answer this before
Every earlier check we ran was confounded. The videos were whatever our own users had linked, and 812 of 907 came from four accounts. A correlation across that pool mostly measures which creator a video belongs to, not whether the score means anything. At that size the honest confidence interval was wide enough to be compatible with a real effect, so the question stayed open rather than answered.
How the study was built
- 300 channels found through the YouTube API across 29 niches, from 569 subscribers to 3.78 million.
- 138 of them sampled, 12 videos each, chosen randomly inside the channel rather than most recent, because recency tracks views and would have narrowed the performance range artificially. 1,655 videos scored.
- Scored blind by the exact production code path, with no views, likes, subscriber count or channel identity visible to the scorer.
- Judged on within channel percentile: every video ranked only against other videos on the same channel, so channel size, niche and audience are held constant by construction rather than adjusted for afterwards.
- Significance by permutation, shuffling scores inside each channel, which is the exact null hypothesis and immune to the clustering that makes naive p values far too small.
The result
Correlation between score and within channel performance, where 0 means no relationship:
- Overall score: +0.019 (p = 0.48)
- Title: +0.011 (p = 0.68)
- Description: +0.040 (p = 0.19)
- Hashtags: +0.006 (p = 0.89)
- Thumbnail: minus 0.005 (p = 0.86)
Nothing rescues it
The same null shows up on engagement rate instead of views, in all five subscriber tiers separately, and for short and long form separately. The 95% confidence interval on the overall figure runs from about minus 0.03 to +0.07, which rules out anything above roughly 0.07: at most half a percent of the variation in how a creator's videos perform.
There is also no hidden subset of creators it works for. 54% of the 138 channels show a positive correlation and the median is 0.060, which is precisely the spread you would expect from noise at 12 videos per channel.
The two checks that make a null result trustworthy
A null is easy to manufacture with a broken measurement, so we tried to break ours twice.
First, a positive control. We ran the identical machinery against dumb mechanical features, and it found them: video duration at minus 0.054, and raw title length at +0.057. The apparatus detects small effects, and the number of characters in your title beat our entire rubric. Neither survives correction for testing seven things at once, which is the real headline: within channel performance is overwhelmingly not a packaging phenomenon.
Second, range restriction. The obvious objection is that scores barely vary inside one creator's catalogue, so there was nothing to correlate. They do vary. The median gap between a channel's best and worst scored video is 24 points, while the median gap between its best and worst viewed video is 28 times. Both sides move. They do not move together.
What this does not say
This is an observational ranking study, and the distinction matters. It establishes that packaging quality as we measure it does not explain which of a creator's videos won. It does not establish that improving your title has no effect. Those are different claims, and a genuine effect could be swamped by topic and by how the platform chose to distribute each video.
It also covers four of the seven things we score. Hook, pace and audio are measured off the video file itself, which the API does not hand over, so this study says nothing about them and should not be quoted as if it did.
What we did about it
We stopped making performance claims, including the softened versions. No number on our site says a higher score earns more views, because we cannot support that and we know it.
What a score can honestly be is craft conformance: a linter for the thing you are about to publish. A linter is genuinely useful and does not promise revenue. So the product leads with specific measured problems now, the kind with a right answer, rather than with a grade.
See what we can actually measure in your video
Upload your video and get a scored breakdown of your title, thumbnail, hook, pace, and audio before you post. Free to start.
Score my video free