Step one: find moments, not intervals
The instinct is to slice the video into even lengths and post the best looking pieces. It does not work, because a Short has no context around it. A viewer arrives with no idea who you are or what came before, so the clip has to contain its own beginning.
What you are looking for is a stretch where something is set up and then paid off inside the clip. A question and its answer. A claim and the proof. A plan and the disaster. If you have to explain the setup in the caption, it is the wrong stretch.
- Good: you say what you are about to try, then it goes wrong.
- Good: a guest is asked something sharp and answers it.
- Bad: the funniest ten seconds, with the setup left in the long video.
- Bad: a stretch that opens halfway through a sentence.
Step two: decide what survives the crop
A 16:9 frame taken to 9:16 loses roughly two thirds of its width. Something has to go, and the question is what you are willing to lose.
For one person talking on a plain background, keep the person and let the background go. For a screen recording, the thing being demonstrated usually matters more than your face. For gameplay or any frame with a facecam plus a second picture, picking one loses the clip: the reaction without the cause, or the cause without the reaction. The answer there is to stack them vertically, face in a band across the top and the other picture whole underneath, which keeps both at the cost of each being smaller.
Step three: captions are not optional
Most Shorts are watched muted. A talking clip with no captions is a silent video of a mouth moving, and it gets scrolled. Burn them in rather than relying on platform captions, so they survive being downloaded and reposted, and time them to the voice so they read as speech rather than as blocks of text.
One thing to avoid: adding captions to a video with no speech in it. It sounds obvious, and it is a common bug in automated tools, which will happily render empty or garbled caption boxes over silent footage.
Step four: give the first two seconds a reason
A line of text on screen at the very start, saying what the clip is about to show, measurably changes whether people stay. It is not a title. It is the promise the next thirty seconds keeps. Keep it to a handful of words and put it where it does not cover a face.
Doing it by hand, and doing it automatically
By hand this is real work: watch the long video with a notepad, mark timestamps, cut each one in an editor, reframe each one, generate and correct captions, add the hook, export. Twenty minutes a clip once you are quick, plus the watching.
Automatically, the steps are the same, and the part that varies between tools is step one and step two. Anything can cut a video up and put captions on it. Whether the moment stands on its own, and whether the reframe kept the thing that mattered, is where the difference shows.
What we do with these steps
Uploadworthy transcribes the whole video, reads that transcript end to end to pick moments that contain their own setup, then rebuilds each one at 9:16, stacking a facecam over a second picture where there is one, burning the captions in word by word and putting a hook line on the front. Point it at a YouTube video you own and the first two are free to cut and to watch.
Try the four steps on your own video
Point us at a long video you own and see the moments it picks, reframed and captioned. The first two are free.
Frequently asked questions
›How long should a Short from a long video be?
Long enough to contain its own setup and payoff, which usually lands somewhere between twenty and sixty seconds. Cutting a good moment short to hit a number is worse than letting it run to where it ends.
›How many Shorts can one long video give me?
Fewer than you would hope. A long video typically holds three to five stretches that genuinely stand on their own. Beyond that you are posting filler, and filler trains the feed not to show you.
›Should I post the same clip to several platforms?
Yes, that is most of the point of making it vertical. A 9:16 clip with burned in captions works on Shorts, Reels and TikTok without further work.
›Do I need an editor to do this?
To do it by hand, yes, and the reframing is the slow part. The steps themselves are not complicated, they are just repetitive, which is exactly why it is worth automating.
›Does this work on gameplay?
Yes, with one extra decision: do not centre crop it. Stack the facecam over the gameplay so you keep the reaction and the cause in one frame.