Why a complete thought matters more here than anywhere
A gaming clip can survive starting slightly late, because the picture carries it. A podcast clip cannot. It is two people talking, so if the clip opens mid sentence the viewer has nothing to hold onto and scrolls. The unit that works is a question and its answer, or a claim and the line that backs it up.
So the whole episode is transcribed, and then the transcript is read from start to finish to find those units. What comes back is up to five stretches that each make sense to somebody who has never heard your show, which is the only audience a Short actually has.
Long episodes, in one piece
A pasted YouTube link can be up to four hours, so a long episode goes in whole rather than being split up first. The transcription is done in parallel across the episode, which is why a two hour source does not take two hours to process.
You will wait longer for a long episode than a short one, and we email you when the batch is ready so you can close the tab.
Two people in frame
A lot of podcast video is two cameras in one frame, or a guest window beside a host. Where there are two pictures we stack them for vertical instead of picking one: the top band gets one, the rest gets the other, whole. Where there is a single speaker on a plain background we simply keep the speaker.
Captions, and what happens without audio
Captions are burned in word by word, which for a talking video is most of the reason the clip works at all in a muted feed. If a video has no speech to caption, the captions are left off rather than filled with filler.
Audio only shows
This works on video. If your podcast lives as an audio file, it needs to be a video first, even a static one, and the usual road is the version you already publish to YouTube. Point us at that.
Your last episode has five Shorts in it
Paste the link and find out which five. Your first two videos are free to cut and to watch.
Frequently asked questions
›Can I use an audio file?
Not directly. The pipeline reads video, so an audio only episode needs a video version first. If you already publish the episode to YouTube, paste that link and you are set.
›How many clips from one episode?
Up to five. A long episode does not automatically produce more, because the aim is the five moments worth posting rather than a large pile.
›Does it label who is speaking?
No. Captions are the words, timed to the voice, without speaker names on them.
›Can I choose the moments myself?
You can look at every clip before anything is posted, and rerun a video if a batch was not useful. Hand picking timestamps up front is not something the tool takes today.
›Will it work on an interview with a remote guest?
Yes, and that is one of the better cases. Two windows in one frame is exactly the shape the stacked vertical layout is for.