ClippenCut my clips
Home Podcast to Shorts

Turn a podcast episode into vertical shorts

TL;DR

One episode in, 6 to 12 clips out per hour of recording, each 15 to 60 seconds at 1080×1920 with word-level captions burned in.

Two-speaker shots are stacked in half-frames rather than cropped down to one person, so a guest answer stays a conversation instead of becoming a monologue with an off-screen voice.

A 60-minute episode is finished in about 20 minutes. Paid, from $9.99 a month, or $4.99 a month billed annually. Credits are charged per minute of source video, not per clip.

How many shorts does one episode actually yield?

6 to 12 clips per hour of recording, each 15 to 60 seconds long. The count moves with the episode rather than a quota: a tight interview where the guest lands three good answers in a row produces more usable segments than an hour of scheduling talk, and the selection reflects that instead of padding.

That is the number worth planning around, because it changes what an episode is. A weekly show that publishes one episode publishes one thing a week. The same show clipped publishes one long-form piece and eight to ten short ones, on three platforms where the discovery actually happens. The recording already exists. The guest has already given you the hour. The clips are the part of that hour that reaches anybody who has not already subscribed.

Most podcast archives are unpublished inventory in exactly this sense: a back catalogue of fifty episodes is fifty hours of material that has been paid for once and distributed once. Clipping an hour by hand — finding the moments, cutting, reframing, timing captions — is most of a working day, which is why the back catalogue stays where it is.

Why does a two-speaker shot need a different layout?

This is the single biggest reason generic clippers produce disappointing podcast clips, and it is worth being specific about.

A podcast is usually shot as a two-shot: both people in one 16:9 frame, sitting apart. Crop that to 9:16 around one face and the other person leaves the frame entirely. The result is a clip where a disembodied voice asks a question and a single visible person answers it, with the reaction — the raised eyebrow, the laugh, the interruption — cropped out of existence. The exchange was the content, and the crop deleted half of it.

Clippen's SPLIT layout stacks both speakers in half-frames, one above the other, so the clip is still a conversation. It is gated deliberately: SPLIT only engages when both faces are genuinely visible in the same frame for a sustained part of the scene. That test is what separates a real two-shot from a shot/reverse-shot edit where the camera alternates between the hosts — stacking that would show the same person twice, top and bottom, which looks broken.

Where the episode is cut as alternating single shots, the speaker-cut mode handles it instead: hard cuts to whoever is currently talking, driven by mouth activity normalised per speaker rather than raw motion, because raw motion just hands the scene to whoever is better lit.

Audiograms or real video clips?

An audiogram is a static image with a waveform and subtitles over it. It is what an audio-only podcast can produce, and for an audio-only show it is a reasonable answer.

It is not what Clippen makes. Clippen needs video in, and it gives you the actual footage recomposed vertically: the speaker tracked, the reaction shot kept, the camera pushing in slightly on the beats so a fixed tripod does not produce a fixed-looking clip. On feeds where the content competes with anything else on the platform, a waveform over a cover image is competing with a face.

The practical consequence: if your show is recorded audio-only, the honest answer is that a clipper of this kind has nothing to reframe, and an audiogram tool is the right purchase. If you already record video and only publish the audio — which is most shows that started as audio and added cameras — that video is sitting on a drive doing nothing, and it is the entire input this tool needs.

What comes back, and what does it cost?

Every clip arrives finished rather than as a raw cut:

  • 1080×1920 vertical, with 9:16, 1:1 and 16:9 available from the same run.
  • Word-level captions burned in, with each word highlighted as it is spoken and the language detected automatically.
  • An AI-written hook headline over the opening seconds, where a scrolling viewer decides.
  • A layout chosen per scene: one speaker tracked, two speakers stacked, slides or a shared screen stacked over the presenter.
  • Optional dubbing into 31 languages with the speaker's voice preserved and the captions re-transcribed to match.
  • Download, or post to TikTok, Instagram Reels and YouTube Shorts from the dashboard.

A 60-minute episode is finished in about 20 minutes, because processing runs at 0.30 to 0.36× real time. Start the job when the guest leaves and the clips exist before you have finished writing the show notes.

Clippen is a paid tool and bills by the length of what you feed it: one credit is one minute of source video, and that one charge covers the whole run — every clip the video yields, the vertical reframing, the captions and the hook. A 60-minute podcast costs 60 credits whether it comes back with six clips or twelve. Plans start at $9.99 a month for 50 credits ($4.99 a month billed annually) and run to $199.99 a month for 1200 credits. There is no free tier.

For a weekly hour-long show that is 60 credits a week, known in advance, whether the episode comes back with six clips or twelve. Compare that with tools metered in finished videos per month, where the cost of an episode depends on an output you cannot predict before you run it.

Common questions

How many shorts can I get from one podcast episode?
Between 6 and 12 per hour of recording, each 15 to 60 seconds long. The count scales with what is actually in the episode rather than a fixed quota, so a dense interview yields more usable clips than an hour of housekeeping.
Will both hosts stay in the frame?
Yes, when they are genuinely in the same shot. The SPLIT layout stacks two speakers in half-frames so the exchange survives the crop to 9:16. It only engages when both faces are visible in the same frame for a sustained part of the scene, because stacking a shot/reverse-shot edit would show the same person twice.
Does it work on audio-only podcasts?
No. Clippen reframes real footage, so it needs video in. If your show is audio-only, an audiogram tool is the right purchase instead. If you record video and only publish the audio, that footage is exactly what this is for.
How long does an episode take to process?
Processing runs at roughly a third of real time, so a 60-minute episode is finished in about 20 minutes and a 90-minute one in a little under half an hour. The job runs server-side, so you can close the tab.
Can I publish clips in another language?
Yes. A finished clip can be dubbed into 31 languages with the speaker's voice preserved, and the dubbed audio is re-transcribed so the burned-in captions match the new language rather than the original. Caption rendering has been verified end to end on English, Spanish, German and Russian.
What does clipping a podcast cost?
Clippen is a paid tool and bills by the length of what you feed it: one credit is one minute of source video, and that one charge covers the whole run — every clip the video yields, the vertical reframing, the captions and the hook. A 60-minute podcast costs 60 credits whether it comes back with six clips or twelve. Plans start at $9.99 a month for 50 credits ($4.99 a month billed annually) and run to $199.99 a month for 1200 credits. There is no free tier.
Are podcast clips audiograms or real video?
Real video. Each clip is the actual footage recomposed to 9:16 with the speaker tracked, the reaction kept where both people share the frame, word-level captions burned in and a hook over the opening seconds. There is no waveform-over-cover-art output.

Related comparisons