Pricing
from $9.99/mo · one credit = one minute of source video · no free tier
Real output from the clip generator — same footage, reframed by face tracking.
1 credit = 1 minute of video processing.
Just getting started
Billed monthly · $4.99/mo if you pay yearly
Cancel anytime.
Try it out
Billed monthly · $14.99/mo if you pay yearly
Cancel anytime.
For serious clippers
Billed monthly · $29.99/mo if you pay yearly
Cancel anytime.
For studios and back catalogues
Billed monthly · $119.99/mo if you pay yearly
Cancel anytime.
Speaker tracking with a stabilised camera (TRACK), full-width-over-blur for groups and landscapes (GENERAL), and beat punch-ins. The layout is chosen per scene by the classifier.
Everything in the standard engine plus the layouts that need a second pass over the video: the AI layout router that picks per source, two-speaker split screen, slides and shared screens stacked over the presenter, and hard cuts to whoever is speaking.
Each of these reads the video a second time before rendering — a Gemini pass on sampled frames, a two-face scan, a per-speaker mouth-activity scan. That extra pass is what the tier pays for.
AI Shorts (AI-actor UGC videos) and voice dubbing use premium generation from fal.ai and ElevenLabs. Connect your own keys for those — you're billed by those providers directly, typically around $0.65 to $2 per generated video. Your plan still covers the script and the orchestration around them. Managed credits for these are coming later.
Inside the plan: your credits cover video processing. Titles and descriptions cost no credits; AI thumbnail image generation uses about 3 credits per batch.
One credit is one minute of source video. The cost of a run is shown before it starts. There is no free tier.
The unit
A credit is a minute of what you feed in, not a clip that comes out. Here is everything that one charge covers.
An hour of conversation comes back as 6 to 12 clips, each 15 to 60 seconds. One charge covers the whole run, however many clips it yields.
Processing runs at roughly a third of real time: a 60-minute episode is finished in about 20 minutes, a 20-minute talk in about seven. The job runs server-side — the clips are waiting when you come back.
1080×1920, reframed around whoever is speaking, captions burned in word by word, a hook headline over the opening seconds.
Voice dubbing into 31 languages runs on your own ElevenLabs key, billed by them. Captions are verified end to end in English, Spanish, German and Russian.
The REST API and the MCP server draw on the same credits as the dashboard, so automating a pipeline costs exactly what running it by hand costs.
Credits meter the length of what you feed in, never the number of clips that come out. A long episode costs what it costs whether it yields six clips or twelve.
Paste a link and the cost of that job is on screen before anything starts. Nothing is spent on a video you decide not to process.
Billing runs on Stripe. Change or cancel a plan from your account page; unused top-up credits stay on the account.
Billing · FAQ
One credit is one minute of source video, and that single charge covers the whole run — every clip the video yields, the vertical reframing, the captions and the hook. A 13-minute talk is 13 credits; a 60-minute episode is 60, whether it comes back with six clips or twelve. Plans run from $9.99 a month for 50 credits to $199.99 for 1200, and paying yearly takes up to 50% off. There is no free tier, and the credit cost of a run is shown before it starts rather than after.
Four layouts. Speaker tracking with a stabilised camera (TRACK), full-width-over-blur for groups and landscapes (GENERAL), and beat punch-ins. The layout is chosen per scene by the classifier. That is what every plan runs. Pro and Ultra add the layouts that have to read the video a second time before rendering: the AI layout router that picks per source, split screen for two-speaker shots, slides and shared screens stacked over the presenter, and hard cuts to whoever is speaking. Punch-ins are on the standard side of that line and run on every plan: they ride the audio pass the renderer makes anyway, so they cost no second read of the video. If a run asks for one of the four advanced layouts on a plan that does not carry them, it is refused before anything is charged, and the refusal names the engines to clear so the same video can start on your current plan straight away.
No, on any plan. There is no free tier to hold the watermark back from, so this is not a feature we can sell you: every clip comes out at 1080×1920 with nothing burned over it but your own captions and hook.
Between 6 and 12 from an hour of conversation, each 15 to 60 seconds long. The number is not fixed: it scales with how much of the recording stands on its own, so a dense interview returns more than a slow monologue. You choose which of them are worth posting.
You do. Clippen takes no rights in your footage or in the clips it produces beyond processing the file you submitted, and publishes nothing on your behalf unless you press the button that posts it.
No. API keys are created on your account page, and calls through the REST API or the MCP server draw on the same credits as the dashboard. What the hosted endpoint buys you is that it is always on: an agent or a scheduled pipeline can clip and publish while your own machine is off.
No — those two run on your own fal.ai and ElevenLabs keys and those providers bill you directly, typically around $0.65 to $2 per generated video. Your Clippen plan still covers the script and the orchestration around them. Add the keys on the settings screen.