Use casesPodcasts

Podcast clips, cut from the episode

A two-hour episode contains maybe six minutes anyone will share. Finding them means scrubbing the whole thing again, which is why most episodes get one clip cut in a hurry or none at all. Glyphcast reads the episode and proposes the moments.

Why podcasts are the easiest case

Everything the scorer is good at, a podcast has. There is continuous speech, so the transcript is dense and word-level timings are accurate. There are natural pauses between thoughts, which is where a clip should start and stop. And the interesting moments are audible - a laugh, an interruption, someone raising their voice - which the audio-energy pass picks up independently of what was said.

What it is not good at is a video with no speech in it. A podcast is the opposite of that problem.

Clips start and end where a sentence does

The failure that makes an automatically cut clip unusable is not a bad moment - it is a good moment that starts three words late and ends mid-word. Glyphcast grows a clip outward from the scored peak and then snaps each edge to a pause in the transcript, so the clip opens on the start of a thought and closes after it lands.

You can still move either edge. Trimming a clip re-runs it through the same pipeline rather than cropping the finished file, so the captions re-time to the new cut instead of drifting out of sync with it.

Video podcasts and audio-only episodes

A video podcast is the straightforward case: the source has a picture, and the 9:16 output keeps the whole frame with a blurred copy of itself filling the bars above and below. Nothing is cropped, so a two-person shot stays a two-person shot.

An audio-only episode has no picture to work with. Glyphcast is built for video sources, and an audio file with a static cover image will produce correctly cut, correctly captioned clips of a still image - which is a real format on some platforms, but it is worth knowing that is what you will get rather than discovering it afterwards.

How it works, step by step

  1. 01Upload the episodeUpload the episode. A video podcast is the straightforward case; an audio-only file with a static cover works, with the caveat below.
  2. 02Let it read the episodeSpeech is transcribed with word-level timings, and every moment is scored on what was said, how loud it got, and how much movement there was.
  3. 03Review the proposed clipsEach clip arrives cut to a natural pause, reframed to 9:16 and captioned. Rate them, trim the edges, or delete the ones that missed.
  4. 04DownloadDownload clips individually or take the whole set as a ZIP. They are ordinary MP4 files.

Questions

How long can a podcast episode be?

The ceiling depends on your plan, and it is listed on the pricing page. Longer episodes take longer to process - the transcription pass is the slow part - but length does not change how the moments are chosen.

Will it find moments in a two-person conversation?

Yes, and conversation is the material it reads best: interruptions, laughter and changes in pace all register as signal. What it does not do is crop to whoever is speaking - the 9:16 output keeps the whole frame rather than tracking a face, so both people stay in shot throughout.

Do the clips have captions?

Yes, burned into the picture and generated from the real transcript rather than from a second pass over the clip, so the words match what is actually said and land on the right frame.