Use casesInterviews

Interview clips, chosen from what was said

An interview is a long search for about five good answers. You know they are in there, because you were in the room when they happened; what you do not have is the hour it takes to find them again.

The transcript does most of the work

In an interview, what makes a moment worth clipping is almost entirely what was said. Motion tells you nothing - two people are sitting still - and audio energy only tells you when someone got animated, which correlates with a good answer but does not identify one.

So the scoring leans on the transcript, and the AI pass reads the actual words of the candidate moments rather than looking at thumbnails of them. A moment is proposed because of what is in it, not because something moved.

A clip that starts mid-answer is a wasted clip

This matters more in an interview than anywhere else. A highlight from a talk can survive starting a beat late. An answer cannot: the first six words are usually the ones that make the rest of it make sense, and a clip that opens on the second sentence reads as a fragment of an argument nobody heard the start of.

Clip edges are snapped to pauses in the transcript, so a clip begins where the answer begins.

Both people stay in frame

The vertical output keeps the entire original frame rather than cropping into it, which for a two-shot means both people remain visible for the whole clip. There is no speaker tracking deciding whose face to follow, and so no moment where it follows the wrong one - which in an interview is the failure that shows, because it happens exactly when someone interrupts and the cut jumps to the person who stopped talking.

How it works, step by step

  1. 01Upload the recordingUpload the recording. One frame with both people in it is the usual case and the one the reframe is built for.
  2. 02It transcribes and readsSpeech is transcribed with word-level timings, and the AI scoring pass reads the transcript of each candidate moment alongside its keyframes.
  3. 03Answers become clipsEach proposed clip is grown around a scored peak and snapped to the pauses either side of it, then captioned and reframed to 9:16.
  4. 04Pick the ones you would have pickedRate, trim or delete. Download what survives, individually or as a ZIP.

Questions

Does it label who is speaking?

No. There is no speaker diarisation, so captions are not attributed by name. For a two-person interview shot in one frame this rarely matters, because you can see who is talking.

Will it cut around a question, or just the answer?

It scores moments rather than recognising question-and-answer structure, so a clip may contain the question, the answer, or both depending on where the scoring peaks and where the pauses fall. Trimming lets you include or exclude the question deliberately.

How accurate are the captions on names and jargon?

Speech recognition is weakest exactly there, and a name it gets wrong once it usually gets wrong throughout. Trimming a clip re-runs the transcription over the new boundaries rather than shifting the old text, so a bad cut is at least not compounded by captions drifting out of sync with it.