What to compare, and why it matters
Every tool in this category says an AI finds the best moments. These are the six decisions underneath that sentence where they genuinely differ - what each choice costs, how to test it yourself, and where Glyphcast lands, including where that is a limitation.
Does it crop the frame, or keep it?
Turning a 16:9 video into 9:16 means losing something. A tool either crops a vertical slice out of the middle and scales it up - filling the screen, discarding about 44% of the picture - or it keeps the whole frame at full width and fills the bars above and below.
Neither is wrong, and the right answer depends entirely on your footage. One person centred in frame: crop, and they get bigger. Two people, a screen share, slides, or anything with text across a wide shot: cropping deletes half of what the shot was about, and the output gives no sign that anything was ever there.
This is the single biggest difference between tools in this category and the one least often stated on a pricing page. Ask it first.
Glyphcast: Glyphcast keeps the whole frame, over a blurred copy of itself. Nothing is cropped, and there is no face tracking - tracking exists only to decide where to crop.
Where does it put the start and end of a clip?
A scoring model finds a moment, which is a point in time. Turning that into a clip means choosing two boundaries, and this step is where most automatically cut clips are ruined.
The cheap approach is fixed padding - take the peak, go back fifteen seconds, forward fifteen seconds. It produces clips that open mid-sentence, which no amount of good scoring can rescue. The better approach uses the transcript's word-level timings to move each edge to a real pause in speech.
It is easy to test: run a dense-speech video through and look only at the first word of each clip.
Glyphcast: Clip edges snap to pauses in the transcript. Trimming re-runs the pipeline rather than cropping the rendered file, so captions re-time instead of drifting.
Where do the captions come from?
Captions can be generated from the transcript the tool already produced for scoring, or from a second speech-recognition pass over the finished clip. The first is more accurate and correctly timed; the second re-does work and can disagree with the transcript the clip was chosen from.
The second question is what happens when you trim a clip. A tool that crops the rendered file leaves the captions where they were, so they drift out of sync with the new boundaries. One that re-runs the cut re-times them.
Glyphcast: Captions come from the same transcript used for scoring, burned in with word-level highlighting. Trimming a clip re-runs the pipeline, so they re-time rather than drift.
What is it actually scoring?
"AI picks the best moments" describes every product in this category and distinguishes none of them. The useful question is what the model is shown.
A model shown only keyframes is judging whether a frame looks interesting, which is a poor proxy - the best moment in an interview looks identical to the worst one. A model shown the transcript is judging what was said. Some tools add signals the video itself cannot provide, such as which parts of a published video viewers actually rewatched.
Glyphcast: Scene cuts, audio energy, motion and a local word-level transcript, with the AI pass seeing both keyframes and the real transcript rather than thumbnails alone.
What happens to your video?
Two separate questions, and they are often answered together in a way that obscures both. First: is your content used to train models? Second: how long is it kept, and what is deleted when you cancel?
Where the speech recognition runs matters too. A tool that transcribes locally never sends your audio to a third-party speech API; one that does has added a processor to the list of companies holding your material, which may or may not be disclosed on the page you are reading.
Glyphcast: Videos are never used to train anything, and are not sold or shared. Speech-to-text runs locally on the processing machine. Finished clips are deleted after 90 days; source videos stay until you delete them.
What does the free plan actually run?
A free plan that switches off the scoring model is not a trial of the product - it is a trial of a different, worse product, and it tells you nothing about whether the real one would work on your footage.
Check which of the limits is volume (fewer videos, shorter sources, fewer clips) and which is capability (no AI scoring, no captions, lower resolution). Volume limits are a fair trade. Capability limits mean the free tier cannot answer the question you are using it to answer.
Glyphcast: The free plan runs the same pipeline with the same AI scoring. What it limits is volume, and its exports carry a watermark.
Why there is no comparison table here
A feature matrix published by one of the products in it is an advertisement, and it requires stating facts about other companies' software that go out of date quietly - competitors change pricing and features continuously, and a table written once is wrong within months in whichever direction flatters whoever wrote it.
The questions above are more useful than a verdict would be, because you can take them to any product and get the answers from the source. If one of them cannot be answered from a tool's own documentation, that is itself an answer.