A 16:9 frame is 1.78 times wider than it is tall. A 9:16 frame is 1.78 times taller than it is wide. Getting from one to the other means losing something, and every tool that does it automatically has quietly picked which loss you get.
There are only two answers. Both are defensible. Neither is free, and the marketing copy for these tools almost never says which one it implements.
Crop: fill the frame, lose the edges
Take a vertical slice out of the middle of the landscape frame, scale it up to fill 1080×1920, discard everything either side. The result uses the whole screen, looks native to the platform, and is what most people picture when they imagine a vertical clip.
What it costs is about 44% of the picture. On a centred talking head, that 44% is wall, and losing it is an improvement - the subject gets bigger. This is the case the technique is built for and the case every demo reel shows.
It stops working the moment the frame is composed. Two people in a two-shot: one of them is gone, or both are half-gone. A presenter beside their slides: you get the presenter or the slide, never both. A product on a table, a lower third, text across the bottom of a wide shot, a whiteboard - all outside the slice, all deleted, with nothing in the output to indicate anything was there.
Speaker tracking is the usual mitigation: follow whoever is talking, so the crop is at least aimed at the right part. It genuinely helps for single speakers. In a conversation it introduces a new failure - the crop jumps to the person who just stopped talking, exactly at the interruption, which is the moment you cut the clip for.
Letterbox: keep everything, use less screen
Put the entire landscape frame at full width in the middle of the vertical canvas, and fill the bars above and below. Usually with a blurred, darkened copy of the frame itself, so the result reads as one image rather than as a video with black bars.
Nothing is cropped. The two-shot stays a two-shot, the slide stays beside the presenter, the text at the bottom of the frame is still there.
What it costs is size. The picture occupies the middle third or so of the screen rather than all of it, which means everything in it is physically smaller on a phone. A face is fine. A headline on a slide is fine. A code listing, a spreadsheet, a dense chart - not fine, and no strategy would have saved those, because the alternative was cropping most of them off entirely.
The other cost is aesthetic, and it is real: a full-bleed vertical video looks more made for the platform than a letterboxed one. Some of that is fashion. Some of it is a genuine signal to a viewer that this was shot vertically rather than harvested from something else.
How to tell which one your footage needs
The question is not which technique is better. It is whether your composition survives losing the outer 44%.
Crop is fine when there is one subject, they stay near the middle, and nothing important lives at the edges - a single-camera talking head, a vlog, a piece to camera.
Letterbox is better when the frame contains more than one thing that matters: two or more people, a screen share, slides, on-screen text, a demo of a physical object, or any footage where the framing was a decision.
A useful test: pause the source video at a few points, cover the left and right thirds with your hands, and ask whether the remaining strip still tells the story. If it does, crop. If you keep uncovering an edge to check something, letterbox.
What Glyphcast does
Letterbox. The whole frame, full width, with a blurred copy of itself filling the bars - and therefore no face tracking, because tracking exists only to decide where to crop and nothing is being cropped.
That is a choice rather than a limitation, and it follows from what the tool is pointed at: long videos of people talking. Podcasts, interviews, webinars and conference talks are the formats where a centre crop is most likely to delete the second person or the slide. The case where cropping clearly wins - one person, centred, nothing at the edges - is also the case where letterboxing costs you the least.
More on the rest of the pipeline in how it works, and what this means per kind of source in use cases.