Make a Highlight Reel from a Long Video
A two-hour recording contains about three minutes worth watching. The hard part was never the cutting — it was finding them.
To make a highlight reel from a long video in Valmera, upload the recording and type what the reel is for and how long it should be: "make a 3-minute highlight reel of the best moments from this talk". The AI agent works from three sources at once — a word-level transcript with speaker labels, the measured energy and vocal stress of the audio, and labeled frames sampled across the whole file — picks the moments, assembles them into one piece, renders a preview, and looks at the frames it produced. You then correct it in plain English, because the first pass is a proposal, not a verdict.
Build a Highlight Reel Free
Upload the long recording, say what the reel is for, review the preview. 50 free credits, no credit card.
Start free →How to Make a Highlight Reel from a Long Video
- 1Upload the full recordingDrag in the whole thing — MP4, MOV, MKV or WebM, up to 3 hours. Do not pre-trim it: the parts you would have cut by hand are exactly the parts the selection needs to see in order to rank what is left. Valmera runs a one-time analysis that builds the word-level transcript, detects silences and shot changes, and samples labeled frames across the entire file.
- 2Say what the reel is FOR, and how longType it literally: "make a 3-minute highlight reel of the best moments from this talk, for people who weren't there". Purpose and audience change the picks more than any other instruction — a recap wants payoffs and conclusions, a trailer wants unanswered questions. Give a duration; without one, "highlight reel" has no length.
- 3Argue with the first passWatch the preview and correct it by name and timecode: "the demo at 22:10 is the whole point, give it 40 seconds", "drop the third moment, it's the same idea as the first", "nothing from the last twenty minutes made it — check again". Each correction re-edits the existing reel rather than starting over.
- 4Finish it, then exportAsk for a title card, music ducked under speech, word-accurate captions, fades and one transition style across the cuts, and a 9:16 version if you need one. The export renders from your original file at source quality — the preview proxy is only for iterating.
The analysis runs once per upload, with visible progress; on a multi-hour recording it takes a while. Every request after that is a chat message.
- "make a 3-minute highlight reel of the best moments from this talk"
- "cut a 90-second recap of this 2-hour stream for people who missed it"
- "build a 45-second trailer for this episode — hook first, no spoilers"
- "the demo at 22:10 is the whole point, give it 40 seconds"
- "drop the second moment and find something from the last 20 minutes"
- "add a title card, put music under it and duck it against my voice"
Three Signals, and What Each One Alone Would Miss
The transcript is the cheapest signal to score and the easiest to rank, so it is the one automatic selection usually leans on — and on its own it is blind in a specific, expensive way. Here is the same forty-two minutes read three ways, and the four moments where the readings disagree.
How the Moments Are Actually Found
What the transcript tells you
Indexing produces a word-level transcript with a start and end time on every word, plus speaker labels when more than one person talks. That is the signal that answers what was said and by whom: the claim, the number, the punchline, the question that got a real answer. It is also the only signal that supports a request like "find the part where I talk about pricing" — you are searching language, and language is indexed.
Word timings matter more than the text here. A moment is a range, and a range chosen from a word-level transcript can start on the first syllable of the sentence rather than four hundred milliseconds into it. Cuts can be snapped to those word boundaries, and when a boundary does land inside a word the system flags it rather than letting it through — which is what keeps the joins from clipping the front of a sentence or leaving a stranded half-word behind.
What the audio tells you
The audio is measured, not guessed at. Analysis produces an energy envelope in half-second bins across the whole recording — where the loudest point is, where the quietest is, and the biggest sustained rise, expressed in decibels relative to peak. A rise of eight decibels climbing into a moment is a real thing that happened in the room, and it happens whether or not anybody said a quotable sentence.
Separately, every word carries a measured vocal stress score: the peak of the speech-band envelope during that word, relative to a rolling ceiling over roughly the surrounding ten seconds. The rolling window is the part that matters. Measured against the whole file, a speaker who warms up over an hour has every late word scored as emphatic and every early one as flat, which is a measurement of the microphone gain rather than the speaker. Measured locally, emphasis means louder than the person was thirty seconds ago — which is what emphasis actually is.
What the frames tell you
This is the signal transcript-only tools do not have, and the reason they miss what they miss. The index samples frames across the entire file and lays them out as labeled tiles — a 2×2 grid of stills per tile with the source timestamp printed under each frame, up to 36 tiles, so 144 stills spanning the whole recording. Those tiles are attached to the agent every single turn. It is not a description of the footage written by another model and then read as text; it is the frames themselves, and the agent reads what is on screen the way an editor does — by having looked.
On top of the standing strip, the agent can call for exact frames at any timestamp it wants, as often as it wants, including frames of the assembled reel rather than the source. That is how it checks a boundary before committing to it: the transcript says the sentence ends at 22:41.8, the frame at 22:41.8 shows him mid-blink with his hand across the lens, so the cut moves. Shot detection supplies the other half — the points where the picture actually changes, which are the natural seams a moment should start and end on.
Every reply the agent sends about the reel is checked server-side against the edits actually recorded, so it cannot tell you it kept a moment it did not keep. When you are not going to watch forty-two minutes to audit a two-minute cut, that verification is the part that makes the reel usable.
Length, Pacing, and Whether to Keep Chronology
Selection is only half the job. Six correctly chosen moments cut back to back with no thought about order or rhythm still play as a list, not a reel.
Set length by purpose, not by source
A 3-hour stream and a 40-minute talk usually want the same-length reel, because the limit is the viewer's patience rather than the footage. A trailer runs 30 to 60 seconds. A recap for people who missed the thing runs 90 seconds to 3 minutes. A sizzle reel that has to stand alone as content runs 2 to 5 minutes. Say the number in the request — without one, "highlight reel" has no defined length and you will get an arbitrary one.
Vary the moment lengths
Reels that feel mechanical are usually reels where every moment is the same duration. Real ones breathe: a 40-second passage that earns its length, then three 8-second beats, then a 6-second button at the end. Ask for it directly — "give the demo 40 seconds and keep the rest to under ten each" — because a duration target alone tends to divide evenly, and evenly is the one rhythm no editor would choose.
Chronology is a decision, not a default
A recap should stay chronological: it is an account of what happened, and reordering it misrepresents the event. A trailer should not: it opens on the strongest few seconds no matter where they occurred, because its only job is to make someone watch the full thing. Valmera assembles a reel from the main footage in the recording's own order — that is the right default for recaps, event summaries and conference edits. If you want a later moment first, bring it in as an inserted clip rather than reordering the main timeline; see the inserted-clip workflow.
Finish the joins
Between-moment cuts are visible in a way that within-a-take cuts are not — the setting changes, the framing changes, sometimes the lighting changes. One transition style across all the cuts, a title card, a fade in and out, and music ducked under the speech will do more for how finished the reel feels than another hour of moment-hunting. If the reel is music-led rather than speech-led, snapping the cuts onto the measured beat grid is one request.
Highlight Reel, Clip, Trailer, Recap — Four Different Jobs
These get treated as one thing and are not. Asking for the wrong one is the most common reason a first pass disappoints.
If what you actually want is a steady output of vertical shorts from every episode, that is a different tool category with different tradeoffs — read what a clipping agent does, including how virality scores are really produced.
When It Goes Wrong
Automatic selection fails in patterns, and the patterns are recognisable enough to fix in one follow-up message.
Everything comes from the first twenty minutes
Long recordings frequently front-load their strongest language, and any ranker inherits that bias. Say so directly: "nothing from the last twenty minutes made it — look again after 1:20:00". Worth knowing precisely: the vocal-stress measurement covers the first hour of audio, and words past that point carry a sentinel score of zero rather than a fabricated one. That is deliberate — a made-up score would rank moments on audio nobody measured — but it does mean that on a two-hour recording you should ask explicitly about the second half rather than assume it was weighed.
The moment starts mid-sentence
A moment that begins on "...and that's why we stopped doing it" needs the sentence before it to mean anything. Ask for the run-up: "start that one about eight seconds earlier so the setup is in". The opposite problem — a moment that keeps rolling past its own ending — is fixed the same way.
It picked the loudest, not the best
Energy and stress are strong signals and they are also how you end up with a reel of someone shouting. Correct it toward substance: "pick for what he actually explains, not for volume — I want the two moments where the argument changes".
Two moments make the same point
Redundancy is the failure mode nobody anticipates: both moments are genuinely good, and together they are worse than either alone. "Three and five say the same thing, keep the better one and find something different" is the whole fix.
The frames are wrong even though the words are right
Somebody walks through the shot, the slide has not advanced yet, the screen share is still grey. Tell the agent to look: "check the frames at 31:40, I think the slide is still the old one". It can pull exact frames at any timestamp, and this is exactly what that is for. Nothing is destructive either way — the original upload is never modified, so any moment you disagree about can be restored at full quality.
How People Do This Without Valmera
All of these work, and some of them are the right answer depending on what you have and what you need.
Premiere Pro or DaVinci Resolve, by hand
The classic professional method is a marker pass and a string-out: watch the recording at 1.5× or 2×, drop a marker every time something happens, then pull each marked range onto a timeline as a rough assembly and cut it down. Both applications transcribe footage and support editing from the transcript, and Premiere Pro's Scene Edit Detection will mark the cuts in an already-flattened file for you, so the mechanical parts are faster than they were. What you do not get from either is a ranked "here are the moments worth keeping" pass — that judgement stays yours, and on a two-hour recording the marker pass alone is the better part of an afternoon. Third-party plugins and scripts exist that automate parts of it.
ffmpeg and a script
Entirely doable and genuinely instructive. silencedetect gives you the pauses, ebur128 or astats gives you a loudness curve, the scdet filter or a select expression on the scene score gives you shot changes, and Whisper — with WhisperX or a word-timestamp build if you want per-word times — gives you the transcript. Then you write the scoring function yourself, produce a list of ranges, and assemble with the concat demuxer. The cost is not the ffmpeg commands — those are a weekend — it is that the scoring function is the actual product, and tuning it against your own footage is an open-ended project. It is also the only approach on this list that costs nothing per video and runs entirely on your own machine.
Automatic clipping tools
Opus Clip, Vizard, Klap and similar tools take a long video, pick out the moments they rate highest, and return a batch of vertical clips with captions — several of them also attach some form of predicted-performance score. They are optimised for volume: if your need is ten shorts a week from a weekly podcast, they are built for exactly that and will beat a manual process comfortably. What they return is many candidates rather than one considered piece, and a batch of clips is not the same deliverable as a single edited reel with an arc. Different job, honestly done.
Your own viewers, if the video is already public
If the recording is already on YouTube with meaningful watch time, the most-replayed graph is real audience data about which seconds people went back to, which is better evidence than any model's guess. It only exists after publication, which makes it useless for the first edit and excellent for the second.
Honest Limits
- The reel plays in the recording's own order. Moments selected from the main footage stay chronological. Opening on a later moment means bringing it in as an inserted clip, not reordering the main timeline.
- One deliverable per request. There is no batch mode that returns ten candidate reels with scores. If you want three different reels, that is three requests, and each one gets a full pass.
- No virality score. Nothing here predicts how a reel will perform. Any number that claimed to would be invented, and inventing it is exactly what the honesty layer exists to prevent.
- Vocal-stress analysis covers the first hour. Uploads run to 3 hours, and words past the first hour score zero rather than a guess. Ask explicitly about the later stretches of a long recording.
- The standing frame strip is coarse on long footage. 144 stills across three hours is roughly one frame every 75 seconds. The agent pulls exact frames on demand to close the gap, but it has to be pointed at a region to do that — which is what your corrections are for.
- Captions are burned in. There is no SRT or VTT import or export, so the reel's captions live in the pixels.
- One transition style across all cuts, and no true crossfade or dissolve. You choose a style for the reel, not per join.
- English is the best-tested transcript path, and no direct publishing to YouTube or TikTok — you download the export and post it yourself.
The Rest of the Reel
Selection is the hard part, but the finish is the same conversation. Add word-accurate captions, a title card and lower thirds, punch-ins on the most emphasized lines, a grade, and a 9:16 or 1:1 version for vertical placements. Within a single moment, cutting the dead air and the filler words tightens it further without changing what it says.
If you would rather drive the selection yourself by reading rather than watching, text-based editing and the editable transcript and timeline are both there. Everything on this page is also available to Claude directly over the Valmera MCP server, which exposes the same tool registry the studio agent uses.
Find the Three Minutes Worth Watching
Upload the whole recording. Say what the reel is for. Argue with the first pass.
Start free →Frequently Asked Questions
Make Your Highlight Reel
The agentic AI video editor — it reads the words, measures the audio, looks at the frames, and shows you its picks. 50 free credits, no credit card.
Start editing free →