← Home
WORKFLOW GUIDE

By Valmera Editorial · Published · Updated

Edit Interview Video by Speaker with AI

Valmera edits interviews from a transcript that knows who is speaking. Upload a two-person interview and describe the edit in plain English: keep the guest's answers, drop repeated questions, frame whoever is talking in a vertical cut, add captions. The agent makes the cuts, renders a preview it checks itself, and you revise by chatting.

THE OUTCOME

A tight main interview, speaker-framed vertical clips, and a version history of every change. Your original file is never modified, so any cut can be restored.

How Editing by Speaker Works

When you upload, Valmera transcribes the interview word by word and labels each line by speaker: S0, S1 and so on. Tell the agent who is who once (“S0 is the host, Maya; S1 is the guest, Dev”) and you can edit by person instead of by timestamp. Speaker labels come from the audio, so check the joins where people talk over each other.

Keep or cut by speaker

“Remove the host's questions unless the answer needs them” or “cut everything S1 says before minute three.” The agent finds the lines in the transcript and cuts on word boundaries.

Build the story from complete answers

Give the structure and target length, and the agent keeps the answers that serve it and cuts the rest, in recorded order. Ask it to keep answers whole rather than splicing words into new sentences.

Tighten without flattening

Dead air, ums, false starts and repeated takes come out. Laughter, breaths and pauses that carry meaning stay when you say so, and every cut can be restored.

Related tools: Text-based video editing; Remove repeated takes; Remove silence; Transcript and timeline docs.

Two-Person and Two-Camera Setups

Framing is chosen section by section: in a vertical cut the crop moves to the next speaker at the turn rather than continuously following movement. Pick the setup that matches your recording.

One wide camera

The agent aims a 9:16 crop at whoever is speaking in each section, using the speaker turns and the faces it finds in the frame. For reactions, it keeps the full two-shot inside the vertical frame with room for a headline and captions.

Side-by-side recording

A Zoom, Riverside or StreamYard layout puts both people in one frame. The agent crops to the half that is talking, section by section, so each person fills the vertical frame on their own answers.

Two cameras

Upload camera A as the main video and camera B as a clip, and give the sync point, such as a clap at 0:04 on A and 0:07 on B. The agent can cut to camera B for the sections you choose while the main audio keeps playing, or build a side-by-side panel of up to 20 seconds.

Framing guides: Auto reframe; Reframing docs; Multi-speaker video podcasts.

What Makes This Editing Job Difficult

Meaning can change at a cut

Removing a qualifier, setup or follow-up can make a quote stronger and less true. Every structural cut is an editorial decision.

Two people, one vertical frame

A 9:16 crop holds one face comfortably. The edit has to decide who is on screen in each section, and when a reaction matters more than the answer.

The conversation isn't the story

Questions arrive in recording order. The finished piece needs a hook, an argument and a close, which means choosing the answers that build it and cutting everything else.

What to Prepare Before You Upload

Most of the editorial context lives in your head. Give the agent the names, the story you want and the rules it must not break.

  • The interview recording (MP4, MOV, M4V, MKV or WebM, up to 14 GB and 3 hours) and any second-camera file, with a sync point if you have one.
  • Who is who (“S0 is the host, Maya Chen; S1 is the guest, Dev Patel”), plus spellings of names, companies and terms.
  • Target length, audience and where the cut will be published.
  • Must-keep quotes, anything off the record, and claims that need exact wording.
  • Caption style, logo, b-roll and music you want to use.

How to Auto-Edit a Two-Person Interview in Valmera

Each step is one message in the chat. The agent renders a preview after every change, so approve the story before you spend time on polish.

  1. 1

    Upload and name the speakers

    Upload the interview. When indexing finishes, check the transcript, tell the agent who S0 and S1 are, and correct any misheard names.

  2. 2

    Ask for the story cut

    Describe the structure and target length. The agent selects complete answers, keeps the questions that give them context, and shows you the cut before adding polish.

  3. 3

    Clean up the dialogue

    Remove dead air, ums, false starts and repeated takes. Protect pauses, laughter and qualifiers that change meaning, then listen across the joins.

  4. 4

    Frame the speakers

    Keep 16:9 for the main cut, or ask for 9:16 framed on whoever is speaking, section by section. Use the full two-shot for reactions. With two cameras, give the sync point and the sections for camera B.

  5. 5

    Add captions, names and sound

    Add word-timed captions and a name title on each person's first appearance. Duck any music under speech and master the mix to -14 LUFS.

  6. 6

    Make clips and export

    Turn complete answers into vertical shorts, review each preview, then export full-resolution MP4s.

Prompts You Can Copy

These are production briefs, not magic words. Replace bracketed details with the real audience, duration, platform, claims, and assets for your project.

Edit by speaker

S0 is the host, Maya; S1 is the guest, Dev. Build a [6–8 minute] cut from Dev's answers about [topic]. Keep Maya's questions only where an answer needs them; otherwise show the question as a short on-screen title. Don't combine words from separate answers. Show me the structure first.

Two-person cleanup

Remove pauses longer than [0.7 seconds], ums, false starts and repeated questions. Keep laughter, the pause before Dev's answer about [topic], and every qualifier. Flag any cut that changes what someone meant.

Speaker-framed vertical cut

Make a 9:16 version that frames whoever is speaking and switches at each speaker turn. When they talk over each other or laugh, show the full two-shot with captions below. Use [podcast] captions.

Two-camera cutaways

Camera B is Dev's close-up. The clap is at 0:04 on the main camera and 0:07 on camera B. Cut to camera B during Dev's answers from 08:10 to 11:30, keep the main audio, and add “Dev Patel, Founder” on his first close-up.

One context-safe clip

From [source range], make one 45–60 second vertical clip. Start on the question that sets up the answer, keep [essential qualifier], and end after the payoff. Accurate captions, no music.

Limits

Speaker labels come from transcription and can slip when voices overlap or sound alike; English is the best-tested language. The main recording plays in recorded order: to open on a later moment, upload that moment as a separate clip. Framing is chosen per section, not tracked continuously, and Valmera doesn't sync cameras or switch angles on its own. It doesn't remove background noise or balance two voices recorded into one track, captions are burned in (no SRT or VTT files), and you post the finished MP4 yourself. There are no team seats yet.

Learn more: Valmera docs; Automate with MCP; How the AI video editor works.

Edit the Interview by Speaker

Upload the interview, name the speakers, and ask for the cut you want.

Direct the first edit →
See pricing →

The Review Gate: What a Human Must Check

Quote integrity

Compare every shortened answer with the source. Certainty, chronology and conditions must survive the cut.

Who is on screen

In vertical cuts, check each speaker turn: the right face, no one cut in half, and reactions shown where they matter.

Speech and captions

Listen on headphones and a phone speaker. Check names, numbers and terms in the captions against the final audio.

Deliverables to Request

  • Main interview MP4 in 16:9 with captions and name titles.
  • Speaker-framed 9:16 clips, each built around one complete answer.
  • A caption-free version if your platform adds its own captions.
  • A short teaser cut from the approved interview.

When an Agentic Editor Fits—and When It Does Not

A strong fit

  • Speech-led interviews with clear audio and identifiable speakers.
  • Teams that know the story they want and the quotes that must survive.
  • Shows that need both a long cut and vertical clips from one recording.

Use another workflow

  • Multicam productions that need automatic sync and angle switching.
  • Recordings that need noise repair before anyone can understand them.
  • Investigative or legally sensitive work without a human editor reviewing every cut.

The Operating Principle

Decide what the conversation says first, then how viewers see and hear it. A transcript with speaker labels makes the first decision fast: you edit people and answers, not timestamps. The agent handles the cutting, framing and captions; you own the meaning.

New to the model? Read the complete guide to AI video editing or start with production-ready editing prompts.

Frequently Asked Questions

Yes. Valmera labels each transcript line by speaker (S0, S1, …). Tell the agent who is who, then ask in plain English: keep only the guest's answers, cut the host's restarts, or drop every answer that's off topic. It cuts on word boundaries and every cut can be restored.
Upload the recording, name the speakers, and describe the cut: target length, structure, what to remove and what to protect. The agent builds the cut, renders a preview and checks it. Revise by chatting, then add captions, framing and export.
Yes, section by section. In a vertical cut the agent aims the crop at whoever is speaking and switches at speaker turns, or shows the full two-shot when both reactions matter. It does not track movement continuously within a section.
No. Upload one camera as the main video and the other as a clip, and give the sync point (for example, the clap time on each). The agent can then cut to the second camera for the sections you choose or build a short side-by-side panel.
No. Valmera has no noise removal and can't balance two voices inside one mixed track. It can raise or lower the original audio over a time range and master the final mix to -14 LUFS. Repair noisy audio before upload.
No. Captions are burned into the MP4, and there is no SRT or VTT import or export. If a platform needs a separate caption file, create it from the finished edit in another tool.
It depends on the job. Descript suits editing the transcript by hand and offers automatic multicam switching. Riverside suits recording remote interviews and clipping them in one place. Valmera suits directing the whole edit in plain English, by speaker, and getting both the long cut and vertical clips.

Direct the Interview Edit

Give the agent the footage, the speakers and the story. Review a rendered cut instead of starting from an empty timeline.

Edit this workflow with Valmera →
See pricing →

Related Articles

AI Podcast Editor and Clip Maker
Multi-speaker episodes, speaker framing and shorts.
Edit a Podcast with AI
A step-by-step conversation-editing workflow.
Auto Reframe
Vertical framing for speakers and screens.
Remove Repeated Takes
Keep the best version of each line.