← Home
FEATURE

Timeline, Transcript & Player

Agent-first, hands-on when you want it. The chat does the work — these are your direct controls.

Agent-First, Hands-On When You Want It

In Valmera, the AI agent is the editor: you describe the edit in plain English and it makes the cuts, adds the captions, and mixes the audio. But you are never locked out of the edit. The Studio gives you three direct surfaces — a visual program timeline, an editable transcript, and a custom player — so you can see exactly what the agent did and adjust it by hand.

The rule of thumb: describe intent in chat ("tighten the intro, add captions"), use your hands for placement ("this music block starts two seconds earlier"). Both write to the same edit, so you can switch freely.

The Program Timeline

The timeline shows the video the agent actually built — not the raw upload. Kept footage appears as blocks, with notches marking every cut the agent made, so a "remove the silences" request becomes something you can see at a glance.

Kept-footage blocks with cut notches

Each block is a stretch of footage that survived the edit. Notches show where cuts landed — silences, filler words, and repeated takes the agent removed.

Draggable insert, music, and voiceover blocks

Extra clips, images, music tracks, and voiceover all appear as blocks you can drag to reposition. Nudge a music cue or move an insert without typing a word.

Drop a file to insert it

Drag a clip or image from your computer straight onto the timeline and it is inserted at that point — clips up to 500MB, images up to 10MB.

Playhead scrub

Drag the playhead to jump anywhere in the edit instantly. Scrubbing is the fastest way to verify a cut landed exactly where you wanted.

Canvas Projects: Start Without a Main Video

A project does not have to begin with a video. In a canvas project you start with an empty canvas in the aspect ratio you choose, and every asset you add — images, clips, generated media, music — becomes a block on a sequential timeline. It is how you build a photo slideshow, an image-plus-voiceover video, or a short assembled entirely from pieces you bring or generate inside the edit.

Any asset can be the first one

Drop in an image, a clip, or a track and the canvas builds around it. Assets play in sequence in the order they sit on the timeline, and you can drag them to reorder.

Any aspect ratio

Pick 16:9, 9:16, 1:1, or 4:5 for the canvas itself, so a slideshow for Reels is portrait from the first frame rather than cropped later.

Text, overlays, music and motion still apply

On-screen text templates, image and video overlays, Ken Burns motion on stills, background music with ducking, and sound effects all work exactly as they do on a normal project.

What a canvas project cannot do, because there is no main video to analyze:

  • No transcript captions — there is no speech transcript to generate them from. Use on-screen text templates instead.
  • No speed changes on canvas projects.
  • No censor regions (blur, pixelate, black-out) and no reframe of an existing edit — you choose the canvas ratio up front.
  • It is a sequential timeline, not a freeform multi-video grid. Two videos on screen at once means one of them is a picture-in-picture overlay.

The Editable Transcript

When you upload footage, Valmera builds a word-level transcript of everything spoken. That transcript powers the agent's speech edits — silence removal, filler-word cuts, repeated-take detection — and it drives the auto captions.

It is also yours to correct. If a name or technical term was misheard, click the sentence in the transcript panel, fix the word, and the captions re-render with your correction. One wrong word never means redoing the edit.

The Player and Ratio Pills

The custom player previews the current edit, and the ratio pills above it switch the frame between 16:9, 9:16, 1:1, and 4:5 — YouTube, Shorts and Reels, square feed, and portrait feed. Reframing works by crop, pad, or blurred-pad; there is no automatic subject tracking, so for tight vertical crops tell the agent what to keep in frame.

Previews render from a fast proxy so iteration stays quick even on long uploads (main videos can be up to 2GB — see Uploading Footage). The final export always renders from the original full-quality file.

Chat or Hands? When to Use Which

  • Use chat for judgment calls. Pacing, what to cut, caption style, effects, and audio mixing — anything where the agent's read of the footage does the heavy lifting.
  • Use the timeline for placement. Dragging a block two seconds left is faster than describing it. Drop files where they belong instead of narrating positions.
  • Use the transcript for words. Fixing a misheard term in the transcript beats asking the agent to guess which caption was wrong.
  • Trust the replies. Valmera's honesty layer checks every agent reply against what the system actually did — if the reply says a cut was made, it was.

Common Questions

Do I have to use the timeline to edit in Valmera?

No. Every edit in Valmera can be done entirely through chat — cuts, captions, music, reframes, effects. The timeline is an optional layer of direct control: it shows you exactly what the agent built, and lets you drag inserts, music, and voiceover blocks by hand when a request is easier to do than to describe.

Can I fix a misheard word in the captions?

Yes. The transcript panel is editable — click the sentence, correct the word, and the captions re-render with your fix. You never have to re-run the whole edit or re-upload anything to repair one misrecognized name or term.

Can I add a clip by dropping it onto the timeline?

Yes. Drop a video file or an image directly onto the timeline and it becomes an insert block at that point. Insert clips can be up to 500MB and images up to 10MB, and once placed, the block stays draggable so you can fine-tune the position.

Can I make a video without uploading a main video?

Yes — that is a canvas project. You pick an aspect ratio and build a sequential timeline out of images, clips, generated media, and music, which is how slideshows and image-plus-voiceover videos get made. Because there is no main video to analyze, canvas projects have no transcript captions, no speed changes, and no censor regions or reframing; on-screen text, overlays, Ken Burns motion, music, and sound effects all work.

Which aspect ratios does Valmera support?

16:9, 9:16, 1:1, and 4:5 — switchable from the ratio pills above the player. Reframing works by crop, pad, or blurred-pad; Valmera does not do automatic subject tracking, so for tight crops tell the agent what to keep in frame.

Is the preview the same quality as the export?

No, by design. Previews render from a fast proxy so you can iterate quickly, while the final export renders from your original full-quality file — up to 2GB uploads are supported. What you check in the player is the same edit; the export is simply full resolution.

Next Steps