← Home
USE CASE

Published · Updated

AI Video Editing for SaaS Product Demos

A demo video is the shortest path between "what does this do" and "I want it". It is also the piece of marketing most likely to sit unmade, because making one properly means recording the product, cutting it, redacting a customer name, enlarging a cursor nobody can follow, and doing it again vertically.

Valmera does that as a conversation. You point it at a URL or upload a screen recording, and an agent records the product, splices the capture in, pushes the frame in on every click, redacts what should not ship, enlarges the cursor, wraps the picture in a floating window, and cuts a vertical version from the same file. This page is the playbook: what the opening seconds have to do, the structure that works, how long to cut for each channel, the exact requests to type, and the things it cannot do.

Make Your Demo Video

Point it at your product, describe the walkthrough, review the cut. 50 free credits, no credit card.

Start Free →
See pricing →

What the First Five Seconds Have to Do

Almost every bad SaaS demo is bad in the same place: the beginning. A logo animation, a slow fade, a title slide, a founder introducing himself. None of it is the product, and the product is the only reason the viewer pressed play. On a landing page the video autoplays muted into a viewport the visitor is already scrolling past; in a feed it is competing with a thumb. The opening is not an introduction, it is a claim.

Four things belong in the first five seconds, and nothing else does.

The product, on screen, doing something
Open on real UI mid-action — a cursor moving toward a button, a result appearing. Not a home page. Not a logo. If the first frame could belong to any company, it is wasted.
A legible screen
The frame the viewer sees is 400 pixels wide on a phone. A full 1440-wide dashboard shrunk into it is texture, not information. Push in on the part that matters and enlarge the cursor, or the opening is an abstract pattern.
One sentence of claim
Text on screen, because the sound is off. A single line that names the outcome — not the feature. "Ship your first project in 90 seconds" beats "Introducing Workspaces".
A reason to keep watching
Show the end state early — the thing the product produces — then go back and show how. Withholding the payoff for 40 seconds is how demos lose the people who would have bought.

Two of those four are editing problems rather than scripting problems, and they are the two that usually go unfixed. Pushing in on the part of the screen that matters is a zoom aimed at a coordinate; making the cursor readable is a repaint of the pointer. Both are one sentence each in a recorded walkthrough or in a screen recording you upload.

The Structure That Works: Problem, Product, Proof

Every demo that holds attention is the same three movements in the same order. The lengths change per channel; the order does not.

1

Problem — the state of the world without you

Ten to fifteen percent of the runtime, and it is the part everyone over-runs. You are not educating the viewer about their own job; you are proving in one line that you understand it. Best delivered as a line of on-screen text over the product already moving, not as a separate scene.

2

Product — one path, all the way through

Sixty to seventy percent. Pick a single job the product does and follow it end to end without branching. Every click gets a push-in. The temptation to show a second feature is the single most reliable way to make a demo forgettable — a walkthrough that clicks everything shows nothing.

3

Proof — the result, and the next step

Twenty to twenty-five percent. The output exists: the report generated, the invite sent, the dashboard populated. Hold on it long enough to be read, then a title card with one instruction. One call to action, not three.

The reason to write the walkthrough before recording anything is that the walkthrough is the middle movement. When you hand Valmera a URL you also hand it the steps — click this, type that, scroll there — and four to ten steps that tell one story produce a usable capture. Twenty steps produce footage you then have to edit down, which is the work you were trying to skip.

How to Make a SaaS Demo Video with Valmera

  1. 1
    Write the walkthrough as steps
    Four to ten actions that tell one story: open the page, click the primary button, type into the field that matters, scroll to the result. Name the visible button labels — that is how the recorder finds them.
  2. 2
    Record the product, or upload your recording
    Give Valmera the URL and the steps and a headless browser performs them with a visible cursor while capturing video, returning an event track of every click and scroll with its position in the frame. If the product is behind a login, record it yourself and upload the file (MP4, MOV, MKV or WebM, up to 14GB or 3 hours).
  3. 3
    Cut it like a product video
    One request places the capture and writes a travelling zoom across each run of clicks, aimed by the event track. Add cursor enhancement if the pointer is small, and the floating window treatment if you want the standard product-video look.
  4. 4
    Clean the frame
    Erase a customer name or an email address so it is gone rather than covered; black-bar anything that must be a true redaction. The regions are measured off real frames and follow the footage through every later cut and reframe.
  5. 5
    Narrate, then cut one version per channel
    Add a voiceover or music with ducking, caption it, and export. Then ask for the vertical cut, the 30-second feed cut and the 15-second Shorts cut from the same upload — each one a separate request, each rendered from your original file at source quality.

Analysis runs once per upload with visible progress. Everything after that is a chat message, and every reply is verified server-side against the edits actually recorded.

The Requests, In Order

This is a real sequence, not a menu. Each message builds on the edit the last one produced, and any of them can be corrected in the next message rather than started over.

Record the product
Record a demo of https://app.example.com — scroll to the features section, click "New project", type "Q3 launch" into the name field, press Enter, then click "Invite team". Landscape.
Cut it like a product video
Place that capture at the start and cut it like a product demo: push in on each click, and make the cursor twice as big so it reads on a phone.
Clean the frame before anyone sees it
The customer name in the top-left of the dashboard and the email in the invite dialog both have to go — erase them, not blur them. Black-bar the billing number in the sidebar.
Give it the product-video look
Put the whole thing in a floating rounded window on a dark navy-to-purple gradient, add a title card that says "Ship your first project in 90 seconds", and master the audio.
Cut the vertical
Now make a 20-second vertical version of just the invite flow. 9:16, karaoke captions on my voiceover, hook first.

Notice what the third request does not say: it does not give coordinates. The agent looks at real frames of your footage, reads where the name sits, and measures the rectangle off the frame — then looks at the frames it rendered to check the result. That is also why the reply you get back is trustworthy: every reply is verified server-side against the edits actually recorded, so the agent cannot tell you it redacted something it did not.

Three Layers, and Why the Order Matters

A demo edit is unusual because it stacks operations that live at genuinely different depths. Some repaint the source pixels. Some rearrange time. One wraps everything that came before it. Knowing which is which is the difference between a clean edit and one where a redaction slides off the thing it was covering.

The three layers of a SaaS demo editOn the left, a demo frame drawn as nested rectangles: a gradient backdrop, a rounded floating window inset inside it, and within that window a browser capture containing a black redaction bar, an enlarged cursor and a dashed zoom rectangle aimed at a button. On the right, three stacked cards naming the layer each of those belongs to. Layer one is the source frame: cursor enhancement, erasing and censoring, all anchored in source coordinates so later cuts keep them. Layer two is the edit decision list: the splice, the travelling zoom per run of clicks, cuts, captions and audio, and it is the only layer that changes timing. Layer three is the finished picture: the floating window and mid-video aspect shifts, applied to everything at once and changing no timing at all. Dashed leader lines connect each element on the left to its layer on the right.ONE DELIVERED FRAMEREDACTEDzoom on click runCreatethe window is inset in the backdrop; captions and overlays live inside it1 · THE SOURCE FRAMEenhance_cursor  ·  erase_region  ·  blur_regionmeasured as fractions of the source frame, soevery later cut keeps it and no timestamp moves2 · THE EDIT DECISION LISTsplice  ·  travelling zoom per click runcuts  ·  captions  ·  voiceover  ·  musicthe only layer that changes timing — versioned,and every cut restorable3 · THE FINISHED PICTUREset_screen_frame  ·  add_aspect_shiftwraps everything above it at once, and changesno timing at all — add or remove it any timeTHE ORDER: redact and fix the pointer first, cut second, frame the window last.The original upload is never modified — every layer is re-derived from it on export.

The practical consequence of layer one is the useful one. Because cursor enhancement and erasing are anchored to the source frame rather than to the finished program, you can redact a customer name once and then cut, reorder, reframe and re-export as many times as you like without the black bar drifting off the thing it was covering. It is also why you should do it early: the pass that repaints pixels re-encodes the file, and doing it once beats doing it after every re-cut.

Layer three is the opposite kind of tool. set_screen_frame applies to the whole finished picture, which is exactly why captions and overlays scale with it — they are inside the window. It changes no timing, so adding it or taking it away is free at any point, and you can decide the look after the edit is right rather than before.

The Tools That Compound Here

Most video editors are built for footage of people. A SaaS demo is footage of a screen, which needs a different set of operations — and this is the one use case where all of them land in the same edit.

Record the product from a URL
A headless browser opens your public page with a drawn cursor and performs steps you write — click, type, hover, scroll, press, navigate. It returns an event track: every action timestamped and positioned in the frame. How it works.
Cut it like a product video
One call splices the clip in and writes a continuous zoom that travels between the buttons across each run of clicks, aimed by the event track rather than by a guess. On a recording you made yourself, pass the click times and it does the same job with eased centre punches instead of travel.
Make the pointer readable
The pointer is located in the source frames, the original repainted out, and a new one drawn up to 4x bigger along a tremor-filtered path. It reports the fraction of frames it actually found the cursor in. Cursor enhancement.
The floating-window look
Inset the picture, round the corners, drop a shadow, float it on a solid colour or a two-colour gradient in four directions. The standard treatment for a screen recording, and it costs no timing.
Erase a customer name
A rectangle repainted out of the source pixels with the background reconstructed inward — the name is gone, not covered. Several marks go in one call. Excellent on thin text over a steady screen. What it can and cannot rebuild.
Redact test data visibly
Blur, pixelate or a black bar over a rectangle, when you want the viewer to see that something was covered. A black bar is the only one of the three that is a true redaction. The difference explained.
Take a wide dashboard vertical
Morph the visible frame to 9:16, 1:1, 4:5 or 4:3 partway through and back, inside a fixed canvas — so nothing desyncs. Or reframe the whole video. Mid-video aspect shift.
Narrate and finish
A voiceover layer, music that ducks under speech, AI sound effects generated from a description and placed at exact times, word-accurate captions, and one-request mastering to a consistent loudness.

The compounding is the point. Recording gives you the event track; the event track aims the zooms; the zooms are what make a 1440-wide dashboard legible in a feed; the enlarged cursor is what makes the zoom readable rather than dizzying. Each one alone is a nice-to-have. Together they are the difference between a screen recording and a product video.

How Long, and for Which Channel

One demo video for every surface is the most common mistake after the slow opening. The same capture should be cut several times, because the viewer's situation is different in each place — whether they chose to watch, whether sound is on, and how much screen the frame gets. These are working defaults, not rules: start here and adjust once you see the finished cut.

Where it playsLengthFrameWhat it must survive
Landing-page hero30–60s16:9Autoplay, muted, no controls — on-screen text carries it
Product Hunt / launch60–90s16:9A skeptical audience that has seen a thousand of these
LinkedIn / X feed30–45s1:1 or 4:5Sound off by default — captions are mandatory
Shorts / Reels / TikTok15–30s9:16One feature only, and a hook before the first second ends
Sales email / outbound45–90s16:9A thumbnail that shows the product, not a play button
In-app onboarding60–120s16:9Someone who already signed up and wants the shortcut
Full YouTube walkthrough3–8 min16:9A viewer who chose it — chapters of the same three movements

Each of those is one request against the upload you already analyzed, so the second cut costs a fraction of the first — the expensive part, indexing the footage, happened once. Ask for one deliverable per message: "now a 20-second vertical of just the invite flow, karaoke captions, hook first". Batch output is not supported, and asking for four cuts in one message gets you one.

Redaction Is Not Optional, and Blur Is Not Redaction

Every real product recording contains something that should not be published: a customer's company name in the account switcher, a colleague's email in an invite dialog, an API key in a settings pane, seeded test data that reads as a real person. A demo video is the most-shared asset a SaaS company makes, and it is also the one most likely to leak something, because the recording was made in a real account.

There are three honest options and they are not interchangeable.

Erase — the pixels are repainted
The rectangle is painted out and the background reconstructed inward from the surrounding pixels, with no generative synthesis, so nothing is hallucinated into the frame. On thin text over a steady screen recording — which is exactly what a dashboard is — this is the right answer, and the result is measured and reported back before anything is promised.
Black bar — the only true visible redaction
A solid bar removes the information completely and tells the viewer it was removed. Use it where the fact of redaction is fine or even desirable: billing figures, a real customer's name in a case-study demo.
Blur or pixelate — a cover, not a removal
Softens or mosaics the region. It reads as censored and is fine for incidental detail, but it is a cover: a heavy blur over short, high-contrast text is weaker than a bar. If the thing genuinely must not be recoverable, use a bar or an erase.

All three are placed as fractions of the source frame, measured off real frames rather than estimated, and all three follow the footage through later cuts and reframes. What none of them do is track motion. The region is an axis-aligned rectangle that stays put, so an element that moves — a name in a list that scrolls, a tooltip that follows the pointer — needs either a rectangle large enough to contain its whole path, a window limited to the seconds it is visible, or a cut that avoids it. That is a real constraint and it is better to design the walkthrough around it than to discover it in review.

The Honest Limits

Everything below is a thing Valmera will not do for a SaaS demo. It is here because finding out in the middle of a launch is worse than finding out now.

It will not invent your product
This is not a text-to-video generator. There is no prompt that produces a demo of an interface that does not exist. It records a real page or edits a real recording. Generated stills and short generated clips can decorate the edit; they cannot be the product.
It cannot record behind a login
The browser is a fresh anonymous session with no cookies and no access to your logged-in browser, and it refuses outright to type into anything that looks like a password or payment field. For an authenticated product tour, use a public demo or sandbox account, or record it yourself and upload the file.
Large sites answer a datacenter IP with a bot wall
A capture of a site that shows a CAPTCHA or a consent screen succeeds technically and contains a CAPTCHA. That is detected and reported as a problem rather than handed back as clean b-roll, because retrying will not help.
The capture is silent, and there is no sound-effects library
No page audio is recorded, ever. Narration and music are added afterwards in the same chat, and there is no bundled pack of clicks or whooshes to pick from — each accent is an AI sound effect generated from your description and placed at a time you name, one per request.
Erasing is capped, and not magic
The repaint pass re-encodes the whole file, so it is limited to 10 minutes of source and about 900 megapixel-seconds — roughly 7 minutes at 1080p, under 2 minutes at 4K. Every other edit works on the full upload. A large object over a moving, detailed background can leave a soft patch; a small stationary mark over a plain background is reliable.
Captions are burned in
There is no SRT or VTT import or export, so you cannot hand the file to a platform that wants a sidecar subtitle track. What you get is a caption layer rendered into the picture, with an editable transcript that re-renders it.
No publishing, no seats, no brand kits
It does not post to YouTube, LinkedIn or TikTok, there are no team seats or share links, and there are no stored brand kits. The practical workaround for a house style is a standing brief you paste each time — the same sentence that got last month's demo right.
One deliverable per request
There is no batch export of four aspect ratios at once. Each cut is its own message against the same upload.
Fonts and transitions are a fixed set
12 bundled font families, no custom font upload; 7 transition styles, one applied to all cuts, and no true crossfade or dissolve. For a demo this rarely bites — hard cuts are the correct grammar for screen footage — but it is worth knowing before you promise a brand font.

One more, which is a feature rather than a limit: the agent will tell you when it did not do something. Replies are verified server-side against the edits actually recorded, so "I removed the customer name" is a statement the server checked. If a pass came back weak, you are told, and you get the frames to judge it yourself.

Cutting Release Videos From Inside Your Editor

There is a version of this workflow that fits how engineering teams already work. Valmera publishes its full editing toolset as a Model Context Protocol server, so Claude — in the app, in Claude Code, or in Cursor — can record the staging URL, cut the capture and export the file without anyone opening a video editor. For a release video that has to be remade every sprint, the demo becomes something you ask for in the same place you asked for the changelog. In the Claude app it connects over OAuth, so there is no token to copy; Claude Code and Cursor can be attached with a bearer token instead, and the setup page has the exact config per client.

Editing the same project from the studio and over MCP at once is refused in both directions, and slow operations return a job id rather than a fabricated completion — the same honesty rule as everywhere else.

What It Costs to Run This

Start on the free plan: 50 credits, granted once, no credit card. That is enough to record a capture, cut it, redact a name and export a real clip — a full pass at your actual product rather than a preview of somebody else's footage. Free-plan exports carry a small Valmera mark in the corner, and every export closes with a brief end card.

Beyond that, paid plans open with a 3-day free trial and export with no watermark at all. Creator is $30/month for 2,000 credits, Pro is $50/month for 4,000, and Frontier is $100/month for 10,000 on a stronger model for both reasoning and vision. Credits are charged in proportion to the AI work each request actually took, which matters here more than in most use cases: the expensive step is indexing the footage, and it happens once per upload. Every extra cut of the same capture — the vertical, the 30-second feed version, the 15-second Shorts version — is charged as the small edit it is.

If you are new to the workflow, the getting-started docs walk through a first project, and the tools hub has a page for every job on this page. Other roles working the same way are in use casescourse creators in particular share most of the screen-recording problems described here.

Record It, Cut It, Ship It

Point Valmera at your product and describe the walkthrough. 50 free credits, no credit card.

Make a Demo Video →
See pricing →

Frequently Asked Questions

Give Valmera the URL of your product's public page and describe the walkthrough — "open pricing, click Start free trial, type an email into the signup field, scroll to the FAQ". A headless browser opens the site with a visible cursor drawn on screen, glides to each target, clicks, waits for the page to react and types at human speed while the session is captured as video. The capture lands in your project as a normal asset, and it comes back with an event track: every click, scroll and keystroke timestamped and positioned in the frame. That track is what the edit is aimed with. If your product is behind a login, record it yourself and upload the file — every tool on this page works on an uploaded recording too.
It depends entirely on where it plays. A landing-page hero that autoplays muted wants 30 to 60 seconds and has to be legible with no sound. A feed post on LinkedIn or X wants 30 to 45 seconds, square or 4:5, captioned. A Shorts, Reels or TikTok cut wants 15 to 30 seconds vertical, one feature only. A sales email wants 45 to 90 seconds. A full YouTube walkthrough can run 3 to 8 minutes because the viewer chose it. Cut the same capture once per destination rather than shipping one length everywhere.
Yes, two different ways, and the difference matters. Blur, pixelate or a black bar puts a visible cover over a rectangle — the viewer can see something was censored, and a black bar is the only one of the three that is a true redaction. Erasing repaints the pixels and reconstructs the background, so the name is gone rather than covered. Both are placed as fractions of the source frame measured off real frames, and both follow the footage through later cuts and reframes. Erasing is excellent on thin text over a steady screen recording, which is exactly what a dashboard is.
Because a pointer sized for a 27-inch monitor is a few pixels wide on a phone. Valmera finds the pointer in the source frames, repaints the original out and redraws it up to 4x bigger along a path filtered to remove hand tremor, while fast deliberate moves stay sharp. It reports what fraction of frames the pointer was actually found in and refuses on footage with no visible cursor rather than re-encoding your video to change nothing. Click ripples can be added at times you supply — a click cannot be seen in pixels, so those times come from the recorder's event track or from you, never from a guess.
Yes, and there are two ways depending on what you want. Reframing converts the whole video to 9:16, 1:1 or 4:5 by crop, pad or blurred pad, with captions rescaled to the new frame. A mid-video aspect shift keeps the project in one shape and morphs the visible frame to another ratio partway through and back — useful when one feature deserves a vertical moment inside a wide video. The rendered file always keeps one resolution, so the shift is the frame itself closing in, which means it cannot desync audio or move a caption.
No. A browser capture is always silent — the browser is launched muted and no page audio is recorded. Add narration afterwards in the same chat: a voiceover layer, music from the built-in library of 23 CC0 tracks across eight moods or your own file, and AI sound effects generated from a description and placed at exact times. There is no bundled sound-effects library to pick from, so an accent like a click or a whoosh is generated from your description and placed at the moment you name — one request per sound.
No, and this is the sharpest limit on the page. Valmera is not a text-to-video generator. It records a real page or edits a real recording; it will not invent an interface. It can splice in generated stills and short generated clips as decoration around your footage, but the product on screen has to exist and be reachable — a live URL, or a file you record yourself and upload.
Every account starts with 50 free credits, granted once, with no credit card — enough to record a capture, cut it and export a real clip. Paid plans open with a 3-day free trial: Creator is $30/month for 2,000 credits, Pro is $50/month for 4,000, Frontier is $100/month for 10,000 on a stronger model. Credits are charged in proportion to the AI work each request actually took, so a re-cut of a capture you already have costs far less than the first pass. Free-plan exports carry a small Valmera mark in the corner; every paid plan exports with no watermark at all, and every export closes with a brief end card.

Your Product Deserves a Real Demo

The agent records the walkthrough, cuts it like a product video, redacts what should not ship, and exports from your original file. Free plan available.

Start Free →
See pricing →

Related Articles

Record a Website as Video
Give it a URL and it drives the browser — clicks, typing, scrolls — with a visible cursor.
Enhance the Cursor
Find the pointer in the pixels, smooth its path, redraw it up to 4x.
Remove Something From the Frame
Repaint a rectangle out and reconstruct the background behind it.
Change Aspect Ratio Mid-Video
Go vertical for one feature and back, without moving a single timestamp.