← Home
SOURCE-CHECKED PRODUCT REVIEW

By Valmera Editorial · Published · Updated

Captions App Review 2026: Pricing, AI Edit V3, Clips Chat, and the Limits That Matter

Captions has grown far beyond automatic subtitles. It now combines multi-clip AI Edit V3, prompt-based revisions, the separate Cappy SMS editor, long-to-short Clips Chat, Avatar X Twins and actors, localization, teams, API workflows, recording, and a manual editor. This review separates those paths, converts current model rates into capacity, and follows the product through project sync, data terms, delivery, and deletion—not merely the first polished generation.

DIRECT VERDICT

Is Captions App worth it?

Yes—for creators who repeatedly turn short, vertical, presenter-led footage into styled social video, or generate that presenter with an avatar. AI Edit coordinates cuts, captions, B-roll, music, motion, and effects; chat makes targeted revision accessible; Clips Chat handles iterative long-to-short repurposing. Horizontal and multi-clip support broaden the product, but the weaker fit remains bespoke long-form, multicamera, documentary, music-led, or picture-first editing. Credits make cost depend on models, generation, and retries; project sync, plan names, and rollover also have conflicting official pages.

BEST FOR
Vertical presenter video
FLAGSHIP PLAN
$24.99 Max · 500 credits
STANDOUT
Whole-video AI Edit
MAIN LIMIT
Credits + format fit

Captions in 2026: The Facts That Decide the Purchase

CURRENT CORE OFFER

Free, $24.99 Max with 500 credits, and Scale at $69.99/$139.99/$279.99 with 1,400/2,800/5,600; Enterprise is custom

LOWER HELP-CENTER TIERS

$9.99 Basic with 200 credits and Android-only $4.99 Lite remain documented even though the public price card centers Max and Scale

ROLLOVER

Current August credit documentation allows Basic and higher to bank up to 3× the monthly allowance; cancellation and some downgrades forfeit credits

AI EDIT

10–40 credits before model-dependent extras; 116 documented styles, multi-clip support, horizontal output, project duplication, and shot-level reruns

PROMPT TO VIDEO

16 credits per six seconds in the current meter; the web-specific guide says up to 30 seconds, 39 languages, and four-second generated segments

GENERATION RANGE

Current model table spans 1-credit images to 140-credit Kling 2.1 video calls, with several Veo variants at 36–100 credits

CLIPS

Long sources up to three hours or 10GB, 30+ candidates, and August 2026 Clips Chat for requesting more or different cutdowns

PROJECT CONTINUITY

2026 release notes say sync is rolling out across iOS and web; older help still says projects are local and can vanish on logout or cache clearing

TEAMS

Admin and Member roles, shared projects and assets, seat billing, immediate access revocation on removal, and team workspaces by default for new users

API

Mirage-hosted endpoints now expose caption rendering and video generation; current per-organization limits list 100 caption and 2 generation requests per minute

DATA

Standard Terms grant a broad service-improvement license over Input and Output; the Privacy Policy includes AI/ML training uses, while Enterprise advertises training-data exclusion

DELIVERY

Current public overview says 1080p publishing; other platform matrices and export controls expose 4K, HDR, SRT, bitrate, and frame-rate only on qualifying device/project paths

CAPPY

A separate SMS-based editor in beta: send clips, photos, a product link, voice note, article, or idea; receive a draft and revise by texting; free only for a limited trial period with no public ongoing price

MIRAGE AVATAR X

Current first-party pages claim a 225-avatar library, prompt-built custom avatars, and AI Twins from about 10 seconds of footage, with audio, expression, face, hands, and body generated as one performance

VENDOR STRESS TEST

Mirage says its 24-hour synthetic news stream aired 100+ segments but still produced two factual errors, drew about one-minute average watch time, and exposed long-form audio-cadence weaknesses

SECURITY CLAIM

Mirage announced completion of a SOC 2 audit in March 2026; enterprise buyers should request the report, type, period, criteria, scope, exceptions, subservice treatment, and bridge letter

How Much Is the Captions App?

The shortest answer is $24.99 per month for Max, the plan Captions foregrounds for its signature AI workflow. The full answer starts at $0, includes an Android-only $4.99 Lite tier and a $9.99 Basic tier in official help material, then rises through Scale allowances from $69.99 to $279.99 per month.

Prices below are published USD monthly figures checked August 26, 2026. Captions says final pricing can vary by monthly versus annual billing, App Store or Play Store versus web checkout, country, currency, and available account offer. The price shown before purchase is authoritative.

PLANPUBLISHED MONTHLY PRICEAI CREDITSPRACTICAL POSITION
Free$060–200 one-time credits in the current help guide; no recurring allowanceBasic captions, camera, trim, zoom, media, voiceover, transitions, resize, and basic export; current release notes say Free output carries a Captions watermark
Lite (Android)$4.99/moNo published renewable AI-credit bucketAndroid-only manual editing tier; it cannot be restored on desktop or iPhone without a separate Basic, Max, or Scale subscription
Basic$9.99/mo200 credits / monthFormerly Pro; watermark-free Basic projects, with Max-only AI Edit, chat, generated media, and AI Creator excluded
Max$24.99/mo500 credits / monthAI Edit styles, chat editing, full-video creation, digital twins or custom actors, and generated media
Scale$69.99/mo1,400 credits / monthEverything in Max, higher volume, and Captions' more sophisticated generative model tier
Scale 2x$139.99/mo2,800 credits / monthTwice the standard Scale usage for a higher-volume individual workflow
Scale 4x$279.99/mo5,600 credits / monthLargest published self-serve allowance; eligible for top-up credits under the current policy
EnterpriseCustomCustomCustom seats, bulk-credit discounts, onboarding, support, and training-data exclusion

Why two official pages appear to disagree about Free

The public pricing comparison labels Free as having no monthly AI allowance. The more detailed credit and subscription guides describe 60–200 one-time credits that do not refresh. Those statements can both be true: Free has a limited initial grant, not a renewable monthly pool. Treat the live account meter as the final answer for your signup.

What 500 Captions Max Credits Actually Buy

Paid tiers bundle credits at about five cents each, but that does not make each video predictable. The official meter charges the first generation, chat revision, regeneration, and generated asset separately. Credits are consumed even when a project is never exported.

ACTIONPUBLISHED BASE RATEMAX-PLAN EXAMPLE
AI Edit10–40 creditsAbout 12–50 base AI Edit runs from Max's 500 credits
AI Creator / AI Ads / AI Skits1 credit / secondA 60-second generation starts around 60 credits
Prompt to Video16 credits / 6 secondsA 60-second result uses about 160 credits before revisions
Chat-based edit2 per message + 1 per actionA one-action request starts at 3 credits; generated media costs extra
Actor regeneration40 creditsTwelve regenerations use 480 of Max's 500 monthly credits
Generated image1–3 creditsModel choice changes the cost
Generated video1–20 credits in the general tableModel-specific entries can be higher, so inspect the live meter
AI Edit style change1 creditTargeted restyling is much cheaper than regenerating the full edit

These examples are arithmetic, not a promise of output. AI Edit varies with duration, model, and algorithm choices; generated video has model-specific rates that can exceed the general add-on range. Check the meter after a representative project before forecasting production volume.

Need an Agent for the Footage, Not a Style for the Presenter?

Upload footage to Valmera, describe the finished edit, inspect the rendered result, and revise in conversation. Uploads are free; editing requires a subscription.

Create an account →
See pricing →

Is Captions Free? Trial, Export, and Watermark Rules

FREE PLAN

Yes—basic editing

Captions, camera, trim, zoom, sounds, media, voiceover, transitions, resize, and basic exports. Current release notes say Free output includes a Captions watermark.

FREE TRIAL

Checkout-specific

Current billing help explains trial cancellation, but the public pricing and plan pages do not promise one universal trial. Confirm the exact end date and renewal price in checkout.

WATERMARK

Free plus entitlement-sensitive

Current Free output is watermarked. Inactive paid subscriptions or Max-only features used on Basic can also produce a watermark according to the July 2026 troubleshooting guide.

RENEWAL

Automatic until canceled

Cancel through Apple, Google, or Captions' web billing portal at least 24 hours before renewal. Direct monthly web plans are described as non-refundable.

Safest test: make and export one free-only project before activating a paid feature or checkout trial. After adding AI Edit, Co-editor, AI-generated media, or AI Creator, inspect the plan badge before export rather than assuming the project still qualifies for the free path.

What Captions App Actually Is in 2026

“Caption app” now understates the product. Captions is a creator production suite with three different starting points: record a presenter in its camera; import existing footage for manual, automatic, or chat-based editing; or generate a presenter and supporting media from a prompt. A fourth workflow, Clips, mines long recordings for short social moments. Cappy is a fifth, separate surface that accepts media or an idea and returns revisions through SMS.

That breadth is the advantage and the source of confusion. AI Edit is optimized for a short, vertical, single-speaker source. Clips can accept a much longer podcast. AI Twin creates new presenter footage. Translation and Lipdub localize a finished performance. They live in one product, but they do not share the same limits, meter, or editorial purpose.

Captions Features Reviewed by the Job They Perform

01

AI Edit

The purchase-defining feature

Captions applies cuts, captions, B-roll, transitions, music, sound effects, motion graphics, and a chosen style in one pass. AI Edit V3 now supports multiple clips and horizontal output, while current workflow guidance still says short, clean, presenter-led footage is the safest source.

BEST FIT
Short, vertical, single-speaker footage that needs a coordinated visual treatment
CHECK BEFORE BUYING
A style is a strong starting point, not an unlimited editorial brief. The current library has 116 styles, older device limits conflict with newer releases, and the best workflow is to test a 30-second segment before spending credits on the full source.
02

Chat-based editor / Co-editor

Useful for targeted revision

The chat interface accepts natural-language changes and applies actions to the project. You can still open the manual editor afterward, which makes chat a revision layer rather than an irreversible black box.

BEST FIT
Changing caption color, removing a shot, moving text, adjusting music, or requesting B-roll without hunting through controls
CHECK BEFORE BUYING
Every message and action consumes credits, and generated assets add their own cost. Specific commands are more reliable and cheaper than broad requests such as “make it better.”
03

Cappy SMS editor

The lowest-friction outcome interface

Cappy accepts clips, photos, a product link, voice note, article, or description, asks follow-up questions, creates the structure and style, returns a video, and continues revising through SMS without an installed app.

BEST FIT
Phone-first users who want to send media or an idea, receive a draft, and revise it through ordinary text messages
CHECK BEFORE BUYING
It is a limited-time free beta with no public continuing price, credit relationship, file envelope, project handoff, retention schedule, or deterministic action log. Carrier rates and separate messaging-data questions also apply.
04

Clips

A materially broader long-form feature

Captions says Clips accepts an upload or URL up to three hours or 10 GB, detects multi-speaker conversation, reframes speakers, avoids doubling existing captions, scores proposed clips, and can produce 30 or more candidates. August 2026 adds Clips Chat for requesting more or different cutdowns from the same source.

BEST FIT
Podcasts, interviews, and long videos that need many vertical social cutdowns
CHECK BEFORE BUYING
Clip selection is a different workflow from AI Edit. Test whether the chosen moments preserve context, whether scores match your audience, and how much repair each clip needs before comparing it with specialist clippers.
05

AI Twin and actors

Strong for repeatable presenter content

AI Twin creates a reusable digital version of the user with a chosen or cloned voice; the current help center says an account can create up to 30 Twins. Photo, guided recording, and August 2026 video-upload setup are available, alongside reusable Avatar Looks and library actors.

BEST FIT
Ads, course updates, internal messages, and recurring social scripts without a new camera session
CHECK BEFORE BUYING
Synthetic delivery is not a substitute for consent, factual review, brand approval, or a real performance when credibility depends on human presence. Regeneration also consumes credits.
06

Captions, dubbing, and translation

A deep localization bundle

Automatic captions are available across current device platforms. Captions advertises recognition across 100+ languages, while translated captions, dubbing, Prompt to Video, and the 29 listed Lipdub language variants each have narrower, different inventories.

BEST FIT
Stylized captions, translated subtitles, dubbed versions, and lip-synced localization in one project family
CHECK BEFORE BUYING
Language counts differ by operation, and official guidance still recommends native-speaker review for important releases. Clean single-speaker audio produces the safest dubbing result.
07

Camera, teleprompter, and eye contact

Excellent creator utilities

The recording workflow combines a camera and teleprompter, while Eye Contact adjusts gaze after recording. The eye-contact tool is documented across desktop, iOS, and Android.

BEST FIT
People who record, script, caption, correct gaze, and finish talking-head video in one app
CHECK BEFORE BUYING
These utilities increase the value for presenters but do little for documentary, music, multicamera event, or picture-led narrative editing.

How Captions Actually Works: Credits, AI Edit, Avatars, Teams, API, and Data

The following analysis reconciles the live price card, the August 2026 credit guide, the current changelog, feature references, API documentation, and legal pages. Where an older official help article contradicts a newer source, the conflict is named and the safer operating rule is stated.

The live pricing page and the detailed subscription guide serve different purposes. The public page sells Free, Max, three Scale multipliers, and Enterprise, and explicitly says its prices reflect iOS plans. The help center additionally documents Android-only Lite at $4.99 and Basic at $9.99 with 200 credits. A January release note says Pro and Max were renamed Starter and Creator, yet the current August pricing page and credit guide again use Basic and Max. The safe buying rule is to quote the active checkout and entitlement screen, not infer a plan from an older review or name.

Max and every Scale multiplier price included credits at exactly five cents apiece: $24.99 for 500 is just under five cents, while $69.99/1,400, $139.99/2,800, and $279.99/5,600 differ only by rounding. Scale does not buy cheaper allowance; it buys more of it and access to the most sophisticated generative model tier. A buyer should therefore choose Scale for model access or measured volume, not because the headline bundle automatically improves the credit unit economics.

Current August documentation permits Basic and higher to carry unused monthly credits until the balance reaches three times the plan allowance. Max can therefore hold up to 1,500, standard Scale 4,200, Scale 2x 8,400, and Scale 4x 16,800. A downgrade immediately reduces the balance to the new cap and forfeits the excess. Cancellation lets credits remain through the paid period and then removes them. Upgrade timing, downgrade timing, and the planned generation backlog are part of the purchase decision.

A March troubleshooting article still says credits do not roll over. The canonical Understanding Credits page, last modified August 18, 2026, and the live pricing FAQ both say they do, capped at three times the monthly amount. The newer, more specific sources should govern this review. The contradiction is still operationally relevant because a support agent or product surface could be in transition. Preserve a screenshot of the balance and rollover language before making a downgrade or allowing a large saved balance to cross renewal.

AI Edit costs 10–40 credits in the general table before generated-media additions, but the range is not a quoted price for a finished video. Duration, style, model choices, algorithmic decisions, reruns, chat, images, video, music, and sound can all change the total. Max's 500 credits theoretically fund 12 full 40-credit base runs or 50 ten-credit runs. Neither number includes rejected edits or generated B-roll, so production capacity should be measured from the credit history after several representative jobs.

The current style guide documents 116 AI Edit styles: 21 Premium and 95 Basic. A style controls caption typography, motion graphics, B-roll strategy, music and sound, pacing, and the overall aesthetic. It is closer to a coordinated art-direction preset than a simple caption theme. Captions lets users preview, change the style on one shot, apply a custom brand color, replace media, edit words and timing, or request changes through chat. That editability is the difference between a useful starting system and a locked template.

AI Edit's source envelope is broader than the older vertical-single-clip description, but it remains format-sensitive. Release notes document 16:9 horizontal output, multi-clip sessions, multi-clip merging, importable voiceover in the edit plan, and long-to-short suggestions. The current workflow guide still says one clean speaker, vertical footage, and roughly 60–90 seconds produce the safest results. Product support and product design center are not the same: horizontal and multiple clips now work, but talking-head social footage remains the best benchmark.

The older device matrix still lists a one-minute web and two-minute iOS AI Edit ceiling. Newer release notes add multi-clip and long-to-short behavior without publishing one replacement maximum for every account. That makes duration an in-product entitlement rather than a reliable universal number. Test the exact device, project type, source count, combined duration, orientation, and plan. Do not turn a successful three-hour Clips upload into a claim that ordinary AI Edit accepts a three-hour timeline.

AI Edit V3 can create a first assembly of transcript cuts, captions, generated or selected B-roll, music, sound effects, transitions, zooms, motion, and visual styling. Afterward, individual shots can be rerun or preserved, and duplicated projects make alternate direction safer. This is meaningful automation. It does not establish multicamera synchronization, professional color management, nested timelines, source relinking, track locking, AAF/XML interchange, or a deterministic offline workflow. Those remain separate evaluation questions.

Chat-Based Editing charges two credits for the user's message, one for every action the system takes, and model costs for any generated asset. A message that changes one caption color starts at three credits; a broad request that replaces several shots and adds generated media can cost much more. The efficient pattern is to inspect first, ask for one bounded outcome, state which existing elements must remain, and verify the action list. A vague instruction is both harder to reproduce and more expensive to debug.

Prompt to Video is a different product path from AI Edit. It builds a synthetic talking performance, optional B-roll, music, transitions, sound, captions, cuts, and zooms from a prompt, selected actor, voice, and style. The plan can be reviewed before generation; a voiceover-only option removes the face; custom media can replace generated insertions. Changing the prompt creates a new plan, selecting several actors produces separate videos, and adding generated media or AI Edit styles increases the credit charge.

The current credit meter charges Prompt to Video at 16 credits per six seconds, about 2.67 credits per output second. A 30-second web result is therefore roughly 80 base credits before actors, edits, or additional media; Max can fund about six such base generations. The web-specific guide says a final video can be up to 30 seconds, a segment up to four seconds, the prompt or audio boundary is 30 seconds or 1,000 characters, and 39 languages are supported. Other device matrices describe one-minute creator projects, so the active path matters.

AI Creator, AI Ads, and AI Skits use the simpler published rate of one credit per output second. A 60-second generation starts at 60 credits. Actor regeneration costs 40 credits independently, so one 60-second video plus three full actor retries can reach 180 before model add-ons. AI Ads can ingest a product URL according to current release material, but the user must still verify price, offer, availability, trademarks, product appearance, disclosures, testimonials, and every claim the synthetic presenter makes.

The standalone model table makes 'generated video' a poor capacity unit. Ray 2 Flash and Hailuo 2.3 Fast are currently eight credits, Grok Imagine Video 11, Sora 2 and MAGI Distilled 16, Pika 20, Seedance variants 48–61, Veo variants 36–100, and Kling 2.1 140. One Kling 2.1 call consumes 28% of Max. Model duration and resolution are not uniform in the summary, so compare the actual quoted generation, accepted seconds, and retries rather than dividing credits by a generic clip count.

Image generation is cheaper but still model-dependent. Several current options cost one credit, common models cost two or three, and Nano Banana Pro costs six. A 20-shot edit with one two-credit image per shot is 40 credits; three variants per scene would be 120. Generated music ranges from two or three credits for current listed options to ten for ElevenLabs Music, while audio isolation is two credits per minute and several sound-effect models cost one to four. A polished automatic pass is the sum of many small meters.

Credits are charged when the AI work happens, even if the project is never exported or saved. Unlimited eligible exports therefore do not mean unlimited creation. Failed creative ideas, abandoned ads, actor retries, duplicate prompt plans, and chat actions still consume the balance. On iOS, the credit history exposes the transaction record. Use it as a production ledger: record source, requested operation, model, charge, accepted output, and whether the asset reached a published video.

When the balance reaches zero, current troubleshooting says ordinary AI Edit and related AI features become unavailable until renewal or upgrade. AI Creator can continue through a slower queue with waits up to one hour. Individual top-ups are not offered on normal personal plans; current top-ups are reserved for existing Teams or Business customers and Scale 4x. Scale 4x is therefore the only published self-serve tier with a path to more credits without a full plan change, and that entitlement should be verified in the account.

Clips is the long-form extraction product. It accepts an upload or URL up to three hours or 10GB, can identify multiple speakers, reframe for social, avoid doubling existing burned-in captions, score candidates, and return 30 or more options. August 2026 adds Clips Chat, letting the user request more or different shorts from the same source without restarting. That makes the extraction loop conversational, but the score remains a ranking hypothesis. Every candidate needs context, crop, caption, speaker, hook, and rights review.

Captions' product vocabulary is shifting from Clips toward AI Shorts and Long to Short in newer help material. The functional center remains the same: turn a long upload or YouTube link into social derivatives. It should not be conflated with AI Edit's automatic finishing or Prompt to Video's synthetic production. A podcast team may use all three—Clips to select moments, AI Edit to style an accepted short, and chat to repair it—but each stage can impose its own limits and credits.

AI Twin currently supports up to 30 digital clones. Setup can begin from a photo, a guided recording, or—since August 2026—a video upload, and users can create saved Avatar Looks for later projects. A Twin can be paired with a selected voice or a voice clone and reused in Prompt to Video. Deleting the actor is documented, but deletion does not remove already exported media. Consent, employment, campaign, and post-relationship use need a policy outside the feature itself.

The voice clone is created through the Twin setup and can narrate editor voiceovers, dubbing, and Prompt to Video. Captions recommends a quiet one-minute calibration at a natural pace, because cadence, emotion, reverb, and microphone distance affect the clone. The availability of a voice in the app does not settle whether it may be used for ads, regulated communications, impersonation, politics, or after consent is withdrawn. Preserve the approval and create an asset-removal process.

Lipdub changes mouth and facial movement to align with translated audio. The current help page lists 29 language variants, while automatic captions support a much larger language inventory and ordinary translation or dubbing has a different list. '100+ languages' describes caption recognition, not one universal feature matrix. For every target language, test transcription, translation, pronunciation, voice identity, mouth motion, onscreen text, timing, and cultural adaptation with a fluent reviewer.

Camera and Teleprompter make the iOS experience especially coherent for presenters. A script can be written or generated, shown in continuous or scene-by-scene scroll mode, recorded in HD or 4K at 30 or 60fps where supported, and then passed through Eye Contact and captioning. This saves round trips for solo creators. Web and Android do not mirror every camera, keyframe, voiceover, project, or HDR control, so a creator-suite review must remain device-specific.

The current export story has multiple layers. The product overview recommends 1080p and direct social publishing; the older device matrix lists up to 4K on iOS and web and 2K on legacy desktop; advanced controls can include frame rate, bitrate, Smart HDR, and SRT depending on source, device, plan, and project level. Free users can export with a watermark after the September 2025 policy change, while eligible paid projects can export without one. Run the final target settings, not a generic export test.

Project continuity is in transition. February release notes announced real-time sync, April said iOS-to-web was available to all users with the reverse direction still rolling out, June added cloud-backed caption templates, and July added native video uploads to keep projects synchronized. Older March and April help articles still warn that projects are local, disappear on logout or browser-cache clearing, and do not appear on another device. The safest interpretation is that sync exists for newer supported paths but is not universal across legacy desktop, web, devices, imports, and account states.

The web migration adds another continuity risk. Current sign-in help distinguishes legacy desktop.captions.ai from the new captions.ai web experience and warns that some sign-in paths may not load the expected subscription, projects, or workflows during the transition. July disabled new legacy-desktop signups and restored a macOS app representing the web experience. Existing desktop users should bookmark their correct surface, confirm the account method, export masters, and validate a copied project before clearing a browser or switching products.

Teams are now an actual collaboration layer rather than a future enterprise promise. New users enter a team workspace by default, Admins can manage people, billing, roles, and all team projects, while Members can create, edit, and share without administrative control. Removing someone immediately revokes access to projects and assets and can free a paid seat. Before offboarding, transfer ownership, preserve reusable media, verify Twin and voice-clone access, and export anything that must survive the account relationship.

Captions' API has moved toward Mirage branding and current endpoints at api.mirage.app. The video-captions API accepts vertical MP4 or MOV up to 50MB and five minutes, requires a caption-template ID, returns an asynchronous job, and exposes status and content retrieval. Current rate limits are 100 caption requests and two video-generation requests per minute per organization. These constraints are not the same as the consumer app's import, duration, project, or export rules.

Older API documentation advertises 360 initial credits, $0.05 per credit, purchases from 100 to 12,000 credits, one credit per second for AI Creator and Ads, and five requests per minute. The newer rate-limit page now says two video-generation requests per minute, and the product is branded Mirage in current examples. Treat the dashboard and current contract as authoritative for price, free credit, endpoint, model, rate, data, and rights. Do not project consumer-plan credits onto API calls.

Standard legal terms say the user owns Input and Output as between the parties, subject to Captions and licensor rights embedded in Captions Content. They also grant Captions a limited, nonexclusive, royalty-free, sublicensable license to host, store, modify, adapt, publish, translate, create derivatives, distribute, provide, improve, and develop services from that Input and Output, with the license surviving termination. Commercial buyers should read this alongside model-specific terms, asset licenses, and the Acceptable Use Policy.

The standard Privacy Policy states that information can be used to train AI or machine-learning models, improve services, infer interests, and—under described circumstances—be disclosed to third parties for their business purposes, including training algorithms. The public Enterprise offer explicitly lists training-data exclusion, which suggests standard and contracted treatment can differ. A company with confidential recordings, biometric likeness, unreleased products, children, health data, legal privilege, or client restrictions should obtain the applicable DPA and written data settings before upload.

Deleting a Captions account is permanent and removes in-app projects, but the current deletion guide warns that it does not cancel an active subscription. Apple, Google, and web billing have separate cancellation routes. That separation can leave a recurring charge after access is destroyed. Cancel with the actual payment provider, record the effective date, export projects and transcripts, remove team seats and API keys, delete the account only after verification, and retain the cancellation receipt.

A serious Captions evaluation should run four different tests rather than one polished demo: a 60-second presenter clip through AI Edit; a long real conversation through Clips and Clips Chat; a 30-second Prompt to Video using the intended actor and voice; and one manual or chat repair. Record device, plan, source, duration, model, credits, render time, sync behavior, export settings, watermark, human correction time, and accepted result. This converts a sprawling feature list into a defensible monthly capacity decision.

Cappy is a separate outcome-level editor, not another name for the in-project Co-editor. The official workflow starts in a phone's messaging app: a user can send clips, photos, a product-page link, voice note, article, or only an idea; Cappy asks questions, makes structure and style decisions, returns a finished draft, and accepts follow-up text revisions. The delivered video can be saved back to the phone. This is strategically important because it moves Captions from a tool with chat controls toward an editor available as a conversation channel.

Cappy's convenience also creates a separate buying boundary. Captions calls it a beta that is free to try for a limited time, says most videos arrive within minutes, and provides no public continuing price, credit allowance, file limit, resolution, retention period, supported country list, or revision meter. Carrier message and data rates may apply, and texting enrolls the number for automated replies until the user sends STOP. A buyer should not infer that an existing Max or Scale subscription funds Cappy or that a beta conversation inherits the app's project, export, and team controls.

Cappy and Captions' in-app Co-editor solve different problems. Co-editor modifies an existing structured project, exposes manual fallback and undo, and has documented message/action credit rates. Cappy accepts a broader starting object and returns a video over SMS, but its public pages do not document an editable project handoff, action log, exact meter, model selection, or deterministic undo. Use Cappy for a low-risk phone-first brief; use the app when review, exact correction, reusable project state, and known credit accounting matter.

Mirage Avatar X is now the named model behind Captions' synthetic presenters. Current first-party pages say it generates voice, face, micro-expressions, eye contact, hands, and body as one performance rather than applying lip sync to a still. The public library page says 225 ready-made avatars, also summarized as 200+, and allows prompt-built custom characters. That is a much broader casting surface than the older generic actor description, but every realism and fidelity statement remains a vendor claim until tested with motion, emotion, difficult phonemes, continuity, framing, and the target export.

The current AI Twin landing page says Avatar X needs about ten seconds of footage and supports uploaded video, a fresh recording, or a selfie, while the help guide also documents a photo-based setup and up to 30 saved Twins. These paths may trade speed for fidelity: the company itself recommends video for the strongest identity and motion match. Ten seconds is an input claim, not proof of perfect ownership, consent, accent, voice, hand movement, emotional range, or long-form identity stability. Test the same script twice and compare face, cadence, gestures, outfit, background, and cross-shot continuity.

Mirage's own 24-hour News Network experiment is more useful than a curated avatar reel because it publishes operational evidence and failures. The company says one machine and roughly $50,000 in tokens produced more than 100 original segments around licensed newswire reporting. Scripts were mapped to shot lists; automated checks covered frozen frames, A/V sync, speech tails, and asset rights; real guests gave written permission and approved talking points; every frame carried an AI-generated disclosure. That is a strong example of the controls required around synthetic media, not evidence that an ordinary Captions subscription can reproduce the system or cost.

The same postmortem documents the limit clearly. Mirage says two factual errors aired, including a wrong identity caption, and both needed public correction. It also reports overly uniform cadence and poorly placed pauses in long-form audio. More than 50,000 people tuned in and the stream generated over three million impressions, but average watch time was around one minute. Those vendor-reported results demonstrate three separate facts: generated presentation can scale, truth still needs editorial verification, and technical realism does not guarantee sustained attention.

Mirage announced in March 2026 that it completed a SOC 2 audit with independent auditors. That is materially stronger than an unverified security slogan, but the public post does not substitute for the report. It does not state in the article whether the report is Type I or Type II, the exact review period, Trust Services Criteria, system boundary, exceptions, complementary user controls, or subservice-organization method. Enterprise procurement should request the current report under NDA, bridge coverage, penetration-test summary, DPA, subprocessors, retention, deletion, incident terms, and the precise products included.

Avatar X also raises a provenance distinction the UI must preserve. A library avatar is a licensed synthetic presenter, a prompt-built avatar is a generated character, and an AI Twin represents a real person. Captions says its ready-made avatars can be used commercially subject to the content rules, but that does not authorize a customer's script, product claim, trademark, music, reference asset, cloned voice, or real-person likeness. Keep the consent record, prompt, source, actor type, generated version, disclosures, approvals, and final export attached to each campaign.

What Captions Does Exceptionally Well

AI Edit has become more flexible without losing its design center

Horizontal output, multiple clips, project duplication, voiceover import, long-to-short suggestions, shot-level reruns, and new styles expand the workflow while the clean talking-head path remains easy to understand.

The style system is unusually explicit

Captions documents 116 coordinated AI Edit styles and explains which visual, caption, B-roll, music, motion, and pacing decisions each style controls. Buyers can preview, customize, or mix them by shot.

Credit rates are published down to individual models

Current documentation lists plan pools, feature charges, image models, video models, music, isolation, and sound. That enables real capacity math instead of treating every generation as one vague AI action.

Clips Chat makes long-to-short iterative

A user can ask for more or different short candidates from the same processed source in August 2026 rather than uploading again. This is a useful conversational refinement loop for podcast and interview repurposing.

Generated and recorded presenter workflows meet in one editor

A real camera performance, AI Twin, library actor, imported voiceover, cloned voice, Prompt to Video project, or voiceover-only production can all reach captions, media, sound, manual correction, and export.

Localization reaches beyond subtitle replacement

Caption recognition, subtitle translation, dubbing, voice cloning, and Lipdub provide several levels of localization. Teams can choose the least invasive operation appropriate to the release.

Cross-device and team infrastructure is actively improving

Project sync, cloud-backed templates, native uploads, shared workspaces, real-time co-editing, Admin and Member roles, and macOS support reduce the isolation of the original mobile-only workflow.

The API now covers both caption rendering and video generation

Organizations can add stylized captions to vertical video or generate presenter media programmatically, with asynchronous job status, downloadable output, API keys, and published rate limits.

Rollover rewards uneven schedules

Current Basic-and-higher accounts can bank two extra months of unused allowance, giving Max a ceiling of 1,500 credits and standard Scale 4,200 without losing every quiet month's balance.

The free surface can validate basic editing

Trimming, transitions, a caption template, resizing, and other basic tools let a buyer test import, transcription, timing, interface, and device compatibility before making the flagship generative features the purchase decision.

Human review remains possible after automation

Captions exposes caption correction, timing, shot style, media replacement, color, intensity, chat, timeline controls, exports, and original or unedited synthetic takes. The first pass is not the only available pass.

Enterprise states the key data difference plainly

Training-data exclusion is listed as an Enterprise entitlement rather than being implied for every customer. That makes procurement's next question clearer, even though the contract still needs review.

Cappy tests a genuinely low-friction editing channel

Clips, photos, a product link, voice note, article, or idea can enter through ordinary texting, and the user can iterate without learning an editor. It is a meaningful outcome-level interface, not just another toolbar.

Avatar X joins casting, voice, expression, and body motion

Captions now names the model behind a 225-avatar library, prompt-built characters, and short-recording AI Twins. Recorded and synthetic presenters can still reach the same correction and export environment.

Mirage published unusually candid long-run evidence

The News Network postmortem names its controls, token spend, two factual errors, audio weakness, and short average watch time instead of presenting only a highlight reel. Buyers can use those failures to design review gates.

The SOC 2 audit claim improves the enterprise starting point

Mirage says independent auditors completed a SOC 2 audit. Procurement still needs the actual report and scope, but the claim is more concrete than a generic enterprise-security badge.

The Limits That Matter More Than the Feature Count

AI Edit has a narrow ideal source

Official guidance favors one person speaking, vertical 9:16 footage, little or no prior editing, and a short duration. That is an excellent social format and a real boundary.

Platform parity is incomplete

Web, iOS, and Android are all supported, but trial, Lite-plan, duration, resolution, and feature availability can differ. A review based on one device does not settle another.

Credit cost appears after the headline

Generation, chat, actions, retries, and assets can each deduct credits. The subscription does not provide unlimited use of its flagship AI features.

Styles can converge on a house look

One of 116 style systems makes speed and consistency possible. It can also make the output feel template-led unless the user customizes captions, color, B-roll, pacing, and sound.

Synthetic video still needs review

An avatar can save a shoot, and localization can save a re-record. Neither removes the need to check likeness, consent, delivery, facts, translation, product details, and brand safety.

Export is not the only unit of value

A technically unlimited export allowance does not matter if credits run out, the source misses AI Edit's format, or every result needs extensive correction. Measure finished videos per month.

Current plan names conflict with the changelog

The live page sells Max while a January 2026 release says Max became Creator; the current credit guide calls the $9.99 tier Basic, formerly Pro or Starter. Checkout is authoritative, but comparisons become easy to misread.

The public pricing page reflects iOS

Captions explicitly says its displayed prices and features reflect iOS plans. Web, Google Play, legacy desktop, country, currency, tax, annual offers, and existing-account entitlements can differ.

Scale does not reduce the nominal price per credit

Max and Scale bundles all sit around five cents per included credit. Scale earns its premium through volume and model access, not a materially better allowance exchange rate.

Rollover documentation contradicts an older help page

The August authority permits balances up to 3×, while a March troubleshooting page says no rollover. The newer rule should win, but subscribers should preserve the account meter before a high-stakes billing change.

Downgrading can destroy saved credits immediately

A lower plan's 3× ceiling applies at downgrade and excess balance is forfeited. A user who banked Scale credits can lose them before the next normal renewal if timing is not planned.

Cancellation forfeits the remaining balance

Credits stay usable only through the paid term and disappear at expiry. An exported project does not preserve unused generation capacity, and an annual commitment does not make credits transferable.

AI Edit's 10–40 credit range is incomplete cost

Generated B-roll, images, video, music, sound, chat, style changes, and reruns can sit on top of the base charge. The monthly video count cannot be known from the plan table alone.

A premium video model can consume 28% of Max

The current Kling 2.1 entry costs 140 credits per call. Several Veo variants cost 36–100. A few rejected synthetic clips can crowd out an entire month of normal editing.

Multi-clip support does not equal long-form NLE support

AI Edit can merge and process several clips, but that does not prove multicamera sync, nested timelines, relink, interchange, advanced color, audio buses, or hour-long deterministic finishing.

Device matrices lag release notes

Older official tables still describe vertical-only, local projects, and different duration or feature limits even after 2026 releases added horizontal AI Edit, sync, multi-clip work, and web insertions.

Prompt to Video limits vary by surface

The web guide says a 30-second final with four-second segments, while other matrices describe one-minute creator projects. Actor count, prompt length, upload size, language, and available plan also differ.

Every selected actor creates another paid video

Choosing several actors in the web prompt workflow produces separate outputs and uses additional credits rather than giving one free comparison grid. Casting experiments should be budgeted as generations.

Actor regeneration is expensive

One retry costs 40 credits, so a few delivery, face, motion, or continuity problems can exceed the base price of the finished synthetic speech. A short casting proof is essential.

Synthetic ads can invent commercial facts

URL-based and prompt-driven ads still need product, price, availability, testimonial, disclosure, logo, demonstration, and jurisdiction review. A plausible synthetic presenter is not an approved claim source.

Virality and clip scores are hypotheses

Clips can rank many candidates and Chat can request alternatives, but no score knows the actual audience, campaign objective, context, legal sensitivity, or channel performance before publication.

AI Twin scale increases consent risk

Thirty reusable clones, upload-based creation, saved Looks, and cloned voice simplify production and make unauthorized reuse easier. Access, campaign, duration, geography, and post-employment rules need explicit ownership.

Language counts are not one entitlement

Captions recognition, translated subtitles, dubbing, Lipdub, and Prompt to Video have different supported lists. Marketing a 100-language caption tool does not promise 100-language lip synchronization or generation.

Project sync is not universal enough to trust blindly

Release notes describe staged iOS/web rollout and native-upload requirements, while older help still warns of local loss. Legacy desktop, cache, browser, account method, import path, and device direction can change continuity.

The web transition can split projects and subscriptions

Legacy desktop and the new captions.ai experience coexist. Signing in through a different method or surface can expose the wrong account, missing projects, or missing entitlement during migration.

Team removal revokes access immediately

Offboarding before asset and ownership transfer can strand a project, prompt, media file, actor, or clone. A seat-saving click is also a production-access event.

Free export wording changed

A newer release says Free exports carry a Captions watermark, while older article versions described clean output for free-only projects. Test the active project badge and a real download instead of carrying forward the old promise.

4K is not a universal export promise

One matrix exposes 4K on qualifying iOS and web paths, while the current overview recommends 1080p and legacy desktop lists 2K. HDR, source, project level, device, plan, bitrate, and frame rate affect delivery.

The consumer and API products use different limits

API captioning is vertical, MP4/MOV, 50MB, and five minutes with its own template and organization RPM. Consumer upload, Clips, AI Edit, and export ceilings cannot be projected onto it.

API documentation spans old Captions and new Mirage surfaces

Pricing, initial credits, rate limits, endpoint host, and product names have evolved. The dashboard and signed agreement must replace copied integration examples before launch.

Standard data terms are broader than Enterprise positioning

The standard Terms and Privacy Policy permit service improvement and AI/ML training uses; Enterprise advertises training-data exclusion. Confidential footage needs a written entitlement, not an assumption based on the sales page.

Deleting an account does not cancel billing

Captions explicitly warns that deletion can erase projects while leaving an Apple, Google, or Stripe subscription active. Billing cancellation must be completed separately before destructive account removal.

Input and Output ownership has qualifications

Users own their material as between the parties, subject to Captions Content and a broad surviving license granted for operating, improving, and developing services. Commercial and regulated teams should review the full contract.

Automatic polish can converge on a Captions look

The same styles, typography, generated B-roll logic, sounds, and pacing recur across many creators. A brand must replace assets, control intensity, correct captions, and define a repeatable treatment to remain distinctive.

Cappy has no public continuing economics

The SMS editor is free only for a limited beta period. Public pages do not disclose the eventual price, included jobs, revision meter, relationship to Max or Scale credits, or failure-refund rule.

SMS is a separate data and delivery surface

Phone number, carrier metadata, message history, attachments, generated drafts, and automated replies pass through a messaging workflow. Teams should clarify retention, deletion, subprocessors, country support, STOP handling, and whether projects enter the normal workspace.

Avatar realism can increase deception impact

A model designed to preserve identity, voice, expression, hands, and body can make an unauthorized or incorrect statement more persuasive. Consent, disclosure, claim review, access control, and source preservation become more important as realism improves.

The vendor's own broadcast produced factual errors

Two errors reached air despite licensed reporting, scripted shot maps, automated checks, written guest approval, and disclosure. Synthetic presentation needs human fact and identity review even in a heavily controlled pipeline.

SOC 2 scope is not visible in the announcement

The public post does not specify Type I or II, audit period, criteria, system boundary, exceptions, complementary controls, or subservice treatment. The report—not the badge—must answer whether the required Captions and Mirage surfaces are covered.

Who Should Choose Captions—and Who Should Not?

Choose Captions when

  • Your recurring source is a person speaking to camera.
  • The desired output is a short, vertical, polished social video.
  • A strong automatic style is more valuable than designing every decision.
  • You want avatars, digital twins, eye contact, captions, and localization together.
  • You will reuse a long podcast through the separate Clips workflow.
  • You can measure credit use on representative projects before committing annually.

Choose another editor when

  • Your edit is long-form, multicamera, documentary, montage, or music-led.
  • You need exact timeline interchange, collaborative project files, or professional finishing.
  • The editor must interpret an unusual open-ended brief rather than apply a creator style.
  • You require predictable unlimited AI generation without a usage meter.
  • Platform parity across a mixed web, iOS, and Android team is mandatory.
  • Your organization cannot use synthetic likeness or cloud generation under its policy.

Captions App Alternatives by Workflow

No single product is “Captions without the limits.” Each alternative changes the center of the workflow. Start with what enters the system and what must come out.

ALTERNATIVECHOOSE IT FORHOW THE WORKFLOW CHANGESDETAILED GUIDE
ValmeraOpen-ended editing of uploaded footageDescribe the outcome, inspect a rendered preview, and revise the edit in conversationCompare →
DescriptPrecise transcript-led podcasts and interviewsEdit words and structure directly through a document-like transcriptCompare →
OpusClipDedicated long-to-short repurposingOptimize clip extraction, reframing, hooks, captions, and batch social outputCompare →
SubmagicRepeatable captioned shortsUse recognizable animated captions, B-roll, hooks, cleanup, and social finishingCompare →
CapCutHands-on mobile and trend-led editingChoose a broad manual social editor with templates, effects, and many utility AI toolsCompare →
VEEDCollaborative browser editing and brand workflowsKeep timeline, transcript, subtitles, generation, translation, stock, and review in one web workspaceCompare →

Captions vs Valmera in one sentence

Captions is the stronger creator suite when a ready-made style, synthetic presenter, eye contact, and localization drive the purchase; Valmera is the stronger fit when uploaded footage and an open-ended editorial brief drive the purchase.

How to Test Captions App Before Paying for a Year

  1. 1
    Choose a representative source, not a demo-friendly clip
    Use a real 30–60 second vertical talking-head recording with a difficult name, one pause, imperfect audio, and at least one moment where relevant B-roll would help. This tests Captions' design center without wasting credits.
  2. 2
    Make a free-plan export first
    Add captions, trim, resize, and use only features listed in the free plan. Export before touching a paid feature so you can verify the free and watermark behavior on your exact device and account.
  3. 3
    Run two contrasting AI Edit styles
    Preview styles first, then generate the same short segment with one restrained and one expressive style. Compare caption accuracy, cut timing, B-roll relevance, sound, motion, brand fit, and distraction—not just polish.
  4. 4
    Repair one result through chat and manually
    Request three precise changes through Co-editor, record the credits used, then correct a caption and visual element by hand. Measure whether the product saves time after the first pass, not only during it.
  5. 5
    Test your second most important workflow
    If you publish podcasts, run Clips on a real episode. If you localize, dub one minute and have a fluent speaker review it. If avatars matter, generate the same script twice and inspect continuity, delivery, and regeneration cost.
  6. 6
    Calculate a full month before subscribing
    Multiply your videos by their observed AI Edit, chat, asset, dubbing, and regeneration usage. Add the cost of human correction. Choose Max or Scale only if the measured month fits the allowance with room for retries.

A focused test takes about two hours plus generation time. The purchase decision should use exported files and observed credit history from your actual device, footage, and workflow.

How This Captions App Review Was Researched

We reviewed Captions' current product, pricing, help, feature, billing, and app-store material on August 26, 2026. We did not assign a fictional numerical score or claim hands-on laboratory testing that did not occur. Every material number on this page is tied to publisher-owned documentation; calculations are labeled as calculations.

  • Price: current published USD monthly rates; platform and checkout variability disclosed.
  • Free: ongoing features separated from a one-time credit grant and any checkout-specific trial offer.
  • Cost: action-level credit rates converted into transparent examples without promising a fixed number of videos.
  • Fit: official format guidance used to separate AI Edit, Clips, avatar, localization, and manual workflows.
  • Conflict: Valmera publishes this review and competes for AI video-editing customers; Captions' advantages are stated because the right answer is sometimes Captions.
  • Commercial policy: no affiliate links, referral parameters, sponsored ranking, or invented user review quotes.

Read the Valmera editorial policy. Product pages and the final account checkout remain authoritative at purchase time.

Official Sources Checked

Captions product overviewCaptions pricingSubscriptions and plansCredit allowances and ratesManage subscription and billingWatermark troubleshootingImport, formats, and limitsDevice feature availabilityAI Edit product pageAI Edit referenceChat-based Co-editorText-based video editorClips product pageClips workflow referenceExport settings and SRTVideo translationAI dubbing and LipdubEye Contact referenceAutomatic captions referenceApple App Store listingCaptions Help — August 2026 Product ChangelogCaptions Help — Current Product and API OverviewCaptions Help — Current Caption StylesCaptions Help — Prompt to VideoCaptions Help — AI Twin Creation and DeletionCaptions Help — Lipdub LanguagesCaptions Help — Dubbing and TranslationCaptions Help — Current Product IntroductionCaptions Help — iOS AppCaptions Help — Android AppCaptions Help — Current Credit TroubleshootingCaptions Help — Team Members, Roles, and RemovalCaptions Help — Current Project ManagementCaptions Help — Lost Project WarningCaptions Help — Permanent Account DeletionCaptions Help — Web and Legacy Desktop TransitionCaptions Help — Sign-In and Web MigrationCaptions Help — API Rate LimitsCaptions Help — Video Captions APICaptions Help — API ReferenceCaptions — API TermsCaptions — Terms and Input/Output RightsCaptions — Privacy Policy and Model TrainingCaptions Help — Cappy SMS Video EditorCaptions — Cappy Launch and Beta TermsCaptions — Mirage Avatar X and 225-Avatar LibraryCaptions — Mirage Avatar X AI Twin WorkflowCaptions — Mirage News Network 24-Hour PostmortemCaptions — Mirage SOC 2 Audit Announcement

Edit Real Footage by Describing the Result

Valmera turns an open-ended brief into a rendered edit you can review and revise. Create an account and upload for free; editing requires a subscription.

Try Valmera →
See pricing →

Frequently Asked Questions

Checked August 24, 2026: Basic is $9.99 per month for 200 credits, Max is $24.99 for 500, Scale is $69.99 for 1,400, Scale 2x is $139.99 for 2,800, and Scale 4x is $279.99 for 5,600. Android also has a $4.99 Lite tier. Annual, platform, region, currency, tax, and checkout prices can differ.
Yes. The current Free plan supports basic editing and export and includes 60–200 one-time credits that do not renew. September 2025 release notes say Free exports carry a Captions watermark, while most AI-powered project levels require a paid subscription. Test the current account because older help described different free-export behavior.
Captions' current billing help explains how to cancel an active trial, but the current public pricing and subscription pages do not promise one universal trial length or platform offer. Treat any trial shown in Apple, Google, or web checkout as account- and platform-specific; confirm its end date and renewal price before starting it.
Yes on current Free exports according to the September 2025 release notes. Paid exports can also show a watermark when the subscription is inactive or a Basic project uses Max-only AI Edit, Co-editor, generated media, or AI Creator. Verify the project level, account, and downloaded file before publication.
Max is Captions' $24.99-per-month flagship individual plan with 500 monthly credits. It includes curated AI Edit styles, full-video creation, custom actors or digital twins, chat-based changes, and generated B-roll, images, music, and sound effects. Credits—not exports alone—are the practical monthly limit.
There is no honest single number. The official guide prices AI Edit at roughly 10–40 credits, AI Creator at one credit per second, Prompt to Video at 16 credits per six seconds, actor regeneration at 40 credits, and chat at two credits per message plus one per action and any generated media. Max's 500 credits could cover about 12–50 base AI Edit runs or roughly three 60-second Prompt to Video generations before revisions.
Eligible Basic-and-higher subscribers can roll unused credits forward up to a total balance of three times the monthly allowance. A downgrade immediately applies the new lower cap, and cancellation forfeits the remaining balance at the end of the billing period. Individual top-ups are generally unavailable; current top-ups are limited to specified business tiers and Scale 4x.
It is worth testing for frequent short, vertical, presenter-led video when AI Edit styles, chat revisions, avatars, captions, eye contact, and localization replace several separate tools. It is less compelling for long-form narrative, multicamera, documentary, music-led, or highly bespoke editing, or when generation and retries exceed the plan's credits.
Captions documents workflows on iOS, web, and Android, with a separate legacy desktop experience still appearing in its availability table. Feature parity is incomplete: AI Edit's current guide recommends at most two minutes on iOS and one minute on web, while Lite is Android-only. Test the exact device and workflow you will publish from.
The answer depends on the workflow. AI Edit is optimized for short vertical talking-head video and its official reference lists a one-minute desktop or two-minute iOS limit. The separate Clips feature accepts sources up to three hours or 10 GB for long-to-short repurposing. That does not make AI Edit a general three-hour timeline editor.
Yes. Captions says subscriptions can be canceled and remain usable through the paid period. App Store and Play Store purchases are canceled through those stores; web subscriptions are canceled through Captions' Stripe portal. Plans renew automatically unless canceled at least 24 hours before renewal, and current documentation says direct monthly website plans are non-refundable.
Captions now calls its $9.99 plan Basic. Current subscription documentation says Basic was formerly Pro. Some older third-party reviews use Starter, but that is not the current official plan name; use Basic and verify the live entitlement list.
Lite is a $4.99-per-month Android-only editing plan. Captions says it can be restored and used only on Android; signing into the same account on desktop or iPhone does not transfer the Lite subscription there. Cross-platform users need a separate Basic, Max, or Scale subscription.
The limit depends on the workflow. Current standard project imports can be up to 60 minutes. AI Edit recommends at most two minutes on iOS or one minute on web. Clips accepts a source up to three hours or 10 GB. These limits are not interchangeable.
Captions' current import guide lists MP4, MOV, and M4V video plus MP3, M4A, AAC, and WAV audio. Projects can also contain images, and iOS or web can import from a YouTube or other public URL. Device and workflow limits still apply after import.
Captions says its Clips workflow can produce 30 or more candidates from a source up to three hours or 10 GB. It ranks moments, reframes footage, tightens dead air, and supports podcast, gaming, webinar, sports, and already-captioned sources. Review every proposed clip for context and crop quality.
Captions exposes advanced resolution, frame-rate, bitrate, supported Smart HDR, and separate SRT controls according to plan, device, source, and project level. Its availability table lists 4K on iOS and web but 2K for the legacy desktop experience, and the web editor does not support HDR video.
Yes. Co-editor accepts natural-language changes, image or video attachments, chained requests, and follow-up instructions, then lets the user review or undo the applied edit. Each user message costs two credits, each action costs one, and generated media adds its own model charge.
Captions says exports are unlimited within the project level included by the plan: Free for Basic-feature projects, Basic for Basic projects, Max for Basic and Max projects, and Scale for all corresponding levels. Credits, duration, device capability, and watermark rules can still constrain the workflow before export.
Valmera is the stronger alternative for open-ended agentic editing of uploaded footage. Descript is better for exact transcript-led work, OpusClip for specialist long-to-short clipping, CapCut for hands-on mobile editing, Submagic for templated captioned shorts, and VEED for a collaborative browser workspace. The best replacement depends on the output, not the number of AI features.
The live public price card centers Free, Max at $24.99 with 500 credits, Scale at $69.99/$139.99/$279.99 with 1,400/2,800/5,600, and Enterprise. Detailed help also lists Basic at $9.99 with 200 credits and Android-only Lite at $4.99. Checkout and device are authoritative.
Official materials span several naming periods. The current credit guide calls the $9.99 tier Basic, formerly Pro or Starter. A January 2026 changelog says Pro and Max became Starter and Creator, yet the current pricing page again sells Max. Match the price, credits, device, and entitlement—not only the name.
Not necessarily. The public pricing page explicitly says its displayed prices and features reflect iOS plans. Captions also says billing length, App Store, Play Store, Stripe, country, currency, tax, and available account offers can change checkout.
Max and Scale bundles work out to roughly five cents per included credit. $24.99/500 is just under five cents; $69.99/1,400, $139.99/2,800, and $279.99/5,600 differ only by rounding. Output per credit varies radically by feature and model.
No meaningful discount appears in the published monthly bundles. Scale mainly adds capacity and access to the more sophisticated generative-model tier. Choose it for measured volume or model entitlement, not because the nominal credit exchange rate improves.
Yes according to the canonical credit guide updated August 18 and the live pricing FAQ. Basic and higher can hold the current month plus up to two extra months, capped at 3× allowance. A stale March troubleshooting page says no rollover; the newer sources supersede it.
The current rollover ceiling is 3× monthly allowance: Basic 600, Max 1,500, Scale 4,200, Scale 2x 8,400, and Scale 4x 16,800. Top-up treatment for eligible Teams, Business, and Scale 4x accounts can be separate.
The balance is immediately capped at three times the new lower monthly allowance and excess credits are forfeited. Use or document a large saved balance before downgrading and preserve the account meter in case the transition is disputed.
They remain available through the end of the already-paid billing period and are forfeited when the subscription expires. Cancellation does not convert them to API credits, cash, or a permanent balance.
Individual top-ups are not available on standard personal plans. Current documentation reserves top-ups for existing Teams or Business customers and Scale 4x users. Other subscribers must wait for renewal or upgrade.
AI Edit and ordinary AI features become unavailable until refresh or upgrade. Current troubleshooting says AI Creator can still generate in a slower queue with waits up to one hour. Manual and already-entitled export behavior should be verified in the project.
The published base range is 10–40 credits, but actual usage varies with duration, model, and algorithm choices. Generated B-roll, images, video, music, sound, chat actions, style changes, and reruns can add more.
Its 500 monthly credits theoretically cover 12 runs at a 40-credit base or 50 at ten credits. Real capacity is lower when a project uses generated media, chat, retries, actor work, music, sounds, or other AI features.
The current official style guide documents 116: 21 Premium and 95 Basic. Styles coordinate captions, graphics, B-roll, music, sound, pacing, and aesthetic, and can be previewed or changed after generation.
Yes. Current release notes say Horizontal AI Edit supports 16:9 output. The workflow is still optimized for talking-head social content, and an older feature matrix remains vertical-only, so test the current account and exact style at 16:9.
Yes. 2026 releases added multi-clip sessions, merging, and multi-clip support in the edit plan. This helps assemble several recordings but does not establish multicamera synchronization, professional track interchange, or unrestricted long-form finishing.
The older matrix still lists one minute on web and two on iOS. Newer releases add multi-clip and long-to-short behavior without one replacement universal maximum. The active device, project type, source count, orientation, plan, and edit-plan screen are authoritative.
Yes. April 2026 release notes added AI Edit project duplication, letting users branch before trying a different direction. Duplicating preserves the original project, but any new AI generation still consumes credits.
Yes. The current workflow supports changing or rerunning an individual shot's style and preserving existing images. This is more efficient than rebuilding the full video when one segment fails, though the targeted generation can still consume credits.
Each user message costs two credits, each action taken costs one, and any generated media adds its model charge. A one-action text request starts at three credits. Broad multi-action revisions cost more.
The documented Co-editor accepts natural-language changes and image or video attachments, can chain requests, and lets the user review or undo the edit. State one bounded objective and what must remain to control cost and interpretation.
It creates a synthetic talking video from a prompt, actor or AI Twin, voice, optional AI Edit style, generated or uploaded B-roll, captions, music, sound, transitions, and cuts. The plan can be reviewed before generation and the final project remains editable.
The current meter lists 16 credits per six seconds, about 2.67 per output second. A 30-second web result is roughly 80 base credits before additional actors, AI Edit styling, generated media, or reruns.
The current web-specific guide says up to 30 seconds with generated segments up to four seconds. An older cross-device matrix lists one-minute AI Creator projects. Treat the active iOS or web workflow as authoritative.
The current Prompt to Video FAQ says 39 languages. That number is separate from Captions' larger caption-recognition inventory and the smaller Lipdub list.
The web flow lets a user select several actors, but each actor produces a separate video and consumes additional credits. It is a paid comparison or variant workflow, not one multi-character scene by default.
Yes. Prompt to Video includes a voiceover-only option and can combine the narration with custom or generated media. Generated inserts and optional AI Edit styling can add credits beyond the base voice-led project.
The current general meter lists one credit per output second for each. A 60-second base generation starts at 60 credits before actor regeneration, generated media, chat, or other additions.
Forty credits. Twelve actor regenerations consume 480 of Max's 500 monthly credits, so test a short script, actor, voice, framing, and delivery before committing to a long or high-volume series.
Yes according to current release notes. Captions can use a product URL in its AI Ads path. The user must verify product visuals, price, offer, availability, claims, endorsements, disclosures, trademarks, and regional legality.
The current meter lists Flux variants, GPT-4o Image, Ideogram, Imagen, Nano Banana, Photon, Recraft, and Stable Diffusion variants at one to six credits. Model access depends on plan and can change.
Current documentation lists Veo, Grok Imagine Video, Hailuo, Kling, MAGI, Minimax, Pika, Ray, Seedance, Sora, and Seedream entries. Published charges range from eight to 140 credits, but duration and output properties are not uniform.
Kling 2.1 is currently 140 credits in the official meter, or 28% of Max's monthly allowance per call. Google Veo 2 is 100. Rates and availability can change, so inspect the live generation quote.
The current model table lists Soundraw at two credits, Lyria 3 Pro at three, and ElevenLabs Music at ten. The general GenAI Music range elsewhere says one to two, showing why model-specific entries should override summaries.
Current entries list audio isolation at two credits per minute, common sound-effect generation at two, Stable Audio Open at one, and MMAudio at four. A full AI Edit can invoke audio alongside other metered assets.
Yes. Credits are deducted when AI generation runs whether or not the result is saved, exported, or published. Abandoned prompts, retries, chat actions, and unused generated assets still count.
An August 2026 update lets users ask for more or different short clips from the same processed long video without starting over. It makes candidate discovery iterative, but every proposed clip still needs contextual and visual review.
Newer product material uses Clips, AI Shorts, and Long to Short around the long-source-to-social workflow. The center is extractive repurposing, distinct from AI Edit's styling and Prompt to Video's synthetic creation.
The current Clips guide says sources can be up to three hours or 10GB and may return 30 or more candidates. Import and URL availability, processing, crop, context, captions, speaker detection, and export should still be tested.
Captions says it detects multi-speaker conversation and reframes speakers in the Clips workflow. Review shot switching, active-speaker crop, names, overlap, context, and captions on a real panel or podcast.
The current AI Twin guide says up to 30. A Twin can be deleted, but exported videos remain. Organizations should define who can create, use, share, and remove likeness and voice assets.
Yes. August 2026 release notes add video-upload creation alongside the photo and guided-recording paths. The source subject's permission and the intended scope of reuse remain the user's responsibility.
The August 2026 release lets users create and save alternate looks for an avatar and reuse them in Prompt to Video and other avatar projects. Each look still needs identity, wardrobe, brand, and continuity review.
It is created during AI Twin setup from a recommended quiet one-minute calibration and appears as My Voice for editor narration, dubbing, and Prompt to Video. Recording quality, pace, emotion, and consent affect both output and acceptable use.
Yes. The voice clone can be selected for supported dubbed languages so the translation retains a version of the user's vocal identity. A fluent native speaker should review pronunciation, meaning, pace, and disclosure.
The current page lists 29 language variants, including simplified and traditional Chinese and Hinglish. This is smaller than the 100+ caption-recognition claim, so feature-specific language lists matter.
The caption workflow includes a Multiple Languages source option and a large recognition list. Translation, dubbing, and Lipdub have narrower target-language lists and should be tested separately.
Yes, especially on iOS. It supports script generation or paste, continuous or scene-by-scene teleprompter modes, camera settings, countdown, recording, Eye Contact, captions, and export. Device parity is incomplete.
The current camera guide describes HD or 4K and 30fps or 60fps where the device supports it. That recording option does not make every editor, generated-video, web, or export path a universal 4K/60 entitlement.
Newer 2026 release notes say yes for supported rollout paths: real-time sync, iOS-to-web availability, rolling web-to-iOS, native uploaded media, and cloud-backed templates. Older help still says no. Export masters and test your exact two-device direction before relying on it.
Older and still-live help warns that local projects can disappear after logout, cache clearing, browser changes, or device switches. New sync is staged and may not cover legacy or non-native paths. Contact support and preserve exports before changing sessions.
Captions is migrating from a legacy desktop web product to a newer captions.ai experience. New legacy signups are disabled, but existing users may still need desktop.captions.ai for their original projects and subscription during transition.
Yes again as of July 2026, according to the changelog. It brings the current web experience to macOS. Captions does not describe it as a full native professional NLE, and legacy desktop project behavior still requires care.
Yes. Captions has team workspaces, real-time co-editing, shared projects and assets, Admin and Member roles, seat management, and in-app invitations. New accounts start in a team workspace by default as of July 2026.
Admins can add or remove members, change roles, manage billing, access all team projects, and adjust team-wide settings. Members can create, edit, and share but cannot manage billing or roles.
Their access to team projects, assets, and billing is revoked immediately and the seat can be freed. Transfer necessary ownership and preserve avatar, voice, prompt, media, and project assets before removal.
Yes. Current Mirage-backed API documentation covers stylized video captions and video generation through API keys and asynchronous jobs. Consumer-plan credits, input limits, and export entitlements should not be assumed to apply.
Current examples use api.mirage.app with an x-api-key header, reflecting Captions' parent Mirage platform. Use the current dashboard and reference rather than older copied Captions endpoints.
The current per-organization rate-limit page lists 100 requests per minute for Video Captions and two per minute for Video Generation. Enterprise customers can ask an account manager for higher limits.
Current docs require vertical 9:16 MP4 or MOV, up to 50MB and five minutes, plus a caption-template ID. Jobs return queued, processing, complete, failed, or cancelled states before content download.
An older official getting-started page lists 360 initial credits and $0.05 per credit with purchases from 100 to 12,000. Because endpoint branding and rate limits have changed, verify current dashboard pricing and contract terms before forecasting.
Do not assume so. The consumer app bundles monthly credits with its plans; the API has a separate dashboard, key, purchase model, limits, and terms. Confirm whether any organization contract pools them.
The standard Terms say users own their Input and Output as between the parties, subject to Captions and licensor rights in Captions Content included in the output. Users remain responsible for rights, claims, consent, and publication.
The standard Privacy Policy includes using information to train AI or machine-learning models, and the Terms grant a service-improvement and development license. Enterprise publicly advertises training-data exclusion. Obtain the exact contract and data settings for sensitive footage.
The current public pricing page lists training-data exclusion among Enterprise benefits. Procurement should obtain the actual scope, effective date, models, subprocessors, retention, deletion, and DPA in writing rather than relying only on the card.
Yes—and that is a risk. Captions warns that permanent account deletion removes in-app projects but does not automatically cancel an Apple, Google, or web subscription. Cancel billing separately first and keep the receipt.
No. The deletion guide says videos already exported to the user's device remain there. In-app projects and the account are permanently removed and cannot be recovered.
Cappy is Mirage's separate SMS-based AI video editor. Send clips, photos, a product page, voice note, article, or just an idea; it makes a draft and accepts follow-up edits through ordinary text messages.
No. Co-editor changes an existing Captions project with documented message and action credits, undo, and manual fallback. Cappy starts and delivers through SMS and does not publish the same project-state or metering details.
No. Captions says Cappy works through text messaging and requires nothing to download. A phone, supported messaging route, and the media or idea are enough to begin.
Official pages list video clips, photos, product-page links, voice notes, articles, scripts, and text descriptions. Cappy can also start from an idea without an uploaded video.
Captions calls Cappy free to try for a limited time while it is in beta. It does not publish the continuing price, included jobs, revision charge, file limits, or relationship to Captions plan credits.
Captions says most videos are ready within a few minutes and more complex edits may take longer. That is vendor guidance, not a service-level guarantee; source size and revision scope are not publicly quantified.
Yes. Follow-up messages can request pacing, captions, cuts, music, hooks, or other changes. Public pages do not document exact revision cost, action history, deterministic undo, or how long the conversation remains available.
Yes. It can start from an idea, article, voice note, photos, or description and create supporting visuals and a finished draft. Facts, rights, product terms, and generated-media provenance still need review.
The public Cappy pages do not say. Do not assume the beta SMS product shares Max or Scale credits, rollover, projects, team roles, export entitlements, or data settings until the account or written terms confirm it.
It is Captions' proprietary model for AI avatars and Twins. Captions says it generates voice, face, expression, eye contact, hands, and body together as one performance rather than merely lip-syncing a still image.
The current avatar page says 225 ready-made avatars and summarizes the library as 200+. Users can also prompt a custom synthetic character or create an AI Twin based on a real person.
The current landing page says about ten seconds of footage. Captions also supports a fresh recording, uploaded video, or selfie and recommends video when the strongest identity and motion match matters.
Yes. The Avatar X page says a user can describe a custom synthetic presenter and save it alongside library avatars. A prompt-built character is distinct from an AI Twin that represents a real person.
Captions says its avatars can be used for personal or professional commercial projects subject to its terms and content rules. That does not grant rights to a script, product claim, music, trademark, reference asset, or real person's likeness.
Mirage reported that Avatar X scaled across 100+ segments, but two factual errors aired, long-form synthetic audio sounded rhythmically flat, and average watch time was around one minute despite more than three million impressions.
Mirage says the 24-hour experiment used roughly $50,000 in tokens on a single machine. That was a custom orchestrated broadcast with licensed reporting and many controls, not Captions subscription pricing.
The presentation was synthetic, including AI-generated versions of real guests. Mirage says each real guest gave written permission and approved talking points, and every frame carried an AI-generated disclosure.
Yes. Mirage says two errors reached air, including a wrong identity caption, and were corrected publicly within an hour. This occurred despite licensed reporting, scripted shot mapping, and automated technical and rights checks.
Mirage announced in March 2026 that independent auditors completed a SOC 2 audit. The public announcement does not provide the full report scope, type, period, criteria, exceptions, or exact products covered.
Request the current SOC 2 report under NDA, type and review period, Trust Services Criteria, system boundary, exceptions, subservice treatment, complementary user controls, bridge letter, DPA, subprocessors, retention, deletion, and incident terms.
Captions is better for a standardized creator workflow centered on presenter footage, AI styles, synthetic actors, localization, long-to-short, and mobile recording. Valmera is better when unique uploaded footage and an open-ended editorial brief are the center of the job.

Related Articles

Valmera vs Captions
Open-ended editing agent versus creator-focused AI suite.
Best AI Video Editors
Compare agents, creator suites, clippers, and manual editors.
Edit Talking-Head Videos
A complete pacing, caption, audio, and visual workflow.
AI Video Editing Cost
How credits, subscriptions, and repair time change the price.