Captions App Review 2026: Pricing, AI Edit V3, Clips Chat, and the Limits That Matter
Captions has grown far beyond automatic subtitles. It now combines multi-clip AI Edit V3, prompt-based revisions, the separate Cappy SMS editor, long-to-short Clips Chat, Avatar X Twins and actors, localization, teams, API workflows, recording, and a manual editor. This review separates those paths, converts current model rates into capacity, and follows the product through project sync, data terms, delivery, and deletion—not merely the first polished generation.
Is Captions App worth it?
Yes—for creators who repeatedly turn short, vertical, presenter-led footage into styled social video, or generate that presenter with an avatar. AI Edit coordinates cuts, captions, B-roll, music, motion, and effects; chat makes targeted revision accessible; Clips Chat handles iterative long-to-short repurposing. Horizontal and multi-clip support broaden the product, but the weaker fit remains bespoke long-form, multicamera, documentary, music-led, or picture-first editing. Credits make cost depend on models, generation, and retries; project sync, plan names, and rollover also have conflicting official pages.
Captions in 2026: The Facts That Decide the Purchase
Free, $24.99 Max with 500 credits, and Scale at $69.99/$139.99/$279.99 with 1,400/2,800/5,600; Enterprise is custom
$9.99 Basic with 200 credits and Android-only $4.99 Lite remain documented even though the public price card centers Max and Scale
Current August credit documentation allows Basic and higher to bank up to 3× the monthly allowance; cancellation and some downgrades forfeit credits
10–40 credits before model-dependent extras; 116 documented styles, multi-clip support, horizontal output, project duplication, and shot-level reruns
16 credits per six seconds in the current meter; the web-specific guide says up to 30 seconds, 39 languages, and four-second generated segments
Current model table spans 1-credit images to 140-credit Kling 2.1 video calls, with several Veo variants at 36–100 credits
Long sources up to three hours or 10GB, 30+ candidates, and August 2026 Clips Chat for requesting more or different cutdowns
2026 release notes say sync is rolling out across iOS and web; older help still says projects are local and can vanish on logout or cache clearing
Admin and Member roles, shared projects and assets, seat billing, immediate access revocation on removal, and team workspaces by default for new users
Mirage-hosted endpoints now expose caption rendering and video generation; current per-organization limits list 100 caption and 2 generation requests per minute
Standard Terms grant a broad service-improvement license over Input and Output; the Privacy Policy includes AI/ML training uses, while Enterprise advertises training-data exclusion
Current public overview says 1080p publishing; other platform matrices and export controls expose 4K, HDR, SRT, bitrate, and frame-rate only on qualifying device/project paths
A separate SMS-based editor in beta: send clips, photos, a product link, voice note, article, or idea; receive a draft and revise by texting; free only for a limited trial period with no public ongoing price
Current first-party pages claim a 225-avatar library, prompt-built custom avatars, and AI Twins from about 10 seconds of footage, with audio, expression, face, hands, and body generated as one performance
Mirage says its 24-hour synthetic news stream aired 100+ segments but still produced two factual errors, drew about one-minute average watch time, and exposed long-form audio-cadence weaknesses
Mirage announced completion of a SOC 2 audit in March 2026; enterprise buyers should request the report, type, period, criteria, scope, exceptions, subservice treatment, and bridge letter
How Much Is the Captions App?
The shortest answer is $24.99 per month for Max, the plan Captions foregrounds for its signature AI workflow. The full answer starts at $0, includes an Android-only $4.99 Lite tier and a $9.99 Basic tier in official help material, then rises through Scale allowances from $69.99 to $279.99 per month.
Prices below are published USD monthly figures checked August 26, 2026. Captions says final pricing can vary by monthly versus annual billing, App Store or Play Store versus web checkout, country, currency, and available account offer. The price shown before purchase is authoritative.
| PLAN | PUBLISHED MONTHLY PRICE | AI CREDITS | PRACTICAL POSITION |
|---|---|---|---|
| Free | $0 | 60–200 one-time credits in the current help guide; no recurring allowance | Basic captions, camera, trim, zoom, media, voiceover, transitions, resize, and basic export; current release notes say Free output carries a Captions watermark |
| Lite (Android) | $4.99/mo | No published renewable AI-credit bucket | Android-only manual editing tier; it cannot be restored on desktop or iPhone without a separate Basic, Max, or Scale subscription |
| Basic | $9.99/mo | 200 credits / month | Formerly Pro; watermark-free Basic projects, with Max-only AI Edit, chat, generated media, and AI Creator excluded |
| Max | $24.99/mo | 500 credits / month | AI Edit styles, chat editing, full-video creation, digital twins or custom actors, and generated media |
| Scale | $69.99/mo | 1,400 credits / month | Everything in Max, higher volume, and Captions' more sophisticated generative model tier |
| Scale 2x | $139.99/mo | 2,800 credits / month | Twice the standard Scale usage for a higher-volume individual workflow |
| Scale 4x | $279.99/mo | 5,600 credits / month | Largest published self-serve allowance; eligible for top-up credits under the current policy |
| Enterprise | Custom | Custom | Custom seats, bulk-credit discounts, onboarding, support, and training-data exclusion |
Why two official pages appear to disagree about Free
The public pricing comparison labels Free as having no monthly AI allowance. The more detailed credit and subscription guides describe 60–200 one-time credits that do not refresh. Those statements can both be true: Free has a limited initial grant, not a renewable monthly pool. Treat the live account meter as the final answer for your signup.
What 500 Captions Max Credits Actually Buy
Paid tiers bundle credits at about five cents each, but that does not make each video predictable. The official meter charges the first generation, chat revision, regeneration, and generated asset separately. Credits are consumed even when a project is never exported.
| ACTION | PUBLISHED BASE RATE | MAX-PLAN EXAMPLE |
|---|---|---|
| AI Edit | 10–40 credits | About 12–50 base AI Edit runs from Max's 500 credits |
| AI Creator / AI Ads / AI Skits | 1 credit / second | A 60-second generation starts around 60 credits |
| Prompt to Video | 16 credits / 6 seconds | A 60-second result uses about 160 credits before revisions |
| Chat-based edit | 2 per message + 1 per action | A one-action request starts at 3 credits; generated media costs extra |
| Actor regeneration | 40 credits | Twelve regenerations use 480 of Max's 500 monthly credits |
| Generated image | 1–3 credits | Model choice changes the cost |
| Generated video | 1–20 credits in the general table | Model-specific entries can be higher, so inspect the live meter |
| AI Edit style change | 1 credit | Targeted restyling is much cheaper than regenerating the full edit |
These examples are arithmetic, not a promise of output. AI Edit varies with duration, model, and algorithm choices; generated video has model-specific rates that can exceed the general add-on range. Check the meter after a representative project before forecasting production volume.
Need an Agent for the Footage, Not a Style for the Presenter?
Upload footage to Valmera, describe the finished edit, inspect the rendered result, and revise in conversation. Uploads are free; editing requires a subscription.
Create an account →Is Captions Free? Trial, Export, and Watermark Rules
Yes—basic editing
Captions, camera, trim, zoom, sounds, media, voiceover, transitions, resize, and basic exports. Current release notes say Free output includes a Captions watermark.
Checkout-specific
Current billing help explains trial cancellation, but the public pricing and plan pages do not promise one universal trial. Confirm the exact end date and renewal price in checkout.
Free plus entitlement-sensitive
Current Free output is watermarked. Inactive paid subscriptions or Max-only features used on Basic can also produce a watermark according to the July 2026 troubleshooting guide.
Automatic until canceled
Cancel through Apple, Google, or Captions' web billing portal at least 24 hours before renewal. Direct monthly web plans are described as non-refundable.
Safest test: make and export one free-only project before activating a paid feature or checkout trial. After adding AI Edit, Co-editor, AI-generated media, or AI Creator, inspect the plan badge before export rather than assuming the project still qualifies for the free path.
What Captions App Actually Is in 2026
“Caption app” now understates the product. Captions is a creator production suite with three different starting points: record a presenter in its camera; import existing footage for manual, automatic, or chat-based editing; or generate a presenter and supporting media from a prompt. A fourth workflow, Clips, mines long recordings for short social moments. Cappy is a fifth, separate surface that accepts media or an idea and returns revisions through SMS.
That breadth is the advantage and the source of confusion. AI Edit is optimized for a short, vertical, single-speaker source. Clips can accept a much longer podcast. AI Twin creates new presenter footage. Translation and Lipdub localize a finished performance. They live in one product, but they do not share the same limits, meter, or editorial purpose.
Captions Features Reviewed by the Job They Perform
AI Edit
The purchase-defining featureCaptions applies cuts, captions, B-roll, transitions, music, sound effects, motion graphics, and a chosen style in one pass. AI Edit V3 now supports multiple clips and horizontal output, while current workflow guidance still says short, clean, presenter-led footage is the safest source.
Chat-based editor / Co-editor
Useful for targeted revisionThe chat interface accepts natural-language changes and applies actions to the project. You can still open the manual editor afterward, which makes chat a revision layer rather than an irreversible black box.
Cappy SMS editor
The lowest-friction outcome interfaceCappy accepts clips, photos, a product link, voice note, article, or description, asks follow-up questions, creates the structure and style, returns a video, and continues revising through SMS without an installed app.
Clips
A materially broader long-form featureCaptions says Clips accepts an upload or URL up to three hours or 10 GB, detects multi-speaker conversation, reframes speakers, avoids doubling existing captions, scores proposed clips, and can produce 30 or more candidates. August 2026 adds Clips Chat for requesting more or different cutdowns from the same source.
AI Twin and actors
Strong for repeatable presenter contentAI Twin creates a reusable digital version of the user with a chosen or cloned voice; the current help center says an account can create up to 30 Twins. Photo, guided recording, and August 2026 video-upload setup are available, alongside reusable Avatar Looks and library actors.
Captions, dubbing, and translation
A deep localization bundleAutomatic captions are available across current device platforms. Captions advertises recognition across 100+ languages, while translated captions, dubbing, Prompt to Video, and the 29 listed Lipdub language variants each have narrower, different inventories.
Camera, teleprompter, and eye contact
Excellent creator utilitiesThe recording workflow combines a camera and teleprompter, while Eye Contact adjusts gaze after recording. The eye-contact tool is documented across desktop, iOS, and Android.
How Captions Actually Works: Credits, AI Edit, Avatars, Teams, API, and Data
The following analysis reconciles the live price card, the August 2026 credit guide, the current changelog, feature references, API documentation, and legal pages. Where an older official help article contradicts a newer source, the conflict is named and the safer operating rule is stated.
The live pricing page and the detailed subscription guide serve different purposes. The public page sells Free, Max, three Scale multipliers, and Enterprise, and explicitly says its prices reflect iOS plans. The help center additionally documents Android-only Lite at $4.99 and Basic at $9.99 with 200 credits. A January release note says Pro and Max were renamed Starter and Creator, yet the current August pricing page and credit guide again use Basic and Max. The safe buying rule is to quote the active checkout and entitlement screen, not infer a plan from an older review or name.
Max and every Scale multiplier price included credits at exactly five cents apiece: $24.99 for 500 is just under five cents, while $69.99/1,400, $139.99/2,800, and $279.99/5,600 differ only by rounding. Scale does not buy cheaper allowance; it buys more of it and access to the most sophisticated generative model tier. A buyer should therefore choose Scale for model access or measured volume, not because the headline bundle automatically improves the credit unit economics.
Current August documentation permits Basic and higher to carry unused monthly credits until the balance reaches three times the plan allowance. Max can therefore hold up to 1,500, standard Scale 4,200, Scale 2x 8,400, and Scale 4x 16,800. A downgrade immediately reduces the balance to the new cap and forfeits the excess. Cancellation lets credits remain through the paid period and then removes them. Upgrade timing, downgrade timing, and the planned generation backlog are part of the purchase decision.
A March troubleshooting article still says credits do not roll over. The canonical Understanding Credits page, last modified August 18, 2026, and the live pricing FAQ both say they do, capped at three times the monthly amount. The newer, more specific sources should govern this review. The contradiction is still operationally relevant because a support agent or product surface could be in transition. Preserve a screenshot of the balance and rollover language before making a downgrade or allowing a large saved balance to cross renewal.
AI Edit costs 10–40 credits in the general table before generated-media additions, but the range is not a quoted price for a finished video. Duration, style, model choices, algorithmic decisions, reruns, chat, images, video, music, and sound can all change the total. Max's 500 credits theoretically fund 12 full 40-credit base runs or 50 ten-credit runs. Neither number includes rejected edits or generated B-roll, so production capacity should be measured from the credit history after several representative jobs.
The current style guide documents 116 AI Edit styles: 21 Premium and 95 Basic. A style controls caption typography, motion graphics, B-roll strategy, music and sound, pacing, and the overall aesthetic. It is closer to a coordinated art-direction preset than a simple caption theme. Captions lets users preview, change the style on one shot, apply a custom brand color, replace media, edit words and timing, or request changes through chat. That editability is the difference between a useful starting system and a locked template.
AI Edit's source envelope is broader than the older vertical-single-clip description, but it remains format-sensitive. Release notes document 16:9 horizontal output, multi-clip sessions, multi-clip merging, importable voiceover in the edit plan, and long-to-short suggestions. The current workflow guide still says one clean speaker, vertical footage, and roughly 60–90 seconds produce the safest results. Product support and product design center are not the same: horizontal and multiple clips now work, but talking-head social footage remains the best benchmark.
The older device matrix still lists a one-minute web and two-minute iOS AI Edit ceiling. Newer release notes add multi-clip and long-to-short behavior without publishing one replacement maximum for every account. That makes duration an in-product entitlement rather than a reliable universal number. Test the exact device, project type, source count, combined duration, orientation, and plan. Do not turn a successful three-hour Clips upload into a claim that ordinary AI Edit accepts a three-hour timeline.
AI Edit V3 can create a first assembly of transcript cuts, captions, generated or selected B-roll, music, sound effects, transitions, zooms, motion, and visual styling. Afterward, individual shots can be rerun or preserved, and duplicated projects make alternate direction safer. This is meaningful automation. It does not establish multicamera synchronization, professional color management, nested timelines, source relinking, track locking, AAF/XML interchange, or a deterministic offline workflow. Those remain separate evaluation questions.
Chat-Based Editing charges two credits for the user's message, one for every action the system takes, and model costs for any generated asset. A message that changes one caption color starts at three credits; a broad request that replaces several shots and adds generated media can cost much more. The efficient pattern is to inspect first, ask for one bounded outcome, state which existing elements must remain, and verify the action list. A vague instruction is both harder to reproduce and more expensive to debug.
Prompt to Video is a different product path from AI Edit. It builds a synthetic talking performance, optional B-roll, music, transitions, sound, captions, cuts, and zooms from a prompt, selected actor, voice, and style. The plan can be reviewed before generation; a voiceover-only option removes the face; custom media can replace generated insertions. Changing the prompt creates a new plan, selecting several actors produces separate videos, and adding generated media or AI Edit styles increases the credit charge.
The current credit meter charges Prompt to Video at 16 credits per six seconds, about 2.67 credits per output second. A 30-second web result is therefore roughly 80 base credits before actors, edits, or additional media; Max can fund about six such base generations. The web-specific guide says a final video can be up to 30 seconds, a segment up to four seconds, the prompt or audio boundary is 30 seconds or 1,000 characters, and 39 languages are supported. Other device matrices describe one-minute creator projects, so the active path matters.
AI Creator, AI Ads, and AI Skits use the simpler published rate of one credit per output second. A 60-second generation starts at 60 credits. Actor regeneration costs 40 credits independently, so one 60-second video plus three full actor retries can reach 180 before model add-ons. AI Ads can ingest a product URL according to current release material, but the user must still verify price, offer, availability, trademarks, product appearance, disclosures, testimonials, and every claim the synthetic presenter makes.
The standalone model table makes 'generated video' a poor capacity unit. Ray 2 Flash and Hailuo 2.3 Fast are currently eight credits, Grok Imagine Video 11, Sora 2 and MAGI Distilled 16, Pika 20, Seedance variants 48–61, Veo variants 36–100, and Kling 2.1 140. One Kling 2.1 call consumes 28% of Max. Model duration and resolution are not uniform in the summary, so compare the actual quoted generation, accepted seconds, and retries rather than dividing credits by a generic clip count.
Image generation is cheaper but still model-dependent. Several current options cost one credit, common models cost two or three, and Nano Banana Pro costs six. A 20-shot edit with one two-credit image per shot is 40 credits; three variants per scene would be 120. Generated music ranges from two or three credits for current listed options to ten for ElevenLabs Music, while audio isolation is two credits per minute and several sound-effect models cost one to four. A polished automatic pass is the sum of many small meters.
Credits are charged when the AI work happens, even if the project is never exported or saved. Unlimited eligible exports therefore do not mean unlimited creation. Failed creative ideas, abandoned ads, actor retries, duplicate prompt plans, and chat actions still consume the balance. On iOS, the credit history exposes the transaction record. Use it as a production ledger: record source, requested operation, model, charge, accepted output, and whether the asset reached a published video.
When the balance reaches zero, current troubleshooting says ordinary AI Edit and related AI features become unavailable until renewal or upgrade. AI Creator can continue through a slower queue with waits up to one hour. Individual top-ups are not offered on normal personal plans; current top-ups are reserved for existing Teams or Business customers and Scale 4x. Scale 4x is therefore the only published self-serve tier with a path to more credits without a full plan change, and that entitlement should be verified in the account.
Clips is the long-form extraction product. It accepts an upload or URL up to three hours or 10GB, can identify multiple speakers, reframe for social, avoid doubling existing burned-in captions, score candidates, and return 30 or more options. August 2026 adds Clips Chat, letting the user request more or different shorts from the same source without restarting. That makes the extraction loop conversational, but the score remains a ranking hypothesis. Every candidate needs context, crop, caption, speaker, hook, and rights review.
Captions' product vocabulary is shifting from Clips toward AI Shorts and Long to Short in newer help material. The functional center remains the same: turn a long upload or YouTube link into social derivatives. It should not be conflated with AI Edit's automatic finishing or Prompt to Video's synthetic production. A podcast team may use all three—Clips to select moments, AI Edit to style an accepted short, and chat to repair it—but each stage can impose its own limits and credits.
AI Twin currently supports up to 30 digital clones. Setup can begin from a photo, a guided recording, or—since August 2026—a video upload, and users can create saved Avatar Looks for later projects. A Twin can be paired with a selected voice or a voice clone and reused in Prompt to Video. Deleting the actor is documented, but deletion does not remove already exported media. Consent, employment, campaign, and post-relationship use need a policy outside the feature itself.
The voice clone is created through the Twin setup and can narrate editor voiceovers, dubbing, and Prompt to Video. Captions recommends a quiet one-minute calibration at a natural pace, because cadence, emotion, reverb, and microphone distance affect the clone. The availability of a voice in the app does not settle whether it may be used for ads, regulated communications, impersonation, politics, or after consent is withdrawn. Preserve the approval and create an asset-removal process.
Lipdub changes mouth and facial movement to align with translated audio. The current help page lists 29 language variants, while automatic captions support a much larger language inventory and ordinary translation or dubbing has a different list. '100+ languages' describes caption recognition, not one universal feature matrix. For every target language, test transcription, translation, pronunciation, voice identity, mouth motion, onscreen text, timing, and cultural adaptation with a fluent reviewer.
Camera and Teleprompter make the iOS experience especially coherent for presenters. A script can be written or generated, shown in continuous or scene-by-scene scroll mode, recorded in HD or 4K at 30 or 60fps where supported, and then passed through Eye Contact and captioning. This saves round trips for solo creators. Web and Android do not mirror every camera, keyframe, voiceover, project, or HDR control, so a creator-suite review must remain device-specific.
The current export story has multiple layers. The product overview recommends 1080p and direct social publishing; the older device matrix lists up to 4K on iOS and web and 2K on legacy desktop; advanced controls can include frame rate, bitrate, Smart HDR, and SRT depending on source, device, plan, and project level. Free users can export with a watermark after the September 2025 policy change, while eligible paid projects can export without one. Run the final target settings, not a generic export test.
Project continuity is in transition. February release notes announced real-time sync, April said iOS-to-web was available to all users with the reverse direction still rolling out, June added cloud-backed caption templates, and July added native video uploads to keep projects synchronized. Older March and April help articles still warn that projects are local, disappear on logout or browser-cache clearing, and do not appear on another device. The safest interpretation is that sync exists for newer supported paths but is not universal across legacy desktop, web, devices, imports, and account states.
The web migration adds another continuity risk. Current sign-in help distinguishes legacy desktop.captions.ai from the new captions.ai web experience and warns that some sign-in paths may not load the expected subscription, projects, or workflows during the transition. July disabled new legacy-desktop signups and restored a macOS app representing the web experience. Existing desktop users should bookmark their correct surface, confirm the account method, export masters, and validate a copied project before clearing a browser or switching products.
Teams are now an actual collaboration layer rather than a future enterprise promise. New users enter a team workspace by default, Admins can manage people, billing, roles, and all team projects, while Members can create, edit, and share without administrative control. Removing someone immediately revokes access to projects and assets and can free a paid seat. Before offboarding, transfer ownership, preserve reusable media, verify Twin and voice-clone access, and export anything that must survive the account relationship.
Captions' API has moved toward Mirage branding and current endpoints at api.mirage.app. The video-captions API accepts vertical MP4 or MOV up to 50MB and five minutes, requires a caption-template ID, returns an asynchronous job, and exposes status and content retrieval. Current rate limits are 100 caption requests and two video-generation requests per minute per organization. These constraints are not the same as the consumer app's import, duration, project, or export rules.
Older API documentation advertises 360 initial credits, $0.05 per credit, purchases from 100 to 12,000 credits, one credit per second for AI Creator and Ads, and five requests per minute. The newer rate-limit page now says two video-generation requests per minute, and the product is branded Mirage in current examples. Treat the dashboard and current contract as authoritative for price, free credit, endpoint, model, rate, data, and rights. Do not project consumer-plan credits onto API calls.
Standard legal terms say the user owns Input and Output as between the parties, subject to Captions and licensor rights embedded in Captions Content. They also grant Captions a limited, nonexclusive, royalty-free, sublicensable license to host, store, modify, adapt, publish, translate, create derivatives, distribute, provide, improve, and develop services from that Input and Output, with the license surviving termination. Commercial buyers should read this alongside model-specific terms, asset licenses, and the Acceptable Use Policy.
The standard Privacy Policy states that information can be used to train AI or machine-learning models, improve services, infer interests, and—under described circumstances—be disclosed to third parties for their business purposes, including training algorithms. The public Enterprise offer explicitly lists training-data exclusion, which suggests standard and contracted treatment can differ. A company with confidential recordings, biometric likeness, unreleased products, children, health data, legal privilege, or client restrictions should obtain the applicable DPA and written data settings before upload.
Deleting a Captions account is permanent and removes in-app projects, but the current deletion guide warns that it does not cancel an active subscription. Apple, Google, and web billing have separate cancellation routes. That separation can leave a recurring charge after access is destroyed. Cancel with the actual payment provider, record the effective date, export projects and transcripts, remove team seats and API keys, delete the account only after verification, and retain the cancellation receipt.
A serious Captions evaluation should run four different tests rather than one polished demo: a 60-second presenter clip through AI Edit; a long real conversation through Clips and Clips Chat; a 30-second Prompt to Video using the intended actor and voice; and one manual or chat repair. Record device, plan, source, duration, model, credits, render time, sync behavior, export settings, watermark, human correction time, and accepted result. This converts a sprawling feature list into a defensible monthly capacity decision.
Cappy is a separate outcome-level editor, not another name for the in-project Co-editor. The official workflow starts in a phone's messaging app: a user can send clips, photos, a product-page link, voice note, article, or only an idea; Cappy asks questions, makes structure and style decisions, returns a finished draft, and accepts follow-up text revisions. The delivered video can be saved back to the phone. This is strategically important because it moves Captions from a tool with chat controls toward an editor available as a conversation channel.
Cappy's convenience also creates a separate buying boundary. Captions calls it a beta that is free to try for a limited time, says most videos arrive within minutes, and provides no public continuing price, credit allowance, file limit, resolution, retention period, supported country list, or revision meter. Carrier message and data rates may apply, and texting enrolls the number for automated replies until the user sends STOP. A buyer should not infer that an existing Max or Scale subscription funds Cappy or that a beta conversation inherits the app's project, export, and team controls.
Cappy and Captions' in-app Co-editor solve different problems. Co-editor modifies an existing structured project, exposes manual fallback and undo, and has documented message/action credit rates. Cappy accepts a broader starting object and returns a video over SMS, but its public pages do not document an editable project handoff, action log, exact meter, model selection, or deterministic undo. Use Cappy for a low-risk phone-first brief; use the app when review, exact correction, reusable project state, and known credit accounting matter.
Mirage Avatar X is now the named model behind Captions' synthetic presenters. Current first-party pages say it generates voice, face, micro-expressions, eye contact, hands, and body as one performance rather than applying lip sync to a still. The public library page says 225 ready-made avatars, also summarized as 200+, and allows prompt-built custom characters. That is a much broader casting surface than the older generic actor description, but every realism and fidelity statement remains a vendor claim until tested with motion, emotion, difficult phonemes, continuity, framing, and the target export.
The current AI Twin landing page says Avatar X needs about ten seconds of footage and supports uploaded video, a fresh recording, or a selfie, while the help guide also documents a photo-based setup and up to 30 saved Twins. These paths may trade speed for fidelity: the company itself recommends video for the strongest identity and motion match. Ten seconds is an input claim, not proof of perfect ownership, consent, accent, voice, hand movement, emotional range, or long-form identity stability. Test the same script twice and compare face, cadence, gestures, outfit, background, and cross-shot continuity.
Mirage's own 24-hour News Network experiment is more useful than a curated avatar reel because it publishes operational evidence and failures. The company says one machine and roughly $50,000 in tokens produced more than 100 original segments around licensed newswire reporting. Scripts were mapped to shot lists; automated checks covered frozen frames, A/V sync, speech tails, and asset rights; real guests gave written permission and approved talking points; every frame carried an AI-generated disclosure. That is a strong example of the controls required around synthetic media, not evidence that an ordinary Captions subscription can reproduce the system or cost.
The same postmortem documents the limit clearly. Mirage says two factual errors aired, including a wrong identity caption, and both needed public correction. It also reports overly uniform cadence and poorly placed pauses in long-form audio. More than 50,000 people tuned in and the stream generated over three million impressions, but average watch time was around one minute. Those vendor-reported results demonstrate three separate facts: generated presentation can scale, truth still needs editorial verification, and technical realism does not guarantee sustained attention.
Mirage announced in March 2026 that it completed a SOC 2 audit with independent auditors. That is materially stronger than an unverified security slogan, but the public post does not substitute for the report. It does not state in the article whether the report is Type I or Type II, the exact review period, Trust Services Criteria, system boundary, exceptions, complementary user controls, or subservice-organization method. Enterprise procurement should request the current report under NDA, bridge coverage, penetration-test summary, DPA, subprocessors, retention, deletion, incident terms, and the precise products included.
Avatar X also raises a provenance distinction the UI must preserve. A library avatar is a licensed synthetic presenter, a prompt-built avatar is a generated character, and an AI Twin represents a real person. Captions says its ready-made avatars can be used commercially subject to the content rules, but that does not authorize a customer's script, product claim, trademark, music, reference asset, cloned voice, or real-person likeness. Keep the consent record, prompt, source, actor type, generated version, disclosures, approvals, and final export attached to each campaign.
What Captions Does Exceptionally Well
AI Edit has become more flexible without losing its design center
Horizontal output, multiple clips, project duplication, voiceover import, long-to-short suggestions, shot-level reruns, and new styles expand the workflow while the clean talking-head path remains easy to understand.
The style system is unusually explicit
Captions documents 116 coordinated AI Edit styles and explains which visual, caption, B-roll, music, motion, and pacing decisions each style controls. Buyers can preview, customize, or mix them by shot.
Credit rates are published down to individual models
Current documentation lists plan pools, feature charges, image models, video models, music, isolation, and sound. That enables real capacity math instead of treating every generation as one vague AI action.
Clips Chat makes long-to-short iterative
A user can ask for more or different short candidates from the same processed source in August 2026 rather than uploading again. This is a useful conversational refinement loop for podcast and interview repurposing.
Generated and recorded presenter workflows meet in one editor
A real camera performance, AI Twin, library actor, imported voiceover, cloned voice, Prompt to Video project, or voiceover-only production can all reach captions, media, sound, manual correction, and export.
Localization reaches beyond subtitle replacement
Caption recognition, subtitle translation, dubbing, voice cloning, and Lipdub provide several levels of localization. Teams can choose the least invasive operation appropriate to the release.
Cross-device and team infrastructure is actively improving
Project sync, cloud-backed templates, native uploads, shared workspaces, real-time co-editing, Admin and Member roles, and macOS support reduce the isolation of the original mobile-only workflow.
The API now covers both caption rendering and video generation
Organizations can add stylized captions to vertical video or generate presenter media programmatically, with asynchronous job status, downloadable output, API keys, and published rate limits.
Rollover rewards uneven schedules
Current Basic-and-higher accounts can bank two extra months of unused allowance, giving Max a ceiling of 1,500 credits and standard Scale 4,200 without losing every quiet month's balance.
The free surface can validate basic editing
Trimming, transitions, a caption template, resizing, and other basic tools let a buyer test import, transcription, timing, interface, and device compatibility before making the flagship generative features the purchase decision.
Human review remains possible after automation
Captions exposes caption correction, timing, shot style, media replacement, color, intensity, chat, timeline controls, exports, and original or unedited synthetic takes. The first pass is not the only available pass.
Enterprise states the key data difference plainly
Training-data exclusion is listed as an Enterprise entitlement rather than being implied for every customer. That makes procurement's next question clearer, even though the contract still needs review.
Cappy tests a genuinely low-friction editing channel
Clips, photos, a product link, voice note, article, or idea can enter through ordinary texting, and the user can iterate without learning an editor. It is a meaningful outcome-level interface, not just another toolbar.
Avatar X joins casting, voice, expression, and body motion
Captions now names the model behind a 225-avatar library, prompt-built characters, and short-recording AI Twins. Recorded and synthetic presenters can still reach the same correction and export environment.
Mirage published unusually candid long-run evidence
The News Network postmortem names its controls, token spend, two factual errors, audio weakness, and short average watch time instead of presenting only a highlight reel. Buyers can use those failures to design review gates.
The SOC 2 audit claim improves the enterprise starting point
Mirage says independent auditors completed a SOC 2 audit. Procurement still needs the actual report and scope, but the claim is more concrete than a generic enterprise-security badge.
The Limits That Matter More Than the Feature Count
AI Edit has a narrow ideal source
Official guidance favors one person speaking, vertical 9:16 footage, little or no prior editing, and a short duration. That is an excellent social format and a real boundary.
Platform parity is incomplete
Web, iOS, and Android are all supported, but trial, Lite-plan, duration, resolution, and feature availability can differ. A review based on one device does not settle another.
Credit cost appears after the headline
Generation, chat, actions, retries, and assets can each deduct credits. The subscription does not provide unlimited use of its flagship AI features.
Styles can converge on a house look
One of 116 style systems makes speed and consistency possible. It can also make the output feel template-led unless the user customizes captions, color, B-roll, pacing, and sound.
Synthetic video still needs review
An avatar can save a shoot, and localization can save a re-record. Neither removes the need to check likeness, consent, delivery, facts, translation, product details, and brand safety.
Export is not the only unit of value
A technically unlimited export allowance does not matter if credits run out, the source misses AI Edit's format, or every result needs extensive correction. Measure finished videos per month.
Current plan names conflict with the changelog
The live page sells Max while a January 2026 release says Max became Creator; the current credit guide calls the $9.99 tier Basic, formerly Pro or Starter. Checkout is authoritative, but comparisons become easy to misread.
The public pricing page reflects iOS
Captions explicitly says its displayed prices and features reflect iOS plans. Web, Google Play, legacy desktop, country, currency, tax, annual offers, and existing-account entitlements can differ.
Scale does not reduce the nominal price per credit
Max and Scale bundles all sit around five cents per included credit. Scale earns its premium through volume and model access, not a materially better allowance exchange rate.
Rollover documentation contradicts an older help page
The August authority permits balances up to 3×, while a March troubleshooting page says no rollover. The newer rule should win, but subscribers should preserve the account meter before a high-stakes billing change.
Downgrading can destroy saved credits immediately
A lower plan's 3× ceiling applies at downgrade and excess balance is forfeited. A user who banked Scale credits can lose them before the next normal renewal if timing is not planned.
Cancellation forfeits the remaining balance
Credits stay usable only through the paid term and disappear at expiry. An exported project does not preserve unused generation capacity, and an annual commitment does not make credits transferable.
AI Edit's 10–40 credit range is incomplete cost
Generated B-roll, images, video, music, sound, chat, style changes, and reruns can sit on top of the base charge. The monthly video count cannot be known from the plan table alone.
A premium video model can consume 28% of Max
The current Kling 2.1 entry costs 140 credits per call. Several Veo variants cost 36–100. A few rejected synthetic clips can crowd out an entire month of normal editing.
Multi-clip support does not equal long-form NLE support
AI Edit can merge and process several clips, but that does not prove multicamera sync, nested timelines, relink, interchange, advanced color, audio buses, or hour-long deterministic finishing.
Device matrices lag release notes
Older official tables still describe vertical-only, local projects, and different duration or feature limits even after 2026 releases added horizontal AI Edit, sync, multi-clip work, and web insertions.
Prompt to Video limits vary by surface
The web guide says a 30-second final with four-second segments, while other matrices describe one-minute creator projects. Actor count, prompt length, upload size, language, and available plan also differ.
Every selected actor creates another paid video
Choosing several actors in the web prompt workflow produces separate outputs and uses additional credits rather than giving one free comparison grid. Casting experiments should be budgeted as generations.
Actor regeneration is expensive
One retry costs 40 credits, so a few delivery, face, motion, or continuity problems can exceed the base price of the finished synthetic speech. A short casting proof is essential.
Synthetic ads can invent commercial facts
URL-based and prompt-driven ads still need product, price, availability, testimonial, disclosure, logo, demonstration, and jurisdiction review. A plausible synthetic presenter is not an approved claim source.
Virality and clip scores are hypotheses
Clips can rank many candidates and Chat can request alternatives, but no score knows the actual audience, campaign objective, context, legal sensitivity, or channel performance before publication.
AI Twin scale increases consent risk
Thirty reusable clones, upload-based creation, saved Looks, and cloned voice simplify production and make unauthorized reuse easier. Access, campaign, duration, geography, and post-employment rules need explicit ownership.
Language counts are not one entitlement
Captions recognition, translated subtitles, dubbing, Lipdub, and Prompt to Video have different supported lists. Marketing a 100-language caption tool does not promise 100-language lip synchronization or generation.
Project sync is not universal enough to trust blindly
Release notes describe staged iOS/web rollout and native-upload requirements, while older help still warns of local loss. Legacy desktop, cache, browser, account method, import path, and device direction can change continuity.
The web transition can split projects and subscriptions
Legacy desktop and the new captions.ai experience coexist. Signing in through a different method or surface can expose the wrong account, missing projects, or missing entitlement during migration.
Team removal revokes access immediately
Offboarding before asset and ownership transfer can strand a project, prompt, media file, actor, or clone. A seat-saving click is also a production-access event.
Free export wording changed
A newer release says Free exports carry a Captions watermark, while older article versions described clean output for free-only projects. Test the active project badge and a real download instead of carrying forward the old promise.
4K is not a universal export promise
One matrix exposes 4K on qualifying iOS and web paths, while the current overview recommends 1080p and legacy desktop lists 2K. HDR, source, project level, device, plan, bitrate, and frame rate affect delivery.
The consumer and API products use different limits
API captioning is vertical, MP4/MOV, 50MB, and five minutes with its own template and organization RPM. Consumer upload, Clips, AI Edit, and export ceilings cannot be projected onto it.
API documentation spans old Captions and new Mirage surfaces
Pricing, initial credits, rate limits, endpoint host, and product names have evolved. The dashboard and signed agreement must replace copied integration examples before launch.
Standard data terms are broader than Enterprise positioning
The standard Terms and Privacy Policy permit service improvement and AI/ML training uses; Enterprise advertises training-data exclusion. Confidential footage needs a written entitlement, not an assumption based on the sales page.
Deleting an account does not cancel billing
Captions explicitly warns that deletion can erase projects while leaving an Apple, Google, or Stripe subscription active. Billing cancellation must be completed separately before destructive account removal.
Input and Output ownership has qualifications
Users own their material as between the parties, subject to Captions Content and a broad surviving license granted for operating, improving, and developing services. Commercial and regulated teams should review the full contract.
Automatic polish can converge on a Captions look
The same styles, typography, generated B-roll logic, sounds, and pacing recur across many creators. A brand must replace assets, control intensity, correct captions, and define a repeatable treatment to remain distinctive.
Cappy has no public continuing economics
The SMS editor is free only for a limited beta period. Public pages do not disclose the eventual price, included jobs, revision meter, relationship to Max or Scale credits, or failure-refund rule.
SMS is a separate data and delivery surface
Phone number, carrier metadata, message history, attachments, generated drafts, and automated replies pass through a messaging workflow. Teams should clarify retention, deletion, subprocessors, country support, STOP handling, and whether projects enter the normal workspace.
Avatar realism can increase deception impact
A model designed to preserve identity, voice, expression, hands, and body can make an unauthorized or incorrect statement more persuasive. Consent, disclosure, claim review, access control, and source preservation become more important as realism improves.
The vendor's own broadcast produced factual errors
Two errors reached air despite licensed reporting, scripted shot maps, automated checks, written guest approval, and disclosure. Synthetic presentation needs human fact and identity review even in a heavily controlled pipeline.
SOC 2 scope is not visible in the announcement
The public post does not specify Type I or II, audit period, criteria, system boundary, exceptions, complementary controls, or subservice treatment. The report—not the badge—must answer whether the required Captions and Mirage surfaces are covered.
Who Should Choose Captions—and Who Should Not?
Choose Captions when
- Your recurring source is a person speaking to camera.
- The desired output is a short, vertical, polished social video.
- A strong automatic style is more valuable than designing every decision.
- You want avatars, digital twins, eye contact, captions, and localization together.
- You will reuse a long podcast through the separate Clips workflow.
- You can measure credit use on representative projects before committing annually.
Choose another editor when
- Your edit is long-form, multicamera, documentary, montage, or music-led.
- You need exact timeline interchange, collaborative project files, or professional finishing.
- The editor must interpret an unusual open-ended brief rather than apply a creator style.
- You require predictable unlimited AI generation without a usage meter.
- Platform parity across a mixed web, iOS, and Android team is mandatory.
- Your organization cannot use synthetic likeness or cloud generation under its policy.
Captions App Alternatives by Workflow
No single product is “Captions without the limits.” Each alternative changes the center of the workflow. Start with what enters the system and what must come out.
| ALTERNATIVE | CHOOSE IT FOR | HOW THE WORKFLOW CHANGES | DETAILED GUIDE |
|---|---|---|---|
| Valmera | Open-ended editing of uploaded footage | Describe the outcome, inspect a rendered preview, and revise the edit in conversation | Compare → |
| Descript | Precise transcript-led podcasts and interviews | Edit words and structure directly through a document-like transcript | Compare → |
| OpusClip | Dedicated long-to-short repurposing | Optimize clip extraction, reframing, hooks, captions, and batch social output | Compare → |
| Submagic | Repeatable captioned shorts | Use recognizable animated captions, B-roll, hooks, cleanup, and social finishing | Compare → |
| CapCut | Hands-on mobile and trend-led editing | Choose a broad manual social editor with templates, effects, and many utility AI tools | Compare → |
| VEED | Collaborative browser editing and brand workflows | Keep timeline, transcript, subtitles, generation, translation, stock, and review in one web workspace | Compare → |
Captions vs Valmera in one sentence
Captions is the stronger creator suite when a ready-made style, synthetic presenter, eye contact, and localization drive the purchase; Valmera is the stronger fit when uploaded footage and an open-ended editorial brief drive the purchase.
How to Test Captions App Before Paying for a Year
- 1Choose a representative source, not a demo-friendly clipUse a real 30–60 second vertical talking-head recording with a difficult name, one pause, imperfect audio, and at least one moment where relevant B-roll would help. This tests Captions' design center without wasting credits.
- 2Make a free-plan export firstAdd captions, trim, resize, and use only features listed in the free plan. Export before touching a paid feature so you can verify the free and watermark behavior on your exact device and account.
- 3Run two contrasting AI Edit stylesPreview styles first, then generate the same short segment with one restrained and one expressive style. Compare caption accuracy, cut timing, B-roll relevance, sound, motion, brand fit, and distraction—not just polish.
- 4Repair one result through chat and manuallyRequest three precise changes through Co-editor, record the credits used, then correct a caption and visual element by hand. Measure whether the product saves time after the first pass, not only during it.
- 5Test your second most important workflowIf you publish podcasts, run Clips on a real episode. If you localize, dub one minute and have a fluent speaker review it. If avatars matter, generate the same script twice and inspect continuity, delivery, and regeneration cost.
- 6Calculate a full month before subscribingMultiply your videos by their observed AI Edit, chat, asset, dubbing, and regeneration usage. Add the cost of human correction. Choose Max or Scale only if the measured month fits the allowance with room for retries.
A focused test takes about two hours plus generation time. The purchase decision should use exported files and observed credit history from your actual device, footage, and workflow.
How This Captions App Review Was Researched
We reviewed Captions' current product, pricing, help, feature, billing, and app-store material on August 26, 2026. We did not assign a fictional numerical score or claim hands-on laboratory testing that did not occur. Every material number on this page is tied to publisher-owned documentation; calculations are labeled as calculations.
- Price: current published USD monthly rates; platform and checkout variability disclosed.
- Free: ongoing features separated from a one-time credit grant and any checkout-specific trial offer.
- Cost: action-level credit rates converted into transparent examples without promising a fixed number of videos.
- Fit: official format guidance used to separate AI Edit, Clips, avatar, localization, and manual workflows.
- Conflict: Valmera publishes this review and competes for AI video-editing customers; Captions' advantages are stated because the right answer is sometimes Captions.
- Commercial policy: no affiliate links, referral parameters, sponsored ranking, or invented user review quotes.
Read the Valmera editorial policy. Product pages and the final account checkout remain authoritative at purchase time.
Official Sources Checked
Edit Real Footage by Describing the Result
Valmera turns an open-ended brief into a rendered edit you can review and revise. Create an account and upload for free; editing requires a subscription.
Try Valmera →