# Clipping-agent review worksheet

Companion guide: https://valmera.io/clipping-agent
Revised: September 8, 2026

This is a blank evaluation template, not a completed benchmark. Use recordings you have permission to process. Keep private source URLs and client details out of public copies.

## 1. Define the trial

- Reviewer and date:
- Product, interface, plan, and version/model if shown:
- Source identifier, duration, and file size:
- Processing timeframe and selection prompt:
- Requested count, duration range, and aspect ratio:
- Intended audience and destination:
- Required source passage and original timestamps:
- Qualification that must remain:
- Correct spelling of the chosen name:
- Chart, demonstration, or subject that must remain visible:
- Required final format, subtitle/project handoff, and branding treatment:
- Trial cost allocation rule:

Use the same source and acceptance requirements when comparing products. Record any differences the tools require instead of presenting different tasks as identical.

## 2. Register candidates before judging

| Candidate ID | Source range | Score and scale, if supplied | Vendor's stated reason | Score hidden during first review? |
| --- | --- | --- | --- | --- |
| | | | | |
| | | | | |
| | | | | |
| | | | | |

Add a row for every candidate in the fixed test set, including rejected candidates. If you already saw the score, mark the review as unblinded. Keep the original score with the original candidate; distinguish later revisions.

## 3. Copy this block for each candidate

Candidate ID:
Project/version before revision:
Review decision: acceptable / repairable / unsuitable

- Meaning and required qualification:
- Opening and ending; missing or cut-off speech:
- Caption spelling and timing:
- Frame, chart, subject, and overlay placement:
- Speech, music, and cut continuity:
- Other reason for the decision:

Required repair:
Controls or follow-up prompt used:
Extra attempts or processing:
Hands-on correction minutes:
Did the repair meet the requirement?:
Accepted project/version, if any:

A complete word or sentence boundary does not prove that the excerpt preserves the speaker's meaning. Check the relevant surroundings in the source.

## 4. Verify delivery

- Candidate and accepted version:
- Export/job identifier:
- Final file retrieved at:
- File opened outside the editor:
- Complete runtime, including ending or audio tail:
- Dimensions and aspect ratio:
- Captions and sound checked through the ending:
- Overlay watermark or end card:
- Separate subtitles/project package tested, if required:
- Correct destination/account/post status, if publishing was authorized:
- Final outcome: delivered / needs repair / not delivered

A processing success response, preview URL, or accepted edit is not itself proof that the agreed final file was delivered. Keep a stable copy while the download link is valid.

## 5. Account for the work

| Stage | Hands-on minutes | Elapsed minutes | Actual charge or usage | Notes and retries |
| --- | --- | --- | --- | --- |
| Upload and generation | | | | |
| First review | | | | |
| Corrections and rechecks | | | | |
| Export and file inspection | | | | |
| Required external finishing/publishing | | | | |

- Candidates generated:
- Candidates reviewed in the fixed set:
- Accepted after review/correction:
- Checked final files delivered:
- Total hands-on minutes:
- Total elapsed minutes:
- Actual trial cost, with currency and allocation rule:
- Usage units and included allowance used:

Acceptance rate = accepted / reviewed.
Hands-on time per accepted clip = total hands-on minutes / accepted.
Allocated cost per accepted clip = trial cost / accepted.

If the denominator is zero, report the counts and total effort/cost; the ratio is undefined. Do not substitute requested or generated clips for accepted clips. Record delivery separately. Elapsed time can overlap across jobs and should not be added as if all tasks ran sequentially.

## 6. Assess scores and limits

Reveal any hidden scores after the initial editorial review.

- Were acceptable clips near the top of the ranking?
- Which highly scored clips were unsuitable, and why?
- Which lower-scored clips were acceptable?
- Did revisions materially change the original scored clip?
- What did this source fail to test, such as accents, noisy audio, screen shares or multiple speakers?

Do not compare different vendors' numbers as equivalent probabilities. This trial measures your chosen source, task and workflow; it does not establish a universal winner.

## 7. Optional publishing observations

For each published version, record the platform, account, publication time, exact metric definition, observation window, and changes to the hook, caption, thumbnail, promotion or distribution.

Keep editorial acceptance separate from audience metrics. An observational comparison is not a controlled causal test and does not prove that a score caused a result.
