How we test AI actors
An AI actor is AI-generated on-camera talent for video. This is the fixed protocol behind the Playcut AI Actor Benchmark: 12 platforms, 34 metrics across six tiers, run the same way every time — instruments, a disclosed LLM judge, and policy audits.
Playcut builds one of the platforms in this benchmark
The protocol is fixed before any scoring run. All takes are published for every platform including Playcut's failures; every score is reproducible from the scripts and raw takes; nothing is retracted or re-rolled. We publish the metrics where Playcut loses. That structure — not a claim of neutrality — is the answer to the obvious conflict of interest.
The honest-numbers rule
A published figure is a real, sourced value or it is absent. Vendor claims are labeled as claims; anything we cannot verify is left blank rather than estimated to look complete. Wherever a measured number lands — including where Playcut loses — it ships. Every cell on a PAB page carries one of these classes:
The v0.9 release covers the verified-facts and pricing layer. The instrument and LLM-judge tiers below run under the same protocol and publish with their raw data as the lab run completes.
The fixed protocol
- • The same 10 scripts, the same reference faces, and the same product references across every platform.
- • Five takes per platform per task, so a single lucky output can't set a score.
- • Each row names the model and version tested — competitors by their productized model name; Playcut as the Playcut Actor Engine or Playcut Voice Engine with a month/year.
- • A dated test window on every row, and four benchmark languages (ES / DE / FR / JA) for the multilingual metrics.
- • Original-resolution files only — instruments run on the delivered output, never a re-encode.
Three method tiers
Every metric is produced one of three ways, and every published column says which. Mixing methods honestly — rather than pretending everything is "measured" — is the point.
The 34 metrics
Tier A — Identity & consistency
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| Identity consistency | Face-embedding cosine (ArcFace buffalo_l) of the actor vs each take. Academic name CSIM. | Higher | Instrument |
| Cross-surface portability ⭐ | The same face-embedding across still, talking-head, UGC, on-product and outfit-variant — the mean of 5 surfaces. No other benchmark measures this. | Higher | Instrument |
| Temporal identity drift | Per-frame face cosine across a clip; variance and minimum. Catches a face that melts mid-take. | Lower | Instrument |
| Wardrobe persistence | Clothing-region image-embedding similarity across takes, cross-checked by the judge. | Higher | Instrument + judge |
| Multi-character hold | Identity + voice consistency across a multi-scene timeline. Not yet scored — a v2 addition. | Higher | v2 |
Tier B — Audio & speech
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| Lip-sync | SyncNet LSE-C / LSE-D. Paired with a judged lip-sync score and human spot-check, because these metrics correlate only loosely with human perception. | Higher / Lower | Instrument |
| Script fidelity (WER) | Whisper large-v3 transcript vs the input script, word error rate. | Lower | Instrument |
| Voice naturalness | UTMOS automatic mean-opinion-score (1–5), with DNSMOS as a secondary read. | Higher | Instrument |
| Voice consistency | Speaker-embedding cosine (ECAPA-TDNN) of the voice across takes. | Higher | Instrument |
| Loudness compliance | Integrated LUFS and true-peak vs the named delivery spec (e.g. TikTok ≈ −14 LUFS). Compliance = distance to spec. | Closer | Instrument |
Tier C — Visual quality & stability
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| Flicker index | Variance of per-frame optical-flow energy. Lower is steadier. | Lower | Instrument |
| Sharpness | Laplacian variance across sampled frames; mean and minimum (min catches momentary blur). | Higher | Instrument |
| Color stability | Consecutive-frame RGB histogram correlation; drift = 1 − mean correlation. Catches exposure pumping. | Lower | Instrument |
| Delivered vs claimed specs | ffprobe-measured resolution / fps / bitrate vs marketing claims. The gap is the story; the field read is disclosed. | Match | Instrument |
Tier D — LLM-as-judge (published rubric)
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| Artifact rate | Fixed defect taxonomy (hands, teeth, eyes, hair, jewelry, extra fingers) → rate per take. Gates usable-take rate. | Lower | LLM-judge |
| Prompt adherence | Brief and script vs the delivered frames and transcript, 1–5. | Higher | LLM-judge |
| Emotion range | Same script in five requested emotions, scored 1–5 each. | Higher | LLM-judge |
| Gesture naturalness | Robotic loops, dead hands, clipping, 1–5. | Higher | LLM-judge |
| UGC-nativeness | Framing, energy and camera grammar vs native TikTok/Reels, 1–5. Buyers' most-cited sort criterion. | Higher | LLM-judge |
| On-product fidelity | Label, shape and color of a held/shown SKU vs the reference, 1–5. | Higher | LLM-judge |
| Background coherence | Warped geometry, melting text, 1–5. | Higher | LLM-judge |
| Skin-texture realism | Plastic-skin / uncanny index, 1–5. | Higher | LLM-judge |
| Gaze stability | Wandering or dead eyes, 1–5. | Higher | LLM-judge |
| Multilingual delivery | Judged delivery plus WER per language (ES / DE / FR / JA in v1). | Higher | Judge + instrument |
| Actor-library diversity | A catalog census — counts of age band, ethnicity and body type on a stated date. Counted, not judged. | Counted | Census |
| Brand-safety robustness | Allow/refuse consistency across a standardized risky-script set. Consistency is scored, not ideology. | Consistent | LLM-judge |
Tier E — Economics & ops
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| Usable-take rate | Share of takes shippable without a retry, defect-gated by the artifact threshold. The hit rate buyers feel but rarely name. | Higher | Instrument |
| Cost per usable video | Dated list price ÷ usable-take rate, per 30-second video. The number buyers actually sort by. Feeds the Pricing Index. | Lower | Instrument + facts |
| Generation latency | Wall-clock p50 and p95 from submission to finished file, queue conditions disclosed. | Lower | Instrument |
| Queue reliability | Failure and timeout rate across the full run. | Lower | Instrument |
| Pricing transparency | Rubric audit of hidden caps, watermark fine print and credit-math obfuscation. Playcut scores itself here — checkable via the published rubric. | Higher | Policy audit |
Tier F — Trust, rights & provenance
| Metric | What it measures | Direction | Method |
|---|---|---|---|
| AI-detectability | Reframed: manifest/watermark presence is scored; there is no headline 'AI pass-rate' because no credible, free, reproducible general AI-video detector exists. | — | Code + note |
| C2PA / watermark | Content-credential manifest present or absent (c2patool + metadata inspection). Who signs vs who ships naked files. | Present | Instrument |
| Commercial rights | Terms-of-service audit of licensing model, ad-usage rights and exclusivity, with the ToS retrieval date published. | Higher | Policy audit |
| Voice-clone consent | Whether a consent flow gates voice/face cloning, evidence linked. | Gate exists | Policy audit |
| API availability | First-party checklist: public API, docs, auth model, webhooks vs polling, MCP/agent-callable. | Higher | Facts |
How the judge is disclosed
"Judge: claude-opus-4-8 (pinned model ID). Scores are returned as structured output against a fixed JSON schema; the rubric is byte-identical across items. Sampling-temperature control is not available on this model generation, so run-to-run determinism is measured rather than assumed: every item is scored in three independent runs, and we publish exact-agreement percentage and Krippendorff's α per metric. Rubrics and prompts are published in the repository. 10% of judged items are human-verified per release, and the agreement number is published."
Items are scored blind: metadata is stripped and files are renamed so the platform identity never enters a prompt. Disagreements greater than one point go to a human review queue. Newly added platforms are marked provisional until the full protocol completes.
Reproducibility
Each release ships its scripts, judge rubrics, pinned environment lockfiles, a dated pricing snapshot, and the raw takes, so any number can be recomputed from source. The identity instrument uses ArcFace buffalo_l face embeddings; lip-sync uses SyncNet; transcription uses Whisper large-v3; delivered specs use ffprobe; provenance uses c2patool and metadata inspection. The public data package and repository publish alongside the lab-phase results.
Methodology FAQ
What is an AI actor? +
An AI actor is AI-generated on-camera talent for video — a synthetic presenter or performer, not a stock clip of a real person and not a software agent. PAB tests platforms that generate this kind of on-camera talent for UGC, ads, and presenter video.
Can Playcut rank #1 on its own benchmark? +
Playcut builds one of the platforms tested and is labeled the vendor throughout. The protocol is fixed before any scoring run, every take is published including Playcut's failures, and every score is reproducible from the scripts and raw takes. We publish the metrics where Playcut loses. That structure — not a promise — is what keeps a vendor-run benchmark honest.
Which metrics are automated versus judged? +
Three method tiers, and every published column is labeled with which one produced it. Instrument metrics (identity, lip-sync, WER, loudness, flicker, sharpness, delivered specs, latency) run on code. Judge metrics (artifact rate, UGC-nativeness, emotion, gesture and more) use a pinned LLM against a published rubric. Policy audits (rights, consent, pricing transparency) use a rubric against terms and pricing pages.
How is the LLM judge kept honest? +
The judge is a pinned model (claude-opus-4-8) run against a fixed JSON-schema rubric that ships in the repo. Sampling-temperature control is not available on this model generation, so run-to-run determinism is measured rather than assumed: every item is scored in three independent runs and we publish exact-agreement percentage and Krippendorff's α per metric. Items are scored blind, and 10% are human-verified per release with the agreement number published.
Why is there no AI-detectability score? +
Because there is no credible, free, reproducible detector for general AI video. Publishing a 'percent detected' column would imply a rigor that does not exist. Instead we score what is deterministic — whether a content-credential manifest or watermark is present — and say plainly why a pass-rate column is not offered.
What is the honest-numbers rule? +
A published figure is a real, sourced value or it is absent. Vendor claims are labeled as claims; anything unverified is left blank rather than estimated to look complete. Wherever a measured number lands — including where Playcut loses — it ships. This v0.9 release covers the verified-facts and pricing layer; the instrument and judge tiers publish with their raw data as the lab run completes.
How often is it re-run? +
Quarterly, on one URL, versioned in place. Newly added platforms are marked provisional until the full protocol completes; retired versions stay visible and marked rather than being silently removed; every row is date-stamped. Corrections are logged, never made silently.