Playcut v2 is coming MCP · Actor v2 · & much more
PAB protocol · v0.9 · 2026-07-17

How we test AI actors

An AI actor is AI-generated on-camera talent for video. This is the fixed protocol behind the Playcut AI Actor Benchmark: 12 platforms, 34 metrics across six tiers, run the same way every time — instruments, a disclosed LLM judge, and policy audits.

Playcut builds one of the platforms in this benchmark

The protocol is fixed before any scoring run. All takes are published for every platform including Playcut's failures; every score is reproducible from the scripts and raw takes; nothing is retracted or re-rolled. We publish the metrics where Playcut loses. That structure — not a claim of neutrality — is the answer to the obvious conflict of interest.

The honest-numbers rule

A published figure is a real, sourced value or it is absent. Vendor claims are labeled as claims; anything we cannot verify is left blank rather than estimated to look complete. Wherever a measured number lands — including where Playcut loses — it ships. Every cell on a PAB page carries one of these classes:

Measured — produced by an instrument on real output.
Vendor claim † — the vendor's figure, not independently verified.
Editorial / audit — a rubric score against published criteria.
Blank — not published, not verified. Never guessed.

The v0.9 release covers the verified-facts and pricing layer. The instrument and LLM-judge tiers below run under the same protocol and publish with their raw data as the lab run completes.

The fixed protocol

  • • The same 10 scripts, the same reference faces, and the same product references across every platform.
  • • Five takes per platform per task, so a single lucky output can't set a score.
  • • Each row names the model and version tested — competitors by their productized model name; Playcut as the Playcut Actor Engine or Playcut Voice Engine with a month/year.
  • • A dated test window on every row, and four benchmark languages (ES / DE / FR / JA) for the multilingual metrics.
  • • Original-resolution files only — instruments run on the delivered output, never a re-encode.

Three method tiers

Every metric is produced one of three ways, and every published column says which. Mixing methods honestly — rather than pretending everything is "measured" — is the point.

Instrument — code runs on the file (face embeddings, lip-sync, transcription, ffprobe).
LLM-judge — a pinned model against a published rubric, N=3 runs, agreement reported.
Policy audit — a rubric against terms of service and pricing pages.

The 34 metrics

Tier A — Identity & consistency

Metric What it measures Direction Method
Identity consistency Face-embedding cosine (ArcFace buffalo_l) of the actor vs each take. Academic name CSIM. Higher Instrument
Cross-surface portability ⭐ The same face-embedding across still, talking-head, UGC, on-product and outfit-variant — the mean of 5 surfaces. No other benchmark measures this. Higher Instrument
Temporal identity drift Per-frame face cosine across a clip; variance and minimum. Catches a face that melts mid-take. Lower Instrument
Wardrobe persistence Clothing-region image-embedding similarity across takes, cross-checked by the judge. Higher Instrument + judge
Multi-character hold Identity + voice consistency across a multi-scene timeline. Not yet scored — a v2 addition. Higher v2

Tier B — Audio & speech

Metric What it measures Direction Method
Lip-sync SyncNet LSE-C / LSE-D. Paired with a judged lip-sync score and human spot-check, because these metrics correlate only loosely with human perception. Higher / Lower Instrument
Script fidelity (WER) Whisper large-v3 transcript vs the input script, word error rate. Lower Instrument
Voice naturalness UTMOS automatic mean-opinion-score (1–5), with DNSMOS as a secondary read. Higher Instrument
Voice consistency Speaker-embedding cosine (ECAPA-TDNN) of the voice across takes. Higher Instrument
Loudness compliance Integrated LUFS and true-peak vs the named delivery spec (e.g. TikTok ≈ −14 LUFS). Compliance = distance to spec. Closer Instrument

Tier C — Visual quality & stability

Metric What it measures Direction Method
Flicker index Variance of per-frame optical-flow energy. Lower is steadier. Lower Instrument
Sharpness Laplacian variance across sampled frames; mean and minimum (min catches momentary blur). Higher Instrument
Color stability Consecutive-frame RGB histogram correlation; drift = 1 − mean correlation. Catches exposure pumping. Lower Instrument
Delivered vs claimed specs ffprobe-measured resolution / fps / bitrate vs marketing claims. The gap is the story; the field read is disclosed. Match Instrument

Tier D — LLM-as-judge (published rubric)

Metric What it measures Direction Method
Artifact rate Fixed defect taxonomy (hands, teeth, eyes, hair, jewelry, extra fingers) → rate per take. Gates usable-take rate. Lower LLM-judge
Prompt adherence Brief and script vs the delivered frames and transcript, 1–5. Higher LLM-judge
Emotion range Same script in five requested emotions, scored 1–5 each. Higher LLM-judge
Gesture naturalness Robotic loops, dead hands, clipping, 1–5. Higher LLM-judge
UGC-nativeness Framing, energy and camera grammar vs native TikTok/Reels, 1–5. Buyers' most-cited sort criterion. Higher LLM-judge
On-product fidelity Label, shape and color of a held/shown SKU vs the reference, 1–5. Higher LLM-judge
Background coherence Warped geometry, melting text, 1–5. Higher LLM-judge
Skin-texture realism Plastic-skin / uncanny index, 1–5. Higher LLM-judge
Gaze stability Wandering or dead eyes, 1–5. Higher LLM-judge
Multilingual delivery Judged delivery plus WER per language (ES / DE / FR / JA in v1). Higher Judge + instrument
Actor-library diversity A catalog census — counts of age band, ethnicity and body type on a stated date. Counted, not judged. Counted Census
Brand-safety robustness Allow/refuse consistency across a standardized risky-script set. Consistency is scored, not ideology. Consistent LLM-judge

Tier E — Economics & ops

Metric What it measures Direction Method
Usable-take rate Share of takes shippable without a retry, defect-gated by the artifact threshold. The hit rate buyers feel but rarely name. Higher Instrument
Cost per usable video Dated list price ÷ usable-take rate, per 30-second video. The number buyers actually sort by. Feeds the Pricing Index. Lower Instrument + facts
Generation latency Wall-clock p50 and p95 from submission to finished file, queue conditions disclosed. Lower Instrument
Queue reliability Failure and timeout rate across the full run. Lower Instrument
Pricing transparency Rubric audit of hidden caps, watermark fine print and credit-math obfuscation. Playcut scores itself here — checkable via the published rubric. Higher Policy audit

Tier F — Trust, rights & provenance

Metric What it measures Direction Method
AI-detectability Reframed: manifest/watermark presence is scored; there is no headline 'AI pass-rate' because no credible, free, reproducible general AI-video detector exists. Code + note
C2PA / watermark Content-credential manifest present or absent (c2patool + metadata inspection). Who signs vs who ships naked files. Present Instrument
Commercial rights Terms-of-service audit of licensing model, ad-usage rights and exclusivity, with the ToS retrieval date published. Higher Policy audit
Voice-clone consent Whether a consent flow gates voice/face cloning, evidence linked. Gate exists Policy audit
API availability First-party checklist: public API, docs, auth model, webhooks vs polling, MCP/agent-callable. Higher Facts

How the judge is disclosed

"Judge: claude-opus-4-8 (pinned model ID). Scores are returned as structured output against a fixed JSON schema; the rubric is byte-identical across items. Sampling-temperature control is not available on this model generation, so run-to-run determinism is measured rather than assumed: every item is scored in three independent runs, and we publish exact-agreement percentage and Krippendorff's α per metric. Rubrics and prompts are published in the repository. 10% of judged items are human-verified per release, and the agreement number is published."

Items are scored blind: metadata is stripped and files are renamed so the platform identity never enters a prompt. Disagreements greater than one point go to a human review queue. Newly added platforms are marked provisional until the full protocol completes.

Reproducibility

Each release ships its scripts, judge rubrics, pinned environment lockfiles, a dated pricing snapshot, and the raw takes, so any number can be recomputed from source. The identity instrument uses ArcFace buffalo_l face embeddings; lip-sync uses SyncNet; transcription uses Whisper large-v3; delivered specs use ffprobe; provenance uses c2patool and metadata inspection. The public data package and repository publish alongside the lab-phase results.

Methodology FAQ

What is an AI actor? +

An AI actor is AI-generated on-camera talent for video — a synthetic presenter or performer, not a stock clip of a real person and not a software agent. PAB tests platforms that generate this kind of on-camera talent for UGC, ads, and presenter video.

Can Playcut rank #1 on its own benchmark? +

Playcut builds one of the platforms tested and is labeled the vendor throughout. The protocol is fixed before any scoring run, every take is published including Playcut's failures, and every score is reproducible from the scripts and raw takes. We publish the metrics where Playcut loses. That structure — not a promise — is what keeps a vendor-run benchmark honest.

Which metrics are automated versus judged? +

Three method tiers, and every published column is labeled with which one produced it. Instrument metrics (identity, lip-sync, WER, loudness, flicker, sharpness, delivered specs, latency) run on code. Judge metrics (artifact rate, UGC-nativeness, emotion, gesture and more) use a pinned LLM against a published rubric. Policy audits (rights, consent, pricing transparency) use a rubric against terms and pricing pages.

How is the LLM judge kept honest? +

The judge is a pinned model (claude-opus-4-8) run against a fixed JSON-schema rubric that ships in the repo. Sampling-temperature control is not available on this model generation, so run-to-run determinism is measured rather than assumed: every item is scored in three independent runs and we publish exact-agreement percentage and Krippendorff's α per metric. Items are scored blind, and 10% are human-verified per release with the agreement number published.

Why is there no AI-detectability score? +

Because there is no credible, free, reproducible detector for general AI video. Publishing a 'percent detected' column would imply a rigor that does not exist. Instead we score what is deterministic — whether a content-credential manifest or watermark is present — and say plainly why a pass-rate column is not offered.

What is the honest-numbers rule? +

A published figure is a real, sourced value or it is absent. Vendor claims are labeled as claims; anything unverified is left blank rather than estimated to look complete. Wherever a measured number lands — including where Playcut loses — it ships. This v0.9 release covers the verified-facts and pricing layer; the instrument and judge tiers publish with their raw data as the lab run completes.

How often is it re-run? +

Quarterly, on one URL, versioned in place. Newly added platforms are marked provisional until the full protocol completes; retired versions stay visible and marked rather than being silently removed; every row is date-stamped. Corrections are logged, never made silently.