# Tovaki AI Video Model Blind-Benchmark Methodology — 2026 R1

Status: **prepared_not_executed**
Published: 2026-08-25
Reviewed: 2026-08-25
Canonical HTML: https://tovaki.com/research/ai-video-model-benchmark-methodology
Machine-readable JSON: https://tovaki.com/data/ai-video-model-benchmark-methodology.json

## Evidence boundary

This record describes a planned Tovaki integration benchmark. No provider generation, human review, quality result or model ranking is represented as complete.

This is a public projection of a private, frozen protocol. It is not a result page. No provider generation, human score, winner, or quality ranking is available.

## Protocol commitment

- Algorithm: SHA-256
- Digest: `3efa03ed5fb305054069571bee501e0a06cb247e3334530c9dd7820cff8840d7`
- Covers: The complete private protocol file, including its exact prompts, model roster, scoring weights, retry policy and publication gate.
- Interpretation: A matching digest can prove that a later disclosed protocol is byte-for-byte identical to the version committed here; the digest does not prove that the benchmark was executed.

## Fixed scope

- 4 Tovaki model integrations
- 6 task lanes
- 24 first-take cells
- 3 blinded reviewers
- 6 seconds per request
- 9:16 aspect ratio
- Tovaki 720p resolution tier
- Audio: Native synchronized audio expected for every accepted sample; optional native audio is enabled and always-on audio remains enabled.
- Review proxy: 720x1280 H.264 at 24 fps with AAC 48 kHz audio, normalized with one locked transcode recipe while originals are preserved.

## Public task lanes

- Relationship handoff
- Dialogue reaction
- Moving action geography
- Material contact and recovery
- Continuous camera and identity
- Bright surreal causality

Exact prompts and the model roster remain withheld until all reviewer scorecards are locked.

## Scoring

Scale: 0-5 in 0.5 increments. Weights total 100.

| Dimension | Weight |
|---|---:|
| Prompt adherence | 15/100 |
| Temporal stability | 15/100 |
| Subject identity | 15/100 |
| Motion contact | 15/100 |
| Camera geography | 10/100 |
| Performance reaction | 10/100 |
| Sound completion | 10/100 |
| Usable edit rate | 10/100 |

Close-result rule: If the top two aggregate scores differ by less than 3 points out of 100, report that they are not distinguishable in this sample.

## Blind-review controls

- Every candidate receives the same exact prompt and settings inside a task.
- Anonymous labels are A-D, independently randomized inside each task.
- Each reviewer receives an independently randomized order.
- Model and vendor identity remain hidden until the three scorecards are hash-locked.
- Accepted review proxies require SHA-256 for every accepted review proxy.
- Final proxy audio must be directly monitored; a waveform or track-presence check is insufficient.
- Identity reveal requires an explicit unlock after score lock.

## Retry and budget gates

- Maximum retries per cell: 1.
- Allowed technical reasons: provider_error, timeout, corrupt_or_unplayable_file, wrong_duration_or_resolution_from_provider, missing_required_audio_track.
- Selection rule: A visually weak, surprising or unattractive result is evidence, not a reason to reroll.
- Planned first pass: 1620 Tovaki credits.
- Absolute protocol ceiling: 3240 Tovaki credits.
- Generation is not authorized. A named owner must approve a credit limit before any provider request.

## Disclosure contract

Public now:
- scope and fixed settings
- task-lane coverage without exact prompts
- score dimensions and weights
- randomization, retry, locking and publication rules
- protocol SHA-256 commitment

Withheld until scores are locked:
- exact prompts
- model roster for this round
- task-local anonymous mapping

Published with any result:
- exact prompts and settings
- accepted sample hashes
- technical failures and retries with sensitive request identifiers redacted
- anonymous raw reviewer scores
- actual Tovaki credits and limitations

Remains private:
- reviewer personal identities unless each reviewer consents
- provider credentials and request tokens
- private provider metadata not needed to reproduce the result

## Publication gate

- Results available: no.
- Ranking available: no.
- A universal-best claim is forbidden.
- Any later conclusion must use the scope: “in this six-task Tovaki integration sample.”
- Cost, resolution, duration, or model availability cannot substitute for a quality result.

## Methodological references

- [VBench: Comprehensive Benchmark Suite for Video Generative Models](https://openaccess.thecvf.com/content/CVPR2024/papers/Huang_VBench_Comprehensive_Benchmark_Suite_for_Video_Generative_Models_CVPR_2024_paper.pdf) — Separating temporal quality, frame-wise quality and condition consistency rather than treating a polished frame as complete video quality.
- [TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation](https://arxiv.org/abs/2406.08656) — Testing whether object attributes and relations change in the requested order while identities remain traceable through time.
- [Video-Bench: Human-Aligned Video Generation Benchmark](https://openaccess.thecvf.com/content/CVPR2025/papers/Han_Video-Bench_Human-Aligned_Video_Generation_Benchmark_CVPR_2025_paper.pdf) — Keeping human judgment attached to explicit dimensions and publishing the limits of an aggregate score.
- [Artificial Analysis video generation benchmarking methodology](https://artificialanalysis.ai/video/methodology) — Treating audio-enabled generation as its own comparison condition and declaring fixed settings and endpoint scope.

## Citation and reuse

Cite this record as: Tovaki. “AI Video Model Blind-Benchmark Methodology — 2026 R1.” Reviewed 2026-08-25. https://tovaki.com/research/ai-video-model-benchmark-methodology

When summarizing this record, preserve the prepared-not-executed status, the Tovaki-integration scope, and the no-result/no-ranking boundary.
