Research methodology · 2026 R1

How Tovaki will compare AI video models without picking the winner first

This is the public, machine-readable contract for a planned four-model test. It fixes what will be generated, how failures count, what reviewers can see, and what must be disclosed before any quality claim is allowed.

Prepared, not executed2026-08-25 · Tovaki Research

No provider generation, human score, result, or ranking has been completed. The page publishes the method before the outcome.

4model integrations
6distinct task lanes
24first-take cells
3blind reviewers

What does this benchmark measure?

It tests whether short AI-generated shots preserve people and objects, complete a requested change in the right order, keep camera geography readable, show physical contact and reaction, and deliver directly monitored synchronized sound. It is a six-task Tovaki integration sample, not a universal vendor leaderboard.

Protocol version commitment

The digest below covers the complete private protocol byte for byte, including exact prompts, the round roster, weights, retry rules, and publication gate. A matching later disclosure proves the file did not change; the hash alone does not prove execution.

SHA-256 · Protocol digest

3efa03ed5fb305054069571bee501e0a06cb247e3334530c9dd7820cff8840d7
Open the machine-readable JSONOpen the agent-readable Markdown

What stays fixed

Every candidate receives the same creative requirement inside a task. Differences are not repaired with model-specific prompting or a prettier reroll.

Six seconds

Every request uses the same requested duration.

9:16 · Tovaki 720p tier

One vertical format and one declared integration tier.

Native audio for every cell

Optional native audio is switched on; always-on audio remains on. Silent and audio-enabled conditions are not mixed.

One exact prompt per task

No vendor-specific rewrite, hidden enhancement, or creative retry.

One normalized review proxy

Originals stay preserved; reviewers receive the same H.264/AAC delivery recipe.

Task-local A–D labels

The same anonymous letter does not permanently identify one model across tasks.

Six tasks, not one beauty shot

The task set deliberately spans bright, warm, comic, kinetic, tactile, observational, and poetic work so one narrow aesthetic cannot stand in for filmmaking ability.

  1. Relationship and object handoff
  2. Dialogue and listener reaction
  3. Moving action geography
  4. Material contact and recovery
  5. Continuous camera and identity
  6. Bright surreal causality

Eight visible scoring dimensions

Reviewers score 0–5 in 0.5 steps. Weights total 100; a polished still frame cannot hide broken time, contact, sound, or story causality.

Prompt adherence

15/100

Requested people, objects, actions, order, setting, and prohibitions are present.

Temporal stability

15/100

Motion develops without flicker, sudden replacement, or unexplained temporal discontinuity.

Subject and object identity

15/100

The same people and traceable objects persist through movement and occlusion.

Motion and physical contact

15/100

Approach, contact, force, transfer, release, and physical response remain readable.

Camera and geography

10/100

Movement reveals the event while screen direction and spatial relationships stay understandable.

Performance and reaction

10/100

Gaze, pauses, listening, and reactions follow the event that causes them.

Sound completion

10/100

Speech and effects are audible, synchronized, perspectivally coherent, and directly monitored.

Usable edit value

10/100

The shot can perform its stated narrative job without hiding the important event.

Blind-review workflow

  1. Freeze and approve

    Verify the protocol digest and obtain a named credit limit before any request.

  2. Generate and fingerprint

    Record every first take, technical failure, retry, actual credit, timestamp, and SHA-256.

  3. Normalize and listen

    Create one anonymous proxy per cell and directly monitor the complete audio track.

  4. Score blind

    Three reviewers use different randomized orders without model names, vendors, prices, metadata, or the private mapping.

  5. Lock, then reveal

    Hash-lock all scorecards before identity is revealed; a changed scorecard invalidates analysis.

A bad-looking result is not a retry reason

Each cell may retry once only after a provider error, timeout, corrupt file, wrong returned duration or resolution, or missing required audio track. A weak composition, strange performance, or failed story beat remains part of the evidence.

Planned first pass: 1620 Tovaki credits. Absolute protocol ceiling if every cell has one allowed technical retry: 3240. Generation remains unauthorized until a named owner approves a limit.

Public evidence without leaking the blind test

Public now

Scope, settings, task categories, score weights, retry rules, lock rules, budget ceiling, publication language, and the protocol digest.

Released only after score lock

Exact prompts, the round roster, accepted sample hashes, redacted failure ledger, anonymous raw scores, actual credits, and limitations.

Remains private

Reviewer personal identities without consent, provider credentials or tokens, and private provider metadata that is unnecessary for reproduction.

Publication gate

  • There is no quality result yet. Cost and technical limits are not used as a substitute ranking.
  • Any later conclusion must say “in this six-task Tovaki integration sample.”
  • If the top two totals differ by less than 3 points out of 100, the result must say they are not distinguishable in this sample.
  • “Best AI video model,” “winner for every use,” and “viral model” remain forbidden claims.

Methodological foundations

These sources inform specific controls; they do not endorse Tovaki or turn this small integration study into their benchmark.