Six seconds
Every request uses the same requested duration.
Research methodology · 2026 R1
This is the public, machine-readable contract for a planned four-model test. It fixes what will be generated, how failures count, what reviewers can see, and what must be disclosed before any quality claim is allowed.
No provider generation, human score, result, or ranking has been completed. The page publishes the method before the outcome.
It tests whether short AI-generated shots preserve people and objects, complete a requested change in the right order, keep camera geography readable, show physical contact and reaction, and deliver directly monitored synchronized sound. It is a six-task Tovaki integration sample, not a universal vendor leaderboard.
The digest below covers the complete private protocol byte for byte, including exact prompts, the round roster, weights, retry rules, and publication gate. A matching later disclosure proves the file did not change; the hash alone does not prove execution.
SHA-256 · Protocol digest
3efa03ed5fb305054069571bee501e0a06cb247e3334530c9dd7820cff8840d7Every candidate receives the same creative requirement inside a task. Differences are not repaired with model-specific prompting or a prettier reroll.
Every request uses the same requested duration.
One vertical format and one declared integration tier.
Optional native audio is switched on; always-on audio remains on. Silent and audio-enabled conditions are not mixed.
No vendor-specific rewrite, hidden enhancement, or creative retry.
Originals stay preserved; reviewers receive the same H.264/AAC delivery recipe.
The same anonymous letter does not permanently identify one model across tasks.
The task set deliberately spans bright, warm, comic, kinetic, tactile, observational, and poetic work so one narrow aesthetic cannot stand in for filmmaking ability.
Reviewers score 0–5 in 0.5 steps. Weights total 100; a polished still frame cannot hide broken time, contact, sound, or story causality.
Requested people, objects, actions, order, setting, and prohibitions are present.
Motion develops without flicker, sudden replacement, or unexplained temporal discontinuity.
The same people and traceable objects persist through movement and occlusion.
Approach, contact, force, transfer, release, and physical response remain readable.
Movement reveals the event while screen direction and spatial relationships stay understandable.
Gaze, pauses, listening, and reactions follow the event that causes them.
Speech and effects are audible, synchronized, perspectivally coherent, and directly monitored.
The shot can perform its stated narrative job without hiding the important event.
Verify the protocol digest and obtain a named credit limit before any request.
Record every first take, technical failure, retry, actual credit, timestamp, and SHA-256.
Create one anonymous proxy per cell and directly monitor the complete audio track.
Three reviewers use different randomized orders without model names, vendors, prices, metadata, or the private mapping.
Hash-lock all scorecards before identity is revealed; a changed scorecard invalidates analysis.
Each cell may retry once only after a provider error, timeout, corrupt file, wrong returned duration or resolution, or missing required audio track. A weak composition, strange performance, or failed story beat remains part of the evidence.
Planned first pass: 1620 Tovaki credits. Absolute protocol ceiling if every cell has one allowed technical retry: 3240. Generation remains unauthorized until a named owner approves a limit.
Scope, settings, task categories, score weights, retry rules, lock rules, budget ceiling, publication language, and the protocol digest.
Exact prompts, the round roster, accepted sample hashes, redacted failure ledger, anonymous raw scores, actual credits, and limitations.
Reviewer personal identities without consent, provider credentials or tokens, and private provider metadata that is unnecessary for reproduction.
These sources inform specific controls; they do not endorse Tovaki or turn this small integration study into their benchmark.
Separating temporal quality, frame quality, and condition consistency.
Testing requested state and relationship changes through time.
Keeping human judgments attached to explicit dimensions and limitations.
Separating audio-enabled conditions and declaring fixed settings and endpoint scope.