How We Test Source Fidelity in AI-Generated Explainer Videos
Test AI explainer videos from source to script to scene. Includes a severity rubric, worked NIST example, review workflow, copyable checklist, and FAQs.

A video can repeat every sentence correctly and still teach the wrong thing. The narrator may preserve the words while the scene reverses a relationship, turns a possibility into a certainty, or makes six concurrent activities look like a checklist. Source fidelity therefore has four layers: source, script, scene, and final experience.
Publishability rule
A polished render is not a passed review. The video passes only when its required facts, relationships, qualifications, visible labels, narration, and scene implications agree with the approved source.
The four-layer review
1. Source
Is it current, approved, complete enough, and free of unresolved contradictions?
2. Script
Are names, numbers, steps, caveats, quotations, and conclusions preserved?
3. Scenes
Do diagrams, icons, labels, scale, and sequence imply the same meaning?
4. Final experience
Can the real audience hear, read, and understand it on the intended device?
Worked example: NIST CSF 2.0
The source says GOVERN informs the other five Functions and that all six should be addressed concurrently. That gives the review two traps: omitting GOVERN from the center, or animating the framework as a one-time sequence.
| Required source element | What the reviewer checks | Severity if wrong |
|---|---|---|
| Six exact Function names | Narration and visible spelling | Major; critical if omission changes the model |
| GOVERN informs the other five | Center placement and narrated relationship | Major |
| Functions are concurrent | No staircase or “finish one, then next” implication | Major |
| Framework is outcome-oriented | No invented certification or prescribed tool list | Critical |
| Source identity | NIST CSWP 29 and date visible or linked | Minor to major depending on context |
Severity beats raw error count
Ten punctuation issues are not equivalent to one reversed safety instruction. Use four levels:
- Critical: changes an action, obligation, identity, amount, safety decision, or central conclusion. Block publication.
- Major: materially distorts a relationship, qualification, or required section. Repair and re-review.
- Minor: does not change meaning but reduces precision or clarity. Repair when practical.
- Editorial: style preference with no factual effect. Do not let it hide higher-risk issues.
A copyable reviewer checklist
Where Golpo helps
Use document input when the source must ground the explanation, Script Mode when wording is approved, the pre-render script review when the plan supports it, Video Instructions for visual constraints, and inserted media for literal evidence. None of those controls removes the need for review; they make the review traceable and corrections more targeted.
Resolve reviewer disagreement explicitly
Two reviewers may agree that every sentence is technically true and still disagree about whether the video is safe to publish. Record the disputed source passage, video timestamp, proposed severity, audience consequence, and final owner decision. The highest-consequence qualified reviewer should decide factual risk; the editorial reviewer decides clarity only after that risk is resolved.
Automation can preflight exact terms, quantities, required sections, and prohibited phrases. It can also compare a generated script with a source ledger. It should not auto-approve diagrams, implied causality, respectful representation, or the decision to omit a caveat. Those judgments depend on how a real viewer will interpret the complete scene.
Create a source-grounded explainer in Golpo, then review the first output with the checklist above.
Frequently asked questions
How do I stop an AI explainer from hallucinating?
Use an approved source, narrow the requested outcome, require a pre-render script review when available, prohibit unsupported additions, and verify both narration and visuals before publishing.
Is an accurate script enough if the visuals are wrong?
No. A scene can imply the wrong sequence, quantity, relationship, or level of certainty even when the spoken sentence is accurate.
How should regulated content be reviewed?
An accountable subject-matter owner should approve the source, script, visual implications, required qualifications, and final delivery. Tool output is not a compliance determination.
Can source fidelity be measured automatically?
Automation can compare names, numbers, terminology, and coverage. Human review is still needed for visual implications, misleading emphasis, nuance, and whether an omission changes meaning.
What makes an error critical?
An error is critical when it could change a required action, safety decision, legal meaning, financial interpretation, identity, amount, or central conclusion.
Continue from here
Tags


