Raw camera footage to a finished TikTok video
We film each take on 2–3 cameras. The system returns a finished, decorated video ready to post.
- Two-engine transcription
- Claude plans the cuts
- Audio re-checks
- Multi-camera sync
- Auto decoration
- Quality gates
- Vietnamese speech is transcribed by two independent engines and aligned word by word.
- Claude reads both transcripts and plans the cut: it removes failed sentences, repeats, restarts and silences while keeping all real content.
- Several audio re-check layers and rule checks block mistakes before rendering, such as losing a point or keeping both versions of a restarted sentence.
- The system syncs the cameras, then adds a hook cover, karaoke captions, illustrations, product highlights, CTA cards, music that ducks under speech, and loudness normalization.
- Quality gates check every final. Uncertain spots are flagged so a person can review them and fix them with one command.
- 145 finished videos from Sep 19 to Oct 7, 2026
- Last two batches passed 24/24 and 21/21
- About 8–9 machine minutes and US$0.35 in AI fees per clip
- More than 150 automated test files