Media
GriotVid
An idea in, a narrated video out — in your language.
A griot is a West African storyteller. GriotVid takes a topic or a script and returns a finished, narrated, captioned video: the script and scene plan are written in the chosen language, footage is pulled from Pexels and Pixabay, narration is generated as one continuous take, music is picked by mood, and FFmpeg assembles the result on the GPU.
Narration uses Edge-TTS for English and several African languages, and YarnGPT for Yoruba, Igbo and Hausa. Paste your own script and it is narrated word for word — the model only picks footage and a title, and never rewrites your words.
How it is built
- Narration rendered as one continuous take rather than stitched per scene, so the pacing holds.
- Word-by-word karaoke captions burned in, with the music ducked under the voice.
- Ken Burns motion fitted to the narration length rather than to fixed clip durations.
- Audience context switches both the script framing and the footage search terms.
- Final render on the GPU through NVENC, in 16:9, 9:16 or 1:1.