whisperX (open source, runs on your computer)
Make captions or subtitles with word-level timings from a video or audio file, and tell speakers apart
Make captions or subtitles with word-level timings from a video or audio file, and tell speakers apart
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization) Open source on GitHub under BSD-2-Clause. It runs on the person's own computer: run_tool hands back how to get it and run it.
Connect the AI you already use to Trillion once. Then ask in your own words; your AI finds this tool by what it does and runs it.
Say to your AI: “Use Trillion to make captions or subtitles with word-level timings from a video or audio file, and tell speakers apart”
What your AI calls: run_tool {"tool_id":"tool:github:m-bain-whisperx"}