whisperX (open source, runs on your computer)

Make captions or subtitles with word-level timings from a video or audio file, and tell speakers apart

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization) Open source on GitHub under BSD-2-Clause. It runs on the person's own computer: run_tool hands back how to get it and run it.

What it asks for

  • input (Optional, words): The file or web address on the person's computer to work on, put where {input} is in the command
  • output (Optional, words): Where the result goes on the person's computer, put where {output} is in the command

Use it from your AI

Connect the AI you already use to Trillion once. Then ask in your own words; your AI finds this tool by what it does and runs it.

Say to your AI: “Use Trillion to make captions or subtitles with word-level timings from a video or audio file, and tell speakers apart”

What your AI calls: run_tool {"tool_id":"tool:github:m-bain-whisperx"}