Skip to main content

Export transcript as word-level JSON with per-word timestamps

The transcribe tool already exports as text, subtitles, PDF, plain text, Markdown, SRT, and VTT — that covers most of what I need. The one format I'd love to see added is a word-level JSON export, where every single word has its own start/end timestamp, rather than grouping multiple words or sentences into one timestamp block (like SRT/VTT do).

Adobe Premiere Pro already supports this kind of export. Having it in Vowel would help in two main cases:

  • YouTube chapter/timestamp generation. Feeding an AI the current subtitle export gives it far fewer timestamp anchors, since each block often spans multiple sentences. A word-level JSON gives the AI enough granularity to place chapter markers exactly when a topic starts — instead of chapters that start or end mid-sentence because the underlying block was too coarse.

  • Motion graphics / animation timing. When converting a design into a motion graphic to sync with a video, having every word's exact timestamp means I know precisely when to trigger each animation, instead of estimating from a block-level subtitle.

This is achievable today by exporting from Premiere Pro instead, but since I'm already using Vowel's transcription for the whole session, having this JSON export available directly from Vowel would save a full extra step.

3 comments

Log in to comment and vote

Comments3

  • Dio

    •

    Sep 12

    @Vowen would this be possible to implement? It would be a game changer in my way of working with transcripted materials

    • Vowen

      Team•

      Sep 13

      We have a draft version of this internally. The issue at the moment is that some providers support word-level timestamps, and some don’t. So we’re thinking of how to resolve the gap. We’ll try adding it as an experimental feature in one of the next updates.

      • Dio

        •

        Sep 17

        That would be great!!

        Thank you so much, and keep up the amazing work c: