Skip to main content

Optimize Whisper build for Windows for faster transcription

Was ecstatic to find an open source wispr, and love the product so far but my only issue is that it can take a lot longer to process, I believe wispr cuts up longer transcripts into smaller pieces automatically and might connect it with a last call, would be great if that could be implemented, would also love to collaborate on it if there is a repo somewhere.

Status: Completed6 comments

Log in to comment and vote

Comments6

  • Gabriel N

    •

    Dec 27, 2025

    I'm on Windows 10 on the base model. Maybe my intuition is wrong. Having used vowen and Whisper, it just seems to be a lot faster, but obviously that might not be the case. So let me know.

    • Vowen

      Team•

      Dec 27, 2025

      I see, since the transcription is happening locally on your device, I think there is a noticeable difference in the speed especially on Windows, we’ve had others reach out to us as well regarding the speed on Windows. If we were to do it on the server with GPU acceleration, it would be lot faster, but we are currently not considering doing it since we wanted the processing to happen locally. I think there might be other optimizations we can consider that makes local transcription much faster, we’ll get back to you on this once we’ve made some improvements for Windows.

    • Vowen

      Team•

      Dec 27, 2025

      Also curious to know if your system has any NVIDIA/AMD GPU? We are trying to rebuild the binary to add support for local GPU acceleration if any GPU is present. This change should definitely speed things up. Let me know.

      • John White

        •

        Jan 21

        Yes! please do this on windows. I’m happy to beta test the nvidia build

  • Vowen

    Team•

    Dec 25, 2025

    @Gabriel N I’m assuming you are using the larger models, as it can be a bit slow, but the response will be more accurate. So far, we haven’t noticed that much of a difference between the base.en model and the other models in terms of accuracy though. Only in specific cases when I want the transcription to be really accurate, I go for the other models, but for everyday use, the base.en works fast. I’ll look into it in more detail and get back to you, thanks once again.

  • Vowen

    Team•

    Dec 25, 2025

    Thanks for the feedback @Gabriel N , we are already processing the data in chunks periodically. I’ll look into making it faster. Can you tell me which model and OS you are running it on? And do you know how long recordings are that is taking a lot of time? We’ll try to reproduce the same issue and address it shortly.