Backlog
54Under consideration
android support in future?
will this be available to android users in future as out of united states, many users use android like me in india form mumbai, i use a android and ios both but primary is android
Voice-Powered Integrations & Workflow Automation
Overview Automate your workflow using natural voice commands. Chain multiple actions together and let your voice assistant handle complex tasks with context from your files, photos, and videos. Key Features Tool Integrations GitHub Issues: "Create a GitHub issue about the bug in LoginScreen.tsx" - automatically pulls context from the file Linear Tasks: "Add a Linear task for the UI redesign based on this screenshot" - attaches image context Smart Context: Reference files, images, videos, or documents in your commands and the assistant extracts relevant information automatically Example Use Cases "Create a GitHub issue with the error logs from error.txt and add priority high" "Make a Linear task from this design mockup with implementation details" "Open Slack, navigate to the engineering channel, and remind the team about standup" "Launch my morning routine: open Calendar, check today's tasks in Linear, and play focus music" "Add calendar events for all the deadlines mentioned in this email thread"
Live Transcription with Inline Editing
Firstly, this product is awesome so thanks for making it. Here’s one issue I would love to resolve: I’m using text to speech and try to say “feature-2.md” but end up with “feature2.md” or “feature dash 2.md” etc. There are many examples like this and in most cases there simply isn’t enough context in what I’m saying to know what I want as the final text. Idea: when using live transcription (e.g. with soniox) I want to see what I’m saying as I am speaking and I want to be able to edit as I go. So if I see a typo I can just fix it right there and if I’m about to try to say something “~/.gitconfig” which I’m sure will be transcribed incorrectly I can just type it myself and then continue to speak. Another idea could come out of this: you log the corrections I make and use that somehow to improve future transcriptions.
ANDROID?
When is Android coming out?
voice recognition
Recognize my own voice within transcription to always label my own voice/speech as me/you. Alternatively this probably could run based on audio channel or similar.
Automatic Microphone Muting Option
I’d love a feature that automatically mutes my mic in other apps while I’m talking in Vowen. Basically when I’m using Vowen and speaking for a while I don’t want people in other apps to hear me or get disturbed. It would be super useful if Vowen could mute my mic system wide or only in specific apps I choose. That way I don’t have to keep muting and unmuting myself all the time manually.
Larger custom dictionary, voice fine-tuning, and free system instruction for the custom API
The dictionary limit is the main thing holding the free tier back I'm a Vowen user, and the 50-word dictionary cap is honestly the only free-tier restriction that gets in my way. I'd happily see a much larger dictionary as a paid feature - it would be genuinely valuable for me and several people around me. For reference, wisprflow.ai lets you add thousands of words and tunes the system around them (optimal LLM prompts, etc.). I've set up a small model that runs locally on my laptop and proofreads my transcribed text. It comfortably handles far more than 50 words using the language model alone - I'd estimate 500 - 1000 words without trouble on my hardware, and I'd expect the same via API requests. So the technical ceiling is much higher than the current limit suggests. The feature I'd really love: personal fine-tuning What would be amazing is a way to fine-tune the model specifically for me - for my voice and my vocabulary. If there's any way I can help test or contribute to this, I'm in. Feedback: please don't paywall the system instruction on the custom API Locking the LLM system instruction behind a paywall doesn't make much sense. It's trivial to work around - with LM Studio, Llama, or any other provider, I can set a system instruction at the provider's own level when connecting the model. So a paid system instruction just pushes me to configure it elsewhere. Given that it used to be free, the better move is the opposite: keep the system instruction as a free feature.
Allow different Provider for Transcript and AI Mode to combine speed in Transcription and quality in AI
My issue is that I would love the speed of Groq and the AI mode quality of Sonnet. As Groq does not support Sonnet, I need to decide between the two of those. So I use Ppenrouter for transcribtion as well as AI Mode. Would be great if I could select Groq for transcription and OpenRouter for AI Mode.
Post-Transcription Automations (Auto-Copy & Auto-Export)
As a premium user who transcribes multiple videos daily, my current workflow involves a repetitive, manual loop once a video finishes transcribing: Wait for the "transcription ready" status. Open the completed transcript. Click the three-dots menu. Click Copy. Paste the text into a Large Language Model (LLM) to generate video descriptions and metadata. When processing multiple videos back-to-back, these manual steps add unnecessary friction to an otherwise seamless experience. Proposed Solution Introduce an Automation Toggle or an "On Completion" setting within the manual transcription menu. When a transcription finishes, Vowen would automatically trigger a pre-selected action based on the user's preference: Option A: Auto-Copy to Clipboard (with Smart Restore) — The full transcript text is instantly copied to the user's clipboard the second it's ready, mirroring the smooth workflow of the existing voice transcript feature. To make this even more seamless, it could include a "smart clipboard" behavior: the transcript temporarily occupies the clipboard, but once the user pastes it into their LLM or document, the system automatically restores their previous clipboard data so nothing is permanently lost. Option B: Auto-Trigger Export Window — Automatically opens the export menu immediately upon completion. Option C: Preemptive Default Export — Allows users to set a default format preference (e.g., Markdown, Plain Text, or PDF) so the file is generated and saved automatically without extra clicks. Why This Is Relevant For power users, content creators, and marketers who use Vowen as a baseline reference for AI workflows, transcription is rarely the final step—it is just the catalyst for the next task. Automating the copy/export bridge would eliminate tedious micro-tasks, speed up content distribution workflows, and significantly increase daily productivity.
Hide notification "Transcription cancelled"
When I cancel hands-free mode via ESC, the notification always appears, which is very annoying. It would be great if there was a simple toggle in the settings to stop the notification from appearing at all! Thanks!
Next up
5Committed and queued
Voice-activated time tracker
Lemme say it first: love you guys, love the product! ❤️ I'd love a voice-activated time tracker! I could say, like, I'm gonna start [task] finish [task] break , and the app would log it with timestamps. Calculating totals isn't essential; I can do that in Google Sheets. Cheers!
Restrict Language Detection to Selected Languages
Auto-detect sometimes selects unrelated languages. Let users choose the languages they actually speak so detection is limited to those options, improving accuracy and avoiding incorrect transcription.
Allow for combination mapping of mouse buttons and keyboard shortcuts
Hi! I have been trying to combine keyboard shortcuts with mouse button 4 and mouse button 5, but it doesn't seem to register correctly. It records the mapping but it doesn't get applied.
Team Functions / shared dictionary / shared threats
Organization → invite team members Share dictionary & threat/default system settings
Custom Start/Stop Sound Effects
Would love the ability to choose the sound effect (whether it’s from a predetermined list or one of my own that I upload) that is played when starting and stopping dictation. Thanks!
In Progress
8Actively being built
iOS & iPadOS Support
Complete voice transcription and AI assistance System-wide keyboard integration Native iOS/iPadOS experience All AI models and capabilities available Native mobile application Custom keyboard extension for any app Share sheet integration Optimized for touch and Apple Pencil (iPad) Works across all iOS/iPadOS applications Context-aware AI responses Local processing for privacy
Linux Support
Complete voice transcription and AI assistance System-wide hotkey support Native Linux integration All AI models and capabilities available Native Linux application Support for major distributions (Ubuntu, Fedora, Arch, Debian) System tray integration Works with popular desktop environments (GNOME, KDE, XFCE) Works across all Linux applications Context-aware AI responses Local processing for privacy Auto-updates and background operation
INTEL Mac support please
I want really try this App, but I can’t becoz no INTEL support
macOS Shortcuts actions
A Shortcuts action that uses Vowen, and the model downloaded, to use in an automated workflow. This would allow using a superior model than Apple’s built-in dictation action. Alternatively, use Vowen as is, but it triggers a macOS Shortcut* after it finishes transcribing. This way, once the text is transcribed and copied to the clipboard, the Shortcut can take over. I would imagine this is yet another keyboard hotkey** to trigger this option. *an option could be that instead of triggering a Shortcut, it triggers an app. You can save Shortcuts as apps, which run on launch and could grab the contents of the clipboard to use in the workflow. **seeing as this would be more keyboard hotkeys, perhaps a single hotkey that when triggered presents the user with a small window to select which option is required, transcribe, transcribe with AI, or transcribe with Shortcuts. I hope all this makes sense. Thanks
Transcribe audio or video file using AI Enhancement
I’ve been fine-tuning my 'Enhance transcription with AI' prompt by comparing sample text from Parakeet with and without enhancement. I then ran some tests using AI-generated audio file samples, but it does not appear to use the 'Enhance with AI' feature; is this correct?
User selected post-process with AI
Hi, I’ve tried sooo many transcription apps that it’s getting out of hand. I’ve tried Handy, Handless, Pindrop, Spokenly, FluidVoice, Pipit, WhisperShortcut, Ottex, Snaply, to name just a few. But I keep coming back to Vowen. I will admit that some of the ones named above have so very nice features not presently available in Vowen, but none are dealbreakers for me. Once again, the app is excellent, and congratulations. So, to the feature request, which I’ve not seen any other app offer. Firstly, create additional Enhance transcription with AI prompts. The could be Clean up, Rewrite, Make markdown note, Email, and so on. Next, once the model transcribes the text, the user is presented with these options, which, when one is selected, then passes the transcription to the AI model to enhance using the respective prompt. This way, the user doesn’t need to decide which prompt to use first, they just speak, and then select what to do with it. Now, I did make a similar-sounding feature request about triggering a Shortcut after transcription to automate what to do with the text. For example, just speak, without first selecting an app to paste the text into, and once the Shortcut is triggered, it could grab the text (clipboard) and present to the user options what to do with it, such as adding it to an existing note. Or using Private Cloud Compute to enhance. I think these two features could co-exist. Thank you.
Spotify auto Duck in while using Vowen
a simple duck in while listening to music instead of having to pause it every time ☺️
Transcript doesn't render
Quite frequently (easily a couple of times a day) I’ll be trying to write something and all that will be captured is the first word. Very frequently this gets rendered as something like - Perspective or - Tests or something like that (with the hyphen being something that is very often present). I’m not sure why this is occurring but it happens with a degree of frequency that means I have worry about whether what I’ve just said will actually be captured and rendered. Macbook M5, 128GB - Small whisper model, and then mistral running locally as custom AI pathway.
Done
122Recently shipped
Speaker Diarization
Ability to distinguish between speakers in the transcripts that are generated
Sync transcriptions across devices
Add support for GPU Acceleration on Windows
Optimize Whisper build for Windows for faster transcription
Was ecstatic to find an open source wispr, and love the product so far but my only issue is that it can take a lot longer to process, I believe wispr cuts up longer transcripts into smaller pieces automatically and might connect it with a last call, would be great if that could be implemented, would also love to collaborate on it if there is a repo somewhere.
Share AI meeting notes with others
Tones Mode
Your AI assistant automatically adapts its writing style to match the context of where you're writing. Whether it's a casual text, professional email, or work message, get perfectly toned responses every time. Key Features Automatically detects the app you're writing in Adjusts tone based on conversation context Seamlessly switches between formal and casual styles No manual switching needed
Add support for nvidia/parakeet-tdt-0.6b-v3 (Windows)
https://huggingface.co/spaces/hf-audio/open_asr_leaderboard Feedback Suggestion Please add support for nvidia/parakeet-tdt-0.6b-v3 in the local transcription software. This model is listed on the Open ASR Leaderboard and demonstrates stronger multilingual performance. It achieves higher transcription accuracy than openai/whisper-large-v3. It is also more than 16× faster than openai/whisper-large-v3-turbo, making it highly efficient. Additionally, since we have already configured the “Configure Your AI – Bring your own API key to power Vowen’s AI features”, we suggest expanding functionality with: Translation feature: Transcribed content could be directly translated into a target language. Real-time transcription + real-time translation: Making the software more suitable for global use cases and multilingual workflows. Voice Activity Detection (VAD): Automatically segment transcription based on natural speech pauses instead of manual hotkeys. Speaker Diarization: Distinguish between different speakers in conversations, improving readability and usability in meetings or multi-speaker recordings.
Add Support for Exporting Transcription Results to Subtitle Formats (VTT, SRT, etc.)
Currently, the app transcribes audio/video files into plain text only. It would be useful to allow exporting transcripts as subtitle files such as.vtt,.srt, or other common formats. Benefits: Use transcripts directly in video editing Save time by avoiding manual timestamp formatting Improve accessibility and multilingual workflows Suggested Behavior: After transcription, allow download as.txt,.vtt, or.srt Optionally, enable subtitle customization (line length, line break rules, timestamp format) So adding subtitle export will integrate well with existing capabilities.
Ability to pause screen recording while taking meeting notes
Allow to call MCP Server in Agent Mode
It would be great if the agent mode could connect to an MCP server. The MCP could then, for instance, call my SuperBase for FAQ knowledge or trigger actions in my Webtool. That would be a total game changer. ❤️ Probably also requires Single Sign On with my Google Account for authentication, though.