VoskFlow
- 5 Devlogs
- 27 Total hours
A free alternative to Wisprflow that has almost all features, is locally private, and fast.
A free alternative to Wisprflow that has almost all features, is locally private, and fast.
Building VoskFlow — an entirely offline, privacy-first voice dictation engine with an Apple/VOID-inspired minimalist HUD that works across any Windows app.
Here’s everything built, debugged, and polished over the latest development sprint! 🚀
#DCA54C) dash docked right above the Windows taskbar.Ctrl + Space), and the dash smoothly expands into a floating pill with live symmetrical audio waveform visualizer bars.Ctrl to lock into hands-free streaming dictation.tiny.en, base.en, small.en).Queue worker loop.float32 Universal Compute Engine: Standardized compute format to ensure compatibility across all CPU architectures (Intel & AMD) without quantization crashes.GetDC (Invalid window handle) errors..exe installer.100% offline. Zero telemetry. Your voice stays on your machine. 🔒
Spent the last sprint squashing bugs and polishing VoskFlow (100% offline, privacy-first voice-to-text) for its V1 release!
What got done:
faster-whisper from pinging HuggingFace API on startup. If local weights exist, it loads straight from disk—zero network required.#1C1917 + Warm Ochre #DCA54C), serif typography, and a live status indicator.keyboard.hook), fixing combo hotkeys like Ctrl+Space without crashes..spec bundling the models, CTranslate2 DLLs, and web UI into a clean, windowed .exe (no ugly terminal popup).Next up: final .exe compilation and shipping! Don’t type, just speak. ✨
Finaaly trinaed a llm and connetced it
Today we overhauled the core engine for commercial-grade speed and accuracy. We migrated to Faster-Whisper, which gives us flawless auto-punctuation and capitalization. To achieve a zero-latency feel, we built a background worker that streams and processes audio chunks every 500ms while you hold the hotkey, meaning the final text drops instantly on release. Finally, we engineered a dynamic disfluency filter that automatically strips out filler words like umms and ahhs before typing. Up next: connecting a local LLM to turn spoken commands into OS-level keyboard shortcuts!
Ngl, the landing page we just dropped absolutely ate and left no crumbs. Added screenshots of our black glassmorphism pill and the bouncing voice bubbles, and the aesthetic is heavily based. Wispr Flow is literally shaking rn. No cap, the UI is serving looks. Now that the site is up, we’re ready to plug in the Vosk ML model and actually let it cook. Huge W. 🔥 https://evinjsubin.github.io/VoskFlow/landing/