An ultra-low latency, privacy-focused Windows desktop voice transcription engine, neural speech assistant, and AI prompt refiner powered by Qwen3-ASR ONNX (DirectML GPU / AVX2), Whisper Large V3, Silero VAD v5, Qwen3-TTS, and Google Gemini AI.
Key Features • Architecture • AI Post-Processing • IPC Bus • Quick Start • License
0xVoice2Text is a high-performance desktop voice dictation engine and AI speech assistant crafted for software engineers, power users, and writers on Windows. It eliminates the friction of voice typing by capturing microphone audio, isolating speech with Silero Voice Activity Detection (VAD v5), transcribing locally at ultra-high speed via Qwen3-ASR 1.7B ONNX (accelerated in GPU VRAM via DirectML / DirectX 12) or cloud Whisper Large V3, and automatically injecting clean, punctuated text directly into your active code editor or browser.
Beyond simple speech-to-text, 0xVoice2Text features an intelligent AI Post-Processing Pipeline (powered by Gemini 3.6 Flash and Gemma) that cleans verbal debris ("um", "like", "you know"), reformats stream-of-consciousness dictation into clean code or prompts, and provides zero-latency neural voice feedback via Qwen3-TTS and Microsoft Edge Neural voices.
- ⚡ Local DirectML GPU Acceleration & Cloud STT Fallback
- Powered by Qwen3-ASR 1.7B ONNX with native DirectX 12 DirectML hardware acceleration (tested on AMD Radeon RX 7800 XT and Intel/NVIDIA) with CPU AVX2 multi-threading and Groq Cloud Whisper Large V3 fallback.
- 🎯 Silero Voice Activity Detection (VAD v5)
- Real-time neural audio stream analysis automatically detects natural speech pauses and finalizes dictation without requiring manual hotkey release.
- 🗣️ Wake Word & Voice Macro Automation
- Continuous low-power listening via Vosk for wake phrases ("Джарвис" / "Jarvis") to activate dictation hands-free, plus custom voice macros for launching applications.
- 🧠 Multi-Tier AI Post-Processing Pipeline
- Toggle seamlessly between raw verbatim output (
DIRECT), automatic speech sanitation (CLEANvia Gemma), and smart code/prompt refactoring (SMARTvia Gemini Flash) directly from the desktop widget.
- Toggle seamlessly between raw verbatim output (
- 🔊 Neural TTS Feedback (Local & Cloud)
- Dual-engine vocal responses supporting local Qwen3-TTS-12Hz zero-shot cloning and Microsoft Edge Neural voices (
ru-RU-SvetlanaNeural,ru-RU-DmitryNeural) played instantly via native Windows MCI (winmm.dll).
- Dual-engine vocal responses supporting local Qwen3-TTS-12Hz zero-shot cloning and Microsoft Edge Neural voices (
- 🎨 Futuristic Cyberpunk UI & Floating Overlays
- Radial Mouse HUD: Neon status circle tracking the cursor with animated recording rings.
- Glassmorphic Desktop Pill: Compact floating widget with interactive mode toggles and audio waveform visualizer.
- History Window: Dedicated searchable log with instant JSON export and replay capabilities.
- 🛰️ Universal IPC Event Bus
- Emits structured JSON events (
~/.0xvoice2text/last_event.jsonandevents.log) enabling live integration with IDEs, bots, and external automation scripts.
- Emits structured JSON events (
┌──────────────────────────────────────────────────────────────────┐
│ Microphone Audio Input Stream │
└─────────────────────────────────┬────────────────────────────────┘
│ 16kHz PCM Audio Stream
┌─────────────────────────────────▼────────────────────────────────┐
│ Silero Voice Activity Detection v5 ONNX Engine │
│ (Filters silence & background noise, detects natural pauses) │
└─────────────────────────────────┬────────────────────────────────┘
│ Buffered Speech Chunk
┌─────────────────────────────────▼────────────────────────────────┐
│ STTEngine Facade (src/core/stt/) │
│ ├── [Adapter 1] Qwen3ONNXAdapter (DirectML GPU / AVX2 CPU) │
│ └── [Adapter 2] GroqSTTAdapter (Whisper Large V3 Cloud Fallback)│
└─────────────────────────────────┬────────────────────────────────┘
│ Raw Transcribed Text
┌─────────────────────────────────▼────────────────────────────────┐
│ AI Post-Processing Pipeline (Gemini/Gemma) │
│ │
│ ┌────────────────────────┐ ┌────────────────────────────────┐ │
│ │ Direct Mode (Verbatim) │ │ Clean Mode (Strip Fillers) │ │
│ └────────────────────────┘ └────────────────────────────────┘ │
│ ┌────────────────────────┐ ┌────────────────────────────────┐ │
│ │ Smart Mode (AI Refine) │ │ Neural TTS Feedback (Qwen3/Edge│ │
│ └────────────────────────┘ └────────────────────────────────┘ │
└─────────────────────────────────┬────────────────────────────────┘
│ Clean Formatted Text
┌─────────────────────────────────▼────────────────────────────────┐
│ Windows Keystroke Injector & IPC Event Bus Dispatcher │
│ (Active Window Focus • ~/.0xvoice2text/last_event.json) │
└──────────────────────────────────────────────────────────────────┘
| Mode | Engine | Purpose | Output Example |
|---|---|---|---|
| ⚡ DIRECT | Qwen3-ASR / Whisper V3 | Instant verbatim output with natural punctuation | "найди в интернете инфу про танк тигр 2" |
| ✨ CLEAN | Gemma 4 / Flash Lite | Strips verbal debris, hesitations ("э-э-э", "ну", "типа"), fixes syntax | "Найди информацию про танк Tiger II." |
| 🤖 SMART | Gemini 3.6 Flash | Converts dictated thoughts into clean prompts, structured specs, or code | "Собери подробную справку по танку Tiger II: история создания, компоновка трансмиссии и бронирование." |
Every completed voice event is instantly broadcast to ~/.0xvoice2text/last_event.json and appended to events.log:
{
"timestamp": "2026-08-25T21:30:00+0300",
"unix_timestamp": 1786752900,
"engine": "qwen3-asr-onnx-directml",
"ai_mode": "smart",
"language": "ru",
"text": "Refactor auth controller into modular middleware pipeline.",
"char_count": 59,
"word_count": 8,
"status": "success"
}import os, json, time
EVENT_FILE = os.path.expanduser("~/.0xvoice2text/last_event.json")
last_mtime = 0
while True:
if os.path.exists(EVENT_FILE):
mtime = os.path.getmtime(EVENT_FILE)
if mtime > last_mtime:
last_mtime = mtime
with open(EVENT_FILE, "r", encoding="utf-8") as f:
event = json.load(f)
print(f"[{event['timestamp']}] ({event['ai_mode']}) {event['text']}")
time.sleep(0.05)git clone https://github.com/T58574/0xVoice2Text.git
cd 0xVoice2Textpython -m venv venv
venv\Scripts\activate
pip install -r requirements.txtCopy .env.example to .env and insert your credentials:
cp .env.example .env# Optional: Fallback cloud transcription (https://console.groq.com/keys)
GROQ_API_KEY=gsk_your_groq_api_key_here
# Optional: Required for Clean & Smart AI modes (https://aistudio.google.com/app/apikey)
GEMINI_API_KEY=your_gemini_api_key_here# Standard background launch
run.bat
# Or launch in interactive debug mode with live terminal logs
run.bat --debug0xVoice2Text/
├── main.py # Master application orchestrator & Qt event bridge
├── run.bat # Intelligent one-click Windows launcher (venv auto-detect)
├── requirements.txt # System dependencies
├── .env.example # Template for API credentials
├── GEMINI.md # System architecture and hardware specifications
├── README.md # Project documentation
├── src/
│ ├── config.py # JSON-backed dynamic configuration manager
│ ├── core/ # Core Audio, Speech & AI Services
│ │ ├── stt/ # Pluggable STT Subsystem
│ │ │ ├── base.py # BaseSTTAdapter ABC
│ │ │ ├── factory.py # Dynamic STT adapter factory
│ │ │ ├── qwen3_onnx_adapter.py # Qwen3-ASR ONNX DirectML/AVX2 engine
│ │ │ └── groq_adapter.py # Groq Cloud Whisper Large V3 engine
│ │ ├── stt_engine.py # Thread-safe unified STT facade
│ │ ├── vad_engine.py # Silero VAD v5 ONNX engine
│ │ ├── audio_recorder.py # sounddevice recorder & RMS audio visualizer
│ │ ├── wake_word.py # Vosk Wake Word & Silero VAD session manager
│ │ ├── ai_engine.py # Gemini & Gemma AI post-processing pipeline
│ │ ├── history.py # SQLite / JSON transcription history manager
│ │ ├── ipc_bus.py # IPC event bus dispatcher (~/.0xvoice2text)
│ │ └── logger.py # Rotating file & console logging system
│ ├── services/ # System & OS Integration
│ │ ├── tts/ # Pluggable TTS Subsystem
│ │ │ ├── base.py # BaseTTSAdapter ABC
│ │ │ ├── factory.py # Dynamic TTS adapter factory
│ │ │ ├── qwen3_tts_adapter.py # Qwen3-TTS-12Hz zero-shot engine
│ │ │ ├── edge_adapter.py # Microsoft Edge Neural TTS engine
│ │ │ └── service.py # JarvisVoiceService (MCI audio & anti-echo guard)
│ │ ├── hotkeys.py # Global Windows keyboard hooks (Toggle / PTT)
│ │ ├── injector.py # Active window keystroke & clipboard injector
│ │ └── macros.py # Voice command macro dispatcher
│ └── ui/ # PyQt6 Desktop User Interface
│ ├── error_dialog.py # Diagnostic error modal with resolution hints
│ ├── history.py # Searchable transcription history drawer
│ ├── mouse_hud.py # Radial neon cursor HUD overlay ring
│ ├── settings.py # Multi-tab settings configuration dialog
│ ├── tray.py # Windows notification area system tray icon
│ └── widget.py # Glassmorphic floating desktop status pill
├── test_full_system.py # Comprehensive full-system validation suite
├── test_stt_adapter.py # STT adapter unit tests
├── test_tts_adapter.py # TTS adapter unit tests
└── LICENSE # MIT License
Distributed under the MIT License. See LICENSE for details.