Skip to content

About

Real-time voice AI assistant & desktop dictation engine (<500ms Whisper Large V3, Silero VAD, Edge-TTS, Gemini/Gemma smart refiner, PyQt6)

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

🎙️ 0xVoice2Text — Real-Time Voice AI Assistant & Desktop Dictation Hub

Python PyQt6 Qwen3-ASR Whisper Silero VAD Gemini License: MIT

An ultra-low latency, privacy-focused Windows desktop voice transcription engine, neural speech assistant, and AI prompt refiner powered by Qwen3-ASR ONNX (DirectML GPU / AVX2), Whisper Large V3, Silero VAD v5, Qwen3-TTS, and Google Gemini AI.

Key Features • Architecture • AI Post-Processing • IPC Bus • Quick Start • License


📖 Overview

0xVoice2Text is a high-performance desktop voice dictation engine and AI speech assistant crafted for software engineers, power users, and writers on Windows. It eliminates the friction of voice typing by capturing microphone audio, isolating speech with Silero Voice Activity Detection (VAD v5), transcribing locally at ultra-high speed via Qwen3-ASR 1.7B ONNX (accelerated in GPU VRAM via DirectML / DirectX 12) or cloud Whisper Large V3, and automatically injecting clean, punctuated text directly into your active code editor or browser.

Beyond simple speech-to-text, 0xVoice2Text features an intelligent AI Post-Processing Pipeline (powered by Gemini 3.6 Flash and Gemma) that cleans verbal debris ("um", "like", "you know"), reformats stream-of-consciousness dictation into clean code or prompts, and provides zero-latency neural voice feedback via Qwen3-TTS and Microsoft Edge Neural voices.


✨ Key Features

  • ⚡ Local DirectML GPU Acceleration & Cloud STT Fallback
    • Powered by Qwen3-ASR 1.7B ONNX with native DirectX 12 DirectML hardware acceleration (tested on AMD Radeon RX 7800 XT and Intel/NVIDIA) with CPU AVX2 multi-threading and Groq Cloud Whisper Large V3 fallback.
  • 🎯 Silero Voice Activity Detection (VAD v5)
    • Real-time neural audio stream analysis automatically detects natural speech pauses and finalizes dictation without requiring manual hotkey release.
  • 🗣️ Wake Word & Voice Macro Automation
    • Continuous low-power listening via Vosk for wake phrases ("Джарвис" / "Jarvis") to activate dictation hands-free, plus custom voice macros for launching applications.
  • 🧠 Multi-Tier AI Post-Processing Pipeline
    • Toggle seamlessly between raw verbatim output (DIRECT), automatic speech sanitation (CLEAN via Gemma), and smart code/prompt refactoring (SMART via Gemini Flash) directly from the desktop widget.
  • 🔊 Neural TTS Feedback (Local & Cloud)
    • Dual-engine vocal responses supporting local Qwen3-TTS-12Hz zero-shot cloning and Microsoft Edge Neural voices (ru-RU-SvetlanaNeural, ru-RU-DmitryNeural) played instantly via native Windows MCI (winmm.dll).
  • 🎨 Futuristic Cyberpunk UI & Floating Overlays
    • Radial Mouse HUD: Neon status circle tracking the cursor with animated recording rings.
    • Glassmorphic Desktop Pill: Compact floating widget with interactive mode toggles and audio waveform visualizer.
    • History Window: Dedicated searchable log with instant JSON export and replay capabilities.
  • 🛰️ Universal IPC Event Bus
    • Emits structured JSON events (~/.0xvoice2text/last_event.json and events.log) enabling live integration with IDEs, bots, and external automation scripts.

🏗️ Architecture

┌──────────────────────────────────────────────────────────────────┐
│                   Microphone Audio Input Stream                  │
└─────────────────────────────────┬────────────────────────────────┘
                                  │ 16kHz PCM Audio Stream
┌─────────────────────────────────▼────────────────────────────────┐
│           Silero Voice Activity Detection v5 ONNX Engine         │
│    (Filters silence & background noise, detects natural pauses)  │
└─────────────────────────────────┬────────────────────────────────┘
                                  │ Buffered Speech Chunk
┌─────────────────────────────────▼────────────────────────────────┐
│         STTEngine Facade (src/core/stt/)                         │
│  ├── [Adapter 1] Qwen3ONNXAdapter (DirectML GPU / AVX2 CPU)      │
│  └── [Adapter 2] GroqSTTAdapter (Whisper Large V3 Cloud Fallback)│
└─────────────────────────────────┬────────────────────────────────┘
                                  │ Raw Transcribed Text
┌─────────────────────────────────▼────────────────────────────────┐
│             AI Post-Processing Pipeline (Gemini/Gemma)           │
│                                                                  │
│  ┌────────────────────────┐  ┌────────────────────────────────┐  │
│  │ Direct Mode (Verbatim) │  │ Clean Mode (Strip Fillers)     │  │
│  └────────────────────────┘  └────────────────────────────────┘  │
│  ┌────────────────────────┐  ┌────────────────────────────────┐  │
│  │ Smart Mode (AI Refine) │  │ Neural TTS Feedback (Qwen3/Edge│  │
│  └────────────────────────┘  └────────────────────────────────┘  │
└─────────────────────────────────┬────────────────────────────────┘
                                  │ Clean Formatted Text
┌─────────────────────────────────▼────────────────────────────────┐
│        Windows Keystroke Injector & IPC Event Bus Dispatcher     │
│        (Active Window Focus • ~/.0xvoice2text/last_event.json)   │
└──────────────────────────────────────────────────────────────────┘

🧠 AI Intelligence Modes

Mode Engine Purpose Output Example
⚡ DIRECT Qwen3-ASR / Whisper V3 Instant verbatim output with natural punctuation "найди в интернете инфу про танк тигр 2"
✨ CLEAN Gemma 4 / Flash Lite Strips verbal debris, hesitations ("э-э-э", "ну", "типа"), fixes syntax "Найди информацию про танк Tiger II."
🤖 SMART Gemini 3.6 Flash Converts dictated thoughts into clean prompts, structured specs, or code "Собери подробную справку по танку Tiger II: история создания, компоновка трансмиссии и бронирование."

🛰️ Universal IPC Event Bus

Every completed voice event is instantly broadcast to ~/.0xvoice2text/last_event.json and appended to events.log:

{
  "timestamp": "2026-08-25T21:30:00+0300",
  "unix_timestamp": 1786752900,
  "engine": "qwen3-asr-onnx-directml",
  "ai_mode": "smart",
  "language": "ru",
  "text": "Refactor auth controller into modular middleware pipeline.",
  "char_count": 59,
  "word_count": 8,
  "status": "success"
}

Python Real-Time Event Consumer

import os, json, time

EVENT_FILE = os.path.expanduser("~/.0xvoice2text/last_event.json")
last_mtime = 0

while True:
    if os.path.exists(EVENT_FILE):
        mtime = os.path.getmtime(EVENT_FILE)
        if mtime > last_mtime:
            last_mtime = mtime
            with open(EVENT_FILE, "r", encoding="utf-8") as f:
                event = json.load(f)
                print(f"[{event['timestamp']}] ({event['ai_mode']}) {event['text']}")
    time.sleep(0.05)

🚀 Quick Start

1. Clone the Repository

git clone https://github.com/T58574/0xVoice2Text.git
cd 0xVoice2Text

2. Set Up Virtual Environment & Dependencies

python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt

3. Configure API Keys (Optional)

Copy .env.example to .env and insert your credentials:

cp .env.example .env
# Optional: Fallback cloud transcription (https://console.groq.com/keys)
GROQ_API_KEY=gsk_your_groq_api_key_here

# Optional: Required for Clean & Smart AI modes (https://aistudio.google.com/app/apikey)
GEMINI_API_KEY=your_gemini_api_key_here

4. Run Application

# Standard background launch
run.bat

# Or launch in interactive debug mode with live terminal logs
run.bat --debug

📁 Project Structure

0xVoice2Text/
├── main.py                     # Master application orchestrator & Qt event bridge
├── run.bat                     # Intelligent one-click Windows launcher (venv auto-detect)
├── requirements.txt            # System dependencies
├── .env.example                # Template for API credentials
├── GEMINI.md                   # System architecture and hardware specifications
├── README.md                   # Project documentation
├── src/
│   ├── config.py               # JSON-backed dynamic configuration manager
│   ├── core/                   # Core Audio, Speech & AI Services
│   │   ├── stt/                # Pluggable STT Subsystem
│   │   │   ├── base.py         # BaseSTTAdapter ABC
│   │   │   ├── factory.py      # Dynamic STT adapter factory
│   │   │   ├── qwen3_onnx_adapter.py # Qwen3-ASR ONNX DirectML/AVX2 engine
│   │   │   └── groq_adapter.py # Groq Cloud Whisper Large V3 engine
│   │   ├── stt_engine.py       # Thread-safe unified STT facade
│   │   ├── vad_engine.py       # Silero VAD v5 ONNX engine
│   │   ├── audio_recorder.py   # sounddevice recorder & RMS audio visualizer
│   │   ├── wake_word.py        # Vosk Wake Word & Silero VAD session manager
│   │   ├── ai_engine.py        # Gemini & Gemma AI post-processing pipeline
│   │   ├── history.py          # SQLite / JSON transcription history manager
│   │   ├── ipc_bus.py          # IPC event bus dispatcher (~/.0xvoice2text)
│   │   └── logger.py           # Rotating file & console logging system
│   ├── services/               # System & OS Integration
│   │   ├── tts/                # Pluggable TTS Subsystem
│   │   │   ├── base.py         # BaseTTSAdapter ABC
│   │   │   ├── factory.py      # Dynamic TTS adapter factory
│   │   │   ├── qwen3_tts_adapter.py # Qwen3-TTS-12Hz zero-shot engine
│   │   │   ├── edge_adapter.py # Microsoft Edge Neural TTS engine
│   │   │   └── service.py      # JarvisVoiceService (MCI audio & anti-echo guard)
│   │   ├── hotkeys.py          # Global Windows keyboard hooks (Toggle / PTT)
│   │   ├── injector.py         # Active window keystroke & clipboard injector
│   │   └── macros.py           # Voice command macro dispatcher
│   └── ui/                     # PyQt6 Desktop User Interface
│       ├── error_dialog.py     # Diagnostic error modal with resolution hints
│       ├── history.py          # Searchable transcription history drawer
│       ├── mouse_hud.py        # Radial neon cursor HUD overlay ring
│       ├── settings.py         # Multi-tab settings configuration dialog
│       ├── tray.py             # Windows notification area system tray icon
│       └── widget.py           # Glassmorphic floating desktop status pill
├── test_full_system.py         # Comprehensive full-system validation suite
├── test_stt_adapter.py         # STT adapter unit tests
├── test_tts_adapter.py         # TTS adapter unit tests
└── LICENSE                     # MIT License

📜 License

Distributed under the MIT License. See LICENSE for details.

About

Real-time voice AI assistant & desktop dictation engine (<500ms Whisper Large V3, Silero VAD, Edge-TTS, Gemini/Gemma smart refiner, PyQt6)

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages