A real-time speech-to-text and translation system with live streaming, multi-user meetings, video/audio processing, meeting history, and RAG chat.
- ποΈ Real-time Streaming: Live microphone transcription with instant translation
- πΉ Meeting Rooms: Multi-user meetings with translation per participant
- Individual Device Mode: Each person uses their own microphone
- Shared Room Mode: Multiple speakers on one mic with AI speaker identification
- π¬ Video Translation: Upload videos for transcription, translation, and TTS audio replacement
- π΅ Audio Recording: Upload audio files with speaker diarization support
- Support for 10+ languages (English, Arabic, Urdu, Spanish, French, German, Chinese, Japanese, Korean, Hindi)
- Auto-detect source language
- Real-time parallel translation to multiple languages
- Voice cloning for TTS (experimental)
- Speaker Diarization: Automatic speaker identification and labeling
- Real-time Collaboration: Multiple users in shared meeting rooms
- Progress Tracking: WebSocket-based progress updates for long operations
- Audio Enhancement: Optional noise reduction for uploaded files
- Transcript Export: Download meeting transcripts in multiple languages
- Meeting History: Account-scoped history with meeting detail views
- RAG Chat: Ask questions about meeting transcripts
- Meeting Minutes: Auto-generated participants, key points, action items, decisions, and summary
- Frontend: Feature-based web UI
- Backend: Go server with WebSocket support and REST API
- ASR Service: Python FastAPI + Whisper for speech recognition
- Translation Service: Python service using Google Translate API
- TTS Service: XTTS v2 for text-to-speech with voice cloning
- Embedding + LLM: RAG pipeline for meeting Q&A
- Database: PostgreSQL for meeting data and participants
- Video Processing: FFmpeg for audio/video manipulation
Browser Mic β WebSocket (16kHz PCM) β Faster-Whisper β Partial Captions
β
VAD
β
Final Segments
β
Translation
β
Translated Output
- Partial captions appear immediately
- Final segments are emitted after silence detection
- Translation runs on finalized segments only
web/
βββ index.html # Landing page
βββ assets/ # Shared resources
β βββ css/ # Modular stylesheets
β βββ js/ # Shared utilities
β βββ images/ # Static images
βββ components/ # Reusable UI components
βββ features/ # Feature modules
βββ home/ # Home page
βββ streaming/ # Live streaming
βββ recording/ # Audio upload
βββ video/ # Video upload
βββ meeting/ # Meeting rooms
βββ history/ # Meeting history + chat
- Go 1.16+ (backend server)
- Python 3.8+ (AI services)
- Docker & Docker Compose (recommended)
- FFmpeg (video/audio processing)
- PostgreSQL 15+
# Ubuntu/Debian
sudo apt install ffmpeg
# macOS
brew install ffmpeg# 1. Clone the repository
git clone <your-repo>
cd audio-translator
# 2. Create .env file
cp .env.example .env
# Edit .env and add your HuggingFace token for speaker diarization
# 3. Start all services
./start-services.shWhat this does:
- β Checks and builds Docker images if needed
- β Starts PostgreSQL, ASR, Translation, TTS, Embedding, and LLM services
- β Runs database migrations
- β Builds and starts the Go web server
Service URLs:
- π Web UI: http://localhost:8080
- π€ ASR Service: http://localhost:8003
- π Translation: http://localhost:8004
- π TTS Service: http://localhost:8005
- π§ Embeddings: http://localhost:8006
- π¬ LLM: http://localhost:8007
# Start services
docker compose up -d
# Run database migrations
cat migrations/*.sql | docker exec -i audio-translator-postgres-1 psql -U audio_translator -d audio_translator
# Start Go server
go build -o bin/server cmd/server/main.go
set -a && source .env && set +a # Load environment variables
./bin/server# Stop everything
docker compose down && pkill -f bin/server
# Stop just Docker containers
docker compose down
# Stop just Go server
kill $(cat bin/server.pid)- Go to http://localhost:8080
- Click "Streaming Translation"
- Select source and target languages
- Click Start and grant microphone permission
- Speak into your microphone
- See real-time transcription and translation
- Download transcript when done
- Go to http://localhost:8080/meeting.html
- Choose Individual Devices or Shared Room
- Share the room code with participants
- Join, select language, and grant microphone permission
- Host can end the meeting for everyone
- Go to http://localhost:8080/features/history/meetings-history.html
- Sign in (Keycloak) to view account-scoped history
- Open a meeting to view minutes and full transcript
- Use the chat panel to ask questions about the meeting
- Go to http://localhost:8080/video.html
- Upload a video file
- Select source/target languages
- Optional: enable "Generate translated audio" and "Clone original voice"
- Process and download results
- Go to http://localhost:8080/recording.html
- Upload an audio file
- Optional: enable Speaker Diarization or Audio Enhancement
- Process and view results
# HuggingFace token (required for diarization)
HF_TOKEN=your_huggingface_token_here
# CORS (optional - leave empty for development)
ALLOWED_ORIGINS=
# Diarization tuning
SPEAKER_SIM_THRESHOLD=0.82
MIN_EMBED_DURATION=0.8
SPEAKER_OVERLAP_RATIO_THRESHOLD=0.25
SPEAKER_CONFIDENCE_THRESHOLD=0.55
SPEAKER_PROFILE_TTL_SECONDS=3600
SPEAKER_PROFILE_CLEANUP_INTERVAL_SECONDS=300
SPEAKER_PROFILE_STORE_URL=
SPEAKER_PROFILE_PERSIST_INTERVAL_SECONDS=15
SPEAKER_PROFILE_DB_TTL_SECONDS=86400
SPEAKER_PROFILE_DB_CLEANUP_INTERVAL_SECONDS=300
# Keycloak JWT verification
KEYCLOAK_ISSUER=
KEYCLOAK_JWKS_URL=
KEYCLOAK_AUDIENCE=
# Backend service URLs
ASR_BASE_URL=http://127.0.0.1:8003
TRANSLATION_BASE_URL=http://127.0.0.1:8004
TTS_BASE_URL=http://127.0.0.1:8005
EMBEDDING_BASE_URL=http://127.0.0.1:8006
LLM_BASE_URL=http://127.0.0.1:8007
OLLAMA_MODEL=llama3.2:3b{
"keycloak": {
"issuer": "http://localhost:8180/realms/audio-transcriber",
"clientId": "audio-translator-client",
"scope": "openid profile email"
},
"services": {
"asrBaseUrl": "http://localhost:8003",
"translationBaseUrl": "http://localhost:8004",
"ttsBaseUrl": "http://localhost:8005",
"embeddingBaseUrl": "http://localhost:8006",
"llmBaseUrl": "http://localhost:8007"
}
}Edit services/asr_py/app.py:
MODEL_NAME = "base" # Options: tiny, base, small, medium, largeEdit cmd/server/main.go:
srv := session.NewServer(session.Config{
PollInterval: 800 * time.Millisecond, // ASR polling frequency
WindowSeconds: 8, // Audio buffer size
FinalizeAfter: 500 * time.Millisecond, // Text stabilization time
})- The TTS service starts in gTTS fallback mode while XTTS v2 loads
- XTTS v2 enables higher quality and voice cloning
- If XTTS is unavailable, the system automatically falls back to gTTS
Check status:
curl http://127.0.0.1:8005/health- XTTS v2 supports voice cloning for multiple languages
- Long text is chunked; if token limits are exceeded, it falls back to gTTS
- First run downloads ~1.8GB model (cached in Docker volume)
- Create a realm (e.g.
audio-transcriber) - Create a public client (e.g.
audio-translator-client) - Set valid redirect URIs to
http://localhost:8080/* - Set web origins to
http://localhost:8080 - Enable self-registration if desired
- Update:
KEYCLOAK_ISSUERin.envweb/config.jsonkeycloak.issuer
Meeting history and chat are account-scoped and require login.
Minutes are generated automatically after a meeting ends.
To backfill minutes for existing meetings:
# Requires LLM_BASE_URL and OLLAMA_MODEL in .env
go run cmd/backfill-minutes/main.go- Check browser microphone permissions (click π in address bar)
- Open browser console (F12) and check for errors
- Ensure you're using HTTPS or localhost (WebRTC requirement)
- Ensure PostgreSQL container is running:
docker ps | grep postgres - Check
.envfile has correct credentials - Verify migrations ran:
docker exec -it audio-translator-postgres-1 psql -U audio_translator -d audio_translator -c "\dt"
- Check browser console for JavaScript errors
- Ensure server is loading
.envvariables - Verify database tables exist (run migrations)
- Add HuggingFace token to
.envfile - Get token from: https://huggingface.co/settings/tokens
- Accept terms for pyannote models: https://huggingface.co/pyannote/speaker-diarization
- Check service is running:
curl http://localhost:8003/health - View logs:
docker compose logs asr -f - Verify Whisper model downloaded successfully
- First run downloads 1.8GB model (3-5 minutes)
- Falls back to gTTS while downloading
- Check status:
curl http://localhost:8005/health
- Docker images run as non-root and use BuildKit cache mounts for faster rebuilds
- Resource limits are set in
docker-compose.ymlto avoid a single service starving the host - ASR uses CUDA runtime images for smaller footprints
- HTTPS/TLS via reverse proxy
- API authentication + rate limiting
- Secrets management (vault or Docker secrets)
- Integration tests + structured logging
- Monitoring (Prometheus/Grafana)
- OpenAI Whisper for speech recognition
- XTTS v2 for text-to-speech
- Pyannote for speaker diarization
- Google Translate API for translations