Skip to content

About

On-device voice dictation, command routing, and AI-assisted rewriting for macOS — a standalone, host-app-agnostic Swift package.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Latest commit

 

History

8 Commits

Folders and files

Repository files navigation

AiVoiceKit

Swift 6.0 iOS 18+ macOS 15+ visionOS 2+ GPLv3 License SPM

A Swift package providing on-device voice dictation, command routing, and AI-assisted rewriting for macOS (iOS/visionOS platform stubs included for future work).

Originally ported from FluidVoice by Aether AI Studio / altic-dev, and adapted for use as a standalone, host-app-agnostic package. AiVoiceKit has no dependency on Alric or any other host application — the host wires it up via a plain callback-based VoiceEngine protocol.

Sponsor NerdSnipe-Inc

AiVoiceKit is free and open source. If it saved you time, sponsoring NerdSnipe Inc pays for the maintenance, bug fixes and new releases that keep it working.

Install

// Package.swift
dependencies: [
    .package(url: "https://github.com/NerdSnipe-Inc/AiVoiceKit.git", branch: "master"),
],
targets: [
    .target(
        name: "YourApp",
        dependencies: [
            .product(name: "AiVoiceKit", package: "AiVoiceKit"),
        ]
    ),
]

Use branch: "master", not from: "<version>". AiVoiceKit pins its own FluidAudio and DynamicNotchKit dependencies to exact git revisions (see Package.swift), which SwiftPM treats as unstable. A tagged release like from: "1.1.1" forces strict semantic-version resolution and fails with no versions of 'aivoicekit' match the requirement ... depends on an unstable-version package 'fluidaudio'. This will be fixed once FluidAudio/DynamicNotchKit cut tagged releases that include the APIs this package needs.

What it does

  • Local ASR via Apple's built-in SFSpeechRecognizer/SpeechAnalyzer (zero download), or downloadable on-device models: Whisper (Tiny–Large via SwiftWhisper), Parakeet Flash/TDT v2/v3, Nemotron Speech, and Cohere Transcribe (all via FluidAudio, CoreML, arm64 macOS)
  • Global hotkey activation — hold-to-record, toggle, or double-tap, with configurable shortcuts for dictation, command mode, edit mode, cancel, and paste-last-transcription
  • Command routing — text prefixed with the wake word "Alric" or "Hey Alric" (e.g. "Alric, summarize this"; a comma, period or space may follow the name) is routed to a host-supplied callback instead of being typed into the frontmost app. The wake word is currently fixed in CommandModeRouter, not configurable per host app.
  • Edit mode — rewrite selected text in the frontmost app via a host-supplied AI callback
  • Notch/bottom overlay — live transcript, waveform, and mode indicator via DynamicNotchKit
  • History, stats, and custom dictionary persistence, file-based (no external database)
  • AI post-processing — punctuation, capitalization, and provider-based (Apple Intelligence / OpenAI / Groq / custom) dictation cleanup

Integration

The host app owns zero AiVoiceKit-specific UI logic beyond wiring two callbacks:

let voiceEngine = VoiceEngineMacOS(
    onCommandReceived: { text in
        await MyApp.shared.handleVoiceCommand(text)
    },
    onEditRequested: { selectedText, instruction in
        return await MyApp.shared.rewrite(selectedText, instruction: instruction)
    }
)

VoiceEngine is a plain AnyObject, ObservableObject protocol — no host framework types appear in the package's public API.

Example app

AiVoiceKit's dictation and voice-command flow is used end-to-end in AICompleteChat, a full-source, production-quality on-device AI chat app built on DesignFoundation and DesignFoundationPro — NerdSnipe Inc's Swift design-system packages. DesignFoundation is the free, MIT-licensed core (tokens, primitives, layout shells); DesignFoundationPro builds on it with complete, drop-in application verticals (chat, CRM, project management, analytics, and more) so a real app's UI is assembled from proven components rather than built from scratch. AICompleteChat is the reference for wiring AiVoiceKit's VoiceEngine into a DesignFoundationPro AIChat vertical — voice input, persistent memory (AiPersona), and on-device MLX inference, all running locally with no cloud dependency.

Package layout

Sources/
  AiVoiceKit/
    Public/     — protocols, types, enums (VoiceEngine, ASRModel, VoiceSettings)
    macOS/      — CoreAudio, NSEvent, Accessibility, overlay (macOS only)
    Shared/     — ASR catalog, history, stats, settings persistence (cross-platform)
  CoreAudioCaptureSupport/ — C target for low-level CoreAudio capture
Tests/
  AiVoiceKitTests/

Requirements

  • macOS 15+ (primary target). iOS 18+/visionOS 2+ platforms are declared but currently compile no macOS-only code paths in — full support is planned but not yet implemented.
  • Swift 6.0 toolchain
  • Microphone access; Accessibility permission for global hotkeys and typed output

License

AiVoiceKit is a derivative work of FluidVoice (GPLv3, © altic-dev contributors) and is distributed under the same license — see LICENSE.

Third-party dependencies and their licenses are listed in-app via the Open Source Licenses sheet, and include:

Dependency License Purpose
FluidAudio MIT Parakeet/Nemotron/Cohere CoreML ASR
SwiftWhisper MIT Whisper inference
DynamicNotchKit MIT Notch overlay window management
PromiseKit MIT Legacy async utilities (updater path)

Support this project

AiVoiceKit is built and maintained by NerdSnipe Inc, a small independent studio in Ottawa. Sponsorship funds new ASR model support.

About

On-device voice dictation, command routing, and AI-assisted rewriting for macOS — a standalone, host-app-agnostic Swift package.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages