AI Engine & Customization

Models Library & Speech Engine

Explore local Whisper GGUFs, Cloud ASR, LLM refinement, quantization, and RAM optimization.

Dicty features a modular AI architecture supporting on-device hardware-accelerated models and ultra-accurate cloud intelligence.


Local Speech Models (ASR)#

Dicty Models Library & Speech Recognition

Local models run 100% offline on your Mac using whisper.cpp optimized with Apple Metal GPU shaders.

ModelSizeSpeedAccuracyBest Used For
Nano (EN / Multi)~75 MB⚡️⚡️⚡️ 1.0★★☆☆☆ (0.4)Ultra-light background typing on older machines.
Base (EN / Multi)~140 MB⚡️⚡️⚡️ 0.8★★★☆☆ (0.6)Fast baseline for clear speakers.
Fast (EN / Multi)~460 MB⚡️⚡️ 0.6★★★★☆ (0.8)Sweet spot: Fast response and high everyday accuracy.
Pro (EN / Multi)~1.5 GB⚡️ 0.4★★★★★ (0.9)Complex technical vocabulary and code names.
Large V3 Turbo~1.6 GB⚡️⚡️ 0.8★★★★★ (0.98)Recommended: Near Large V3 accuracy at 4x inference speed.
Large V3~3.1 GB⚡️ 0.2★★★★★ (1.0)Studio-grade acoustic precision and heavy accents.
Distil Large V3~1.5 GB⚡️⚡️⚡️ 1.0★★★★★ (0.98)Distilled model for English dictation.

Model Specifications & Quality Inspection#

Click the info badge next to any installed speech model to inspect its technical profile, installed quantization tier, and relative benchmarking stats:

Installed Model Specifications Card
  • Status Indicator: Shows Ready • Offline when the weights are validated and available for instant offline transcription.
  • Installed Version: Displays active disk footprint (e.g., 140 MB for Base Q5).
  • Benchmark Gauges: Visual meters for real-time inference speed (e.g., 80%) and acoustic accuracy (e.g., 60%).

Specialized Regional Models#

  • BELLE (Chinese - 3.1 GB): Fine-tuned for Mandarin and Chinese dialects.
  • Large V2 (French - 3.1 GB): Trained on thousands of hours of French audio to resolve silent letter ambiguities.
  • Large V3 Turbo (Japanese - 1.6 GB): Fine-tuned to eliminate Japanese hallucinations.

Understanding Quantization & Download Tiers#

When downloading local models from the library, Dicty provides three precision profiles tailored for different hardware setups:

Model Download Quantization Dialog
  1. Optimized Studio (Q5) — Recommended:
    • Retains ~99% of full 16-bit acoustic accuracy while cutting memory and disk requirements by 50% (e.g. 3.1 GB for Large V1/V3).
    • Provides optimal performance on Apple Silicon unified memory (M1 or later).
  2. Uncompressed (F16) — Maximum Precision:
    • 16-bit uncompressed floating-point weights (~10.2 GB).
    • Best for studio environments, heavy background noise, or subtle phoneme distinctions.
  3. Ultra Light (Q4) — Minimal Footprint:
    • 4-bit integer weights (~2.3 GB).
    • Blazing-fast inference speed and minimal battery usage for laptop multitasking.

LLM Text Refinement & Advanced Model Settings#

Dicty pairs raw speech recognition with local and cloud Large Language Models to format book prose, remove verbal disfluencies, and draft structured notes:

LLM Refinement Models & Advanced Model Settings

LLM Text Refinement Models#

  • Local GGUF Models: Download and run lightweight instruction models locally (e.g. Qwen 2.5 1.5B for fast offline punctuation, Llama 3.2 3B for comprehensive editing).
  • Cloud LLMs: Connect to top-tier reasoning engines (DeepSeek V3, GPT-4o Mini, Qwen 2.5 72B, Claude 3.5 Sonnet) with one click.

Advanced Model Settings#

Control SettingConfigurable OptionsFunctional Description
Translation MethodNone / Built-in / LM Studio / CloudSelects whether translation uses on-device Whisper, local LM Studio, or cloud LLMs.
Spoken Language90+ Languages / Auto DetectSupports 90+ languages (English, Spanish, German, French, Japanese, Russian, Ukrainian, Chinese, etc.). Explicitly selecting your language improves speed and accuracy over Auto Detect.
Translate to EnglishToggle ON / OFFDirectly translates spoken foreign language audio into fluent English text during transcription.
Keep model in RAM (Fast)Toggle ON / OFFKeeps speech recognition weights warm in Apple Silicon unified memory for zero-latency startup. Disable on 8GB RAM Macs to conserve system memory.
Auto-Complete Sentence EndingsToggle ON / OFFUses acoustic context to predict sentence endings and suppresses common Whisper silence hallucinations.