alpharesearch/SmartDictate

C#

0

145 commits

updated Aug 21, 2026

See the code

README

SmartDictate

Getting Started

  1. Open in Visual Studio 2022:
    Open the SmartDictate.sln solution file.

  2. Restore NuGet packages:
    Visual Studio will prompt you to restore packages on first open.

  3. Build and run:
    Press F5 or click Start to build and launch the application.

Key Features

  • Consolidated Tabbed Settings Panel: Access all core and advanced settings (Whisper models, LLM parameters, audio input devices, custom hotkeys, VAD mode, chunking thresholds, prompt profiles, custom vocabulary prompts, and deterministic replacements list) in a single tabbed dialog.
  • Multi-State Visual Pipeline Indicator: View real-time visual status updates (Idle, Listening/Silent, Speech Detected, Processing) through the active color-coded status bar (lblStatusIndicator).
  • Interactive Clipboard Rerun (Rerun LLM button): Dynamically checks for available clipboard text history and enables quick LLM-refinement on-demand.
  • Deterministic Post-Processing & Vocabulary Replacements: Automatically map phonetic or misrecognized terms to exact preferred names (e.g. correcting "smart server" to "Sm@rtServer" or "site top" to "SITOP") using a configurable case-insensitive search-and-replace list.
  • Custom Vocabulary Prompts: Control whether vocabulary-preservation or sounds-like instructions are dynamically injected into the LLM system prompt via simple checkboxes.
  • Integrated Download Shortcuts: Directly click built-in link labels in the Models tab of the settings menu to download recommended high-performance local models.
  • Global Dictation Mode (CTRL + ALT + D): Dictate directly into any active windows application. The app types out your transcribed text right at your cursor.
  • Clipboard Proofreading (CTRL + ALT + P): Instantly refines your currently copied clipboard text using the local LLM to correct grammar, spelling, and punctuation, then auto-pastes it back.
  • Local LLM Refinement: Automatically proofreads and refines your dictations using local GGUF models.
  • Auto-Prompt Formatting: Automatically detects and applies the correct instruction templates for Qwen, Llama, and Gemma models based on the filename. You can manually override this behavior by setting LLMPromptTemplate in the appsettings.json file.
  • Real-time Resource Monitoring: Live tracking of System RAM and GPU VRAM (Application footprint vs. Total system usage).
  • Advanced VAD (Voice Activity Detection): Adjustable sensitivity settings to filter out background noise and handle automated chunking.

Basic Usage

  1. Click the Settings button to open the consolidated configuration panel.
  2. Select your microphone, VAD sensitivity mode, and configure Whisper/LLM models under their respective tabs:
    • Models: Select model file paths (click the built-in recommendation links to download the recommended Whisper Large v3 Turbo or Gemma 4 E4B models if you don't have them yet).
    • Audio & VAD: Select your input microphone, choose a VAD mode, and adjust silence thresholds for normal and dictation modes.
    • LLM & Prompts: Toggle LLM refinement, select from prompt profiles, or modify system and user prompts directly.
    • General: Configure real-time transcription displays and custom global keyboard hotkey overrides.
  3. Click Start on the main window to dictate locally, or use the Global Hotkeys (Ctrl + Alt + D to dictate, Ctrl + Alt + P to proofread clipboard).
  4. Watch the pipeline status indicator show exact states (Idle, Listening, Speech Detected, or Processing) as you speak.
  5. Use the context-aware Rerun LLM button (which is enabled whenever a valid text buffer is available) to re-refine your clipboard text.

Required Models

You must download and provide your own Whisper and LLM models:

Advanced Configuration (appsettings.json)

The application automatically generates an appsettings.json file on first run. You can edit this file to customize advanced behavior:

Customizing Hotkeys

You can change the global shortcut keys used for dictation and proofreading by modifying the Modifiers and Key settings:

"DictationHotkeyModifiers": "Control, Alt",
"DictationHotkeyKey": "D",
"ProofreadHotkeyModifiers": "Control, Alt",
"ProofreadHotkeyKey": "P"

Valid modifiers include Control, Alt, Shift, or combinations separated by commas. Keys can be any standard key like D, F12, NumPad1.

LLM Prompts and Formatting

  • LLMSystemPrompt / LLMUserPrompt: Customize the persona and instructions for the local LLM. The defaults are set up for strict copy editing.
  • LLMPromptTemplate: Leave as "" to use Auto-Prompt Formatting based on the model name. Set to a custom template string (e.g. <|im_start|>system\n{0}...) to manually override.
  • LLMAntiPrompts: A list of stop tokens to prevent the LLM from hallucinating conversational filler or running on indefinitely.
  • LLMContextSize / LLMTemperature / LLMMaxOutputTokens: Fine-tune the underlying local inference parameters to fit your hardware and selected model.

Audio & VAD Tweaks

  • VadGainMultiplier: Boosts the microphone volume only for the Voice Activity Detection analysis (Default: 1.0). Useful if your microphone is too quiet to trigger the VAD.
  • Silence Thresholds: Adjust NormalSilenceThresholdSeconds and DictationSilenceThresholdSeconds to control how long you can pause before the app considers a sentence finished.

Deterministic Replacements and Vocabulary Prompts

  • EnableVocabPrompt1 / EnableVocabPrompt2: Checkboxes to toggle whether vocabulary-preservation and sounds-like instructions are injected into the system prompt.
  • VocabPrompt1Text / VocabPrompt2Text: Editable system prompt templates containing placeholders for custom vocabulary.
  • VocabularyReplacements: A list of search-and-replace pairs applied as a deterministic post-processing pass directly to the LLM refined text.

Project Structure

  • MainForm.cs - Main application logic and UI event handling.
  • MainForm.Designer.cs - UI layout and control definitions.
  • MainForm.resx - Resource file for form localization and assets.

Notes

  • Model files are not included. Download them as described above.
  • Debug output can be enabled for troubleshooting.
  • All processing is done locally; no audio is sent to external servers.

Screenshot

image

License

MIT


Built with ❤️ using .NET 9 and Windows Forms.

icon

alpharesearch/SmartDictate

C#

0

145 commits

updated Aug 21, 2026

See the code

README

SmartDictate

Getting Started

  1. Open in Visual Studio 2022:
    Open the SmartDictate.sln solution file.

  2. Restore NuGet packages:
    Visual Studio will prompt you to restore packages on first open.

  3. Build and run:
    Press F5 or click Start to build and launch the application.

Key Features

  • Consolidated Tabbed Settings Panel: Access all core and advanced settings (Whisper models, LLM parameters, audio input devices, custom hotkeys, VAD mode, chunking thresholds, prompt profiles, custom vocabulary prompts, and deterministic replacements list) in a single tabbed dialog.
  • Multi-State Visual Pipeline Indicator: View real-time visual status updates (Idle, Listening/Silent, Speech Detected, Processing) through the active color-coded status bar (lblStatusIndicator).
  • Interactive Clipboard Rerun (Rerun LLM button): Dynamically checks for available clipboard text history and enables quick LLM-refinement on-demand.
  • Deterministic Post-Processing & Vocabulary Replacements: Automatically map phonetic or misrecognized terms to exact preferred names (e.g. correcting "smart server" to "Sm@rtServer" or "site top" to "SITOP") using a configurable case-insensitive search-and-replace list.
  • Custom Vocabulary Prompts: Control whether vocabulary-preservation or sounds-like instructions are dynamically injected into the LLM system prompt via simple checkboxes.
  • Integrated Download Shortcuts: Directly click built-in link labels in the Models tab of the settings menu to download recommended high-performance local models.
  • Global Dictation Mode (CTRL + ALT + D): Dictate directly into any active windows application. The app types out your transcribed text right at your cursor.
  • Clipboard Proofreading (CTRL + ALT + P): Instantly refines your currently copied clipboard text using the local LLM to correct grammar, spelling, and punctuation, then auto-pastes it back.
  • Local LLM Refinement: Automatically proofreads and refines your dictations using local GGUF models.
  • Auto-Prompt Formatting: Automatically detects and applies the correct instruction templates for Qwen, Llama, and Gemma models based on the filename. You can manually override this behavior by setting LLMPromptTemplate in the appsettings.json file.
  • Real-time Resource Monitoring: Live tracking of System RAM and GPU VRAM (Application footprint vs. Total system usage).
  • Advanced VAD (Voice Activity Detection): Adjustable sensitivity settings to filter out background noise and handle automated chunking.

Basic Usage

  1. Click the Settings button to open the consolidated configuration panel.
  2. Select your microphone, VAD sensitivity mode, and configure Whisper/LLM models under their respective tabs:
    • Models: Select model file paths (click the built-in recommendation links to download the recommended Whisper Large v3 Turbo or Gemma 4 E4B models if you don't have them yet).
    • Audio & VAD: Select your input microphone, choose a VAD mode, and adjust silence thresholds for normal and dictation modes.
    • LLM & Prompts: Toggle LLM refinement, select from prompt profiles, or modify system and user prompts directly.
    • General: Configure real-time transcription displays and custom global keyboard hotkey overrides.
  3. Click Start on the main window to dictate locally, or use the Global Hotkeys (Ctrl + Alt + D to dictate, Ctrl + Alt + P to proofread clipboard).
  4. Watch the pipeline status indicator show exact states (Idle, Listening, Speech Detected, or Processing) as you speak.
  5. Use the context-aware Rerun LLM button (which is enabled whenever a valid text buffer is available) to re-refine your clipboard text.

Required Models

You must download and provide your own Whisper and LLM models:

Advanced Configuration (appsettings.json)

The application automatically generates an appsettings.json file on first run. You can edit this file to customize advanced behavior:

Customizing Hotkeys

You can change the global shortcut keys used for dictation and proofreading by modifying the Modifiers and Key settings:

"DictationHotkeyModifiers": "Control, Alt",
"DictationHotkeyKey": "D",
"ProofreadHotkeyModifiers": "Control, Alt",
"ProofreadHotkeyKey": "P"

Valid modifiers include Control, Alt, Shift, or combinations separated by commas. Keys can be any standard key like D, F12, NumPad1.

LLM Prompts and Formatting

  • LLMSystemPrompt / LLMUserPrompt: Customize the persona and instructions for the local LLM. The defaults are set up for strict copy editing.
  • LLMPromptTemplate: Leave as "" to use Auto-Prompt Formatting based on the model name. Set to a custom template string (e.g. <|im_start|>system\n{0}...) to manually override.
  • LLMAntiPrompts: A list of stop tokens to prevent the LLM from hallucinating conversational filler or running on indefinitely.
  • LLMContextSize / LLMTemperature / LLMMaxOutputTokens: Fine-tune the underlying local inference parameters to fit your hardware and selected model.

Audio & VAD Tweaks

  • VadGainMultiplier: Boosts the microphone volume only for the Voice Activity Detection analysis (Default: 1.0). Useful if your microphone is too quiet to trigger the VAD.
  • Silence Thresholds: Adjust NormalSilenceThresholdSeconds and DictationSilenceThresholdSeconds to control how long you can pause before the app considers a sentence finished.

Deterministic Replacements and Vocabulary Prompts

  • EnableVocabPrompt1 / EnableVocabPrompt2: Checkboxes to toggle whether vocabulary-preservation and sounds-like instructions are injected into the system prompt.
  • VocabPrompt1Text / VocabPrompt2Text: Editable system prompt templates containing placeholders for custom vocabulary.
  • VocabularyReplacements: A list of search-and-replace pairs applied as a deterministic post-processing pass directly to the LLM refined text.

Project Structure

  • MainForm.cs - Main application logic and UI event handling.
  • MainForm.Designer.cs - UI layout and control definitions.
  • MainForm.resx - Resource file for form localization and assets.

Notes

  • Model files are not included. Download them as described above.
  • Debug output can be enabled for troubleshooting.
  • All processing is done locally; no audio is sent to external servers.

Screenshot

image

License

MIT


Built with ❤️ using .NET 9 and Windows Forms.

icon