Changelog

What’s new in DictateSI.

Every shipped release, what changed, why it matters. The source-of-truth document lives in the app repo as CHANGELOG.md; this page mirrors it.

0.7.2

2026-10-07

A stutter is not content. The 0.7.1 guard counted a repeated phrase ("the Third Circuit, um, the Third Circuit") as content and so rejected a correct cleanup; repeats now collapse before counting. Found by a second, held-out evaluation set.

0.7.1

2026-10-07

Never lose a sentence. The first measured pass over the on-device cleanup model caught it returning a paragraph minus its first sentence. Cleanup now has a content guard on every backend.

  • • Content guard. A model's output is compared with what you said, counting content rather than words: fillers and stutters do not count, and a spoken number and its written form ("four thousand two hundred seventeen dollars", "$4,217") are one unit each. An output that keeps too little is discarded and your raw dictation is kept instead. Applies to the on-device model, Ollama, Anthropic, and OpenAI-compatible endpoints.
  • • Measured, in the open. The app repo now carries a 40-case dictation evaluation set with documented conventions and a harness that scores any model through the real pipeline. The shipped 1.5B model scores a word error rate of 0.138 against a careful editor; the 3B model is marginally better on average but changed numbers in two cases, so it stays optional. Method and numbers: Accuracy, measured.

0.7.0

2026-10-07

Better ears. OpenAI's Whisper large-v3-turbo model joins the picker and becomes the default on Apple Silicon Macs with 16 GB or more: the biggest single accuracy step available to an on-device app.

  • • Hardware-tiered default. 16 GB+ Apple Silicon gets large-v3-turbo (547 MB, 5-bit quantized). Intel Macs stay on small.en, 8 GB Macs on base.en. Settings marks the recommended model for your Mac; your choice always wins.
  • • Upgrade without interruption. The app keeps dictating with the model already on disk while the new one downloads in the background, verifies its SHA-256, and takes over at the next idle moment.
  • • Measured, not promised. About half a second per utterance on an M5 Pro, against 0.38 s for small.en. Better sentence punctuation; a spoken "question mark" becomes "?". Turbo writes spoken lists as "Number 1. Banana. 2. Kiwi." and the formatter now reads that form too.
  • • Same rule as every model download: pinned to an immutable revision and SHA-256-verified before install. Turbo is multilingual but DictateSI runs it in English.

0.6.1

2026-10-07

Intent refinements from the first real dictations: "the following items: one apple, two bananas, three mangoes" is now a list even without saying "number one," and spoken punctuation works.

  • • A list introduction relaxes the markers. After "the following," "as follows," "here are," or a colon, "one…, two…, three…" are items, not quantities. Without an introduction, "I bought one apple and two bananas" stays prose.
  • • Spoken punctuation. Say "colon," "comma," "period," "question mark" and you get the mark, not the word. Real words in a sentence ("the trial period") are untouched.
  • • Conjunctions spoken inside lists are dropped from items, and the on-device model is told to keep spoken list markers as spoken.

0.6.0

2026-10-07

The formatting release. Say the structure and you get the structure: "number one, two, three" pastes as a numbered list under its introducing line, "bullet point" makes bullets, and "new paragraph" does what it says. On every backend, including with cleanup off.

  • • Structure comes from code, not from the model. A deterministic formatter lays out spoken lists, bullets, and paragraph breaks and never changes your words. It is conservative on purpose: "I have two cats" and "the first quarter" stay prose.
  • • The small on-device model polishes wording only, block by block, and is never asked about layout. We measured it: given layout instructions it dropped words; given one item at a time it is accurate and fast.
  • • Cloud models get the full brief. Anthropic and OpenAI-compatible backends receive formatting instructions, so connecting a Sonnet-class model gives Sonnet-class documents.
  • • Type-out mode sends real Return keys for line breaks. New General setting to turn formatting off.

0.5.0

2026-10-07

The IT release. On-device stays the default; this gives an institution's IT team real control over it, and gives anyone the option to plug in the AI account they already have.

  • • Managed preferences: configuration profiles from Jamf, Kandji or Intune can pin cleanup to on-device, allow-list backends, force the audit log on, turn local history off, or lock an approved endpoint. Locked controls read "Managed by your organization." Policy can only narrow toward on-device. Sample profiles ship in the deployment kit.
  • • Connect the AI account you already have: a new OpenAI-compatible backend for OpenAI, Azure OpenAI, Google Gemini, Groq, Mistral, OpenRouter, or a gateway your IT team runs. Pick a provider, paste the endpoint and model, add a key if needed, test the connection. Anthropic keeps its own tab.
  • • Honest audit records for it: endpoint runs are logged as endpoint, and the off-device flag comes from the host, so a gateway on this Mac counts as on-device and anything else counts as off-device, even when the request fails.

0.4.0

2026-10-06

The rename release. The product formerly called Ghost is now DictateSI, at dictatesi.com. Nothing about how the app works changed; its identity did, along with a careful upgrade path.

  • • New bundle identifier com.dictatesi.mac. macOS treats this as a new app, so on first launch it asks again for Microphone and Accessibility. Grant both once.
  • • Automatic migration on first launch: settings, downloaded models, and the History and Audit stores move to the new identity. No re-download, and the audit hash chain is untouched.
  • • Not migrated, on purpose: an Anthropic API key stored in the Keychain by the old build. Paste it again under Settings → AI Cleanup if you use the cloud backend.
  • • App icon: the same ink tile and glyph as this site.

0.3.1

2026-07-04

The efficiency release. DictateSI is now a polite guest on 8 GB Macs, with hardware-aware defaults, GPU-accelerated cleanup, and a strict CPU budget.

  • • Hardware-tiered defaults: 8 GB-class Macs default to the lighter Whisper base.en model and free the cleanup model (~1.4 GB) after 5 idle minutes; it reloads in about a second when needed. 16 GB+ Macs keep the fastest setup. A new "Keep model in memory" toggle makes the trade yours.
  • • GPU-accelerated cleanup on Apple Silicon (Metal, unified memory, no extra RAM). Cleanup runs ~5–7× faster and no longer spikes the CPU cores your real work is using.
  • • Burst-thread budget: transcription and cleanup are capped near the performance-core count instead of grabbing every core. No fan spin-ups, no stalling the app you're dictating into.
  • • The optional 3B model is marked "not recommended" on 8 GB Macs.

0.3.0

2026-07-02

The trust release: DictateSI's privacy claims stop being promises and become evidence you can export, verify, and hand to a security reviewer.

Tamper-evident audit log

  • • Settings → Audit Log (off by default): every dictation appends a hash-chained record of which cleanup backend ran and whether anything left the Mac. Metadata only, never the transcript.
  • • Editing, reordering, or deleting a record breaks the chain; the built-in Verify integrity check detects and reports it.
  • • Export the whole log to JSON or CSV with an integrity summary. The headline figure for a clean deployment: off-device events: 0.
  • • The log reflects what cleanup actually did, not what was configured: a failed cloud request still counts as off-device; a local no-op never does.

Verified model supply chain

  • • Every AI model download (Whisper + cleanup) is now pinned to an immutable revision and SHA-256-verified against a published checksum before install. A file that fails verification is deleted, never loaded.
  • • The privacy-invariant test suite (now 31 tests) runs in CI on every commit; no build ships without passing the moat guards.

First-run experience

  • • Onboarding now offers the on-device cleanup model up front (one-time ~1.1 GB download with progress) so it's never a surprise. Enabling is always an explicit tap, and an existing cloud configuration is never overridden.

0.2.3

2026-06-29

The pipeline is now fully local: dictation is transcribed and polished entirely on your Mac, with zero network and zero keys. No competitor offers frictionless end-to-end on-device AI dictation.

On-device AI cleanup

  • • Settings → AI Cleanup → On-device downloads a small open-licensed model once (Qwen2.5-1.5B, ~1.1 GB, Apache-2.0) and runs it in-process via embedded llama.cpp. No Ollama, no API key, works offline. ~1–2 s on Apple Silicon.
  • • On-device is the new default for fresh installs; existing key-holders keep Cloud. "Nothing leaves the machine" is now the default cleanup posture.
  • • Optional higher-quality Qwen2.5-3B download for 16 GB+ Macs; in-Settings on-device test.
  • • Automated moat guard: tests fail the build if any networking code enters the on-device cleanup path.

0.2.1

2026-06-08

DictateSI is now bidirectional voice: dictate text in, and listen to text read back, both entirely on-device. A proofreading and accessibility capability the cloud-only competition structurally can’t match.

Read Aloud: on-device text-to-speech (new)

  • • Select text in any app and press ⌃⌥R to hear it read back; press again or Esc to stop. Rebindable in Settings → Read Aloud.
  • • On-device voices via Apple’s speech synthesizer (default, Enhanced, Premium, Personal Voice): voice picker with live preview, plus adjustable speed and pitch.
  • • A calm teal speaking overlay, distinct from the recording pill.
  • • Auto-stops the read when you start a dictation, so you never dictate over the voice.
  • • Nothing leaves your Mac: the selected text is never transmitted, never written to disk, never saved to history. No new permissions, no new network endpoints.

History controls (compliance hardening)

  • • "Save dictation history" toggle: off means nothing new is recorded.
  • • Retention window (forever / 1 / 7 / 30 / 90 days) with automatic purge. History never leaves the Mac either way; this bounds what exists at rest.

Under the hood

  • • New automated test suite, including a privacy-invariant scan that fails the build if networking or persistence code ever enters the read-aloud path.
  • • Global-hotkey handler now filters by hotkey ID, so a second registered hotkey no longer re-triggers the first.
  • • Type-out output mode no longer blocks the UI on long dictations.
  • • Onboarding reflects your actual hotkeys and introduces Read Aloud.

0.2.0

2026-05-26

The v0.2 release turns DictateSI from a working demo loop into a real product surface: full Settings, AI cleanup, vocabularies, per-app rules, history, and a distribution pipeline.

Distribution & onboarding

  • • First-run onboarding wizard (Welcome → Microphone → Accessibility → Try it → Done).
  • • Full Settings window with sidebar nav: About, Permissions, General, Audio & Model, AI Cleanup, Vocabulary, Per-App Rules, History.
  • • Menu-bar quick settings: mode indicator (🔒 Local / ☁ Cloud cleanup) plus inline toggles for “Polish with AI” and “Show recording overlay.”
  • • Notarized-DMG distribution pipeline producing DictateSI-X.Y.Z.dmg.
  • • Privacy & data flow doc covering exactly what stays local and what goes to the cloud.

Recording

  • • Custom hotkey recorder: bind any key combo or hold any modifier alone (Fn, Right ⌘, etc.).
  • • Double-tap-to-lock for hold-mode hotkeys: quick double-tap latches recording on for hands-free long-form dictation; tap once more to stop.
  • • Notch-style overlay with live level bars, elapsed timer, and lock indicator when latched.
  • • Output mode: paste (⌘V) or type-out per-character keystrokes for password fields and apps that block paste.
  • • Optional start/stop sound effects.
  • • Auto-stop on silence (configurable 1–5 s). Suppressed during locked recording.
  • • Esc cancels mid-recording without pasting.

Transcription

  • • Whisper model picker (tiny.en, base.en, small.en) with in-app download and live model swap (no relaunch).
  • • Domain vocabulary packs: Banking, Insurance, Legal, Real Estate, Title. ~20–25 terms each.
  • • User-editable custom vocabulary, fed to Whisper as initial_prompt for recognition biasing.

AI cleanup (optional, opt-in)

  • • BYOK Anthropic Messages API integration; API key in macOS Keychain only.
  • • Model picker with curated presets (Opus 4.7, Sonnet 4.7, Sonnet 4.6, Haiku 4.5) plus a Custom model ID field for any current Anthropic model string.
  • • Three tone presets (neutral, formal, casual).
  • • Custom system prompt override for vertical-specific tuning.
  • • Hosted-cleanup tier scaffolded in the UI (server arrives in v0.3).
  • • Graceful fallback to raw Whisper on any cleanup error; it never blocks the paste.

Per-app rules

  • • Captures the frontmost app at recording start so rules bind to the target app even if focus shifts.
  • • Per-app overrides for AI cleanup on/off, tone, and vocabulary pack set.

History

  • • Local SwiftData persistence of every successful transcription (timestamp, duration, source app, raw text, polished text, cleanup-used flag).
  • • Searchable list with copy-to-clipboard, delete, clear-all. No syncing; database lives only on this Mac.

Fixed

  • • Critical: audio capture was silently dropping ~99% of recorded audio because the AVAudioConverter input block signaled .endOfStream after every tap callback. Switched to .noDataNow and added converter.reset() at recording start.
  • • Three residual “Quill” strings (pre-rebrand name) cleaned up in AppDelegate and entitlements.

0.1.0

2026-05-26 · pre-rebrand

Project scaffold under the working name “Quill”; rebranded to DictateSI the same day. Never publicly distributed; listed for completeness.

  • • macOS native Swift menu-bar app, XcodeGen + Makefile + stable Apple Developer signing.
  • • ⌃⌥Space global hotkey via Carbon RegisterEventHotKey.
  • • AVAudioEngine recording pipeline (16 kHz mono Float32 PCM).
  • • whisper.cpp xcframework integration with ggml-small.en.
  • • End-to-end loop: hotkey → record → whisper → paste at cursor.