Changelog
What’s new in DictateSI.
Every shipped release, what changed, why it matters. The
source-of-truth document lives in the app repo as
CHANGELOG.md;
this page mirrors it.
A stutter is not content. The 0.7.1 guard counted a repeated phrase ("the Third Circuit, um, the Third Circuit")
as content and so rejected a correct cleanup; repeats now collapse before counting. Found by a second,
held-out evaluation set.
Never lose a sentence. The first measured pass over the on-device cleanup model caught it returning a
paragraph minus its first sentence. Cleanup now has a content guard on every backend.
- • Content guard. A model's output is compared with what you said, counting content rather than words: fillers and stutters do not count, and a spoken number and its written form ("four thousand two hundred seventeen dollars", "$4,217") are one unit each. An output that keeps too little is discarded and your raw dictation is kept instead. Applies to the on-device model, Ollama, Anthropic, and OpenAI-compatible endpoints.
- • Measured, in the open. The app repo now carries a 40-case dictation evaluation set with documented conventions and a harness that scores any model through the real pipeline. The shipped 1.5B model scores a word error rate of 0.138 against a careful editor; the 3B model is marginally better on average but changed numbers in two cases, so it stays optional. Method and numbers: Accuracy, measured.
Better ears. OpenAI's Whisper large-v3-turbo model joins the picker and becomes the default on
Apple Silicon Macs with 16 GB or more: the biggest single accuracy step available to an on-device app.
- • Hardware-tiered default. 16 GB+ Apple Silicon gets large-v3-turbo (547 MB, 5-bit quantized). Intel Macs stay on small.en, 8 GB Macs on base.en. Settings marks the recommended model for your Mac; your choice always wins.
- • Upgrade without interruption. The app keeps dictating with the model already on disk while the new one downloads in the background, verifies its SHA-256, and takes over at the next idle moment.
- • Measured, not promised. About half a second per utterance on an M5 Pro, against 0.38 s for small.en. Better sentence punctuation; a spoken "question mark" becomes "?". Turbo writes spoken lists as "Number 1. Banana. 2. Kiwi." and the formatter now reads that form too.
- • Same rule as every model download: pinned to an immutable revision and SHA-256-verified before install. Turbo is multilingual but DictateSI runs it in English.
Intent refinements from the first real dictations: "the following items: one apple, two bananas,
three mangoes" is now a list even without saying "number one," and spoken punctuation works.
- • A list introduction relaxes the markers. After "the following," "as follows," "here are," or a colon, "one…, two…, three…" are items, not quantities. Without an introduction, "I bought one apple and two bananas" stays prose.
- • Spoken punctuation. Say "colon," "comma," "period," "question mark" and you get the mark, not the word. Real words in a sentence ("the trial period") are untouched.
- • Conjunctions spoken inside lists are dropped from items, and the on-device model is told to keep spoken list markers as spoken.
The formatting release. Say the structure and you get the structure: "number one, two, three"
pastes as a numbered list under its introducing line, "bullet point" makes bullets, and "new paragraph"
does what it says. On every backend, including with cleanup off.
- • Structure comes from code, not from the model. A deterministic formatter lays out spoken lists, bullets, and paragraph breaks and never changes your words. It is conservative on purpose: "I have two cats" and "the first quarter" stay prose.
- • The small on-device model polishes wording only, block by block, and is never asked about layout. We measured it: given layout instructions it dropped words; given one item at a time it is accurate and fast.
- • Cloud models get the full brief. Anthropic and OpenAI-compatible backends receive formatting instructions, so connecting a Sonnet-class model gives Sonnet-class documents.
- • Type-out mode sends real Return keys for line breaks. New General setting to turn formatting off.
The IT release. On-device stays the default; this gives an institution's IT team real control
over it, and gives anyone the option to plug in the AI account they already have.
- • Managed preferences: configuration profiles from Jamf, Kandji or Intune can pin cleanup to on-device, allow-list backends, force the audit log on, turn local history off, or lock an approved endpoint. Locked controls read "Managed by your organization." Policy can only narrow toward on-device. Sample profiles ship in the deployment kit.
- • Connect the AI account you already have: a new OpenAI-compatible backend for OpenAI, Azure OpenAI, Google Gemini, Groq, Mistral, OpenRouter, or a gateway your IT team runs. Pick a provider, paste the endpoint and model, add a key if needed, test the connection. Anthropic keeps its own tab.
- • Honest audit records for it: endpoint runs are logged as
endpoint, and the off-device flag comes from the host, so a gateway on this Mac counts as on-device and anything else counts as off-device, even when the request fails.
The rename release. The product formerly called Ghost is now DictateSI, at dictatesi.com.
Nothing about how the app works changed; its identity did, along with a careful upgrade path.
- • New bundle identifier
com.dictatesi.mac. macOS treats this as a new app, so on first launch it asks again for Microphone and Accessibility. Grant both once. - • Automatic migration on first launch: settings, downloaded models, and the History and Audit stores move to the new identity. No re-download, and the audit hash chain is untouched.
- • Not migrated, on purpose: an Anthropic API key stored in the Keychain by the old build. Paste it again under Settings → AI Cleanup if you use the cloud backend.
- • App icon: the same ink tile and glyph as this site.
The efficiency release. DictateSI is now a polite guest on 8 GB Macs, with
hardware-aware defaults, GPU-accelerated cleanup, and a strict CPU budget.
- • Hardware-tiered defaults: 8 GB-class Macs default to the lighter Whisper base.en model and free the cleanup model (~1.4 GB) after 5 idle minutes; it reloads in about a second when needed. 16 GB+ Macs keep the fastest setup. A new "Keep model in memory" toggle makes the trade yours.
- • GPU-accelerated cleanup on Apple Silicon (Metal, unified memory, no extra RAM). Cleanup runs ~5–7× faster and no longer spikes the CPU cores your real work is using.
- • Burst-thread budget: transcription and cleanup are capped near the performance-core count instead of grabbing every core. No fan spin-ups, no stalling the app you're dictating into.
- • The optional 3B model is marked "not recommended" on 8 GB Macs.
The trust release: DictateSI's privacy claims stop being promises and become
evidence you can export, verify, and hand to a security reviewer.
Tamper-evident audit log
- • Settings → Audit Log (off by default): every dictation appends a hash-chained record of which cleanup backend ran and whether anything left the Mac. Metadata only, never the transcript.
- • Editing, reordering, or deleting a record breaks the chain; the built-in Verify integrity check detects and reports it.
- • Export the whole log to JSON or CSV with an integrity summary. The headline figure for a clean deployment: off-device events: 0.
- • The log reflects what cleanup actually did, not what was configured: a failed cloud request still counts as off-device; a local no-op never does.
Verified model supply chain
- • Every AI model download (Whisper + cleanup) is now pinned to an immutable revision and SHA-256-verified against a published checksum before install. A file that fails verification is deleted, never loaded.
- • The privacy-invariant test suite (now 31 tests) runs in CI on every commit; no build ships without passing the moat guards.
First-run experience
- • Onboarding now offers the on-device cleanup model up front (one-time ~1.1 GB download with progress) so it's never a surprise. Enabling is always an explicit tap, and an existing cloud configuration is never overridden.
The pipeline is now fully local: dictation is transcribed and polished
entirely on your Mac, with zero network and zero keys. No competitor offers frictionless
end-to-end on-device AI dictation.
On-device AI cleanup
- • Settings → AI Cleanup → On-device downloads a small open-licensed model once (Qwen2.5-1.5B, ~1.1 GB, Apache-2.0) and runs it in-process via embedded llama.cpp. No Ollama, no API key, works offline. ~1–2 s on Apple Silicon.
- • On-device is the new default for fresh installs; existing key-holders keep Cloud. "Nothing leaves the machine" is now the default cleanup posture.
- • Optional higher-quality Qwen2.5-3B download for 16 GB+ Macs; in-Settings on-device test.
- • Automated moat guard: tests fail the build if any networking code enters the on-device cleanup path.
DictateSI is now bidirectional voice: dictate text in, and listen to text read
back, both entirely on-device. A proofreading and accessibility capability the cloud-only
competition structurally can’t match.
Read Aloud: on-device text-to-speech (new)
- • Select text in any app and press ⌃⌥R to hear it read back; press again or Esc to stop. Rebindable in Settings → Read Aloud.
- • On-device voices via Apple’s speech synthesizer (default, Enhanced, Premium, Personal Voice): voice picker with live preview, plus adjustable speed and pitch.
- • A calm teal speaking overlay, distinct from the recording pill.
- • Auto-stops the read when you start a dictation, so you never dictate over the voice.
- • Nothing leaves your Mac: the selected text is never transmitted, never written to disk, never saved to history. No new permissions, no new network endpoints.
History controls (compliance hardening)
- • "Save dictation history" toggle: off means nothing new is recorded.
- • Retention window (forever / 1 / 7 / 30 / 90 days) with automatic purge. History never leaves the Mac either way; this bounds what exists at rest.
Under the hood
- • New automated test suite, including a privacy-invariant scan that fails the build if networking or persistence code ever enters the read-aloud path.
- • Global-hotkey handler now filters by hotkey ID, so a second registered hotkey no longer re-triggers the first.
- • Type-out output mode no longer blocks the UI on long dictations.
- • Onboarding reflects your actual hotkeys and introduces Read Aloud.
The v0.2 release turns DictateSI from a working demo loop into a real
product surface: full Settings, AI cleanup, vocabularies, per-app
rules, history, and a distribution pipeline.
Distribution & onboarding
- • First-run onboarding wizard (Welcome → Microphone → Accessibility → Try it → Done).
- • Full Settings window with sidebar nav: About, Permissions, General, Audio & Model, AI Cleanup, Vocabulary, Per-App Rules, History.
- • Menu-bar quick settings: mode indicator (🔒 Local / ☁ Cloud cleanup) plus inline toggles for “Polish with AI” and “Show recording overlay.”
- • Notarized-DMG distribution pipeline producing
DictateSI-X.Y.Z.dmg. - • Privacy & data flow doc covering exactly what stays local and what goes to the cloud.
Recording
- • Custom hotkey recorder: bind any key combo or hold any modifier alone (Fn, Right ⌘, etc.).
- • Double-tap-to-lock for hold-mode hotkeys: quick double-tap latches recording on for hands-free long-form dictation; tap once more to stop.
- • Notch-style overlay with live level bars, elapsed timer, and lock indicator when latched.
- • Output mode: paste (⌘V) or type-out per-character keystrokes for password fields and apps that block paste.
- • Optional start/stop sound effects.
- • Auto-stop on silence (configurable 1–5 s). Suppressed during locked recording.
- • Esc cancels mid-recording without pasting.
Transcription
- • Whisper model picker (tiny.en, base.en, small.en) with in-app download and live model swap (no relaunch).
- • Domain vocabulary packs: Banking, Insurance, Legal, Real Estate, Title. ~20–25 terms each.
- • User-editable custom vocabulary, fed to Whisper as
initial_prompt for recognition biasing.
AI cleanup (optional, opt-in)
- • BYOK Anthropic Messages API integration; API key in macOS Keychain only.
- • Model picker with curated presets (Opus 4.7, Sonnet 4.7, Sonnet 4.6, Haiku 4.5) plus a Custom model ID field for any current Anthropic model string.
- • Three tone presets (neutral, formal, casual).
- • Custom system prompt override for vertical-specific tuning.
- • Hosted-cleanup tier scaffolded in the UI (server arrives in v0.3).
- • Graceful fallback to raw Whisper on any cleanup error; it never blocks the paste.
Per-app rules
- • Captures the frontmost app at recording start so rules bind to the target app even if focus shifts.
- • Per-app overrides for AI cleanup on/off, tone, and vocabulary pack set.
History
- • Local SwiftData persistence of every successful transcription (timestamp, duration, source app, raw text, polished text, cleanup-used flag).
- • Searchable list with copy-to-clipboard, delete, clear-all. No syncing; database lives only on this Mac.
Fixed
- • Critical: audio capture was silently dropping ~99% of recorded audio because the AVAudioConverter input block signaled
.endOfStream after every tap callback. Switched to .noDataNow and added converter.reset() at recording start. - • Three residual “Quill” strings (pre-rebrand name) cleaned up in AppDelegate and entitlements.
0.1.0
2026-05-26 · pre-rebrand
Project scaffold under the working name “Quill”; rebranded
to DictateSI the same day. Never publicly distributed; listed for completeness.
- • macOS native Swift menu-bar app, XcodeGen + Makefile + stable Apple Developer signing.
- • ⌃⌥Space global hotkey via Carbon
RegisterEventHotKey. - • AVAudioEngine recording pipeline (16 kHz mono Float32 PCM).
- • whisper.cpp xcframework integration with ggml-small.en.
- • End-to-end loop: hotkey → record → whisper → paste at cursor.