Built to pass your security review.
DictateSI is the dictation tool a bank, insurer, or law firm's IT team can actually approve, because the architecture, not a promise, keeps voice data on the machine. The mainstream cloud tools send your audio to a server. DictateSI can't, by design.
The whole data path stays on the Mac.
Every arrow below is local. There is no arrow that leaves the device.
You speak
Audio captured to memory, never written to disk.
whisper.cpp
Transcribed on-device on the CPU / Neural Engine.
Pasted at cursor
Text inserted into your app via Accessibility.
Read aloud
Selected text spoken on-device. Optional.
Nothing leaves
No audio, no text, no telemetry transmitted.
AI cleanup runs on-device too (see below), so the default pipeline is local end to end. The only way text ever leaves the machine is if you, or your IT policy, explicitly opt into a cloud cleanup backend: your own Anthropic account, or an OpenAI-compatible provider or gateway. Even then, only text, never audio, and the audit log records every send.
What DictateSI does not do.
Speech is transcribed on-device by whisper.cpp. There is no cloud STT path in the app.
DictateSI sends no usage data, crash reports, or “anonymous” metrics.
DictateSI reads nothing on your screen except the text you explicitly select for Read Aloud.
DictateSI uses Accessibility to insert text, not to read it.
No sign-up, no user database, no email collection.
Updates are user-initiated. Air-gapped installs never call out at all.
Polished on-device by default. Cloud only if you ask.
DictateSI turns raw dictation into clean writing using a small AI model that runs entirely on your Mac: no key, no account, works offline. The whole pipeline (speech → text → polish) stays on the machine. Cloud (Claude) is an optional, deliberate, auditable alternative for those who want it.
On-device by default
A small open-licensed model (Qwen2.5, Apache-2.0) runs in-process via embedded llama.cpp. Nothing is transmitted: not audio, not text.
No key, works offline
New installs default to local cleanup. No Anthropic account, no API key, functions with the network off.
Cloud is opt-in only
Prefer a cloud model? Connect the account you already have: Anthropic natively, or OpenAI, Azure OpenAI, Gemini, Groq, Mistral and any OpenAI-compatible gateway. Keys live in the macOS Keychain. These are the only paths that send text off-device, and only text, never audio.
Fails safe
If cleanup errors, DictateSI pastes the raw local transcript. A local failure never silently escalates to the cloud.
MDM-lockable
Managed preferences, shipped: a configuration profile from Jamf, Kandji or Intune can pin cleanup to on-device, allow-list backends, force the audit log, or lock an approved endpoint. Policy can only narrow toward on-device. Sample profiles are in the deployment kit.
Don't take our word for it. Export the record.
Most privacy claims are marketing. DictateSI's are checkable, by your own team, on your own machines.
Tamper-evident audit log
Turn it on and every dictation appends a hash-chained record: which cleanup backend ran and whether anything left the Mac. Metadata only, never the transcript. Editing or deleting a record breaks the chain, and the built-in integrity check says so. Export the whole log to JSON or CSV and hand your reviewer the headline figure: off-device events: 0. Off by default; built for regulated deployments.
Machine-enforced invariants
"No networking in the local pipeline" isn't a policy. It's a build gate. Automated tests scan the source and fail the build if networking code ever enters the on-device transcription, cleanup, read-aloud, or audit paths, or if content ever enters the audit log. They run in CI on every commit, so no version of DictateSI can ship without passing them.
Verified model supply chain
The AI models DictateSI downloads are pinned to immutable revisions and SHA-256-verified against published checksums before install; a file that fails verification is deleted, never loaded. What you audited is what runs. Air-gapped installs can pre-position the same checksummed files and never touch the network at all.
The numbers behind the model choices.
Every model decision in DictateSI is made against a fixed test with the method written down. We publish our own numbers and how we got them. We do not publish comparisons with other products.
How we test
A set of 40 dictations written the way Whisper actually transcribes speech: fillers, repeated words, false starts, spoken numbers, and real mis-hearings, across banking, legal, clinical, and general business, including spoken lists, paragraph commands, and mid-sentence self-corrections. Each is paired with the text a careful editor would produce, following written conventions: money as $4,217.50, percentages as 18%, times as 2:30 p.m., dosages as 500 mg, and a self-correction keeps only the corrected value.
Every candidate model runs through the shipped pipeline: the same prompts, the same structure formatter, the same content guard. The score is word error rate against the expected text; 0 means every word matched. "Hard" cases need a convention, a mis-hearing fix, or a self-correction, not just filler removal.
What the numbers changed
- • The larger model is not the default. Its average is marginally better, but in two of 40 cases it changed a number: "seventeen dollars" became "fifteen dollars." A dictation tool for banks and law firms must never do that. A slightly rougher sentence is a better failure than a wrong figure.
- • The content guard exists because of a measurement. In one case the default model returned a paragraph minus its first sentence. Since 0.7.1 every backend compares its output with what you said and keeps your words if the model dropped too much.
- • A dictation-trained version of the default model is in evaluation. If it ships, its numbers on this same set appear here first.
| On-device cleanup model | Word error rate, all 40 | Standard cases | Hard cases | Time per dictation |
|---|---|---|---|---|
| Qwen2.5 1.5B (default) | 0.138 | 0.043 | 0.224 | 0.21 s |
| Qwen2.5 3B (optional) | 0.130 | 0.041 | 0.211 | 0.35 s |
Measured 2026-10-07 on an Apple M5 Pro with Metal; your Mac's timing will differ. The expected text is one good answer, not the only one, so a result of 0.05 can be a legitimate variant. Use these numbers to compare our models on the same set, not as an absolute grade. Neither base model applies the number conventions yet, which is most of the "hard" score.
Speech recognition, same Mac: on clean speech, Whisper small.en and large-v3-turbo transcribe the same words; turbo produces better sentence punctuation and takes 0.52 s per utterance against 0.38 s. Real-world speech widens the gap, which is why turbo is the default on 16 GB Apple Silicon Macs since 0.7.0.
Roll it out the way your review allows.
Direct download
PilotsDeveloper-signed binary, drag to Applications. For evaluation and solo practitioners.
MDM-deployable
Managed preferences: shippedShip the signed app via Jamf, Kandji, or Intune, and push a configuration profile that pins cleanup to on-device, allow-lists backends, forces the audit log on, or locks an approved endpoint. Locked controls read "Managed by your organization."
Air-gapped on-prem
RegulatedRuns with zero network access: local model, no telemetry, no update calls. The laptop never sees the internet.
Straight about where we are.
On-device by architecture. Because audio never leaves the Mac, DictateSI sidesteps the data-residency and transmission concerns that block cloud dictation tools in HIPAA, GLBA, and attorney-client-privilege contexts. With AI cleanup off, no dictation content is transmitted anywhere.
No certification we don't have. DictateSI is an early-stage product. We are not claiming SOC 2 or HIPAA certification today. We'll say so plainly when we have them. What we offer now is an architecture your team can verify themselves and a founder who will sit in your security review.
SOC 2 on the roadmap. SOC 2 Type II prep is planned as we move from pilots to broader deployment. If a specific control or questionnaire is blocking you, tell us. That feedback sets our priorities.