Privacy

What Lispr depends on, and why we name it

September 4, 2026 · 7 min read

Most dictation tools describe their infrastructure in one of two ways. They say nothing at all, or they use language like "processed securely on trusted servers." Both tell you approximately the same thing: very little.

This post takes a different approach. It names what Lispr depends on, explains why those dependencies exist, and says what they mean for you. The law does not require it. We do it because you should be able to find out. Everything here describes Lispr as of September 2026, and when the chain changes, this post changes.

The chain, named

When you hold the dictation key and speak, your audio moves through a specific chain of components. Here it is, in order.

Your Mac or PC captures the audio locally. Nothing leaves the device yet.

On key release, the audio travels over an encrypted connection to a Cloudflare edge node. Cloudflare routes the request to an AI inference provider, which runs the speech recognition: Whisper large-v3-turbo, an open-weight model originally from OpenAI, hosted on the provider's hardware. The transcription comes back as text and lands at your cursor.

Two optional steps ride along in the same round trip. Add the translation key and your sentence is rendered into one of 32 target languages before the text returns, which takes the whole thing to roughly 0.7 seconds. Switch on AI formatting and a second pass adds punctuation, paragraphs and lists, or in concise mode cuts filler and repeats. Both of those run on servers rather than on your machine, and neither one is Whisper. They deserve the same treatment this post gives transcription, so they are getting their own post.

The Codebridge relay that sits between your device and the inference provider is stateless. It passes the audio through and keeps nothing: no logs, no transcripts, no record of what was said.

Some of Lispr depends on nothing outside your machine at all. Your personal dictionary is stored on the device. So is your dictation history. With formatting set to Off, you get the raw transcript and no formatting call is made.

Why cloud and not your device

Whisper large-v3-turbo is a large model. Running it locally would mean a download measured in gigabytes, a GPU that gets warm on every dictation, and latency that would make the tool feel like work. The provider's hardware runs the same model and returns a result a median of 346 ms after you release the key. The Lispr download is about 6 MB, roughly 17 MB once installed, and it asks nothing of your GPU.

That is the trade: speed and a small footprint in exchange for audio leaving your machine during transcription. We think it is the right trade for most people. But it should be named as a trade, not hidden inside a reassuring adjective.

What the upstream provider holds

We asked, and we read the policy. Three parties touch the audio, and they hold it for different lengths of time. The Codebridge relay keeps nothing. Cloudflare keeps nothing at the proxy layer. The inference provider holds the audio for up to 30 days for abuse review, then deletes it, and does not use it for training.

When we say elsewhere on this site that your audio is discarded after transcription, we mean discarded by us. Thirty days at the provider is the real number in the chain, and it was not ours to set.

One setting changes what we keep, and you are the one who changes it. Sharing recordings is off unless you switch it on. Turn it on and we keep those recordings and use them to improve transcription. Leave it alone and nothing is kept.

The parts that are not transcription

Transcription is not the only outside service in the picture. Paddle handles checkout and billing as our merchant of record, so your card details sit with Paddle and never reach us. The optional account holds an email address and a plan, and you sign in with a one-time code, which means there is no password for anyone to steal.

None of that is required to dictate. The free plan works with no account and no card.

The lock-in question

Lispr depends on an upstream inference provider for transcription. That is what choosing a cloud architecture means in practice: if the provider changes its terms, adjusts its pricing, or stops offering the service, Lispr's transcription changes. We chose our current provider because it is fast, handles real speech well, and is clear about what it does with audio.

If we switch providers, we will say so, in a post like this one, not buried in a changelog entry.

The alternative is to avoid naming the components and let "trusted infrastructure" carry the weight. Some tools prefer that. We don't.

What this does not change

Naming the dependency chain does not change what the chain does. If you need your audio to never leave your machine by design rather than by policy, Lispr is not the right fit. Applications that run Whisper locally (MacWhisper is the most common example) give you a structural guarantee that a cloud architecture cannot match: the audio stays on your hardware regardless of what any company promises.

We describe that trade-off in more detail in cloud vs on-device transcription.

Why say it plainly

There is a version of the software industry where companies treat their infrastructure choices as confidential. Sometimes for competitive reasons, sometimes because the specifics are not flattering. But "we run Whisper large-v3-turbo via Cloudflare" is not a competitive secret. It is just true, and anyone who thinks to ask already knows enough to verify it.

Tools built on vague language about data handling tend to earn exactly the trust that vagueness deserves. Saying the chain plainly costs us nothing and gives you something you can look up.

Common questions

Does my audio leave my Mac when I use Lispr?

Yes. Your Mac captures the audio locally, and on key release it travels over an encrypted connection to a Cloudflare edge node and on to the inference provider that runs Whisper large-v3-turbo. The text comes back and lands at your cursor. Windows works the same way. If your standard is that the audio must never leave your hardware, an app that runs Whisper locally is the better fit.

How long is my audio kept, and who keeps it?

Only one of the three parties in the chain keeps it at all. The Codebridge relay keeps nothing: no audio, no transcript, no record of what was said. Cloudflare passes the request through at the proxy layer and keeps no audio. The inference provider holds it for up to 30 days for abuse review and then deletes it, which makes thirty days the only real retention window. Your dictation history stays on your own device.

Is my voice used to train an AI model?

No. Whisper large-v3-turbo is a fixed open-weight model, so it does not learn from what you dictate, and the inference provider's terms exclude your audio from training. One setting is the exception. Sharing recordings is off unless you turn it on, and turning it on keeps those recordings so we can improve transcription.

Does Lispr work offline?

No. The speech model runs on the provider's hardware rather than yours, so transcription needs a connection. On a plane with no wifi, Lispr does nothing. If you dictate offline regularly, a local Whisper app is the right tool and this one is not.

If Lispr is free, who pays for the transcription?

Codebridge does, out of Pro subscriptions. The free plan costs nothing forever: 60 dictation minutes a week, resetting every Monday, with no account and no card. Dictation, instant translation into 32 languages, your personal dictionary and your local history are all in the free plan, and they never move behind a paywall. Lispr Pro is $79 a year or $7.99 a month, from $44.90 a year in some countries, and it lifts the minutes cap and adds AI formatting. New installs also get 30 days of Pro with no card. Those subscriptions cover the inference bill for everyone dictating for free, and nothing about the free plan is funded by your audio.

Try Lispr

Voice to text in any app: hold a key, talk, let go. Free, no sign-up.

Download Lispr

Language

Bahasa IndonesiaBahasa MelayuCatalàČeštinaDanskDeutschEnglishEspañolFrançaisHrvatskiItalianoMagyarNederlandsNorsk bokmålPolskiPortuguêsPortuguês (Brasil)RomânăSlovenčinaSuomiSvenskaTiếng ViệtTürkçeΕλληνικάРусскийУкраїнськаעבריתالعربيةहिन्दीไทย한국어日本語简体中文繁體中文