Auri VoiceAuri Keyboard

Speech-to-Text for Sensitive Work (Legal, Medical, HR)

Answer first

Start with the rule your organization already has, then pick tools that fit it. Local recognition keeps audio and text off servers, which removes one category of risk and leaves others: a lost laptop, a cloud backup, an open-plan office, and disclosure obligations you cannot delegate to software. Build two lanes so the sensitive one is a habit rather than a decision made under time pressure. Auri Voice runs recognition and cleanup on device by default, and that is a technical property, not a compliance certification.

Start with the rule, then pick the tool

Sensitive work already has rules: privilege, patient confidentiality, employment law, an approved-tools list. The tool decision follows from those, and a tool cannot change your obligations.

  • Which categories of content are you handling: privileged, clinical, personnel, financial?
  • Does your organization maintain an approved-software list?
  • Are there contractual requirements, such as an agreement a vendor must sign?
  • Where are drafts allowed to live, and for how long?
  • Who signs off on a new tool, and what do they need from you?

Answer those first and the shortlist gets short fast. Many workplaces allow a local dictation app more readily than a cloud service, because the review question changes from data processing agreements to endpoint software.

Write down what you decided and who agreed. A one-paragraph note in your own files is what protects you if the question comes up a year later.

What local recognition solves, and what it does not

On-device recognition means audio is transcribed on your machine rather than uploaded. That is a real reduction in exposure and a narrow one.

RiskLocal recognition helps?What handles it
Audio sent to a serverYesOn-device engine, cloud off
Vendor retention of transcriptsYes, if storage is local tooLocal history, delete path
Lost or stolen laptopNoDisk encryption, screen lock
Transcript in a cloud backupNoBackup scope review
Being overheardNoClosed room, or typing
Disclosure and retention dutiesNoPolicy and legal advice

Keep the two questions apart: where processing happens and where text is stored. An app can transcribe locally and still write searchable history into a folder that syncs.

Also check the fallback path. If a local model is missing or unavailable for a language, some apps quietly use a server instead. Test with the network off so you know what happens when the local path is not available.

Honest limits

No dictation vendor, us included, can hand you compliance. Software provides technical properties. Compliance is a program made of policies, contracts, training, access control, retention schedules, and evidence.

  • On-device processing is a control, not a certificate.
  • Where an agreement with a vendor is required, that is a contract question, not a settings question.
  • Your professional judgment about what to dictate cannot be delegated to a tool.
  • Marketing language about privacy is not a substitute for your security review.

Auri Voice keeps speech to text and text improvement on device by default. That is what it does. It is not a HIPAA certification, and it does not decide whether your workflow is acceptable in your jurisdiction or your firm.

Bring this to the people who own the question: your privacy officer, general counsel, security team, or professional body. Their answer is the one that counts, and getting it in writing takes one email.

Build two lanes so you stop deciding under pressure

The failure mode for sensitive work is a judgment call made while you are already talking and behind schedule. Two prepared lanes remove the decision.

Sensitive lane

  • Local engine only, cloud features off, no server-side rewrite.
  • History off or cleared at the end of the session.
  • Drafting only into approved apps and storage locations.
  • Door closed, headphones off, nobody else in earshot.

Normal lane

  • Everything else: cloud polish, phone keyboard, whatever is fastest.

Make switching physical rather than mental. A second app profile, a different shortcut, a note taped near the camera, a calendar block labeled for the sensitive work. Anything that makes the wrong lane require an extra step.

Then rehearse the sensitive lane once when nothing is urgent. Confirm recognition works with the network off, insertion works in the approved app, and history behaves the way you expect.

Dictate without saying the identifiers

You can often get the drafting speed without putting identifiers into audio or text. Speak the structure and add the details by hand.

  • Use initials, a case number, or a role instead of a full name.
  • Leave a placeholder for dates of birth, account numbers, and identifiers.
  • Dictate the reasoning and paste the facts from the source of record.
  • Skip diagnoses and allegations aloud when a shorthand works.
  • Fill placeholders at the keyboard in the approved system.

This habit also reduces the cost of a mistake. A draft that says the complainant and the March incident is far less damaging in the wrong folder than one with names and dates throughout.

It helps accuracy too. Names and long numbers are exactly what dictation gets wrong most often, so typing them removes both a privacy risk and an editing chore.

Device and disk hygiene

Once text lives on your machine, your machine is the control. These checks take fifteen minutes and cover the failure modes local processing does not touch.

  1. Turn on full disk encryption and confirm the recovery key is stored properly.
  2. Set a short screen lock and require the password immediately.
  3. Review what your backups include, including any folder holding dictation history.
  4. Check clipboard managers, screen recorders, and note apps that sync by default.
  5. Decide a retention rule for drafts, then follow it.
  6. Be careful while screen sharing, since inserted text appears live.

Clipboard history deserves special attention. Utilities that keep everything you copied will happily archive the paragraph you were careful about, and they are easy to forget.

Screen sharing catches people mid-meeting. Text lands at the cursor as you speak, so dictating a sensitive note while sharing a window shows it to the room.

Questions to bring to your reviewer

Security reviews go faster when you arrive with the questions already framed. These map to what a reviewer needs to sign off.

  • Where does recognition run, and what happens with no network?
  • Where is history stored, is it encrypted at rest, and how is it deleted?
  • Which features call a server, and can they be disabled centrally?
  • Which third parties are involved, and under what retention terms?
  • What minimum OS version and permissions does the app need?
  • What does the vendor state in writing, and how is it dated?

For Auri Voice the short answers are: recognition and text improvement run on device by default, cloud features are optional, dictation works with the network off once an engine is downloaded, Apple Watch audio to the Mac is encrypted in transit and transcribed locally, and it requires macOS 15 or later.

Get the rest in writing from us and from any other vendor you are comparing. Then let the people who own the compliance program make the call.

FAQ

Does local dictation make me HIPAA or GDPR compliant?

No. On-device processing reduces where data travels, and compliance depends on your policies, contracts, training, access controls, retention rules, and documentation. Treat a local app as one control among many, and get sign-off from the people responsible for that program.

Can I dictate privileged material at all?

Many people do, with rules: local recognition, cloud features off, a private room, no client names spoken aloud where possible, and drafts stored only in approved systems. Confirm with whoever owns your confidentiality policy before you build the habit.

Is a phone keyboard safe for sensitive text?

Treat it as the wrong lane by default. Auri Keyboard defaults to cloud voice typing on iPhone, with local speech to text available. If sensitive content has to happen on a phone, change that setting first and skip the server-side features.

What about being overheard?

It is the risk local processing does nothing about, and often the most likely one. Open-plan offices, shared homes, cafes, and trains all carry your voice further than you think. Use a closed room, or type.

What should I bring to a security review?

Where recognition runs, what happens with no network, where history is stored and for how long, how deletion works, whether backups capture it, which third parties are involved, and what changes when a cloud feature is enabled. Written answers, not screenshots of marketing.