tl;dr
Say the messy version of your prompt out loud, let Auri's cleanup pass remove the filler and self-corrections, then spend fifteen seconds sharpening the ask before you send. Speaking gets far more context into a prompt than typing does, and context is what most weak prompts are missing. Cover five things: the role, the output you want, the context, the constraints, and one example. Auri Voice does this on Mac, where transcription runs locally by default; Auri Keyboard does it on iPhone.
Speak the messy brief, then fix the prompt
Good prompts are mostly context, and typing makes people stingy with context. The loop takes about a minute.
- Talk for thirty to sixty seconds as if briefing a coworker who will do the task.
- Let the cleanup pass turn that into readable paragraphs.
- Read it once, move the ask to the top, and cut what you talked yourself out of.
Speaking gets the detail in: in a minute you will mention the audience, the deadline, the thing you already tried, and the constraint you would never have typed. The read-through fixes the order, because speech wanders toward the conclusion and models respond better when the request comes before the background.
Why not ChatGPT's built-in mic?
ChatGPT's own voice input works fine inside ChatGPT. If that is the only place you write prompts, it may be all you need.
A dictation app earns its place through scope. One shortcut and one cleanup pass behave the same in ChatGPT, Claude, Gemini, your terminal, and the email where you paste the answer, so you build one habit instead of one per app. Custom vocabulary carries your project's jargon everywhere: add your product and library names once in Auri Voice under Settings > Dictation > Text > Vocabulary and they come out spelled right in every prompt. And Auri Voice transcribes on your Mac by default, so the audio of your prompt stays on the machine.
Five things to say out loud
Hit five slots in any order while talking, then check for them afterwards; no template needed.
- Role or perspective: who should be answering, and for whom.
- Output: what artifact you want, in what shape and length.
- Context: the situation, what you tried, what you know already.
- Constraints: tone, audience, format, hard limits.
- Example: one thing to imitate, or one thing to avoid.
The example slot is the one people skip and the one that changes output most. Ten seconds describing a version you liked does more than three adjectives about tone.
State constraints as concrete limits, positive or negative: under 150 words, plain list, no puns, UK spelling. A number or a named format is easier to say out loud, and easier for the model to follow, than an abstract description of style.
One brief, start to finish
Here is a 45-second take for a real task, and what happens to it.
You said: "okay so I need, um, a landing page headline, no wait, three headline options for our new, uh, budgeting app for freelancers, it's called Ledgerly, the audience is freelancers who hate spreadsheets, keep each one under ten words, no puns, we tried invoicing made easy before and it felt like, you know, every other SaaS site, oh and give them to me as a plain list"
Auri typed: "I need three headline options for our new budgeting app for freelancers. It's called Ledgerly. The audience is freelancers who hate spreadsheets. Keep each one under ten words, no puns. We tried 'invoicing made easy' before and it felt like every other SaaS site. Give them to me as a plain list."
The pass removed the filler, applied the spoken self-correction ("a landing page headline, no wait, three headline options" became "three headline options"), and added punctuation. It did not reorder or add anything; that part is yours.
After fifteen seconds of editing: "Act as a direct-response copywriter. Write three headline options for Ledgerly, a budgeting app for freelancers who hate spreadsheets. Each under ten words, no puns, as a plain list. Avoid anything like 'invoicing made easy'; it reads like every other SaaS site."
Check it against the slots: output (three headlines, plain list), context (Ledgerly, freelancers who hate spreadsheets), constraints (under ten words, no puns), example to avoid (the old headline). The role was the one missing slot, so the edit added the copywriter line up front.
The fifteen-second refinement pass
Scan for four things before sending.
- Is the ask in the first sentence? If not, move it.
- Are the constraints stated as instructions rather than musings?
- Did you contradict yourself mid-take? Keep one version.
- Are names, numbers, and technical terms spelled correctly?
The last check matters most. A misheard library name or version number sends the whole answer somewhere useless, and cleanup cannot catch it because the result is a real word. Add the terms you use weekly to your custom vocabulary so they stop being a problem. When a prompt runs long, put a one-line summary of the request at the very top; the rest then reads as supporting material.
Mac path and iPhone path
Both work. The Mac suits real briefs; the phone suits follow-ups and capture.
| Mac with Auri Voice | iPhone with Auri Keyboard | |
|---|---|---|
| Best for | Long prompts, iteration, pasted context | Follow-ups, ideas on the move |
| How | Global shortcut inserts at the cursor | The keyboard mic inside the chat app |
| Cleanup | On by default | Hold the mic and pick Clean |
| Recognition | Local engine by default | Online engines by default; On-Device is English-only |
On the Mac, focus the chat input, press the shortcut, speak, and the text lands at the cursor. The gain shows most on the long briefs people avoid typing; Auri's stats screen estimates time saved against your own typing speed. On iPhone, Auri Keyboard types the raw transcript out of the box; hold the mic and pick Clean for the cleaned version. Keep phone prompts to a single question, since editing a long prompt on a phone costs more than the speed you gained.
Where things run: Auri Voice transcribes on your Mac by default. The cleanup pass sends the transcribed text (never audio) to the provider you pick — Auri Cloud unless you change it; choose Auri Local in AI settings to keep both steps on the Mac (Apple Silicon).
Working with long pasted context
When the task involves a document, code, or a thread, paste the material first and speak only the instruction. Describing a document you could paste wastes the take and drops detail.
- Paste the source text into the chat field.
- Press your dictation shortcut with the cursor below it.
- Speak the instruction and refer to the material as the text above.
- Name the output shape: a summary, a rewrite, a list of risks, a diff.
Say what to ignore as well. Instructions like skip the changelog or only the second function are quick to speak and prevent an answer that covers everything at low resolution.
For iterations, dictate the correction rather than rewriting the prompt. Two sentences of keep the structure, cut the intro, make the third point concrete is faster than editing the original by hand.
Failure modes to watch
Four things go wrong repeatedly with dictated prompts.
- Buried ask: the question arrives last and the model answers the background instead.
- Answering your thinking: you speculated aloud, and the model treats the speculation as the request.
- Mangled terms: a library, product, or person's name came out as a similar real word.
- Wrong content in a cloud model: local recognition says nothing about where the prompt goes.
No setting fixes the last one. Transcription running on your Mac does not change the fact that sending the prompt puts the text into someone else's service, so decide what belongs there before you start talking. When an answer misses badly, reread your prompt before you blame the model; most of the time the ask was ambiguous in a way the fifteen-second pass would have caught.
FAQ
Why is a spoken prompt better than a typed one?
Because you include more. Typing pushes people toward one-line requests, while forty seconds of speech naturally covers audience, constraints, and what you already tried. Extra relevant context is what usually separates a useful answer from a generic one.
Why not just use ChatGPT's built-in mic?
It works fine inside ChatGPT. A dictation app gives you the same shortcut and cleanup pass in every AI tool and every other app you write in, plus custom vocabulary for your project's terms and, on Mac, transcription that runs locally by default.
Does the filler in my speech confuse the model?
Some, and the cleanup pass removes most of it: filler words, spoken self-corrections, and missing punctuation. The model reads a paragraph instead of a transcript.
Should I send the first take?
Read it once first. The common defect in a spoken prompt is a question buried at the end after a minute of thinking out loud. Move the ask to the top and delete the contradictions you talked yourself out of.
How do I handle a long document I want analyzed?
Paste the text first, then dictate the instruction and refer to it as the text above. Speaking a description of a document you could paste wastes the take and loses detail.
Can I dictate work-sensitive prompts?
Transcription can run locally, and the prompt still goes to whatever model you send it to. Decide what belongs in a cloud model separately from where transcription happens, and keep confidential material out of it.