Auri VoiceAuri Keyboard

Prompting ChatGPT Faster With Voice Typing

Answer first

Say the messy version out loud, let the cleanup pass remove filler and self-corrections, then spend fifteen seconds sharpening the ask before you send. Speaking gets far more context into a prompt than typing does, and context is what most weak prompts are missing. Cover five things: the role, the output you want, the context, the constraints, and one example. This works with Auri Voice on Mac, where recognition and cleanup run on device, or with Auri Keyboard on iPhone.

Speak the messy brief, then fix the prompt

Good prompts are mostly context, and typing makes people stingy with context. The loop that works is three steps, and it takes about a minute total.

  1. Talk for thirty to sixty seconds as if briefing a coworker who will do the task.
  2. Let the cleanup pass turn that into readable paragraphs.
  3. Read it once, move the ask to the top, cut what you talked yourself out of.

The point of speaking is volume of relevant detail. In a minute you will mention the audience, the deadline, the thing you already tried, and the constraint you would never have typed, and those are what change the answer.

The point of the second pass is order. Speech wanders toward the conclusion, and models respond better when the request is stated before the background.

Five things to say out loud

You do not need a template to read from. You need five slots you can hit in any order while talking, then check for afterwards.

  • Role or perspective: who should be answering, and for whom.
  • Output: what artifact you want, in what shape and length.
  • Context: the situation, what you tried, what you know already.
  • Constraints: tone, audience, format, what to avoid, hard limits.
  • Example: one thing to imitate, or one thing to avoid.

The example slot is the one people skip and the one that changes output most. Ten seconds describing a version you liked does more than three adjectives about tone.

Constraints work best as negatives. Saying what you do not want, such as no bullet lists or no more than 150 words, is easier to speak and easier for the model to follow than an abstract description of style.

Let cleanup handle the mess

Speech is full of things that make transcripts hard to read: restarts, half-sentences, filler, missing punctuation. Fixing that by hand is where people give up on voice for prompting.

Auri Voice on Mac improves text on device by default, and the pass covers exactly these problems.

  • Errors and self-corrections, so the earlier wrong version disappears.
  • Punctuation and sentence boundaries.
  • Basic formatting, including lists when you spoke one.
  • Filler words like um and you know.

What arrives is prose you can edit rather than a transcript you have to reconstruct. That difference is what makes the messy-brief approach practical on a busy afternoon.

Cleanup is not restructuring. It will not move your buried question to the top or resolve two contradictory instructions, which is why the fifteen-second read is still part of the loop.

The fifteen-second refinement pass

Scan for four things before sending. This is the difference between a prompt that reads as a clear brief and one that reads as a transcript of someone thinking.

  1. Is the ask in the first sentence? If not, move it.
  2. Are the constraints stated as instructions rather than musings?
  3. Did you contradict yourself mid-take? Keep one version.
  4. Are names, numbers, and technical terms spelled correctly?

That last check matters more than it looks. A misheard library name or version number sends the whole answer somewhere useless, and it is the error cleanup cannot catch because the result is a real word.

When a prompt is long, add a one-line summary of the request at the very top. You already spoke the detail, so the summary costs a few seconds and makes the rest read as supporting material.

Mac path and iPhone path

Both work. The difference is length: the Mac suits real briefs, the phone suits follow-ups and capture.

Mac with Auri VoiceiPhone with Auri Keyboard
Best forLong prompts, iteration, pasted contextFollow-ups, ideas on the move
HowGlobal shortcut inserts at the cursorVoice typing inside the chat app
CleanupOn device, on by defaultGrammar and rewrite styles
RecognitionLocal by defaultCloud by default, local available

On the Mac, focus the chat input, press the shortcut, speak, and the text lands where the cursor is. Speaking runs around five times faster than typing, which is most noticeable on the long briefs people otherwise avoid writing.

On iPhone, dictate the follow-up and use one rewrite pass if the message rambled. Keep phone prompts to a single question, since editing a long prompt on a phone costs more than the speed you gained.

Working with long pasted context

When the task involves a document, code, or a thread, put the material in first and speak only the instruction. Describing a document you could paste wastes the take and drops detail.

  1. Paste the source text into the chat field.
  2. Press your dictation shortcut with the cursor below it.
  3. Speak the instruction and refer to the material as the text above.
  4. Name the output shape: a summary, a rewrite, a list of risks, a diff.

Say what to ignore as well. Instructions like skip the changelog or only the second function are quick to speak and prevent an answer that covers everything at low resolution.

For iterations, dictate the correction rather than rewriting the prompt. Two sentences of keep the structure, cut the intro, make the third point concrete is faster than editing the original by hand.

Failure modes to watch

Four things go wrong repeatedly with dictated prompts. All four are cheap to avoid once you have seen them.

  • Buried ask: the question arrives last and the model answers the background instead.
  • Answering your thinking: you speculated aloud, and the model treats the speculation as the request.
  • Mangled terms: a library, product, or person’s name came out as a similar real word.
  • Wrong content in a cloud model: local recognition says nothing about where the prompt goes.

That last one is a habit, not a setting. Recognition running on your Mac does not change the fact that sending the prompt puts the text into someone else’s service, so decide what belongs there before you start talking.

When an answer misses badly, reread your prompt before you blame the model. Most of the time the ask was ambiguous in a way you would have caught in the fifteen-second pass you skipped.

FAQ

Why is a spoken prompt better than a typed one?

Because you include more. Typing pushes people toward one-line requests, while forty seconds of speech naturally covers audience, constraints, and what you already tried. Extra relevant context is what usually separates a useful answer from a generic one.

Does the filler in my speech confuse the model?

Some, and cleanup handles most of it. Auri Voice improves text on device by default, repairing self-corrections, adding punctuation, applying formatting, and dropping filler, so the model reads a paragraph instead of a transcript.

Should I send the first take?

Read it once first. The common defect in a spoken prompt is a question buried at the end after a minute of thinking out loud. Move the ask to the top and delete the contradictions you talked yourself out of.

How do I handle a long document I want analyzed?

Paste the text first, then dictate the instruction and refer to it as the text above. Speaking a description of a document you could paste wastes the take and loses detail.

Can I dictate work-sensitive prompts?

Recognition can be local, and the prompt still goes to whatever model you send it to. Decide what belongs in a cloud model separately from where transcription happens, and keep confidential material out of it.