Why does ChatGPT's new voice Work tab need a different kind of prompt?
ChatGPT can now take a spoken instruction on your phone and turn it into a finished document, email draft or Slack summary. OpenAI announced the voice-driven Work tab for mobile on 23 September 2026. The catch is that spoken prompts are usually vague, rambling and missing the details a typed prompt would contain, so the output drifts unless you brief it deliberately.
If you have ever dictated a task to ChatGPT while walking to the MTR and got back something generic, you are not doing anything wrong. You are just speaking the way people speak, and an agent that produces real work needs the way people brief.
A voice brief is a spoken instruction structured so an AI agent can complete a task end to end without follow-up questions. It states the outcome you want, the inputs it should use, the shape of the deliverable and the point at which it should stop. It takes about twenty seconds to say.
This article walks through what OpenAI actually shipped, the four-beat structure that makes spoken briefs reliable, a script you can read aloud today, and the places where voice still breaks down.
What did OpenAI actually ship on 23 September 2026?
According to TechCrunch's Ivan Mehta, OpenAI brought voice-based agentic features to the ChatGPT mobile app on 23 September 2026. Plus and Pro subscribers can use the Work tab on their phone by speaking, which lets them create a document, draft an email or summarise Slack messages. Free and Go users get voice access to plugins and connected apps instead.
The same report notes that paid users can go further from the Work tab: build sites, create presentations, use the cloud browser and reach the finances area of ChatGPT. Voice conversations also gained richer text output, and you can switch between typing and speaking mid-task.
The feature that matters most for daily work is handoff. You can start a task by voice on the phone and resume it on desktop. That turns commute time into a place where real work gets kicked off, then reviewed at your desk.
Context helps here. OpenAI launched the GPT-Live conversational voice model in July 2026 and wired it into the desktop app later that month, so voice could drive the Work tab and the Codex tab on a computer. The September update simply brings the same capability to mobile. Anthropic went a different direction two weeks earlier, when it merged Claude's chat and Cowork interfaces, whereas OpenAI still keeps chat and workspaces separate. Source: TechCrunch, 23 September 2026.
Why do spoken prompts produce worse output than typed ones?
Spoken prompts underperform typed prompts for three reasons: people talk in half-finished thoughts, they leave out the format they want because it feels awkward to say, and they never state where the task ends. An agent that can build a whole presentation will happily do so when you only wanted a three-line email.
Consider the difference. Typed, you might write "Draft a 120-word email to the client confirming Thursday's meeting, friendly but professional, bullet the two agenda items." Spoken, the same person usually says "Um, can you write something to the client about Thursday, just confirming and mentioning the agenda." The second version has no length, no tone, no structure and no stopping point.
There is a second, quieter problem. Anthropic's 2026 prompting guidance, and the practitioner write-ups that followed it, note that newer models tend to expand scope when the request is open-ended, adding charts, extra sections or research nobody asked for. Voice makes requests more open-ended by default, so scope creep gets worse, not better.
The fix is not to speak in a robotic way. It is to speak in a fixed order, so that the details a typed prompt would carry are never skipped.
What is the four-beat voice brief?
The four-beat voice brief is a spoken structure with four parts in a fixed order: Outcome (what the finished thing is and who it is for), Inputs (which messages, files or facts to use), Shape (length, format, tone) and Stop (what counts as done and what not to do). Said in that order, it takes 15 to 25 seconds and removes most follow-up questions.
Beat 1, Outcome. Name the deliverable and the reader in one sentence. "A follow-up email to Mandy at the client, confirming Thursday's kickoff." Naming the reader is what lets the model pick the right tone without you describing it.
Beat 2, Inputs. Say where the material comes from. "Use the last five messages in the project Slack channel and the agenda I pasted yesterday." On mobile this is the beat that decides whether the Work tab pulls from a connected app or invents content.
Beat 3, Shape. State the form out loud even if it feels odd. "Under 120 words, two short bullets for the agenda, friendly but professional." Length and format are the two things people most often forget to say.
Beat 4, Stop. Define done and fence off the extras. "Just the draft, no subject line options, no calendar invite, and flag anything you are unsure about instead of guessing." This is the beat that stops a document task from turning into a deck.
The order matters because it mirrors how you would brief a capable colleague in a lift: what, from what, looking like what, and where to stop. Practitioners who already use written output contracts will recognise the same logic; the earlier UD guide on output contracts for consistent AI results covers the typed version.
How do you apply a voice brief to a real workday task?
A realistic scenario: you are a marketing manager leaving a client meeting at 6pm, and three things need to happen before tomorrow. Each one becomes a single four-beat voice brief in the ChatGPT Work tab, spoken on the way to the station, and reviewed on desktop when you get home.
Task one, the recap email. "Outcome: a recap email to the client team after today's meeting. Inputs: the three decisions I am about to list, budget approved at 80,000, launch moved to 15 October, and Sarah owns creative. Shape: under 150 words, one line per decision, warm and direct. Stop: draft only, do not add next steps I did not mention, and leave the subject line blank."
Task two, the Slack summary. "Outcome: a summary of what the design team discussed today, for me only. Inputs: today's messages in the design-review channel. Shape: five bullets maximum, each starting with the person's name. Stop: only things that were decided or blocked, skip general chat, and tell me if the channel had fewer than ten messages."
Task three, the one-page brief. "Outcome: a one-page internal brief for the October launch, for the sales team. Inputs: the recap email you just drafted and the launch plan document in my project. Shape: four headings, target audience, key message, timeline, what sales should say. Stop: one page, no pricing details, and mark any date you could not confirm from the documents."
Notice what each brief has in common. The reader is named. The source of truth is named. The length is a number. And the model is told what to leave out and what to flag. That last instruction is the one that keeps voice-driven agents honest, because it gives the model permission to say "I could not find this" rather than filling the gap.
Where does voice briefing still break down?
Voice briefing fails in five predictable places: tasks that need precise strings such as URLs or product codes, tasks where the input is not connected, noisy environments, anything sensitive said in public, and over-trusting a summary you have not read. Knowing these in advance saves more time than any prompt trick.
Precise strings. Speech-to-text will mangle SKU codes, email addresses and URLs. Speak the brief, then type those details when the draft opens, or reference a pasted note instead of dictating them.
Missing inputs. "Summarise the Slack channel" only works if Slack is a connected app in your ChatGPT account. If it is not, the model may produce a plausible summary of nothing. The Stop beat's "tell me if you cannot find it" line is your safety net.
Plan and rollout limits. Per TechCrunch, the agentic voice features in the Work tab are for Plus and Pro subscribers; Free and Go users get voice with plugins and connected apps. Feature availability can also vary by region and roll out gradually, so check your own app rather than assuming.
Public spaces. Dictating a client's budget on a crowded MTR platform is a confidentiality problem before it is a prompting problem. Save sensitive briefs for the desktop handoff.
Unread summaries. An agent that summarises 80 messages into five bullets will drop things. Treat the summary as a reading guide, not a replacement for the channel, especially when the output goes to someone else.
What voice brief can you try in the next 20 minutes?
Open ChatGPT on your phone, go to the Work tab if your plan includes it, tap the voice control, and read the script below aloud, replacing the bracketed parts. Then open the same conversation on desktop and compare the draft to what you would have written yourself. The whole exercise takes under 20 minutes.
Try this voice brief (read aloud):
Outcome: draft a follow-up email to [reader's name and role] about [topic], sent from me.
Inputs: use the [three or four] points I am about to give you: [point one], [point two], [point three]. Do not use anything else.
Shape: under [120] words, [friendly but professional] tone, one short paragraph then [two] bullets.
Stop: give me the draft only. No subject line options, no calendar invite, no extra suggestions. If any point is unclear, ask me one question instead of guessing.
If your plan does not include the Work tab yet, the same four beats work in ordinary voice mode. You will get the text back in the chat instead of as a finished document, but the discipline is identical, and you will notice the difference in the first reply.
A useful upgrade once the basic brief works: add a fifth line, "Check: before you finish, list any fact in the draft that did not come from my inputs." Newer models are far better at checking a claim than at avoiding one in the first place, and voice-driven tasks benefit from that check more than typed ones.
What is the one thing to remember about briefing AI by voice?
Speak in a fixed order, not in a natural one. Outcome, inputs, shape, stop. The new ChatGPT voice Work tab makes it possible to kick off real deliverables from your phone, but the quality of what comes back is decided in the twenty seconds you spend talking, not in the minutes the agent spends working.
Voice is the most human way to give instructions and, for exactly that reason, the easiest way to give incomplete ones. The four-beat brief closes that gap without making you sound like a machine. We understand AI. We understand you better. With UD by your side, AI doesn't feel cold.
Reviewed by the UD AI team.
Ready to hand real work to an AI colleague?
Briefing an agent well is the first step. The next is giving that agent a defined role in your team, with the inputs, guardrails and review loop already built in. The AI Employee Hub shows what AI staff can take on for marketing, admin, customer service and more, and we'll walk you through every step, from role design to daily operation.