OpenAI has added ChatGPT Voice to its desktop app, so you can now speak commands to control AI agents and run multi-step tasks on your computer.
What Changed
The most direct way to tell a computer what to do is now built into the ChatGPT desktop app. OpenAI announced on July 23 that the app supports ChatGPT Voice, powered by the ChatGPT-Live voice model family it introduced earlier this month.
The feature works with both ChatGPT Work and Codex, and it can tap into computer-use skills to look up websites and open apps on your behalf. On macOS, Appshots lets ChatGPT read what is on your screen, including alt-text, so it can react to what you are actually looking at.
The smartphone version of ChatGPT Voice launched earlier with smoother conversations and better handling when you interrupt the assistant, but it could not take action on the phone. The desktop update is a different animal: you can dictate a complex, many-step instruction and ChatGPT will execute it, then pause to ask you for input or confirmation when it needs it.
In OpenAI’s demo, a developer gave a single spoken command that created a new code thread, opened a pull request, and traced a bug to its root cause. You can also control Codex from the iOS app through remote access, so a voice command on your phone can drive work happening on your desktop.
Who Can Use It
ChatGPT Voice on desktop works with ChatGPT Work and Codex, so it is aimed at people who already use ChatGPT as an active workspace rather than a chat window. On macOS, Appshots extends it further: the app can see what is on your screen, including alt-text, which lets it react to the actual context of your work instead of asking you to describe it.
The iOS remote access route matters for one practical reason: you can dictate a task into your phone and have the desktop agent carry it out. That turns voice control into an async workflow, like leaving an instruction with a colleague rather than operating the machine directly.
Why It Matters
Voice has been the weakest input channel for agents because most implementations still just transcribe text. ChatGPT Voice on desktop closes part of that gap: the model listens, speaks, and coordinates work in the same session, which makes multi-step computer control feel closer to working with a person than with a terminal.
If you are already experimenting with using OpenAI Codex for real tasks, the desktop voice layer is worth testing, because it removes the friction of typing out long agent prompts.
The Wider Picture
OpenAI is not alone here. Anthropic also updated Claude’s voice mode, which can call on Opus, Sonnet, and Haiku models and work inside apps like Gmail, Calendar, Slack, Notion, and Canva. The two companies are converging on the same idea: the agent runs the computer, and you run the agent with your voice.