OpenAI extends ChatGPT Voice to desktop, enabling agent control via voice
The desktop application now supports voice-driven automation and screen access on macOS, marking a significant capability leap from the earlier smartphone-focused launch.

OpenAI has rolled out a global update to its ChatGPT desktop application, introducing ChatGPT Voice to the platform. The update empowers users to control AI agents and execute computer tasks through voice commands, utilising the ChatGPT-Live model family launched earlier this month. This expansion marks a shift from the smartphone version of the feature, which prioritised conversational fluidity and interruption handling but lacked the capacity to take direct action on devices.
The desktop iteration is designed for higher complexity, allowing users to dictate multi-step instructions and coordinate work across agents. The feature integrates with existing OpenAI products, ChatGPT Work and Codex, enabling users to direct multiple agents simultaneously. According to the company, the system can speak, listen, and coordinate tasks within the app at the same time, leveraging computer use skills to navigate websites and applications.
In a demonstration of the new capabilities, OpenAI showcased a developer issuing a single voice command to create a new thread, submit a pull request, and identify the root cause of a software bug. The update also includes Appshots for macOS users, a feature that grants the application access to the screen content, including the ability to generate alt-text for visual elements.
While the desktop version focuses on local execution, OpenAI noted that users can also utilise ChatGPT Voice in Codex from the iOS app through remote access. This hybrid approach allows for flexibility, bridging the gap between mobile convenience and desktop power, although the primary focus of this update remains the enhanced desktop experience.
The move places OpenAI in direct competition with rivals expanding their own voice-driven automation tools. Anthropic recently updated its voice mode for Claude, enabling the model to tap into Opus, Sonnet, and Haiku to perform tasks in productivity applications such as Gmail, Calendar, Slack, Notion, and Canva. As both firms refine their voice interfaces, the focus is shifting towards seamless, multi-step automation across the digital workspace.
