OpenAI has launched GPT-5.4, an advance in its series of generative AI models, featuring significant upgrades that enhance its functionality in professional tasks and its integration with computer systems. GPT-5.4 introduces native computer use capabilities, a milestone that enables the AI model to interact with and manipulate a computer across various applications autonomously. This includes performing complex online tasks and operating software like spreadsheets, documents, and presentations. The new model marks a notable step towards creating AI-powered agents capable of autonomously managing tasks in digital environments, a future that AI companies are diligently working towards.
GPT-5.4 is designed to be integrated across OpenAI’s platforms, including its API and Codex, its AI-powered coding tool. Additionally, a specialized version known as GPT-5.4 Thinking is being rolled out for use in ChatGPT. This variant is tailored to handle tasks that require enhanced reasoning capabilities. GPT-5.4 enhances its predecessors’ abilities by allowing for more sophisticated interactions with computers through writing code, and sending precise keyboard and mouse commands based on screenshots. These capabilities not only extend the scope of tasks GPT-5.4 can perform but also improve the interactivity and efficiency with which these tasks are executed.
The improvements in GPT-5.4 manifold in its browsing capabilities and its ability to engage external tools and APIs to complete complex tasks. Notably, it excels in aggregating information from diverse sources to respond to intricate queries, particularly those where the relevant data is not readily apparent. According to OpenAI, GPT-5.4 has significantly reduced the likelihood of producing factual inaccuracies, claiming a 33% reduction in false information compared to its previous model, GPT-5.2. This enhancement underscores its utility in scenarios that demand high accuracy and reliability in data handling and synthesis.
Inside ChatGPT, the GPT-5.4 Thinking model provides users an outline of its reasoning processes for complex inquiries, adding a layer of transparency and allowing users to modify their queries in real-time to better align with their specific requirements. OpenAI emphasizes that this capability simplifies user interactions with the model by reducing the need for multiple iterations of input to refine the AI’s responses. The feature is currently accessible through the ChatGPT web app and on Android devices, with plans to expand availability to iOS.
The deployment of GPT-5.4 is comprehensive, extending across the different user levels in ChatGPT, including Plus, Team, and Pro categories, each presumably tailored to varying degrees of task complexity and performance requirements. Additionally, a Pro version of the model, designed to deliver peak performance on particularly demanding tasks, is being introduced in the API and for the enterprise and educational users of ChatGPT.
This rollout of GPT-5.4 not only enhances the versatility and efficiency of OpenAI’s offerings but also solidifies its position at the forefront of AI development, particularly in the realm of autonomous agents. As AI technologies continue to evolve, the capabilities demonstrated by GPT-5.4 will likely become foundational elements of more sophisticated digital systems and services. This progression supports the broader goal of realizing fully autonomous digital agents that can effectively and efficiently handle a wide range of tasks across different platforms and industries.
Read the full post on theverge.com


