I Sent ChatGPT Agent Out To Shop For Me And It Couldn’t Finish The Job – theverge.com

I sent ChatGPT Agent out to shop for me and it couldn’t finish the job – theverge.com

OpenAI’s recent launch of ChatGPT Agent, an AI tool designed to perform multifaceted tasks using a “virtual computer,” showcases an innovative yet inconsistent advancement in artificial intelligence. The Verge took the new tool for a test run after purchasing a $200 monthly subscription to ChatGPT Pro, as OpenAI reported that its broader release to Plus and Team users would be delayed due to high demand.

ChatGPT Agent is an extension of OpenAI’s previous projects, Operator and Deep Research, and delves into task completion across various platforms without user interaction on those platforms. It operates by entering an “Agent Mode” prompted by typing “/agent,” and offers assistance in executing specified tasks such as product searches and event planning. However, the experience of using ChatGPT Agent highlighted both its capabilities and limitations.

The process begins with users selecting tasks from suggested examples, such as finding specific items on Etsy or scheduling events through Google Calendar. The agent provides detailed task execution steps, which are visible to the user but focus on the task currently being addressed rather than the overall reasoning chain. When testing the tool to find a vintage-style lamp on Etsy, it displayed each step from setting up desktop operations to filtering search results for specifics like price and item condition.

Despite claiming to add selected items to a shopping cart for user review, the items were not actually added to the user’s real cart but rather to a virtual one inaccessible to the user. This discrepancy in function underscores a vital limitation: ChatGPT Agent operates on a virtual system independent of the user’s actual accounts and browsers. This separation restricts the tool’s ability to perform any action requiring real-time user account interaction, like finalizing purchases or editing personal details.

Moreover, the speed and effectiveness of ChatGPT Agent were found wanting in several aspects. The AI took extended periods—sometimes up to an hour—to complete tasks that a human might manage more swiftly. In a briefing, OpenAI representatives explained that the agent is optimized for handling complex tasks rather than focusing on speed, suggesting it is suitable for running background tasks while users attend to other priorities.

While the tool theoretically handles everyday consumer tasks, it is explicitly restricted in areas involving financial transactions and sensitive personal data. For instance, when attempting to set up an automatic bank transfer, ChatGPT Agent failed to proceed and returned error messages along with explanations clarifying its inability to handle such tasks. This limitation extends to other high-stakes financial operations, where ChatGPT Agent is programmed to avoid executing any tasks that may compromise user security and privacy.

Despite these restrictions, the tool is purportedly capable of managing regular consumer purchases. This claim was tested through a request for ChatGPT Agent to purchase a flower bouquet. Although it could provide options and facilitate decision-making by comparing various offerings and advising on reliability, it could not place the order directly. Users still need to perform critical actions such as finalizing the purchase details and completing transaction forms, highlighting the agent’s role as more of an advisory and research assistant than a transactional agent.

The concept of an AI performing such autonomous tasks is fascinating, and as it stands, ChatGPT Agent does offer substantial help in gathering information, analyzing options, and guiding users through complex decision-making processes. Nevertheless, its practical application is constrained by the current limits of its operational framework and technology, unable to fulfill tasks requiring real interaction with external systems or handling of sensitive information.

In summary, OpenAI’s ChatGPT Agent marks a step forward in AI’s usability for everyday tasks and introduces a potentially powerful tool for users. However, as currently implemented, the agent is more of a personal research and planning assistant rather than a full substitute for human action. It needs further refinement to bridge the gap between automated decision-making support and actual transactional capability. As such, while it showcases potential, it still has significant room for improvement to meet user expectations fully.

Read the full post on theverge.com

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top