Artificial intelligence

OpenAI Bets on AI Agents for Everyday Work: Will Users Trust Them?

OpenAI seeks to move AI agents beyond the programming environment into accounting, investment, medicine, and other office jobs through ChatGPT Work. But expanding usage depends on the product’s ability to manage permissions, explain its steps, and work reliably on tasks that are difficult to evaluate, such as strategy, presentations, and sales.

2026-08-24
6 min read
11 views
فريق تحرير certi.news
OpenAI Bets on AI Agents for Everyday Work: Will Users Trust Them?

OpenAI is pushing AI agents beyond code writing through ChatGPT Work, a version aimed at office workers that the company launched last month and makes available under its lowest subscription tier for $20 per month. The idea is not merely for the model to answer questions, but to execute multistep projects using email, work tools, and cloud platforms on the user’s behalf.

This direction represents OpenAI’s biggest commercial and product bet, but it confronts the company with a question more difficult than improving the model itself: How much control over their digital lives will users be willing to give an automated agent? To complete useful tasks, the agent needs access to sources such as email, Slack, Notion, Figma, and calendars, and may have permission to make changes in them. These permissions increase the tool’s value, but they also increase the risks of exposing private information or carrying out an unintended action.

From Developer Tools to the Rest of the Workforce

ChatGPT Work was built on a modified version of the Codex coding tool. OpenAI says the goal is to give nonengineers the kind of capabilities developers have begun using: a tool that can execute an entire task in a nearly autonomous manner instead of merely producing an answer. Current uses presented in the article include preparing weekly performance reports, turning spreadsheets into planning tools, compiling information about companies into investment memos, and creating dashboards and charts.

The importance of this expansion lies in the fact that AI developers cannot rely on the software sector alone to justify the huge investments in training and computing. Longer tasks consume more tokens, while access to new specialties opens a broader market. At the same time, specialized companies such as Harvey in the legal field and Clay in sales are pursuing the same customers with an approach that is not tied to a single model.

Usage figures inside and outside OpenAI reveal the gap the company is trying to close. A study backed by OpenAI showed that 98% of the company’s employees used Codex in June, compared with just 17% of organizational subscribers and less than 1% of individual subscribers. The company also did not disclose the number of Work users compared with Codex users, while its shared applications are used by about 20 million people, according to the article, compared with more than one billion users whom OpenAI says interact with ChatGPT online.

What Is Changing in Practice?

The agent relies on what engineers call a “harness,” meaning the software layer that determines what information the model sees, which tools it can use, and how the result is presented. In coding tools, the command line was sufficient for some users. A broader audience, however, needs an interface that hides the complexity and clarifies what the agent can do and how to start the task.

The article’s experience illustrates both the benefit and the friction. ChatGPT Work successfully transferred an unusually formatted school calendar from email to Google Calendar, created an automatically updating dashboard for publicly listed companies’ metrics, built a queryable database of space launches, and sent a weekly message about new AI research. But configuring permissions to access cloud storage was confusing, and it was not clear to the user how to grant the tool read-only access; it later emerged that the task required full access.

Some settings are also available only through the web app, forcing the user to move between it and the mobile app. There are other practical limitations, such as the tool’s ability to create events in Google Calendar without creating new calendars. A task may also produce a weak result if the user selects a low effort level, while the reasoning levels do not appear sufficiently understandable to new users, according to an OpenAI engineer’s acknowledgment.

Trust and Evaluation Are the Next Obstacles

The problem is not limited to the interface. Software can often be tested with a straightforward question: Does it work or not? Evaluating a good presentation, business strategy, or sales pitch is less clear and harder to track. For that reason, OpenAI said it uses the GDPval benchmark, which is based on tests for knowledge work in 44 occupations, alongside user feedback. The company still faces the challenge of determining whether it is building the workflow most people need or one suited only to its specialized employees.

Permissions also raise questions about the limits of reliance. A lead engineer at OpenAI tested the application with his email, Slack account, phone, and other applications, but acknowledged that the system could extract information from a private message and then use it in a document that was not supposed to include it. The article’s author, meanwhile, did not give the tool access to their email, bank account, or work drafts, even though they found it useful for less sensitive tasks.

Competition Between the Model and the Harness

ChatGPT Work resembles products such as Claude Cowork and other agent tools targeting ordinary users. The article says that Claude Code preceded Codex in adopting a conversational style that lets users choose a path among several options and receive successive updates, before OpenAI added more interaction points to Codex and developed its desktop and mobile applications. Download statistics indicate that Claude Code was more in demand through April of the current year, before Codex moved ahead by a slight margin.

OpenAI believes the strength of its modern models and their cost-effectiveness are the primary difference, while researchers and users believe that harness design and the ease of requesting reviews and comparisons are no less important. The question remains open: Can a powerful general-purpose model gradually compensate for specialized tools and instructions, or will users continue to need an interface that offers greater explanation and control? The answer will determine whether AI agents move from impressive tools for early adopters to an ordinary layer of office work.

News source
TechCrunch AI
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news