ChatGPT Work Review

3 min read

ChatGPT Work Review

Overview

ChatGPT Work is OpenAI’s bet that the future of its flagship product isn’t answering questions, it’s finishing tasks. Launched July 9 alongside the GPT-5.6 model family, Work turns ChatGPT into an agent that gathers context from your connected apps and files, works through a project independently, sometimes for hours, and hands back a finished spreadsheet, document, deck, or small web app instead of a chat reply you still need to act on. It also absorbed OpenAI’s standalone Codex coding app in the process, so the same workspace now handles both office-style deliverables and multi-file code changes.


The idea is sound and, frankly, overdue. The execution is real but uneven, and OpenAI’s own numbers are honest enough to admit it isn’t the best model on the market for the hardest version of this job.

Performance

Where ChatGPT Work is genuinely strong is orchestration: pulling context from Slack, Google Drive, SharePoint, email, and a directory of more than 1,400 connected apps, then stitching that into a coherent finished artifact without you manually assembling the pieces. The Sites feature, which builds shareable, auto-updating dashboards and trackers, is the most forward-looking part of the product, and Scheduled Tasks turn genuinely tedious recurring work into something that runs unattended.


Where it’s less convincing is raw capability at the hardest end of the spectrum. OpenAI’s own launch benchmarks show GPT-5.6 Sol, the model powering Work on paid plans, trailing Anthropic’s Claude Fable 5 on GDPval AA, the professional-work benchmark most directly relevant to what Work is supposed to do. That’s OpenAI’s own data, not a rival’s spin, and it’s worth taking seriously: this product is a strong generalist, not a category leader on the specific measure that matters most for professional deliverables.


The company’s security claims deserve the same scrutiny. OpenAI says internal red-teaming blocked 100% of attempts to extract protected data through the Compliance API layer an impressive number, entirely self reported, from the same company grading its own homework. Treat it as a claim, not a verified result, until independent testing catches up.

Pros

The core promise, delegate a real project and get back a finished file, genuinely works for the use cases OpenAI demonstrated, and the connector breadth means most common business tools are supported without custom integration work. Folding Codex into the same workspace is a smart consolidation rather than a gimmick; developers get agentic coding and office style automation from a single tool. The governance layer, inherited from ChatGPT Enterprise, gives IT teams real oversight rather than shipping a shadow-IT risk to already nervous security departments.

Cons

The flagship Sol model isn’t the top performer on the benchmark that matters most for this exact product, and OpenAI’s own materials say so. Usage is metered and shared with Codex, so heavy agentic use burns through plan limits faster than casual chat ever did, meaning the “included in your existing plan” pitch understates the real cost for power users. Free and Go tier users get the noticeably weaker Terra model, not Sol, which makes early impressions from that tier a poor proxy for what paid users actually experience. And every governance and security claim in this launch is self reported by OpenAI, with no independent audit yet available to check it against.

Pricing

ChatGPT Work isn’t a separate add-on charge it’s bundled into existing plans, but metered against a usage pool shared with Codex. Pro, Enterprise, and Edu users got access first on July 9; Plus and Business followed within days. Free and Go users get Work as well, but running on Terra rather than Sol. On desktop, it’s available across all plans immediately, since it arrived through the Codex-ChatGPT app merger rather than a staged rollout. There’s no meaningfully cheaper way to “try” Sol powered Work specifically without a paid plan, which limits how easily a skeptical user can kick the tires before committing.

Verdict

ChatGPT Work is a legitimately useful product wrapped in slightly oversold framing. If your work involves recurring, multi tool projects, the kind you currently assemble by hand from three different apps every week, this will save real time and is worth testing on one workflow you already know well. If you need the single best model for the hardest professional grade tasks, OpenAI’s own benchmarks say to look at Claude Fable 5 first, and treat Work as a strong, well integrated generalist rather than the category leader its launch messaging implies. The usage metering fine print and the self reported security claims are the two things I’d want resolved before recommending this for anything genuinely sensitive.


Leave a Reply

Your email address will not be published. Required fields are marked *

Never Miss the Stories Shaping AI

Receive reporting on artificial intelligence, Big Tech, startups, and emerging technology from around the world.