Our complete GPT 5.3 Codex review: A new era for agentic AI

Stevia Putri
Written by

Stevia Putri

Katelin Teen
Reviewed by

Katelin Teen

Last edited February 6, 2026

Expert Verified
Image alt text

On February 5, 2026, OpenAI released GPT-5.3-Codex, its newest coding model. The release coincided with Anthropic's Opus 4.6, highlighting the competitive pace of AI development.

OpenAI is positioning this as more than a minor update. They are shifting Codex from a powerful code generator into a general-purpose agent that can operate a computer and handle professional workflows from start to finish. The concept moves from a tool toward an AI teammate.

This article will break down what’s new, review its performance, and analyze what this means for developers and businesses.

What is GPT 5.3 Codex?

At its core, GPT-5.3-Codex is what OpenAI calls its "most capable agentic coding model to date." It follows GPT-5.2-Codex, but with a significantly expanded scope.

According to OpenAI's official announcement, the new model is built on three main principles:

  1. Top-tier agentic skills: The model is designed to handle long, complex tasks across the software development lifecycle and other professional domains.
  2. Improved efficiency: It is reportedly 25% faster and uses fewer tokens than the previous version, which enhances user experience and reduces operational costs.
  3. Self-improvement: Notably, OpenAI states the model helped "create itself." It assisted engineers with tasks like debugging its own training and managing deployments.
The concept is to provide an interactive partner rather than a tool that simply follows commands. This positions it as a teammate that can be guided in real-time, not just an assistant for task delegation.
An infographic detailing the core principles of the GPT 5.3 Codex review: top-tier agentic skills, improved efficiency, and self-improvement.
An infographic detailing the core principles of the GPT 5.3 Codex review: top-tier agentic skills, improved efficiency, and self-improvement.

New capabilities of GPT 5.3 Codex

Let's get into the details of how this new model performs. We’ve dug into OpenAI's claims and the early analysis to see what’s really going on.

Benchmark performance: A leap in agentic skills

OpenAI backed up its release with new scores on key industry benchmarks. These numbers show a significant jump in what the AI can do on its own.

Here’s a look at the data from their blog post, visualized for clarity:
A bar chart infographic for our GPT 5.3 Codex review, comparing its benchmark scores against GPT-5.2-Codex on SWE-Bench Pro, Terminal-Bench 2.0, and OSWorld-Verified.
A bar chart infographic for our GPT 5.3 Codex review, comparing its benchmark scores against GPT-5.2-Codex on SWE-Bench Pro, Terminal-Bench 2.0, and OSWorld-Verified.
BenchmarkGPT-5.3-CodexGPT-5.2-CodexImprovement
SWE-Bench Pro56.8%56.4%A slight edge in multi-language software engineering.
Terminal-Bench 2.077.3%64.0%A massive leap in command-line proficiency.
OSWorld-Verified64.7%38.2%A huge jump in general computer productivity tasks.

The improvements in Terminal-Bench and OSWorld are significant. This suggests the model has improved capabilities for operating within a digital environment and using tools like a person would.

However, the competitive landscape is strong. Community analysis shows that while Codex's 77.3% on Terminal-Bench 2.0 beats Anthropic's Opus 4.6 (65.4%), the tables turn on OSWorld. There, Opus 4.6 scores 72.7% to Codex's 64.7%. This indicates that neither model currently leads across all agentic skills.

Yes. And this is from someone who has always hated codex and only used 5.2 high and xhigh. But 5.3-codex-xhigh is amazing, I’ve build more in 4 hours than I have in the last week.

From coding assistant to professional collaborator

OpenAI is clearly positioning Codex as more than just a tool for developers. They are showing off its ability to manage entire professional workflows.

For example, they shared demos where Codex created a 10-slide PowerPoint presentation for a financial advisor and built fully functional racing and diving games from scratch. This capability extends far beyond suggesting the next line of code.

Regarding the "built itself" claim, it means the model was powerful enough to accelerate its own development. OpenAI's engineers used it to help data scientists build new data pipelines and even had it dynamically scale GPU clusters during the launch. It is a proof of concept for how agentic AI can accelerate complex technical work.

The practical gap for businesses

This capability is impressive. For many businesses, however, this serves as a foundational technology that requires further development for specific applications.

It still takes a lot of technical know-how and engineering time to turn it into a reliable tool for a specific job, like customer support or sales.

Many companies require AI solutions tailored to specific business functions, such as an AI teammate that can learn their products, understand refund policies, and begin handling support tickets. This highlights the gap between a general-purpose model and a business-ready solution.

User experience and accessibility

Beyond its raw power, how does it feel to use GPT-5.3-Codex? And more importantly, who can get access to it?

A more interactive and steerable AI

One of the notable new features is called "steering." It lets you interact with the model while it's working on a task. You can jump in to ask questions, give feedback, and nudge it in the right direction in real time.

This is a significant shift from the typical "black box" approach where a user provides a prompt and waits for the final output. It adds a layer of transparency and control, letting you see the agent's "thought process" and fix its course before it goes too far down the wrong path. It feels less like giving instructions and more like actual collaboration.

Exactly I wouldn't mind if it needed to work 20 hours instead of 1 hour if it could deliver same quality of code I can write myself.

The biggest limitation: No API access

So, how can you try it out? GPT-5.3-Codex is available through the Codex app, a CLI, IDE extensions, and the web interface for paid ChatGPT users.

However, a significant limitation for businesses is that API access is not yet available. OpenAI says it's "rolling out soon," but for now, that's the main roadblock preventing companies from building this power into their own products or internal workflows. Without an API, it remains a powerful but standalone tool, not a scalable part of your tech stack.

This delay presents a challenge for businesses. While businesses wait for API access to build custom solutions, other platforms offer ready-to-deploy applications. For instance, eesel AI provides an AI teammate designed to integrate with help desks like Zendesk, Gorgias, and Intercom. The eesel AI Agent learns from a company's data and can begin handling customer support issues, without requiring custom development.
A view of the eesel AI Agent, an alternative solution mentioned in this GPT 5.3 Codex review, handling customer support tickets autonomously.
A view of the eesel AI Agent, an alternative solution mentioned in this GPT 5.3 Codex review, handling customer support tickets autonomously.

Pricing and the new cybersecurity model

The last pieces of the puzzle are cost and security.

How much does it cost?

Right now, OpenAI hasn't announced any specific pricing for GPT-5.3-Codex. Access is included with paid ChatGPT plans.

Because there’s no API access yet, there's no API pricing available either. This creates uncertainty for businesses planning their AI initiatives, as the cost at scale is unknown, making budgeting difficult.

Some platforms provide more predictable pricing structures. For example, eesel AI's pricing is based on a pay-per-interaction model. This model is not tied to the number of user seats, which can help businesses forecast costs and calculate ROI as they scale their use of AI for customer support.

A "high capability" model for cybersecurity

OpenAI has labeled GPT-5.3-Codex as a "High capability" model for cybersecurity under its Preparedness Framework. This is because it was trained to find software vulnerabilities, making it a strong tool for security professionals.

To manage the risks, OpenAI has rolled out safety measures like the "Trusted Access for Cyber" program, which gives access to vetted cybersecurity experts, and a $10M grant to speed up cyber defense research.

This level of capability has significant security implications. While it is a powerful tool for defense, it also introduces risks that businesses must manage. A managed platform can help address these concerns by offering built-in security and compliance features. For example, eesel AI states that customer data is isolated and never used for training, providing AI capabilities with established security protocols.

A glimpse into the future

GPT-5.3-Codex is a significant step forward for agentic AI. Its performance, speed, and wider skill set make it a powerful tool for developers and other tech professionals. It offers a glimpse into a future where AI agents are our daily collaborators.

However, for many businesses, its current limitations are significant. The missing API access, unknown costs, and the work required to turn a general model into a specific business tool mean it is more of a preview of future capabilities than a solution for immediate implementation.

To see GPT-5.3-Codex in action and hear more detailed first-hand experiences, the following review provides a comprehensive look at its new features and what they mean for the future of AI-assisted development.

A detailed review of OpenAI's GPT-5.3-Codex, covering its new features, performance benchmarks, and its impact on the software world.

How to deploy an AI agent today

A key challenge is that a powerful foundational model like Codex is the engine, but businesses still need to build the application around it. These models are not designed for direct, out-of-the-box business use.

This is where a platform like eesel AI can provide a complete solution. Instead of setting up a tool, you "hire" an AI teammate. The eesel AI Agent connects to the tools you already use, learns your business in minutes, and starts working with your team to handle customer support tickets on its own.

This allows businesses to start using AI agents without waiting for foundational models to become fully productized. Explore how the eesel AI Agent can be applied to customer service operations.

Frequently Asked Questions

What is the main takeaway from this GPT 5.3 Codex review?
The main takeaway is that GPT-5.3-Codex is a significant step forward for agentic AI, especially for developers. However, its lack of an API and undefined pricing make it more of a future-facing tool than a practical business solution you can implement today.
How does GPT 5.3 Codex compare to Anthropic's Opus 4.6?
The comparison is mixed. Codex beats Opus 4.6 on the Terminal-Bench 2.0 benchmark, showing better command-line skills. But Opus 4.6 scores higher on OSWorld, indicating better performance on general computer tasks. Neither model is the clear winner across the board.
What is the biggest limitation highlighted in the GPT 5.3 Codex review?
The single biggest limitation for businesses is the lack of API access. Without an API, companies can't integrate Codex's capabilities into their own products or internal systems, making it a standalone tool for now.
Who should be most excited about this release?
Developers and technical professionals are the primary audience for this release, given the model's capabilities in coding, debugging, and infrastructure management.
What does the "steering" feature mentioned in this review allow users to do?
"Steering" is an interactive feature that lets you guide the model while it's working. You can ask questions, provide feedback, and correct its course in real-time, making it feel more like a collaborative partner than a black-box tool.

Share this article

Stevia Putri

Article by

Stevia Putri

Stevia Putri is a marketing generalist at eesel AI, where she helps turn powerful AI tools into stories that resonate. She’s driven by curiosity, clarity, and the human side of technology.

Related Posts

All posts →
Slack and Perplexity logos in separate white circles
Guides

Brave Leo vs Perplexity AI (2026): privacy, research, and browsing

Compare Brave Leo and Perplexity AI for private browser help, cited research, data controls, and browser actions, with a practical support workflow.

Stevia PutriStevia PutriOct 26, 2025
Illustration of three people around a laptop beneath a ChatGPT speech bubble
Guides

ChatGPT group chat pricing: what to do after retirement

ChatGPT group chats are winding down. This guide explains why no plan upgrade restores the feature and how to evaluate the cost of a replacement.

Rama Adi NugrahaRama Adi NugrahaJun 9, 2026
Illustration of people collaborating in a ChatGPT group chat
Guides

ChatGPT group chat is retiring: what teams can use instead

ChatGPT group chats are winding down. Learn what happens to existing conversations and how to choose a shared AI workflow for real team work.

Kenneth PanganKenneth PanganNov 18, 2025
Image alt text
Guides

A detailed Claude Cowork review: Features, pricing, and limitations

Anthropic's Claude Cowork brings AI agent capabilities to the desktop, allowing users to automate tasks by managing files and browsing the web. This review explores its features, performance, and limitations.

Stevia PutriStevia PutriFeb 6, 2026
Illustration of image inspection connected to code and a chart
Guides

Gemini Agentic Vision: how it works and how to test it

Learn how Gemini Agentic Vision inspects images, what to verify before using its answers, and how eesel CLI helps test a separate support decision.

Alicia Kirana UtomoAlicia Kirana UtomoJan 30, 2026
A person demonstrating a workflow on their Mac while Codex records it as a reusable skill and an AI agent replays it
Guides

OpenAI Codex record and replay, explained

What OpenAI Codex record and replay actually does: demonstrate a workflow on your Mac once, and Codex turns it into a reusable skill. How it works, its limits, and where it fits.

Alicia Kirana UtomoAlicia Kirana UtomoJun 22, 2026
Four illustrated people in overlapping profile windows on a purple background
Guides

Custom personas for ChatGPT: 8 practical prompts to try

Eight practical custom personas for ChatGPT, with copyable prompts for support, research, writing, analysis, and a safer path for customer-facing work.

Kenneth PanganKenneth PanganAug 28, 2025
Abstract yellow and blue gradient with a central square and black connected-nodes icon
Guides

AgentKit vs Claude: choosing an agent-building approach in 2026

Compare OpenAI’s transitioning Agent Builder, the OpenAI Agents SDK, and Anthropic’s Claude Agent SDK. Learn which layer you are choosing and what production ownership remains.

Kenneth PanganKenneth PanganOct 19, 2025
OpenAI logo surrounded by illustrated dollar bills
Guides

Apps in ChatGPT reviews: what businesses should check in 2026

A practical review of Apps in ChatGPT: plugin discovery, app connections, permissions, workspace controls, and limits for support work.

Stevia PutriStevia PutriOct 8, 2025

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free