A model guide for the GPT‑6 family
Practical tips for getting the best results from GPT‑6 models while managing time and cost
GPT‑6 is our most advanced suite of models yet, and offers you a choice of models for different kinds of work.
Whether you’re turning an idea into a working prototype, building and testing a feature, or orchestrating multi-step workflows across code repositories, databases, and external APIs, this guide explains how to choose a GPT‑6 model, give it effective instructions, manage long-running work, and prepare for production.

Run effectively in production. Use caching(opens in a new window) and compaction(opens in a new window) to manage context and cost. Measure task success and latency, and plan for monitoring and data controls.
Match the model to your workload. Balance capability, cost, and latency by choosing the model, reasoning effort(opens in a new window), and speed that fit the task.
Adjust your prompts and skills. Keep prompts, skills, and repository instructions consistent about what the model should deliver, what it can do independently, and what counts as done.
Keep long-running work on track. Use steering(opens in a new window), async tools(opens in a new window), and delegation(opens in a new window) to handle updates and independent work. Set clear boundaries for when the model should ask for input.
Before deploying, there are several checks and best practices you’ll want to put into place.
Keep efficiency in mind. Cut context the task doesn’t need(opens in a new window) while keeping the evidence it does. Where your application supports it, run independent tasks together(opens in a new window) so one slow step doesn’t hold up unrelated work.
Reuse shared context through prompt caching(opens in a new window) for recurring work. Cached input tokens cost up to 95% less(opens in a new window) than uncached input tokens, depending on the model. Put stable instructions and reference material before changing task details, and keep tool definitions consistent. The caching dashboard(opens in a new window) and diagnostics guide(opens in a new window) help you see where that reuse breaks down. Include cache writes and any long-context rates when estimating the cost of a complete workflow.
For longer conversations, compaction(opens in a new window) reduces context size while preserving the state needed to continue.
Decide how you’ll monitor behavior(opens in a new window) and review the data controls(opens in a new window) for your application.
Test before deploying: Run representative tasks and measure task success, latency, and cost per successful task. Check out our API deployment checklist(opens in a new window).
Think of the model choice and reasoning level as an intelligence/ price tradeoff.
Model:
GPT‑6 Astra(opens in a new window) for the hardest reasoning work where maximum intelligence is needed.
GPT‑6.1 Sol(opens in a new window) for complex coding, research, and computer use.
GPT‑6 Luna(opens in a new window) for focused tasks at scale and everyday, repeated work with a clear goal, such as extracting invoice fields, classifying requests, or producing structured summaries.
When evaluating the best model for the task, compare pricing(opens in a new window) for each model.
Reasoning level: In the API, choose how much effort the model spends on the task.
Low: Routine tasks, such as extracting facts or making small edits.
Medium: Work requiring judgment, such as planning a feature or comparing options.
High: Difficult debugging, deeper analysis, or careful review.
Extra high / Max: Test where supported when High falls short, and keep only if the improvement justifies the added time and cost.
In the API, you can change reasoning effort mid-conversation(opens in a new window) without breaking cache.
In Codex, start with the default reasoning level for that model, then lower it for simpler tasks or increase it for deeper analysis.
Speed:
In the API, use Fast mode(opens in a new window) when response time matters, such as in chat apps or coding tools. It provides faster, more consistent response times at a higher per-token cost than Standard processing.
In Codex and the API, use Ultrafast(opens in a new window) when faster responses are worth the premium, such as rapid coding iterations. It speeds up token generation independently of reasoning effort. Available for GPT‑6 Astra(opens in a new window).

—Eric Provencher, Developer Experience at OpenAI
Start with a clear assignment: the result you want, who it’s for, the relevant context and constraints, and what counts as done. Then review these four areas, summarized from Rethinking skills and prompts for GPT‑6 Astra(opens in a new window), for a deeper dive into updating your instructions:
Create better skills: Keep descriptions short and explicit about when each skill should run, load supporting details only when needed, and replace rigid recipes with guidance suited to the models your team uses.
Update your AGENTS.md: Explain when particular documents and tests are relevant, and explicitly authorize safe routine workflows, such as running local tests with disposable data and no production access.
Set decision boundaries: State which actions can proceed independently and which require approval, replacing blanket “always ask” rules with clear boundaries.
Be prescriptive about persistence: Define what “done” includes—implementing the change, running it, inspecting the result, and fixing failures—and identify any decisions that require your review.
For additional guidance, see reasoning best practices(opens in a new window).
Whether you’re working in Codex or building with the API, specify which decisions the model can make, when it should ask for input, and what a useful response looks like.
Give the model enough direction to keep work moving without guessing at decisions that matter. Tell it which choices it can make and when to ask for your input(opens in a new window), for example, it can choose how to organize a summary, but should check with you before changing the project’s scope. Describe what a useful response looks like(opens in a new window), for example plain language, technical detail suited to your audience, and a short handoff covering what changed, what was checked, and what still needs attention.
With the GPT‑6 family of models, you can now take on tasks that span hours or days. Use the following features to better manage agents on long-running tasks.
In the API, use steering, asynchronous tools, and parallel work to keep long-running tasks moving.
Update instructions during a run: Mid-turn steering(opens in a new window) lets you send a correction through the Responses WebSocket API(opens in a new window) while the model works. Updates are queued; they don’t cancel running tools or undo completed actions.
Keep working while a tool runs: Asynchronous tool calling(opens in a new window) lets the model continue independent work while your app runs a slower task, such as tests. Your app returns the result when it’s ready. Wait for that result before starting work that depends on it.
Delegate independent subtasks: GPT‑6.1 Sol supports multi-agent workflows in the Responses API(opens in a new window). It can assign independent work to subagents (such as investigating different parts of a codebase) and combine their findings into a final response. Multi-agent is currently in beta.
Long-running tasks can uncover decisions you wouldn’t necessarily anticipate in the initial prompt. Use clarification and steering to keep the work on course.
Answer questions as work progresses: With GPT‑6 Astra, Codex can ask for clarification while it works(opens in a new window). Resolve questions that affect the next step, and specify which independent work can continue while you decide. If you’ll be away, tell Codex which tasks can continue and when it should pause for your answer.
Redirect work when requirements change: Steer the active task(opens in a new window) with new information, explaining what should change and what should stay the same. This helps avoid spending more time on an approach that no longer meets your needs.

Computer use(opens in a new window) lets GPT‑6 Astra, GPT‑6.1 Sol, and GPT‑6 Luna interact directly with websites and desktop apps, even applications without an API. For example, you can ask the model to investigate a bug, fix the code, and open your product in a browser to check that the fix works.
Choose the simplest reliable way to do each step:
Use an API or connected tool when it can do the job directly.
Use computer use when the model needs to read a screen, click buttons, or fill in a form for you.
If you’re building computer use into your own app, give the model a tool that can run code to control a browser or desktop. Playwright works with browsers(opens in a new window); PyAutoGUI works with desktop apps.



