Agentic AI Tools

OpenAI's Dots Agents Just Landed: What to Build and What to Buy

Rushil ShahRushil Shah
14 min read
Share

OpenAI shipped Dots at DevDay 2026 alongside computer use in the Agents API. Here is the layer-by-layer audit: trigger, retrieval, tools, handoff, state, verified action and cost ceiling — and which ones you still own.

TL;DR

Dots is a ChatGPT-surface product, not an API primitive. It genuinely absorbs the tool-calling layer, the sandbox layer and a large part of session state. It does not absorb your triggers from systems of record, your access-controlled retrieval, your cross-vendor handoffs, or — the one that matters most — verification of business invariants before a write. OpenAI's Auto-review checks intent and safety, not whether your journal entry balances. Budget for Dots as a knowledge-worker assistant and the Agents API as a harness, and keep coordination logic in something you can export.

What OpenAI actually shipped on 29 September

Three separate things got collapsed into one headline this week, and they sit at different layers of the stack.

The consumer-and-knowledge-worker product is dots: always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser, able to reach over 4,000 apps through the plugin ecosystem. They roll out in ChatGPT to Pro and Business Premium users in eligible markets, with Enterprise, Edu and Healthcare workspaces getting a beta that is off until an admin enables it. Your first dot is included in the plan at no extra cost, with an allowance for deeper work on top.

The developer product is the Agents API, which actually shipped in public beta on 10 September and picked up computer use at DevDay. It exposes the Codex harness — context compaction, tool search, programmatic tool calling, subagents — as a managed service, with the sandbox either OpenAI-hosted, self-hosted, or run through partners including Cloudflare, Modal, E2B, Daytona, Vercel and Oracle.

The third thing is organisational: specialist dots, provisioned with their own identity and credentials, are currently focused enterprise pilots where OpenAI engineering teams work directly with organisations to define responsibilities and approval paths, with a Microsoft Agent 365 integration in progress. That is a professional-services motion, not a self-serve platform, and it should be read as one when you are planning headcount.

Commercially, the same keynote introduced Pro 500 at 25× the Plus allowance, priced at $500 per month, while Altman and Friar fielded IPO questions on the sidelines — Friar told CNBC OpenAI wants to list "when the time is right for our business". Procurement teams reopened agent budgets on the strength of a keynote. The engineering question underneath is narrower and more answerable.

4,000+ apps reachable by a dot through the plugin ecosystem Source: OpenAI, 2026
$10 / $50 gpt-6-astra standard price per 1M input / output tokens, short context Source: OpenAI API pricing, 2026
$60 / $300 top speed tier listed for gpt-6-astra, per 1M tokens Source: OpenAI API pricing, 2026
$10 / 1k web search tool calls, plus search content tokens at model rates Source: OpenAI API pricing, 2026

The seven layers a production agent actually needs

Every agent that survives contact with a real business has the same seven layers, whether or not anyone drew them. Vendor comparisons go wrong because they benchmark the middle three and ignore the outer four, which is where the operational cost lives.

The seven layers, in execution order

1
Trigger

Something in a system of record changes and the agent must know within a bounded time. Webhook, CDC event, queue message, cron.

↓
2
Retrieval

Fetching the right context under the right permissions. The hard part is not embeddings, it is honouring row-level access for the identity the agent is acting as.

↓
3
Tool-calling

Selecting and invoking tools without blowing the context window. Tool search and parallel execution matter once you pass roughly thirty tools.

↓
4
Handoff

Passing work between agents, between an agent and a human queue, and between vendors. Each hop needs an idempotency key and a resume point.

↓
5
State

What the run knows, what it has already done, and what it may safely redo. Distinct from conversational memory.

↓
6
Verified action

The check immediately before a write: does this satisfy the business invariant, not merely the safety policy.

↓
7
Cost ceiling

A hard stop that halts the run when spend or step count exceeds a threshold, independent of the model's own judgement.

Abstract layered cross-section showing a vertical line passing cleanly through central layers but fracturing at the outer layers
Most vendor benchmarks measure the middle bands, where a request becomes a tool call; the fractures at the outer bands are where on-call engineers spend their nights.

Layer by layer: what Dots absorbs and what it leaves with you

Dots is strongest exactly where self-built harnesses are weakest, and silent exactly where compliance reviews concentrate.

LayerDots / Agents API coverageWhat remains on your side
TriggerPartial. Dots run proactive background research using read-only tools, and ChatGPT team tasks fire on a schedule or on events like new mail. DevDay also added support for the proposed MCP Events specification so plugins can start automations.Triggers originating in your ERP, core banking, EHR or warehouse. You still own the webhook receiver, the dedupe table and the replay path when the agent was down.
RetrievalStrong for SaaS. Connected apps inherit the permissions you granted ChatGPT.Permission-aware retrieval over your own data. The agent sees what the connected account sees — which is usually far more than the task requires.
Tool-callingStrong. Tool search loads definitions on demand and programmatic tool calling runs calls in parallel and filters results in code, so only relevant output re-enters context.Tool contracts, error semantics, and what a tool does on partial failure. The harness retries; your API has to tolerate being retried.
HandoffStrong inside the runtime. Subagents each keep their own context while a main agent coordinates.Handoffs that cross vendors or leave AI entirely — to a human queue, an RPA job, a Temporal workflow, a partner's API with its own SLA.
StateStrong within a session. Automatic compaction carries information across context windows; dots persist learned preferences over time.Durable business state you can query, audit and replay. Compaction is lossy by construction and the dropped tokens are not a record you can produce in an audit.
Verified actionSafety-complete, business-incomplete. Auto-review checks planned steps against your instructions, Custom Rules and safety requirements before an email or file change runs, and returns the reason when it blocks.Domain invariants. No external reviewer knows your netting rules, your tax treatment or your refund authority matrix.
Cost ceilingPlan allowances on the ChatGPT side; pure token-and-tool metering on the API side with no additional Agents API fee.Per-run and per-tenant budget enforcement, and the kill switch. This is yours in both products.

The verified-action row is the one worth reading twice. OpenAI has been unusually clear about the boundary: certain steps always come back to you, including changing a password or transferring money between financial accounts, and sharing health data always requires a named recipient. Those are good defaults. They are also generic defaults. Nothing in them knows that your organisation requires dual approval above a currency threshold, or that a credit note against a closed period is forbidden.

!

Treating Auto-review as a business control

Auto-review is genuinely well-designed: OpenAI keeps the controls that enforce it outside the environments dots can change, so an agent cannot disable its own checks. That architecture is exactly right, and it is why teams mistake it for a complete control plane. It is not. It evaluates intent, recipient and policy. It cannot evaluate whether the amount is correct, whether the counterparty is sanctioned, or whether this write would break a reconciliation that runs at 02:00.

Fix: put a deterministic validator in front of every consequential write — schema, invariant, authority limit, duplicate check — and run it in your own code, not in the prompt. The agent proposes a payload; your validator decides whether it executes. Keep that validator in a repository you control so it survives any model or vendor change.

The cost shape nobody models: what happens when an agent loops

Usage-based pricing on a single-turn model call is predictable. Usage-based pricing on an autonomous loop is not, because the loop decides how many turns it takes.

Look at the published numbers. gpt-6-astra is $10 per million input tokens and $50 per million output at standard rates on short context, rising to $20 and $75 on long context. The same page lists a top speed tier for astra at $60 and $300. GPT-6.1 Sol lands at $2 and $10 — which is why OpenAI positions it as near-Astra intelligence at a fifth of the standard token prices, and why the first thing to test on any agentic workload is whether Sol holds your eval.

Now add the parts that scale with autonomy rather than with volume: web search at $10 per 1,000 calls with search content tokens billed at model rates, file search at $2.50 per 1,000 calls plus $0.10 per GB-day of storage, and hosted containers from $0.03 to $1.92 per 20-minute session depending on memory. A long-running agent that browses, re-reads its own artefacts and re-plans will touch all four meters on a single task.

The failure mode in production is not an expensive run. It is a cheap run that repeats. An agent hits a flaky endpoint, the harness retries, compaction re-summarises a growing transcript, a subagent fans out three ways, and the cost of one logical task multiplies by the branching factor. None of that shows up in a demo because demos succeed on the first attempt. Instrument cost per completed task, not cost per call, and alert on the ratio between them.

On the ChatGPT side the shape is different but not simpler. Conversations with your dot do not count toward ChatGPT usage limits, but tasks it starts in Codex or ChatGPT Work do, and OpenAI has said the plan allowance for deeper work carries extended limits only for the first month after launch. Plan a capacity review for late October rather than discovering the steady-state limits during a release week. If you are comparing per-token economics across providers before committing, our model coverage page tracks where each family currently sits.

Abstract branching filament structure where branches grow thicker and brighter with each generation, some looping back on themselves
Retries and subagent fan-out compound multiplicatively rather than additively, which is why cost per call stays flat while cost per completed task climbs.

The switching cost of putting coordination logic inside one lab's runtime

There is a clean, documented precedent for this risk and it is sitting in OpenAI's own docs. Agent Builder, the visual workflow canvas launched at DevDay 2025, now appears under Legacy APIs in the API documentation navigation, with a migration guide that tells you to export to Agents SDK code. That guide is admirably honest: the process "does not convert your workflow graph or guarantee that every behavior transfers unchanged", and under limitations it notes that workflows with strong determinism at their core may not migrate faithfully.

Read that as a general law rather than a complaint about one product. The parts of your system that are deterministic — the graph, the branching, the retries, the approval gates — are the parts that transfer worst between agent runtimes, because every vendor expresses them differently. The parts that are probabilistic — prompts, tool descriptions, evaluation sets — transfer almost for free.

That gives you a portability rule that survives the next three launches: keep determinism outside the lab's runtime and probability inside it. An n8n, Temporal or Airflow layer that owns triggers, idempotency, retries and the human approval queue can call the Agents API for the reasoning step and swap the model underneath without touching the graph. The inverse — coordination expressed as subagent configuration inside a hosted harness — is a rewrite every time the vendor renames the abstraction.

Two things genuinely reduce this risk, and both come from OpenAI. The Agents API is powered by the open-source Codex harness, with the coordinating logic inspectable in a public repository, and the API lets you run the sandbox on your own infrastructure or with a partner rather than OpenAI-hosted. Use both. An agent whose compute you control and whose harness you can read is a materially better lock-in position than one you can only configure through a settings panel.

Two panels showing a rigid node lattice on the left dissolving into scattered particles on the right
Prompts and evaluation sets migrate between runtimes almost intact; the branching, retry and approval structure around them is what has to be rebuilt by hand.

Three workflow shapes where a hosted agent builder is still the wrong home

Regulated write-backs

Anything where a regulator, auditor or insurer can later ask "who authorised this and on what evidence" needs an immutable record of the inputs, the decision rule and the approver — before the write, not reconstructed from a transcript afterwards. Compaction works against you here by design. Keep the proposal in the agent and the commit in a system that writes an audit row in the same transaction as the business record.

Multi-system handoffs with different failure semantics

The moment a workflow spans two vendors with independent retry policies, you need distributed-transaction thinking: idempotency keys, compensating actions, a reconciliation job that catches the cases where system A committed and system B did not. A hosted agent will cheerfully retry a step that already succeeded. This is ordinary integration engineering and no harness removes it; it is most of what real agentic AI implementation work consists of.

Long-running state measured in weeks

Dots are built to persist and learn, and for an individual's working context that is the point. For a case that lives for six weeks across a dozen participants and must be resumable after a thirty-day gap, you want state in a database with a schema, migrations and a query interface — not in an agent's accumulated memory. The test is simple: if you cannot write a SQL query that answers "how many of these are currently stuck at step four", the state is in the wrong place.

A decision table, not a verdict

Map the workflow, not the vendor. Most organisations will end up running all three columns simultaneously, and that is the correct outcome rather than a failure of standardisation.

Workflow shapeRight homeWhy
Individual knowledge work: research, drafting, monitoring a feed, preparing documents for reviewDotsHighest-value use of a persistent assistant with its own computer. No integration work, reversible output, a human reads everything.
Open-ended reasoning inside an app you already run: triage, investigation, document analysis, code tasksAgents API (self-hosted or partner sandbox)You get compaction, tool search and subagents without building a harness, and keep your own environment and UX.
Scheduled or event-driven integration across SaaS, with light judgementn8n or equivalent orchestrator, calling a model per stepDeterministic graph, visible run history, cheap to change, model-agnostic. Reasoning stays a node, not the runtime.
Regulated write-backs, payments, clinical or legal commitmentsCustom, with a deterministic validator and explicit approval queueBusiness invariants and audit evidence cannot be delegated to a vendor's safety reviewer.
Long-running multi-party processes, weeks of state, resumableDurable workflow engine (Temporal and similar) with agent stepsReplayable state and compensating actions are the primitives; the model is one activity among many.
Org-wide roles with their own identity and credentialsSpecialist dots pilot, if you qualify — otherwise waitCurrently a hands-on enterprise programme with OpenAI engineering involvement, not a self-serve product.

The honest summary for anyone mid-way through an orchestration decision: Dots does not replace the stack you were budgeting for, and it was never scoped to. It replaces the assistant layer above it, very well, for a specific population of users. The Agents API replaces the harness you were about to write, which is a real saving and the more consequential launch of the two for engineering teams. Everything in the trigger, verification and durable-state layers is still yours, and the organisations that have shipped agents successfully are mostly the ones that worked out which layer they were actually building before they picked a vendor. If you want that audit run against your specific workflows, start here, or look at how we scope these in our case studies.

Frequently Asked Questions

Can I call an OpenAI Dot from my own application?

Not as a developer primitive. Dots are created and used inside ChatGPT, with messaging through ChatGPT, Slack and Teams. The developer-facing equivalent is the Agents API, which exposes the same Codex harness — subagents, compaction, tool search, computer use — through a normal API call. If you are building a product feature, the Agents API is the correct entry point; Dots is for people, not services.

Does Dots replace n8n or a workflow orchestrator?

No, and they solve different problems. An orchestrator gives you a deterministic graph, run history, retries and idempotency across systems. Dots gives you open-ended judgement and a persistent working context. The durable pattern is an orchestrator owning triggers and commits, calling a model or agent for the steps that genuinely require reasoning. Expressing your branching logic inside a hosted agent is the part that becomes expensive to unwind.

What does Auto-review actually check before an agent acts?

OpenAI describes Auto-review as a separate safety system that checks planned steps — sending email, changing files — against your instructions, your Custom Rules and built-in safety requirements, returning a reason to the agent when it blocks. It covers intent, recipient and policy. It does not know your domain invariants: amounts, authority limits, period locks, sanctions screening. Those checks belong in deterministic code you own and can show an auditor.

How should I budget for agent costs when usage is metered?

Budget per completed task, not per call. Autonomous loops multiply token spend through retries, compaction and subagent fan-out, and they also hit tool meters — web search is billed per thousand calls with content tokens charged at model rates, and hosted containers bill per session. Set a hard step and spend ceiling per run in your own code, log cost against task outcome, and alert when the ratio of attempted to completed runs drifts.

How real is platform lock-in with a hosted agent runtime?

Real for coordination logic, mild for prompts. OpenAI's own Agent Builder migration guide states that exporting does not convert your workflow graph or guarantee unchanged behaviour, and that strongly deterministic workflows may not migrate faithfully. Mitigate it by keeping deterministic control flow in a layer you own, running sandboxes on your own infrastructure or a partner's where the API allows it, and maintaining an evaluation set that can be re-run against any model.

Are specialist dots available to buy now?

Not as a self-serve product. OpenAI describes specialist dots — agents with their own identity, credentials and system access — as focused enterprise pilots where its engineering teams work with organisations to define each dot's responsibilities and review paths, with Microsoft Agent 365 integration in progress. Treat it as a design partnership with a scoping cost, and do not plan a Q4 rollout around availability that has not been announced.

Which workflows should I simply not automate with agents yet?

Anything where an incorrect write is expensive and hard to reverse, where you cannot express a deterministic pre-commit check, or where the evidence trail must satisfy an external party. Payments initiation, clinical decisions, contract execution and anything touching a closed accounting period belong in that group today. Agents can still prepare, reconcile and draft in those domains — the restriction is on the commit, not on the reasoning.

OpenAI DotsDevDay 2026agent orchestrationbuild vs buyAgents APIAI implementationvendor lock-inworkflow automation

Published

AI-assisted writing · Reviewed by the Twarx research team

Share:
Share

Research digest

AI Research Briefing

Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.

  • No hype, just measurable outcomes
  • Read by 2,400+ engineers
  • Unsubscribe anytime

Continue reading

More from Articles

Agentic AI Tools

Five Layers of MCP Design Behind Reliable B2B Agent Workflows

16 min read
Agentic AI Tools

n8n vs Zapier AI Agent Automation: The 2-Layer Stack That Cuts Costs 90%

16 min read
Agentic AI Tools

MCP or LangChain? They Solve Two Different Problems

14 min read
Agentic AI Tools

MCP Integration for Workflow Automation: The 2026 Production Framework

16 min read