OpenAI shipped Dots at DevDay 2026 alongside computer use in the Agents API. Here is the layer-by-layer audit: trigger, retrieval, tools, handoff, state, verified action and cost ceiling — and which ones you still own.
Last Updated: October 2, 2026
Dots is a ChatGPT-surface product, not an API primitive. It genuinely absorbs the tool-calling layer, the sandbox layer and a large part of session state. It does not absorb your triggers from systems of record, your access-controlled retrieval, your cross-vendor handoffs, or — the one that matters most — verification of business invariants before a write. OpenAI's Auto-review checks intent and safety, not whether your journal entry balances. Budget for Dots as a knowledge-worker assistant and the Agents API as a harness, and keep coordination logic in something you can export.
What OpenAI actually shipped on 29 September
Three separate things got collapsed into one headline this week, and they sit at different layers of the stack.
The consumer-and-knowledge-worker product is dots: always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser, able to reach over 4,000 apps through the plugin ecosystem. They roll out in ChatGPT to Pro and Business Premium users in eligible markets, with Enterprise, Edu and Healthcare workspaces getting a beta that is off until an admin enables it. Your first dot is included in the plan at no extra cost, with an allowance for deeper work on top.
The developer product is the Agents API, which actually shipped in public beta on 10 September and picked up computer use at DevDay. It exposes the Codex harness — context compaction, tool search, programmatic tool calling, subagents — as a managed service, with the sandbox either OpenAI-hosted, self-hosted, or run through partners including Cloudflare, Modal, E2B, Daytona, Vercel and Oracle.
The third thing is organisational: specialist dots, provisioned with their own identity and credentials, are currently focused enterprise pilots where OpenAI engineering teams work directly with organisations to define responsibilities and approval paths, with a Microsoft Agent 365 integration in progress. That is a professional-services motion, not a self-serve platform, and it should be read as one when you are planning headcount.
Commercially, the same keynote introduced Pro 500 at 25× the Plus allowance, priced at $500 per month, while Altman and Friar fielded IPO questions on the sidelines — Friar told CNBC OpenAI wants to list "when the time is right for our business". Procurement teams reopened agent budgets on the strength of a keynote. The engineering question underneath is narrower and more answerable.
The seven layers a production agent actually needs
Every agent that survives contact with a real business has the same seven layers, whether or not anyone drew them. Vendor comparisons go wrong because they benchmark the middle three and ignore the outer four, which is where the operational cost lives.
The seven layers, in execution order
Something in a system of record changes and the agent must know within a bounded time. Webhook, CDC event, queue message, cron.
Fetching the right context under the right permissions. The hard part is not embeddings, it is honouring row-level access for the identity the agent is acting as.
Selecting and invoking tools without blowing the context window. Tool search and parallel execution matter once you pass roughly thirty tools.
Passing work between agents, between an agent and a human queue, and between vendors. Each hop needs an idempotency key and a resume point.
What the run knows, what it has already done, and what it may safely redo. Distinct from conversational memory.
The check immediately before a write: does this satisfy the business invariant, not merely the safety policy.
A hard stop that halts the run when spend or step count exceeds a threshold, independent of the model's own judgement.
Layer by layer: what Dots absorbs and what it leaves with you
Dots is strongest exactly where self-built harnesses are weakest, and silent exactly where compliance reviews concentrate.
| Layer | Dots / Agents API coverage | What remains on your side |
|---|---|---|
| Trigger | Partial. Dots run proactive background research using read-only tools, and ChatGPT team tasks fire on a schedule or on events like new mail. DevDay also added support for the proposed MCP Events specification so plugins can start automations. | Triggers originating in your ERP, core banking, EHR or warehouse. You still own the webhook receiver, the dedupe table and the replay path when the agent was down. |
| Retrieval | Strong for SaaS. Connected apps inherit the permissions you granted ChatGPT. | Permission-aware retrieval over your own data. The agent sees what the connected account sees — which is usually far more than the task requires. |
| Tool-calling | Strong. Tool search loads definitions on demand and programmatic tool calling runs calls in parallel and filters results in code, so only relevant output re-enters context. | Tool contracts, error semantics, and what a tool does on partial failure. The harness retries; your API has to tolerate being retried. |
| Handoff | Strong inside the runtime. Subagents each keep their own context while a main agent coordinates. | Handoffs that cross vendors or leave AI entirely — to a human queue, an RPA job, a Temporal workflow, a partner's API with its own SLA. |
| State | Strong within a session. Automatic compaction carries information across context windows; dots persist learned preferences over time. | Durable business state you can query, audit and replay. Compaction is lossy by construction and the dropped tokens are not a record you can produce in an audit. |
| Verified action | Safety-complete, business-incomplete. Auto-review checks planned steps against your instructions, Custom Rules and safety requirements before an email or file change runs, and returns the reason when it blocks. | Domain invariants. No external reviewer knows your netting rules, your tax treatment or your refund authority matrix. |
| Cost ceiling | Plan allowances on the ChatGPT side; pure token-and-tool metering on the API side with no additional Agents API fee. | Per-run and per-tenant budget enforcement, and the kill switch. This is yours in both products. |
The verified-action row is the one worth reading twice. OpenAI has been unusually clear about the boundary: certain steps always come back to you, including changing a password or transferring money between financial accounts, and sharing health data always requires a named recipient. Those are good defaults. They are also generic defaults. Nothing in them knows that your organisation requires dual approval above a currency threshold, or that a credit note against a closed period is forbidden.
Treating Auto-review as a business control
Auto-review is genuinely well-designed: OpenAI keeps the controls that enforce it outside the environments dots can change, so an agent cannot disable its own checks. That architecture is exactly right, and it is why teams mistake it for a complete control plane. It is not. It evaluates intent, recipient and policy. It cannot evaluate whether the amount is correct, whether the counterparty is sanctioned, or whether this write would break a reconciliation that runs at 02:00.
The cost shape nobody models: what happens when an agent loops
Usage-based pricing on a single-turn model call is predictable. Usage-based pricing on an autonomous loop is not, because the loop decides how many turns it takes.
Look at the published numbers. gpt-6-astra is $10 per million input tokens and $50 per million output at standard rates on short context, rising to $20 and $75 on long context. The same page lists a top speed tier for astra at $60 and $300. GPT-6.1 Sol lands at $2 and $10 — which is why OpenAI positions it as near-Astra intelligence at a fifth of the standard token prices, and why the first thing to test on any agentic workload is whether Sol holds your eval.
Now add the parts that scale with autonomy rather than with volume: web search at $10 per 1,000 calls with search content tokens billed at model rates, file search at $2.50 per 1,000 calls plus $0.10 per GB-day of storage, and hosted containers from $0.03 to $1.92 per 20-minute session depending on memory. A long-running agent that browses, re-reads its own artefacts and re-plans will touch all four meters on a single task.
The failure mode in production is not an expensive run. It is a cheap run that repeats. An agent hits a flaky endpoint, the harness retries, compaction re-summarises a growing transcript, a subagent fans out three ways, and the cost of one logical task multiplies by the branching factor. None of that shows up in a demo because demos succeed on the first attempt. Instrument cost per completed task, not cost per call, and alert on the ratio between them.
On the ChatGPT side the shape is different but not simpler. Conversations with your dot do not count toward ChatGPT usage limits, but tasks it starts in Codex or ChatGPT Work do, and OpenAI has said the plan allowance for deeper work carries extended limits only for the first month after launch. Plan a capacity review for late October rather than discovering the steady-state limits during a release week. If you are comparing per-token economics across providers before committing, our model coverage page tracks where each family currently sits.
The switching cost of putting coordination logic inside one lab's runtime
There is a clean, documented precedent for this risk and it is sitting in OpenAI's own docs. Agent Builder, the visual workflow canvas launched at DevDay 2025, now appears under Legacy APIs in the API documentation navigation, with a migration guide that tells you to export to Agents SDK code. That guide is admirably honest: the process "does not convert your workflow graph or guarantee that every behavior transfers unchanged", and under limitations it notes that workflows with strong determinism at their core may not migrate faithfully.
Read that as a general law rather than a complaint about one product. The parts of your system that are deterministic — the graph, the branching, the retries, the approval gates — are the parts that transfer worst between agent runtimes, because every vendor expresses them differently. The parts that are probabilistic — prompts, tool descriptions, evaluation sets — transfer almost for free.
That gives you a portability rule that survives the next three launches: keep determinism outside the lab's runtime and probability inside it. An n8n, Temporal or Airflow layer that owns triggers, idempotency, retries and the human approval queue can call the Agents API for the reasoning step and swap the model underneath without touching the graph. The inverse — coordination expressed as subagent configuration inside a hosted harness — is a rewrite every time the vendor renames the abstraction.
Two things genuinely reduce this risk, and both come from OpenAI. The Agents API is powered by the open-source Codex harness, with the coordinating logic inspectable in a public repository, and the API lets you run the sandbox on your own infrastructure or with a partner rather than OpenAI-hosted. Use both. An agent whose compute you control and whose harness you can read is a materially better lock-in position than one you can only configure through a settings panel.
Three workflow shapes where a hosted agent builder is still the wrong home
Regulated write-backs
Anything where a regulator, auditor or insurer can later ask "who authorised this and on what evidence" needs an immutable record of the inputs, the decision rule and the approver — before the write, not reconstructed from a transcript afterwards. Compaction works against you here by design. Keep the proposal in the agent and the commit in a system that writes an audit row in the same transaction as the business record.
Multi-system handoffs with different failure semantics
The moment a workflow spans two vendors with independent retry policies, you need distributed-transaction thinking: idempotency keys, compensating actions, a reconciliation job that catches the cases where system A committed and system B did not. A hosted agent will cheerfully retry a step that already succeeded. This is ordinary integration engineering and no harness removes it; it is most of what real agentic AI implementation work consists of.
Long-running state measured in weeks
Dots are built to persist and learn, and for an individual's working context that is the point. For a case that lives for six weeks across a dozen participants and must be resumable after a thirty-day gap, you want state in a database with a schema, migrations and a query interface — not in an agent's accumulated memory. The test is simple: if you cannot write a SQL query that answers "how many of these are currently stuck at step four", the state is in the wrong place.
A decision table, not a verdict
Map the workflow, not the vendor. Most organisations will end up running all three columns simultaneously, and that is the correct outcome rather than a failure of standardisation.
| Workflow shape | Right home | Why |
|---|---|---|
| Individual knowledge work: research, drafting, monitoring a feed, preparing documents for review | Dots | Highest-value use of a persistent assistant with its own computer. No integration work, reversible output, a human reads everything. |
| Open-ended reasoning inside an app you already run: triage, investigation, document analysis, code tasks | Agents API (self-hosted or partner sandbox) | You get compaction, tool search and subagents without building a harness, and keep your own environment and UX. |
| Scheduled or event-driven integration across SaaS, with light judgement | n8n or equivalent orchestrator, calling a model per step | Deterministic graph, visible run history, cheap to change, model-agnostic. Reasoning stays a node, not the runtime. |
| Regulated write-backs, payments, clinical or legal commitments | Custom, with a deterministic validator and explicit approval queue | Business invariants and audit evidence cannot be delegated to a vendor's safety reviewer. |
| Long-running multi-party processes, weeks of state, resumable | Durable workflow engine (Temporal and similar) with agent steps | Replayable state and compensating actions are the primitives; the model is one activity among many. |
| Org-wide roles with their own identity and credentials | Specialist dots pilot, if you qualify — otherwise wait | Currently a hands-on enterprise programme with OpenAI engineering involvement, not a self-serve product. |
The honest summary for anyone mid-way through an orchestration decision: Dots does not replace the stack you were budgeting for, and it was never scoped to. It replaces the assistant layer above it, very well, for a specific population of users. The Agents API replaces the harness you were about to write, which is a real saving and the more consequential launch of the two for engineering teams. Everything in the trigger, verification and durable-state layers is still yours, and the organisations that have shipped agents successfully are mostly the ones that worked out which layer they were actually building before they picked a vendor. If you want that audit run against your specific workflows, start here, or look at how we scope these in our case studies.
Frequently Asked Questions
Can I call an OpenAI Dot from my own application?
Not as a developer primitive. Dots are created and used inside ChatGPT, with messaging through ChatGPT, Slack and Teams. The developer-facing equivalent is the Agents API, which exposes the same Codex harness — subagents, compaction, tool search, computer use — through a normal API call. If you are building a product feature, the Agents API is the correct entry point; Dots is for people, not services.
Does Dots replace n8n or a workflow orchestrator?
No, and they solve different problems. An orchestrator gives you a deterministic graph, run history, retries and idempotency across systems. Dots gives you open-ended judgement and a persistent working context. The durable pattern is an orchestrator owning triggers and commits, calling a model or agent for the steps that genuinely require reasoning. Expressing your branching logic inside a hosted agent is the part that becomes expensive to unwind.
What does Auto-review actually check before an agent acts?
OpenAI describes Auto-review as a separate safety system that checks planned steps — sending email, changing files — against your instructions, your Custom Rules and built-in safety requirements, returning a reason to the agent when it blocks. It covers intent, recipient and policy. It does not know your domain invariants: amounts, authority limits, period locks, sanctions screening. Those checks belong in deterministic code you own and can show an auditor.
How should I budget for agent costs when usage is metered?
Budget per completed task, not per call. Autonomous loops multiply token spend through retries, compaction and subagent fan-out, and they also hit tool meters — web search is billed per thousand calls with content tokens charged at model rates, and hosted containers bill per session. Set a hard step and spend ceiling per run in your own code, log cost against task outcome, and alert when the ratio of attempted to completed runs drifts.
How real is platform lock-in with a hosted agent runtime?
Real for coordination logic, mild for prompts. OpenAI's own Agent Builder migration guide states that exporting does not convert your workflow graph or guarantee unchanged behaviour, and that strongly deterministic workflows may not migrate faithfully. Mitigate it by keeping deterministic control flow in a layer you own, running sandboxes on your own infrastructure or a partner's where the API allows it, and maintaining an evaluation set that can be re-run against any model.
Are specialist dots available to buy now?
Not as a self-serve product. OpenAI describes specialist dots — agents with their own identity, credentials and system access — as focused enterprise pilots where its engineering teams work with organisations to define each dot's responsibilities and review paths, with Microsoft Agent 365 integration in progress. Treat it as a design partnership with a scoping cost, and do not plan a Q4 rollout around availability that has not been announced.
Which workflows should I simply not automate with agents yet?
Anything where an incorrect write is expensive and hard to reverse, where you cannot express a deterministic pre-commit check, or where the evidence trail must satisfy an external party. Payments initiation, clinical decisions, contract execution and anything touching a closed accounting period belong in that group today. Agents can still prepare, reconcile and draft in those domains — the restriction is on the commit, not on the reasoning.
Research digest
AI Research Briefing
Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.
- No hype, just measurable outcomes
- Read by 2,400+ engineers
- Unsubscribe anytime






