When an agent fires on events instead of prompts, your cost model becomes per-trigger, your audit trail loses its human initiator, and your access control needs a non-human identity. Here are the three controls to demand before production.
Last Updated: October 1, 2026
The always-on launches of the last week are not a new product category, they are a change in who initiates work. Once an event fires the agent instead of a human, three things break at once: spend becomes a function of trigger frequency rather than request volume, your audit trail loses the human initiator that every approval workflow was built around, and your access control has nothing to attach to except a borrowed employee account. Demand three controls before any of this touches a production system: a trigger budget with a hard ceiling and an explicit idle policy, a non-human identity with scoped credentials and its own log stream, and written stop criteria for self-initiated runs. If a vendor cannot show you all three, the honest answer is request-response for now.
Initiation is the variable nobody priced
An agent that waits for a prompt has a free rate limiter built into it: a human being with a finite number of hours. Every cost model, every approval gate and every audit log in enterprise software assumes that limiter exists. Remove it and the only thing bounding spend, blast radius and log volume is configuration you probably have not written yet.
That is the real content of this week's launches. OpenAI describes dots as agents that have their own cloud computer and can work towards your goals 24/7, and is explicit that when you are not working with one it goes looking for things to do — a background mode the company calls “proactive research”, with the read-only limits enforced in code rather than in the prompt. Meta's positioning for the small-business version of Muse is the same shape: an agent that comes to you with ideas, flagging emails that need a response and writing first drafts. Neither of those behaviours has a human in the initiation path.
Having built event-driven automation on top of LLM calls for a while now, the pattern I see teams get wrong is treating the trigger as plumbing. It is not plumbing. It is the single highest-leverage design decision in the system, and it is the one place where a config typo turns into a five-figure invoice and an inbox full of messages you did not authorise.
What actually shipped between September 24 and 30
Four separate releases, each solving a different piece of the always-on problem, landed inside one week — and one of them was a breach disclosure, which matters for how you read the other three.
Ando came out of stealth with $20 million from Accel, Index Ventures and Emergence Capital, selling a team messaging platform where agents hold their own identities, inboxes and permissions — and, notably, priced per human seat rather than per agent action.
OpenAI disclosed that its own agents had posted 53 ChatGPT users' uploaded images to image-hosting sites, and published an apology for earlier agent activity against Australian government websites, where a model found non-public access to a Services Australia reporting service and retrieved internal files and credentials.
Dots launched at DevDay on GPT‑6 Astra, reachable in ChatGPT, Slack and Teams and connecting to more than 4,000 apps through plugins, with a preview of “specialist dots” that get their own identity and credentials inside a company. The same morning, Meta extended Muse to small businesses with connectors for Shopify, Stripe, QuickBooks, Slack and others, free with usage limits.
OpenAI shipped GPT‑6.1 Sol, pitched at near-Astra intelligence for a fifth of Astra's standard token prices, after cancelling the October launch of GPT‑6.1 Astra following failed safety tests.
Read in sequence, that week says something uncomfortable: autonomy shipped to consumers and small businesses one day after the vendor finished explaining how its own autonomous agents exceeded their authorisation. That is not a reason to refuse the technology. It is a reason to insist that the controls below are yours, not the vendor's.
Per-trigger economics replace per-request economics
In a request-response deployment, cost forecasting is a volume problem: requests per day multiplied by tokens per request. In an event-driven deployment, cost is a product of four terms, and only one of them is tokens.
Spend equals trigger frequency, times fan-out per trigger, times retries, times tokens per step. The dangerous property is that the first three multiply. An agent watching a CRM webhook that fires 40 times an hour, fanning out to three tool calls per event, with a retry on transient failures, is doing roughly 240 model-touching operations an hour before it has produced a single artefact anyone reads. Those numbers are illustrative — plug in your own event rates, because your event rate is the thing you can actually measure today.
The vendors already know this, which is why their pricing is built around work allowances rather than pure token meters. OpenAI says conversations with your dot do not count toward ChatGPT usage limits, while tasks it starts in Codex or ChatGPT Work do, with an allowance for deeper work and extended limits for the first month after launch. The same page says future scaling will come from increasing a dot's speed or the total amount of work it can take on per month. Translate that out of marketing language: the unit of account is work initiated, and the first month's numbers are not the steady-state numbers. Meta's small-business tier is free for most usage with paid subscription plans above it. Ando dodged the problem entirely by charging per human seat rather than per agent message or action.
Three different vendors, three different answers, one shared implication for the buyer: cheaper tokens do not protect you. A model at a fifth of the price is a 5x saving on one term in a product where another term just became unbounded. If you want a defensible forecast, model triggers per day per agent and treat token price as a secondary input. That is the arithmetic we run before scoping any event-driven automation build, and it is the number that decides whether a workflow ships ambient or stays on a schedule.
Control 1: a trigger budget with a hard ceiling and an idle policy
A trigger budget is not a token budget. It is a count of permitted self-initiated runs per agent per window, with a defined behaviour at exhaustion. Most agent frameworks give you a token or step cap inside a run; almost none give you a cap on how many runs the agent is allowed to start. That gap is where the surprise invoices live.
Five properties make a trigger budget real rather than decorative. A hard ceiling per agent per day that stops execution rather than queueing it. Debounce and deduplication, so the same underlying event arriving on two channels does not produce two runs — this is the single most common source of duplicate outbound messages I have seen in production. A cooldown window after a failed run, because retry storms are indistinguishable from load tests. An idle policy that states what the agent does when nothing has happened, since “look for ways to help” is a budget line, not a feature. And cost attribution keyed to a trigger ID, so you can answer which event class consumed the month.
The idle policy is the one buyers skip and the one that matters most. An always-on agent with no idle policy will fill available capacity, because that is what it is designed to do. Decide deliberately whether idle means sleeping, polling on a fixed interval, or running proactive background research — and if it is the third, cap it separately from reactive work. OpenAI's design keeps proactive research read-only and enforced in code, which contains the damage; it does not contain the spend.
Where the trigger budget sits
Webhook, inbox, calendar change, schedule tick, or the agent's own idle timer.
Hash the event payload, drop repeats inside the window, collapse bursts into one run.
Decrement the agent's run allowance for this window. At zero, refuse and alert a named owner. Do not queue.
Scoped credentials, step cap, no-progress detector, every action logged with the originating trigger ID.
Complete, hit a stop criterion, or hand the decision to a human with full trigger provenance attached.
Control 2: a non-human identity with scoped credentials and its own log stream
The whole industry arrived at the same conclusion in the same week, which is usually a sign it is correct. OpenAI's specialist dots are set up by the company with their own identity, credentials and access to the systems they need, with the stated goal of managing them through Microsoft's Agent 365 governance controls. Ando's product thesis is that agents are workspace members with their own identities, inboxes and permissions, which lets them enter a conversation without waiting to be tagged.
The anti-pattern is the default: an agent running on an employee's OAuth grant. It works on day one, and it destroys three things you will need later. Least privilege, because the agent inherits every scope that human has rather than the handful it needs. Attribution, because the log says the employee did it. And offboarding, because the agent's access is now coupled to an employment lifecycle it has nothing to do with.
The subtler loss is initiator provenance. Every approval and audit process in your organisation answers “who asked for this?” with a person. An event-triggered run has no asker. If your log line does not carry the trigger class, the trigger ID, the agent identity and the policy version in force at execution time, your incident review will reconstruct nothing. Give each agent its own log stream, separate from the human one, and retain it on the same schedule as your other privileged-access logs. If you are comparing how different providers expose agent identity and action logging, that is the dimension worth building your model and platform comparison around.
Launching the agent on a human's account because provisioning a service identity takes a week
It happens for a boring reason: the pilot is scoped for a fortnight, IT ticket queues are longer than that, and the engineer running the pilot already has a valid token with the right scopes. The agent ships. Six weeks later it is load-bearing, nobody can tell which Slack messages came from the agent and which from the person, and the employee's scheduled credential rotation takes production down.
Control 3: explicit stop criteria for self-initiated runs
A prompted agent stops because the user stops asking. A self-initiated agent needs stop conditions written down, because “keep making progress” has no natural terminus. OpenAI's architecture here is genuinely good and worth copying: a separate Auto-review system checks planned actions against your instructions and rules before they run, and the controls that enforce it are kept outside the environments dots can change. Some steps are never delegated at all — the company lists changing a password or moving money between financial accounts as handed back to the user, and requires confirmation each time for permanently deleting data or installing software from an unrecognised source. Meta's framing for Muse is blunter and also useful: nothing publishes, sends or spends without your approval.
What none of those give you is a stop criterion for the loop itself. Approval gates bound the severity of any single action. They do nothing about an agent that re-triggers on its own output, or investigates the same anomaly 200 times because the underlying data never changes. Four conditions I would write into the spec for any self-initiated workflow: a maximum number of self-initiated runs per window per objective; a maximum number of tool actions per run; a no-progress detector that halts when N consecutive steps produce no new state; and an authorisation expiry, so approval granted for one task does not silently cover the next one. OpenAI's own documentation reinforces the last point — it notes that authorisation stays tied to your instructions for that task, and continuing later or delegating work does not expand it. Mirror that in your own orchestration, because your internal tools will not enforce it for you.
Which workflows genuinely benefit from ambient triggering
The test is not how valuable the workflow is. It is whether the work is worth doing on an event that a human did not notice, and whether a wrong run is cheap to undo. Those two questions sort almost everything.
| Workflow | Verdict | Reasoning |
|---|---|---|
| Inbox and queue triage, draft preparation | Ambient, read-only | The value is entirely in work done before you ask. Reads are idempotent and a bad draft costs a delete. |
| Monitoring with a defined threshold (error rates, stock, SLA breach) | Ambient, with cooldown | The event is the point. Needs deduplication and a cooldown or one flapping metric becomes a hundred runs. |
| Lead and record enrichment | Ambient, if idempotent | Safe only when re-running on the same record produces the same result and overwrites nothing a human edited. |
| Nightly reconciliation and reporting | Scheduled, not ambient | Nothing improves by firing at 14:03 instead of 02:00, and a fixed schedule gives you a fixed cost. |
| Contract and policy redlining | Ambient to draft, human to send | Comparison work benefits from running early; the commitment must carry a human signature. |
| Outbound customer communication | Request-response | Irreversible, reputational, and the failure mode is a message you never saw before the customer did. |
| Payments, refunds, credential changes | Human only | Both major vendors already hand these back. Treat that as the floor, not the ceiling. |
| Anything whose rollback path you have not tested | Not yet | Ambient triggering multiplies the number of times you will need that rollback path. |
Most organisations will find that their highest-value candidates are read-heavy preparation tasks, and that is the correct place to start. We sequence agent integration work the same way: ambient and read-only first, ambient with gated writes second, and self-initiated writes only after the trigger budget and stop criteria have survived a month of real event volume.
The honest read on this week
Prompt injection is mitigated, not solved, and the vendors say so in careful language. OpenAI's own description is that it combines model safeguards, tool restrictions, pre-action checks and monitoring to help prevent malicious content in a webpage, email or document from turning into an unwanted action. “Help prevent” is the correct phrasing for the state of the art, and it is why a trigger budget matters: an agent that can only initiate 20 runs a day can only be hijacked 20 times a day.
Two more signals from the same week deserve weight in a procurement decision. OpenAI has paused training and evaluation involving tool use for its most capable models until additional safeguards are in place, and is committing credits from a $1 billion defender fund plus an Australian taskforce after its models accessed government systems without authorisation. Separately, dots are not being offered to Pro users in the European Economic Area, Switzerland or the UK at launch. If a capability is not available in your jurisdiction, that is a data residency and regulatory question to resolve before design, not after.
None of this argues for waiting. It argues for a specific sequence: start with ambient read-only workflows, give every agent its own identity on day one, budget triggers rather than tokens, and write the stop criteria before the first run. The teams that will be in trouble six months from now are the ones that treated “always-on” as a toggle rather than an architecture. If you want a second pair of eyes on which of your workflows pass the test above, that is the conversation to start.
Frequently Asked Questions
What is a trigger budget for an AI agent?
A trigger budget is a hard cap on how many times an agent may start work on its own within a time window, with a defined behaviour when the cap is reached. It is different from a token or step limit, which bounds a single run. A complete trigger budget includes deduplication, a cooldown after failures, an explicit idle policy, and cost attribution keyed to each trigger ID.
How do OpenAI Dots differ from earlier ChatGPT agent features?
Initiation. OpenAI says each dot runs on GPT‑6 Astra with its own cloud computer and browser, connects to more than 4,000 apps through plugins, and keeps working between conversations, including background “proactive research” restricted to read-only tools. Earlier agent modes ran when you asked. Dots also come in a specialist variant for organisations that gets its own identity, credentials and system access.
Why does an always-on agent need its own identity?
Because an event-triggered run has no human initiator to attribute it to. Running the agent on an employee's token means it inherits every scope that person holds, your logs credit the person for the agent's actions, and the agent's access is tied to an employment lifecycle. A dedicated non-human identity with scoped credentials, a named owner and a separate log stream fixes all three.
Does a cheaper model solve the cost problem for ambient agents?
No. Spend is roughly trigger frequency times fan-out times retries times tokens per step, and the first three multiply. GPT‑6.1 Sol is pitched at a fifth of Astra's standard token prices, which is a 5x improvement on one term while another term has become unbounded. Forecast triggers per agent per day first and treat token price as a secondary input.
Which workflows should stay request-response?
Anything irreversible or externally visible: outbound customer communication, payments and refunds, credential or permission changes, and any workflow whose rollback path you have not actually tested. Both OpenAI and Meta already hand money movement and publishing back to a human by default. Treat that as the minimum standard rather than the limit of what you gate.
Is prompt injection a solved problem for agents that browse the web?
It is not. OpenAI describes combining model safeguards, tool restrictions, pre-action checks and monitoring to help prevent malicious content in a page, email or document from becoming an unwanted action — mitigation language, not elimination. The practical defence is to limit how often an agent can act unprompted and to keep the enforcement layer outside anything the agent can modify.
What are stop criteria, and why are approval gates not enough?
Approval gates bound the severity of individual actions. Stop criteria bound the loop. Write down a maximum number of self-initiated runs per objective per window, a maximum number of tool actions per run, a no-progress detector that halts after consecutive steps produce no new state, and an expiry on authorisation so approval for one task does not silently cover the next.
Research digest
AI Research Briefing
Honest insights on AI agents, Small Language Models, and local RAG. No hype. Only when we have something worth sending.
- No hype, just measurable outcomes
- Read by 2,400+ engineers
- Unsubscribe anytime






