Imagine a mobile release due Friday. New bug reports arrive overnight, a design decision changes, and the team needs one approved account of what is blocking the launch. This is a more useful way to read a release that spans agents, workspaces, models, and APIs: who tracks the work, where do sources and decisions live, what model handles each task, and who approves a risky action?
Two event reports published on September 30 put Dots, ChatGPT Space, and GPT-6.1 Sol at the center of the story. The second report also describes a live demo that did not run smoothly. Neither stage performance nor launch language is a substitute for testing a real workflow. This guide uses official product and API documentation to turn the announcements into a practical evaluation.
| Layer | Product | Main job | What to test first |
|---|---|---|---|
| Continuous execution | Dots | Give an agent an ongoing, defined responsibility | Access, triggers, approvals, activity logs |
| Shared context | ChatGPT Space | Put people and agents on a common page | Source links, page visibility, file permissions |
| Task reasoning | GPT-6.1 Sol | Handle suitable everyday work at a lower unit price | Pass rate, rework, latency, actual bill |
| Developer infrastructure | Agents API | Manage sessions, sandboxes, tools, and recovery | Region, retention, tool costs, failure handling |
These layers form an evaluation framework, not a prepackaged one-click stack. One important distinction: the Dots overview says Dots is powered by Astra. GPT-6.1 Sol is a separately selectable model. Do not assume every Dot task runs on Sol.
Dots: define the job before granting access
A chat request is usually one exchange. A Dot is better understood as an owner of a continuing responsibility: track a goal, gather updates, and ask for a decision or permission when needed. According to OpenAI’s getting-started guide, each Dot has a separate cloud computer and browser. App connections, computer access, and contact channels are configured separately. A Dot can delegate specific work to Codex or ChatGPT Work.

Launch presentation: the useful question is whether a Dot can manage an ongoing responsibility within clear permissions.
Start with one bounded responsibility. In the desktop app, create a Dot for release triage, connect only the designated issue tracker and test repository, choose a contact channel, and set a recurring trigger. The getting-started guide treats app connections, access to your local computer, and contact methods as separate choices. Local-computer access also requires the computer and desktop app to be online. Do not grant it simply because the Dot can request it.
Goal: Prepare a daily mobile-release bug brief at 09:30 Asia/Shanghai.
Sources: The designated issue tracker and test repository only.
Output: Ticket ID, user impact, reproduction evidence, owner,
proposed next action, and original source link.
Allowed: Read, classify, and draft recommendations.
Approval: Ask the release owner before merging, messaging externally,
changing production settings, or closing a ticket.
Escalate: Missing evidence, conflicting sources, customer data,
or a task that exceeds the agreed time budget.
Stop: Pause the recurring task after seven days for review.
Goal: Prepare a daily mobile-release bug brief at 09:30 Asia/Shanghai.
Sources: The designated issue tracker and test repository only.
Output: Ticket ID, user impact, reproduction evidence, owner,
proposed next action, and original source link.
Allowed: Read, classify, and draft recommendations.
Approval: Ask the release owner before merging, messaging externally,
changing production settings, or closing a ticket.
Escalate: Missing evidence, conflicting sources, customer data,
or a task that exceeds the agreed time budget.
Stop: Pause the recurring task after seven days for review.
Seed the pilot with one incomplete report, one duplicate, and one report containing sensitive information. Check whether the Dot links the original record, stops when uncertain, and avoids noisy notifications. Then inspect its Activity page for delegated work and pending approvals. Ending a chat does not necessarily stop work already assigned to a Dot; cancel the trigger at the end of the pilot. Dots getting started.
Keep data source, allowed action, approval gate, and cost in one permission table. A Dot, work delegated to Codex or ChatGPT Work, and external tools may use different quota or billing rules. Verify the actual account’s usage pages before buying capacity. An app connection is a technical capability, not permission to write to that app.
Space: make the current project state reviewable
If Dots move tasks forward, Space provides a place where the team can examine and revise the result. The Space documentation describes pages that combine writing, files, research, comments, and generated work. The practical benefit is fewer conflicting versions scattered across chat threads, attachments, and temporary summaries.

Launch presentation: the shared page is useful when its sources and decisions remain visible to reviewers.
Start with one page, rather than migrating the whole knowledge base. The Space getting-started guide covers creating a page, adding existing notes or files, asking ChatGPT to organize material, and setting view or edit access. A test page might look like this:
Page: Mobile 2.8 release decision
Goal: Friday 17:00 release; existing crash and payment thresholds apply.
Evidence: Ticket board, latest test report, design file, each with update time.
Open decision: Does P0-138 block release? Product owner decides by Thursday 16:00.
To verify: How many of the 18 reports are duplicates? Link each original.
Approved: Only conclusions confirmed by a human owner.
Page: Mobile 2.8 release decision
Goal: Friday 17:00 release; existing crash and payment thresholds apply.
Evidence: Ticket board, latest test report, design file, each with update time.
Open decision: Does P0-138 block release? Product owner decides by Thursday 16:00.
To verify: How many of the 18 reports are duplicates? Link each original.
Approved: Only conclusions confirmed by a human owner.
Have the Dot classify updates as new, resolved, or still blocking, with ticket IDs and source links. Let QA correct classifications in comments; the product owner moves a conclusion to “Approved.” This makes errors traceable without letting an automated summary overwrite the decision. Test the page through a view-only account before adding private material. The official documentation distinguishes a link to an original file from material copied or summarized into a page: the link keeps its original permissions, while copied content may be visible to page readers. A link to a restricted finance file does not share the file, but copying its revenue numbers into a shared page may reveal them.
GPT-6.1 Sol: unit price is only the start of the cost calculation
The official model pages give the following standard text rates in US dollars per million tokens. They are a budgeting starting point, not the price of a completed task. Sol; Astra.
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | $2.50 | $10 |
| GPT-6 Astra | $10 | $1 | $12.50 | $50 |
Cached-input prices apply only to hits. Above 272K input tokens in one Sol request, the model page says the entire request’s input/cache rates are doubled and output is charged at 1.5 times the listed rate. Fast, Batch, Flex, and processing region also have separate pricing rules. Long-context and agent budgets need the full rule, not just the four numbers in the table.

Launch presentation: evaluate the new model on task success and total cost, not the headline rate alone.
For a text-only example, a run with 100,000 input and 20,000 output tokens would cost 0.1 × $2 + 0.02 × $10 = $0.40 with Sol, or $2.00 with Astra. That fivefold difference can disappear at the level of accepted work. The following pass rates and review times are illustrative assumptions, not measured benchmark results. Suppose 30 runs of the same size cost $12 with Sol and $60 with Astra. Sol produces 23 accepted results and needs 3.5 minutes of human review per task; Astra produces 27 and needs 1 minute. At an internal review cost of $30/hour, excluding tools:
| Model | 30-run model bill | Review labor | Accepted | Cost per accepted result |
|---|---|---|---|---|
| Sol | $12 | $52.50 | 23 | about $2.80 |
| Astra | $60 | $15 | 27 | about $2.78 |
The conclusion can reverse for another task mix. The useful metric is:
Total cost per accepted result = (model + tools/run time + human review/rework) / accepted results
Use 30 historical tasks across synthesis, code changes, long-document analysis, and high-stakes decisions. Fix inputs, acceptance rules, reasoning effort, and allowed tools, then have a reviewer score outputs without knowing which model made them. Record first-pass acceptance, p95 latency, tool costs, retries, and human correction time. The Sol model page directs developers to the Responses API for tool calling; comparing only text outputs from Chat Completions would miss the agent workflow.
OpenAI’s model selection guide offers a starting route: consider Luna for narrow high-throughput tasks, try Sol for everyday work that needs judgment, and evaluate Astra for the hardest tasks. This is a hypothesis, not a permanent rule. If a cheaper model produces more expensive errors in a particular workflow, its lower unit price does not save money.
Agents API and the rest: separate implementation from announcement
Product owners can begin with Dots and Space. A team building agents into its own product should also examine the Agents API overview. It describes a managed Codex execution environment, persistent sessions, sandboxes, tools, and recovery. It reduces work on the execution layer; it does not decide what business data may be read or what result is good enough.
For the release-triage example, draw five explicit states: receive the ticket, retrieve evidence, draft a recommendation, obtain human approval, and execute or return it. Store source links, tool calls, failures, and approver in the same task record. A sandbox constrains execution, but a tool that holds production write access still needs restrictions at the tool layer. Inject timeouts, duplicate tasks, bad citations, and denied permissions into a pilot. See whether the run stops at the right boundary and a person can take over.
Cost has more than one meter: model usage, tools, and managed containers may be billed separately. Set per-task, daily, and retry caps. There is also an immediate deployment constraint: at the time of review, the official documentation says the hosted service is available in the US region and does not yet support zero data retention (ZDR). Resolve strict residency or retention requirements before a pilot. Agents API documentation.
Other launch terms need separate checks:
- UltraFast is a speed mode, not a new model. Official guidance says up to eight times faster; Astra has broader access but low rate limits, and processing is currently limited to US or global regions. Do not infer that every Sol account has the same mode. Measure p95 latency and queueing before paying for speed.
- Sign in with ChatGPT handles identity and consent. Its quickstart separates login, permission, and optional plan usage. Eligibility for plan-backed use is limited to qualifying open-source local apps and selected partners; a user’s subscription does not mean every third-party app gets a free API.
- Decisions API and Marketplace need their own verification. Both appear in event reports, but this review did not find official documentation sufficient to establish public interfaces, availability, and billing. Track them as leads rather than base a contract or roadmap on the names alone.
Check region, eligibility, quotas, and pricing in the actual account before a wider deployment. A stage demo is a clue, not a service agreement.
A seven-day pilot that produces a decision
| Day | Action | Evidence to keep |
|---|---|---|
| 1 | Select 30 historical tickets; have people label correct outcomes and forbidden actions | Samples, reference answers, baseline time, error types |
| 2 | Configure read-only Dot sources, time zone, schedule, contact, and a stop date | Permission list, trigger log, cancellation method |
| 3 | Build one Space page for evidence, pending decisions, and approved conclusions; test with a view-only account | Page access, original links, approval records |
| 4–5 | Compare Sol with the current model on fixed inputs, prompts, tools, reasoning effort, and blind scoring | First-pass rate, p95 latency, tokens, review minutes |
| 6 | Inject network loss, duplicate tasks, bad citations, and an unauthorized request | Errors, notification count, takeover and rollback |
| 7 | Calculate cost per accepted result and decide expand, observe, or stop | Scorecard, actual bill, owner sign-off |
Set thresholds on day one. A starting rule to adjust for your business: at least 27 of 30 tasks meet the existing quality bar; zero unapproved sends, merges, or production writes; every material finding traces to an original source; and cost per accepted result beats the current workflow. Passing 30 examples justifies a wider controlled pilot, not an immediate full rollout. An access violation or sensitive-data leak should stop connections and recurring work immediately rather than wait for day seven.
The final question is whether the team actually saved time. If writing the summary became checking hallucinations and chasing logs, the workflow has not paid off. Put acceptance, review minutes, total cost, and error types side by side before changing prompts, models, or the degree of automation.
For teams managing model calls through ClaudeAPI, record task labels, model IDs, usage, tool charges, and acceptance logs separately before testing Sol. A model announcement does not prove that a gateway supports it yet. Check the platform model list, returned model ID, and actual bill. Those records preserve task definitions, cost baselines, and a rollback path when models or providers change.
Sources and verification
- Event reporting: DevDay summary by 数字生命卡兹克 and live report by 发现明日产品的. This article uses a new structure and an independent pilot framework; the reports’ opinions are not treated as product facts.
- Official documentation: Dots, Dots getting started, Space, GPT-6.1 Sol, model selection, Agents API, UltraFast, and Sign in with ChatGPT. Check current pricing, permissions, and availability at the time of use.



