\n
Tech: The AI Doesn’t Stop When a Single Instruction Ends
“Summarize market trends every Monday morning.”
Until now, AI has mostly been the kind of tool that gives you a one-time answer to a request like this. If you wanted the same task done the following Monday, you had to open a new conversation and repeat the instruction. But OpenAI’s newly unveiled always-on AI agent, dots, aims to change that experience. Even after a user closes the conversation, the AI can remember the goal and keep tracking and carrying out the work it requires.
According to the public description, dots runs on GPT-6 Astra and is less a simple chatbot than a persistent agent that manages long-term goals. For example, if asked to prepare market trend reports, it could monitor relevant news and industry data, prepare summaries when significant changes occur, and deliver reports on a set schedule.
The biggest difference from existing AI is that the work doesn’t stop when the session ends. The agent retains the user’s goals, progress, and the results of previous tasks, then plans what to do next. To do this, agents typically work through a structure like the following:
- Goal management: Stores long-term requests, such as “summarize market trends every week,” and sets priorities.
- Planning: Breaks goals into subtasks, such as gathering data, detecting changes, categorizing key issues, and writing reports.
- Tool and app integration: Finds information and continues working across services such as documents, email, messaging apps, CRMs, and calendars.
- State memory: Uses the key points from last week’s report and the user’s feedback to improve the next one.
- Exception handling: Can ask the user for confirmation when an important change occurs or a task requires authorization.
If this tech shift becomes reality, AI will no longer be just a “tool that answers questions.” It will evolve into a digital work partner that takes on specific goals and keeps working—like a research assistant, operations aide, or development project monitor.
But the advantage of always-on operation also brings new responsibilities. The more an agent moves across apps and data, the more rigorously access permissions, personal information, and unintended automated actions must be managed. Ultimately, the value of agents like dots will depend not only on how much work they automate, but also on what users entrust to them—and how much control they retain.
The Engine Behind Always-On Tech Agents: A System That Remembers, Plans, and Acts
What does it take for AI to keep working while a user is away? The answer isn’t simply “a smarter model.” It takes an entire system that can hold on to goals, track progress, choose the right tools, and put them to work.
An always-on agent like OpenAI’s dots is different from a conventional chatbot that generates an answer in a single conversation. It must maintain users’ long-term goals, move across multiple applications, and decide what to do next at the right moment. This architecture can be understood in four broad layers: memory, planning, execution, and control.
Persistent Memory That Preserves Goals and State
For a typical chatbot, the context of a task can easily disappear when the conversation ends. An always-on agent, by contrast, may need to hold on to a goal for days—for example, “Research competitors and write a report by next week.”
That requires persistent memory that stores task-relevant information in a structured way, rather than simply keeping a chat history.
- Long-term goals: The outcome the user ultimately wants
- Task status: What’s complete and what remains
- User preferences: Report format, approval process, and priorities
- Work context: Connected documents, schedules, projects, and stakeholder information
- Execution history: Which tools were used, when they were used, and what results they produced
For example, if an agent is assigned market research, its job doesn’t end with summarizing the news. It must distinguish between sources already reviewed and those that haven’t been checked, while also taking into account the perspective of previous reports and the user’s feedback. That way, it can deliver more consistent results over time instead of starting the research from scratch every time.
The Role of GPT-6 Astra: The Brain Behind Planning
According to publicly available descriptions, dots is powered by GPT-6 Astra. While the full details of its architecture have not been disclosed, a model like this is likely to serve as the agent’s core brain, handling reasoning, planning, and tool selection.
An agent can’t simply act on a user’s abstract instruction as written. It first has to break the goal down into smaller, verifiable tasks.
“Analyze why new customers are churning and prepare a response plan.”
That request could be broken down into a workflow like this:
- Gather data on recently churned customers from the CRM.
- Analyze customer segments and usage patterns.
- Look for common complaints in customer inquiries, surveys, and support tickets.
- Prioritize the causes.
- Recommend outreach campaigns and product improvements.
- Request review from the appropriate people if needed.
In this process, the model acts not merely as an answer generator, but as a planner. It interprets the goal, determines the order of tasks, and adjusts the plan based on interim results. That’s what can make always-on agents more flexible than traditional automation tools.
An Execution Layer That Connects Multiple Apps
A plan alone doesn’t get the work done. To produce real outcomes, an agent needs to use external systems such as email, document tools, CRMs, databases, and collaboration messengers.
An always-on agent therefore needs an execution layer that turns the model’s decisions into actions. This layer is typically built using APIs, connectors, browser automation, and integrations with internal tools.
An agent might:
- Search a document system for the latest policy files
- Aggregate and analyze figures in a spreadsheet
- Retrieve a list of customers matching specific criteria from a CRM
- Request a review from a stakeholder through Slack or email
- Create a draft meeting invitation in a calendar tool
- Check error logs and change history in a code repository
The key point is that an agent doesn’t get unlimited use of every tool. It needs orchestration to determine which tool to call and when, whether the results are sufficient, and whether it’s ready to move on to the next action. This ability to coordinate across applications could become a key factor in determining which companies lead the race to build tech agents.
A Feedback Loop That Verifies Results and Revises Plans
Real-world work rarely follows the plan perfectly. Data may be missing, access may be denied, or an external system may respond slowly. A good agent shouldn’t treat failure as a reason to simply stop. It should evaluate the results and then plan its next move.
The process can be summarized like this:
Set goal → Plan tasks → Use tools → Check results → Revise plan → Complete or request approval
For example, if there isn’t enough data to write a report, the agent should look for additional sources or ask the user to make a judgment rather than draw conclusions on its own. It’s also safer to build in approval steps before taking consequential actions—for instance, “I’ve prepared a draft. Please review it before I send it.”
Autonomy Depends on Access Controls and Oversight
The more powerful always-on agents become, the more security and control become essential—not optional. An agent that runs for extended periods and can access email, documents, customer information, or deployment systems could cause serious problems even with a small mistake.
In enterprise environments, the following safeguards are especially important:
- Principle of least privilege: Allow access to data and tools only to the extent needed for the task
- Approval gates: Require human confirmation for high-risk actions such as payments, external communications, or data deletion
- Execution logs: Keep an auditable record of what the agent viewed and changed
- Policy-based blocking: Restrict the transfer of sensitive information, privilege escalation, and unauthorized tool use
- Anomaly detection: Monitor repeated actions, unusual data access, and unexpected workflows
Ultimately, the value of always-on agents like dots won’t be determined by a model’s performance alone. It depends on how well persistent memory that retains users’ goals, planning that breaks work into manageable tasks, execution across apps, and governance that limits and verifies actions all work together.
The central question in the age of agents is no longer “What can AI answer?” It is shifting to “How safely and reliably can AI get work done over time?”
Tech: From a Race for Better Models to a Race for AI That Gets the Job Done
Will the winner of the AI race be the company that builds the biggest model—or the one that creates a product capable of seeing a user’s work through to the end?
The question OpenAI’s always-on AI agent, dots, raises is clear: competition in tech no longer stops at unveiling “smarter models.” A new battleground is emerging: the agent experience—AI that understands users’ goals, connects multiple tools, and carries work through over time.
Traditional generative AI has operated one conversation at a time. A user asks a question, and it answers; a user requests a document, and it drafts one. Always-on agents like dots, by contrast, aim to manage goals continuously. If asked to “analyze competitors every week and write a report,” for example, the agent wouldn’t simply deliver a one-time summary. It would keep tracking news and market data, organize the information it needs, and deliver the finished work at the agreed time.
That difference is significant from a technical standpoint, too. An AI that gets the job done needs more than language skills. It must also be able to:
- Maintain long-term memory and task state: Remember what it has already done and how far along the work is.
- Plan and replan: Break a larger goal into smaller tasks, then adjust the sequence when something goes wrong or circumstances change.
- Connect to tools and apps: Safely call APIs across services such as email, documents, CRM systems, calendars, and code repositories.
- Manage permissions and approvals: Require user approval and policy checks before taking high-impact actions, such as sending important emails, making payments, or modifying data.
- Verify results: Go beyond performing the task to confirm that the result meets the goal and its requirements.
Ultimately, model performance is only the starting point. A product’s real competitive edge is likely to depend on the orchestration layer built on top of the model, along with its app ecosystem, user experience, and security framework. Even an exceptional reasoning model will have limited use in the workplace if it can’t connect to a company’s systems, can’t manage permissions reliably, or loses track of a task’s progress.
Seen through this lens, Google’s large-model strategy, OpenAI’s push to turn agents into products, and Nvidia’s agent-safety platform represent different fronts in the competition. Some companies are building more powerful models; others are designing work automation users can experience firsthand; and still others are providing the infrastructure to keep agents within their authorized boundaries.
The most important question in the tech market may soon no longer be, “Which model scored higher on benchmarks?” Instead, it will be: “Which AI can complete our organization’s work more safely, with less supervision, and all the way through?”
The arrival of dots shows that the center of gravity in this competition is shifting from model performance to getting work done end to end—and to the agent ecosystem that makes it possible.
Tech: From Research to DevOps, What You Can Delegate to a Digital Colleague
Asking an agent to “write a report” is very different from having it continually gather the information it needs and update its draft whenever something changes. The first is a conversational request for a one-time deliverable. The second is a long-term mission: work that remembers its goals, schedule, sources, and review criteria as it goes.
That’s why always-on agent systems like dots are drawing attention. They point to the possibility of a “digital colleague” that observes information and keeps work moving within defined boundaries—even when the user isn’t entering prompts each time.
Research: Building Context, Not Just Collecting Information
Research is one of the first areas likely to change. Traditional generative AI has excelled at producing answers at the moment a user asks a question. By contrast, a long-running agent can track a specific topic over time.
Suppose a company sets the goal: “Summarize competitors’ AI product launches and pricing changes every week.” The agent could carry out the work as follows:
- Continuously monitor trusted news sources, regulatory filings, tech blogs, and product update pages.
- Compare newly discovered information with existing research notes.
- Separate out duplicate articles and sources with low credibility.
- Prioritize significant changes, such as price increases, new features, and partnerships.
- At a scheduled time, provide a summary report along with “what’s changed since last week.”
The key isn’t the summary itself. It’s maintaining context as it accumulates over time. Rather than having people repeatedly search for and compare information, the agent manages the history of changes.
External information can, however, contain errors or bias. For research used in important decisions, it’s essential to build in source links, collection timestamps, credibility standards, and a final human review process.
Reports and Operations: A System That Keeps Drafts “Ready”
The role of agents can change in reporting, too. “Write this week’s performance report” is a one-off automation request. A long-running agent, however, can collect the necessary data, look for gaps, and prepare a draft well before the report is due.
For example, a weekly business reporting mission could be set up to:
- Collect key metrics every day from the CRM, advertising dashboards, and customer inquiry systems
- Calculate performance against targets and detect sharp fluctuations
- Connect relevant campaigns, incident records, and customer feedback that may help explain those fluctuations
- Generate both an executive summary and a detailed working draft on Friday
- Request a review from the person in charge before publishing the report if figures fall outside established thresholds
This approach does more than reduce the time spent writing documents. It connects separate tasks—data collection, anomaly detection, organizing possible causes, and report writing—into one workflow. That kind of end-to-end workflow orchestration is precisely why the tech industry is paying attention to agent UX.
Marketing and Customer Operations: Ongoing Observation, Not Just Repeated Execution
In marketing operations, agents could become more like “always-on analytical partners” than campaign execution tools.
If the goal is to increase new-customer conversion, for instance, an agent could be designed to monitor traffic and conversion data by channel, then investigate when a particular segment’s churn rate rises. It could compare factors such as landing-page changes, advertising budget adjustments, and email send schedules to suggest improvements.
But sending messages to customers or changing budgets creates direct business risks. It’s therefore best to divide permissions into levels:
- Observation access: View data and detect anomalies
- Recommendation access: Draft improvements and recommend execution priorities
- Approval-required access: Change budgets, send messages, or modify customer information
- Limited execution access: Automate recurring tasks within pre-approved boundaries
More autonomy isn’t always better. Autonomy and approval steps should be designed according to the potential impact of each task.
Development and DevOps: From Coding Tool to Colleague That Looks After Systems
Development is an area where long-running agents could be especially valuable. Code generation is already a familiar capability, but in real-world development, maintaining a codebase and operational environment over time takes up more effort than writing a single line of code.
An agent could take on long-term missions such as:
- Tracking repository changes and finding areas where test coverage is declining
- Checking dependency packages for security vulnerabilities and needed updates
- Connecting incident alerts, application logs, and infrastructure metrics to suggest possible causes
- Analyzing recurring patterns in deployment failures
- Organizing technical debt and prioritizing it by impact and difficulty of remediation
- Finding discrepancies between documentation and the actual code, then drafting documentation updates
In DevOps environments, separating observation, judgment, and execution is especially important. An agent suggesting that “the database connection pool is likely saturated” after analyzing logs is one thing; restarting a production server or changing its configuration is quite another.
That’s why it’s safer to start with read-only access and a recommendation-focused role. Later, automatic execution permissions can be granted for recurring tasks whose reliability has been validated—with rollback capability and audit logs in place.
What It Takes to Delegate Well: Goals, Permissions, and Review Criteria
Simply telling a digital colleague to “take care of it” isn’t enough to make it useful. A long-term mission requires at least three things.
| Design element | Questions to ask | |---|---| | Goal | What needs to be improved or completed? | | Scope | Which data, apps, and team tools can it access? | | Success criteria | What result, and by when, would count as a job well done? | | Permissions | Can it only make recommendations, or can it take action too? | | Approval process | Which actions must always be confirmed by a person? | | Recordkeeping | Which decisions and actions should be logged? |
Ultimately, the point of adopting agents isn’t to remove people entirely. It’s to create a structure in which people set the goals and boundaries, agents handle ongoing observation and preparation, and important actions go through appropriate approval processes.
From research and reporting to customer operations and DevOps, the tech race ahead may hinge less on how convincing an answer AI can produce in a single shot, and more on how safely and consistently it can keep real work moving.
The Final Threshold of Tech Autonomy: How Much Should We Delegate—and When Should We Stop?
AI’s ability to stay focused on a goal is powerful. But it also means that AI can pursue a badly defined goal for a long time. In the age of always-on agents, the real question is no longer “What can they do?” The question now is “What should they be allowed to do—and when must they stop?”
Agents like dots, which can move across multiple apps to pursue long-term goals, are different from simple chatbots. They can do more than organize schedules or summarize documents: they can affect real workflows by sending emails, updating CRMs, adjusting inventory, or deploying software. In these settings, even a small error in judgment can snowball into serious harm when repeated over and over.
The Longer a Goal Persists, the Greater the Risk
The risk of always-on agents lies less in one-off mistakes than in persistence. For example, an agent given only the goal “cut costs” might overreach in pursuit of short-term gains, scaling back essential functions such as customer support, security reviews, or quality control.
The problem isn’t that AI has malicious intent. It arises when goals are incomplete, success criteria are too simplistic, or the agent’s permissions are too broad. An agent will faithfully optimize for the goal it is given—but we should not expect it to automatically understand an organization’s unspoken context or make ethical judgments on its own.
Particular care is needed in the following situations:
- Ambiguous goals: Instructions open to interpretation, such as “improve performance” or “reduce inquiries”
- Excessive permissions: Granting access all at once to payments, contracts, customer-data deletion, and system deployments
- Contaminated external inputs: Hidden prompt injections in emails, webpages, or documents that alter an agent’s behavior
- Delayed feedback: Errors that go undetected for days or weeks, allowing their effects to accumulate
Separate What Agents Can Do from What Requires Approval
The practical answer is not to “leave AI out of everything.” The key is to sort tasks by risk and clearly define the boundaries of an agent’s autonomy.
| Task type | Agent can act autonomously | Human approval required | |---|---|---| | Information gathering and summarization | Yes | Review only in exceptional cases | | Report drafts and meeting notes | Yes | Review before sharing externally | | Schedule suggestions and reminders | Yes | Approval for changes to important events | | Customer email drafts | Yes | Approval before sending or policy checks | | Payments, contracts, and refunds | Limited | Approval generally required | | Performance reviews and rejecting job candidates | Prohibited or heavily restricted | Human makes the final decision | | Production deployments and data deletion | Limited | Multi-step approval and rollback safeguards |
The principle is simple: automate what’s easy to undo, and make the agent stop before doing what’s hard to undo. Drafting, monitoring, anomaly detection, and proposing options play to an agent’s strengths. But when legal liability, financial loss, personal data, or reputational harm is at stake, humans must retain final control.
Technical Safeguards for Running Agents Safely
The tech industry is focused on more than the safety of the models themselves. The more agents connect to real-world tools and data, the more important it becomes to have a control layer outside the model.
The four most fundamental safeguards are:
The principle of least privilege
Agents should be granted only the permissions they need. For example, an agent that can view customer data should not automatically be able to delete or export it. The key is to separate read, write, approve, and execute permissions.Policy-based execution gates
Certain actions should never be carried out solely on the basis of an AI’s judgment. For tasks such as wire transfers, external emails, deployments, and data deletion, a policy engine should step in to verify the amount, recipient, timing, and conditions.Audit logs and traceability
Organizations should record what information an agent accessed, why it called a tool, and what results it produced. When something goes wrong, “The AI did it” cannot be the end of the story. Reproducible logs are essential for investigating the cause and adjusting permissions.A stop button and rollback plan
Agents need a kill switch that can stop work immediately. Organizations should also have version control, approval queues, deployment rollbacks, and change-history retention in place so completed actions can be reversed where possible.
Efforts to establish an external security layer, separate from the agent itself—as with Nvidia’s agent-safety platform—matter for precisely these reasons. It isn’t enough for AI to decide that it is safe. We need independent systems that can monitor and constrain AI’s actions.
A Good Goal Is an Operating Contract, Not Just an Instruction
When delegating a goal to an agent, it isn’t enough to say only what it should achieve. You also need to define:
- What counts as success?
- What actions are strictly off-limits?
- What are the limits on cost, time, and permissions?
- Under what circumstances should the agent report to a person?
- If goals conflict, which one takes priority?
For example, a better instruction than “Increase revenue” would be:
Propose weekly improvements to reduce churn among existing customers, using customer data only within the approved CRM scope. Issue discount coupons, send external messages, or change prices only after receiving approval from the person responsible. If you detect a privacy or legal issue, stop work immediately and report it.
A well-designed goal like this does not constrain an agent’s capabilities for their own sake. It provides an operating framework that helps the agent deliver results more reliably.
Autonomy Is a Matter of Responsibility, Not a Feature
The next tech race won’t be decided simply by who can automate the longest tasks. A more important competitive advantage will be how safely an organization delegates authority, how clearly it can trace responsibility, and how quickly it can bring an agent to a halt.
AI agents can become outstanding digital colleagues. But just as we define roles, permissions, reporting lines, and boundaries for human colleagues, we need to apply the same principles to agents. Expanding autonomy is not the same as losing control. In the age of agents, the winners won’t be the organizations that hand AI the most work, but those that delegate most wisely—and know exactly how to stop it.
Comments
Post a Comment