Editorial note: Original analysis of the primary source below. No sponsor paid for or reviewed this story.

Microsoft’s September 23 description of AI at work distinguishes a predictable user subscription from usage billing for agent work that can vary greatly in duration and complexity. This is a vendor’s business-model account, not a neutral finding that every customer should adopt its pricing. It raises a practical design question: can the person commissioning a job see what they have access to, what work the agent may perform, how the cost is metered and when to stop it? A clean chat box is insufficient when the unit of work is no longer one conversation.

Seats and work are different objects

A seat tells you who can use a product. A usage meter tells you what execution has consumed. When the two appear in one vague ‘AI included’ label, teams struggle to predict spend. A user may trigger a short summary or a multi-hour research process from the same surface. An administrator needs a budget and a worker needs a meaningful estimate before a job starts. Those are separate user experiences, even if the billing system combines them later.

The design response is to model a job explicitly. Show its owner, purpose, estimated range, current state, connected systems and cancellation path. Show the difference between a permission to start jobs and a quota for work already underway. Users should not have to inspect a financial dashboard to discover that a supposedly passive assistant has been consuming resources overnight.

Turn cost visibility into interaction design

Cost feedback should arrive at the decision point. Before a large run, explain what the agent will do and what could make the work longer. During execution, display progress in units that help a person decide whether to continue. Afterward, attach the spend to an outcome: a report produced, records changed or a task completed. A raw token counter rarely tells a business owner whether the work was worthwhile.

A sensible product can offer stop conditions: maximum spend, maximum duration, required approval before external actions, or an escalation when the plan changes. The exact thresholds depend on the organization; the invariant is agency. If limits exist only in an admin console, the individual starting the work may still misunderstand what is happening. Good cost UX gives administrators policy and workers timely context.

Evaluation has its own ledger

Microsoft also describes continuous improvement loops for agents. That claim invites a separate measurement question: which tasks improved, against what baseline, and at what total cost? Teams should record task success and human correction time alongside bills. A cheap agent that sends incomplete work downstream can be more costly than a more expensive one that resolves the task reliably.

Construct an evaluation ledger with a small set of representative jobs. For each, note expected output, risk level, allowed tools, human acceptance criteria and actual spend. Review failures as product issues, not only model issues. Sometimes the prompt is vague; sometimes the underlying data is stale; sometimes the interface implied the job had authority it did not possess. A single model benchmark will not uncover those differences.

A product team’s decision

When considering a platform, ask whether price controls apply before or after a job starts, whether active work can be stopped, and whether administrators can attribute use to meaningful business outcomes. Look for a clear distinction between included access and metered execution. Ask whether credits expire, whether the meter is observable, and whether people can inspect the actions for which they were charged. These are product questions as much as procurement questions.

The broader shift is from selling AI as an always-available conversation to selling measurable work. That changes onboarding, permission design, usage analytics and trust. Microsoft’s post is one articulation of the shift; another vendor might meter differently. The principle holds across models: users need to understand both what they can ask the system to do and what it costs when the system acts.

What to watch next

Do not take a headline about agent productivity as proof of return on investment. Look for published definitions of success, a credible comparison and the operational costs of supervision. A product that exposes budgets but hides failure patterns is still difficult to manage. A product that exposes beautiful job histories but cannot set spend ceilings is difficult to trust.

The most useful next experiment is a bounded workflow in your organization. Set a budget, define one outcome and compare agent work with the current human process. Document the handoffs, corrections and customer effects. Then decide whether the seat, the usage meter or both belong in the experience. The answer should follow the work, not a platform’s preferred pricing narrative.

Questions to take into your next review

  • Can a worker see an estimate before a long-running task?
  • Can an administrator stop or cap work in progress?
  • Is measured value tied to accepted outcomes rather than activity?

Primary sources and further reading

Continue reading