Editorial note: Original analysis of the primary source below. No sponsor paid for or reviewed this story.

Atlassian’s September account of governed agent loops argues that software agents need shared context, controls over what they can touch and a way to measure results. Its availability notes matter: some features are in open beta, while agent loops, standards and AI review are described as private early access in the announcement. This is a roadmap and product position, not evidence that every team can deploy the whole system today. The useful design question is how a team keeps continuous agent work reviewable without reducing it to a stream of unattended pull requests.

Context is an operational dependency

A coding agent can read a repository and still miss architectural intent. Why does one service own a piece of data? Which migration pattern is prohibited? What incident made a seemingly odd guard necessary? Those answers may live in issue discussions and experienced engineers’ memory rather than code. A shared context layer can help, but only if it points to decisions and their owners rather than dumping every document into a prompt.

Treat context as versioned product infrastructure. Give a standard a scope, date, rationale and example of correct application. Provide a route to challenge a stale rule. When an agent follows a constraint, the reviewer should be able to see which source it used. If the source is wrong, fixing the standard should improve future work more reliably than changing one prompt.

Autonomy requires a map of authority

Agent loops imply repeated planning and execution. That can accelerate maintenance tasks, but the team should decide which steps are automatic, which create a draft and which need a person to approve. The permission boundary should be tied to effect: editing a test file, changing a production configuration and merging a deployment carry different consequences. One ‘agent access’ switch is too broad to explain the risk.

A workable policy distinguishes read, write, propose, merge and deploy. It names the environments and data each level can touch. Reviewers need a compact summary of actions taken since the last checkpoint and the evidence that tests actually exercised the changed behavior. Without that, continuous loops merely move the burden from creating code to understanding surprising code.

Measure delivered quality

Atlassian describes measurement across throughput, quality, adoption and cost. The four dimensions are useful precisely because any one can be misleading. More commits may mean faster delivery or more rework. High adoption may reflect novelty rather than value. Lower per-task model cost may be offset by expensive reviews. Teams should measure accepted work, post-release defects and human intervention against a baseline.

Choose a narrow class of tasks, such as dependency updates or small UI fixes, and create a review rubric. Track elapsed time from issue to accepted change, number of review cycles, test failures, regressions and operational cost. Sample rejected work too. If the agent avoids hard cases, an aggregate success rate will flatter it. A credible metric describes task mix and the acceptance criteria.

Governance cannot be a dashboard alone

Dashboards help leaders see patterns, but an engineer needs actionable feedback inside the workflow. Show when an agent used an outdated standard, which tool was called and which approval was bypassed or requested. Make the trace concise enough to review before merging. If every detail is in a separate analytics product, teams may learn about failure only after release.

Governance also needs an escape hatch. A rule should be overridable with a reason and a reviewer where appropriate, otherwise people will work around it. An override becomes evidence that the standard may need revision. This is a two-way relationship: policies shape agents, and observed agent failures expose weak policies.

The rollout question

Before buying or building a continuous agent loop, ask what happens when the context is stale, the agent is blocked or an evaluation fails halfway. Does the system stop safely? Can a human take over with a clear state? Does the next run know which action has already occurred? These are operational questions, not visual polish. Reliability depends on explicit transitions and recovery.

Atlassian’s announcement marks an industry shift from individual prompt demos to managed execution. The idea is sound, but teams should distinguish the currently available components from early-access promises. The strongest proof will be repeatable, accepted work under visible controls, not the mere presence of agents in a Jira dashboard.

Questions to take into your next review

  • What context source explains each material code decision?
  • Which agent actions require a reviewer before production?
  • How are rejected changes and regressions counted?

Primary sources and further reading

Continue reading