Avatar (Fabio Alessandro Locati|Fale)'s blog

On the autonomy ladder

September 30, 2026

I used to describe my position on AI and infrastructure with a simple rule: AI should never operate your infrastructure directly.

The rule was useful, but I do not think it is the right rule anymore.

Language models are not trustworthy enough to receive unrestricted access to production. That part has not changed.

What has changed is the architecture around them. Automation platforms are increasingly being positioned as a trusted execution layer, separating the probabilistic output of an AI system from approval and deterministic execution. That is a better design than letting a model issue arbitrary commands, but it leaves the important question unanswered.

How much autonomy should we give a machine for a particular operation?

The answer should not be a property of the model. It should depend on the operation, the controls around it, and the consequences of getting it wrong.

The autonomy ladder

I think about autonomy as six levels. They are not steps that every organization should climb until it reaches the top. They are a way to describe what a machine is allowed to do.

L0: Observe

The machine collects information about the system. It can read metrics, logs, configuration, events, and inventory, but it does not recommend or change anything.

This is where I would start when the data or the integrations are new, or when visibility is useful but taking action is too sensitive.

L1: Recommend

The machine proposes a diagnosis or an action. An operator decides whether the proposal makes sense and performs the action themselves.

This is still a useful level for unfamiliar failure modes. It also exposes problems in the observability behind the automation: a recommendation that cannot show which evidence led to it is not very useful.

L2: Prepare

The machine produces a remediation that is ready to run. An operator reviews and approves it before execution.

That is different from L1. The machine is no longer suggesting a general course of action; it is preparing the concrete change in the format the control plane expects.

For infrastructure, this could be a change set, a playbook invocation, or a pull request with the proposed configuration. The human is approving the remediation itself, not just agreeing with the diagnosis.

L3: Execute with approval

The machine can execute the operation after an authorized person approves it. The execution should go through a deterministic control plane, not through free-form commands produced by the model.

This is probably where many teams will want to start with meaningful production changes. The operation is understood and repeatable, but its scope or timing still deserves a human decision.

The approval should show the target, the proposed action, and the evidence behind it. An approval button without that context is just ceremony.

L4: Execute within policy

The machine does not need approval for every individual action. Policies define what it may do, where it may do it, and under which conditions.

A policy might allow a service restart on a non-critical host when health checks fail, while forbidding the same action during a deployment window or on a production database. The machine can act quickly because the organization made the decision in advance.

L4 is not L3 with the approval step hidden. It needs a tested policy, an enforcement point, a clear scope, and an audit trail. Without those, an unattended action is simply an unreviewed action.

L5: Autonomous remediation

The machine detects a problem, reasons about it, executes a remediation, and verifies the result without human intervention in the normal path.

This is the highest level of autonomy, but it is not automatically the best one. It only makes sense when the operation is bounded tightly enough that acting cannot create a worse incident than the one it is trying to fix.

An idempotent service restart could deserve L5 if the detection, action, and verification are reliable. A firewall policy change may stop at L3 because the blast radius is harder to bound. An irreversible data modification may require L2 even when the model is very capable.

The level belongs to the operation, not to the AI product.

Choosing the ceiling

When I assess an operation, I use five dimensions:

blast radius x reversibility x confidence x observability x business criticality

I do not mean this as a universal numerical score. It is a way to force the right questions into the conversation before someone asks whether an AI agent is ready for production.

Blast radius is about how much one wrong decision can affect. An action against one disposable worker is different from the same action against every cluster in a region.

Reversibility is about more than having a backup. If undoing the change does not restore the previous state, the operation is not meaningfully reversible.

Confidence should come from tested signals and bounded behavior, not from a model producing a convincing explanation.

Observability is what lets the system tell that the action worked, partly worked, or made things worse. If the machine cannot verify the result independently, I would not call the operation L5.

Business criticality is the cost of being wrong or being slow. A low-risk restart and a payment authorization may look similar from an infrastructure perspective, but they do not deserve the same autonomy.

High blast radius, poor reversibility, low confidence, weak observability, and high business criticality should push the ceiling down. Small blast radius, easy reversal, strong evidence, good verification, and low business criticality make higher autonomy easier to justify.

The dimensions do not cancel each other out. High confidence does not make an irreversible action safe, and good observability does not make a business-critical outage harmless.

The control plane still matters

This framework does not make the execution layer optional. It makes the boundary between reasoning and execution explicit.

An AI system can inspect evidence, identify a likely cause, and prepare a remediation. The control plane should still enforce identity, authorization, scope, policy, idempotence, validation, and auditability.

That separation is useful even when the reasoning component is not generative AI. The same ladder can describe a monitoring system that restarts services, a script that rotates credentials, or an operator approving a firewall change. It is an architecture framework before it is an AI framework.

It also gives architects better questions to ask about products. Where is policy enforced? Is execution deterministic? How is approval represented? What evidence is recorded? How does the system verify the result?

Whether a vendor calls the feature an agent is much less important.

Do not climb by default

The temptation will be to treat L5 as the destination. That would turn the ladder into another maturity model and create pressure to automate actions that should remain governed.

L4 is often better than L5 when the rules are clear but unusual circumstances still require judgment. L3 is often better than L4 when the action is safe but its timing or scope needs a human decision. L2 is often better than L3 when reviewing the exact proposed change matters more than approving an abstract intent.

The point is not to maximize autonomy. The point is to grant the highest level that the operation has earned, and no higher.

That gives architects a question they can use with Ansible, another automation platform, or no generative AI at all: what is the maximum autonomy this particular operation has earned?