An agent detects a disabled backup job and proposes re-enabling it. The API permission exists. The action is technically straightforward. But the job may have been paused for an approved migration, or its destination may no longer be suitable.
The question is whether this action is authorized for this resource, under the conditions that exist now. A broad instruction to “fix backup drift” cannot carry all of those decisions safely or make them reviewable afterward.
Before enabling write access, define the agent’s operating boundaries in terms that the execution system can enforce. Specify which resources and actions are permitted, the facts that must be true, when a person must approve, and the evidence required before the finding can close.
This article proposes a practical method for doing that. It applies to AI agents and conventional automation alike, because either can execute a technically valid change with an incomplete understanding of its consequences.
Define an action contract before granting write access
An action contract is a documented, enforceable agreement about what an automation may do. The term here is a proposed operating practice, not an industry certification or a named Forttic feature.
The contract should be specific enough for an engineer to evaluate without reading the agent’s full reasoning transcript. It identifies the triggering condition, target resources, allowed change, prerequisites, approval boundary, verification method, and failure route. It also identifies who owns the workflow and when its authority must be reviewed.
Forttic’s article on AI agents and resilience enforcement explains why connected tools alone do not provide a complete resilience operating system. An action contract makes that distinction practical: tool access becomes a narrowly defined capability inside a governed process.
Keep enforcement outside the model’s discretion. Scoped credentials, execution-time policy checks, and approval controls should constrain what the tool can do. Instructions in a prompt can communicate policy, but the workflow should not depend entirely on the model choosing to obey them.
Classify authority by the consequences of the change
A useful starting point is to distinguish read-only investigation, execution under preapproved conditions, and changes requiring explicit approval. These are proposed categories for review, not universal defaults. A seemingly routine action can have material consequences in a particular environment.
| Proposed authority | Example candidate | Conditions to establish |
|---|---|---|
| Read-only investigation | Retrieve job status and inspect configuration | Limited data access, approved scope, traceable observations |
| Preapproved execution | Re-enable one job that deviated from its approved state | No active pause or exception, unchanged destination, bounded load, defined verification |
| Explicit approval | Change retention, destination permissions, or protection scope | Documented impact, named decision-maker, exact resource and parameter review |
| Separate high-impact procedure | Delete recovery points or lock an irreversible setting | Strong authorization, verified prerequisites, explicit acknowledgement of irreversibility |
The difference between enabling and locking a setting is especially important. Microsoft documents that Azure Backup vault immutability can be enabled reversibly, while locking it makes the setting irreversible. That distinction belongs in the decision process before execution.
Do not let a desired protection outcome become permission for every possible method of achieving it. An agent authorized to restore an approved configuration should not infer authority to change the policy itself or delete recovery points to satisfy another check.
Recheck the conditions immediately before acting
The observation that triggered a workflow may already be stale. Another engineer may have corrected the issue, a maintenance window may have begun, or a legitimate exception may have been approved.
For the disabled-job example, read the current job state, policy version, destination, and applicable exception record before execution. Confirm that the resource still belongs to the approved scope. If the relevant facts differ from those used to authorize the action, stop and reassess.
Where supported, use version or conditional-update controls to reduce the risk of acting on a configuration that changed between inspection and execution. If the platform lacks that capability, document the concurrency limitation and choose a narrower scope or stronger coordination mechanism.
Untrusted text should not grant new authority. A resource description, log entry, or retrieved document may inform an investigation; it should not expand credentials, substitute a destination, or remove an approval step. The execution path should validate identifiers and permitted parameters against the contract.
Make approval specific and time-bounded
An approver should see the identified resource, observed gap, proposed action, expected consequences, and verification plan. A request that merely asks whether to “remediate the issue” transfers too much interpretation to the workflow.
Record which version of the plan was approved and the conditions under which that approval expires. If the target, parameters, or relevant state changes, obtain a new decision. An approval given before a migration should not silently authorize a different change after it.
Existing automation systems provide mechanisms for this separation. AWS Systems Manager’s aws:approve action can pause a runbook for designated principals to approve or reject. Its documented limitations include lack of support for multi-account and Region automations; check the fit before choosing it for a cross-environment workflow.
The process also needs a decision when approval does not arrive. Keep the protection finding visible, preserve ownership, and escalate through the agreed route. Silence should not become authorization through an unrecorded timeout behavior.
Set bounds on repetition and scope
One successful change does not establish that the same action should run across every account. A wrong assumption becomes more consequential when applied repeatedly.
Define per-run resource limits, concurrency, retry behavior, and conditions that stop further execution. Use vendor idempotency mechanisms where available, and check whether a timed-out request actually completed before submitting it again. Repeating an operation because its response was lost can create duplicate work or conflicting state.
Begin with a small, representative scope. Inspect the result before expanding. If related workflows can touch the same resource, agree how they coordinate so one process does not undo another’s approved decision.
A stop control should prevent new actions promptly. It may not cancel an operation already accepted by a vendor, so its limitations should be explicit. The operator needs a record of in-flight work and a way to reconcile its eventual outcome.
Verify protection after execution
Separate the state of the automation from the state of the protection requirement. “The request completed” belongs in the execution record. “The required backup exists in the approved destination” is a different claim that needs its own observation.
For a re-enabled job, checking that it is enabled verifies the configuration change. If the underlying finding concerns a missing current recovery point, closure requires the qualifying recovery point and any additional checks specified by that control. Do not silently change the closure condition to match the easiest available evidence.
Use a verification deadline appropriate to the operation. Some results will follow a later scheduled job. Keep the finding in an explicit waiting state while that result is pending. If the deadline passes or verification fails, retain the open finding and route the next decision to its owner.
The verification component should evaluate observable outcomes rather than simply repeat the agent’s success statement. Link the underlying result to the affected resource and the original requirement. Forttic’s verified-closure guide provides the broader operating context for this distinction.
Plan for partial failure before the first run
Consider a workflow that updates a copy configuration but fails to establish destination access. The first step succeeded; the protection outcome did not. A single failed status cannot explain whether a rollback is appropriate or which state now exists.
Retain a result for each consequential step and inspect the current state before taking further action. Where rollback is possible and authorized, document the conditions for using it. Where it is unsafe or impossible, stop and escalate with the partial result intact.
Avoid treating rollback as a universal reset. Restoring a retention value cannot recreate recovery points already deleted. Revoking a permission may disrupt work started by another approved process. Recovery from automation failure needs the same attention to dependencies and consequences as the original action.
If policy permits temporary risk acceptance, record it as a separate decision with an owner and review date. It should remain distinguishable from a corrected and verified control.
Use one small workflow to establish the operating pattern
For an initial pilot, choose one recurring issue whose cause and permitted correction are understood. The disabled-job scenario is a candidate only after the team has established why it occurs and when re-enabling is appropriate.
Document the resource scope, approved baseline, exception lookup, live-state check, execution limit, expected result, and escalation route. Exercise the workflow in a suitable test environment, including stale findings, denied approvals, API timeouts, and unsuccessful verification. The purpose is to learn how it behaves when the preferred path is unavailable.
Track attempts separately from verified corrections. Review out-of-scope requests rejected, actions stopped because conditions changed, and findings still awaiting proof. These measures can reveal whether the boundaries work; a high volume of automated actions alone cannot.
Forttic’s agentic enforcement architecture describes policy knowledge, operational context, triggers, and execution capabilities supporting resilience decisions. Its CRE model connects those decisions to governed remediation, verification, and evidence. The action-contract method in this article offers a practical way to evaluate that operating discipline without assuming every proposed field is a prebuilt product control.
Before expanding an agent’s permissions, ask its owner to explain one completed decision from trigger to verified outcome, including what would have stopped it. If that explanation is clear and supported by evidence, the team has a concrete basis for reviewing the next scope of authority.
Take the Forttic CRE assessment to identify where your current process needs stronger links between authorized action, verification, and evidence.
Bound the action. Verify the outcome.
Write access is not a general instruction to fix backup drift. Give the workflow a specific contract, recheck live conditions, and close the finding only when the protection requirement is observed again.
Frequently asked questions
What are guardrails for AI backup automation?
Guardrails are enforceable limits on which resources an automation can access, what it can change, the conditions required before acting, when approval is necessary, and how failures are handled. They should constrain execution as well as guide the agent’s reasoning.
Which backup tasks can run without individual human approval?
Tasks may qualify when the organization has preapproved a specific action, resource scope, prerequisites, and verification method. Suitability depends on impact, reversibility, uncertainty, and current conditions. There is no universally safe list based on task names alone.
Is a prompt enough to govern an agent’s permissions?
No. A prompt can describe policy, but execution should also be constrained by scoped credentials, parameter validation, policy checks, and approval controls. Retrieved text should not be able to expand the agent’s authority.
What should happen when automated verification fails?
Keep the protection finding open, preserve the executed steps and observed state, and follow the defined failure or escalation route. Retry or roll back only when the contract permits it and the current conditions support that decision.
Can every backup change be rolled back?
No. Some operations are irreversible, and reversing a configuration does not necessarily reverse its consequences. Identify those limits before execution and require the appropriate authorization for high-impact decisions.
How should teams evaluate a backup automation pilot?
Evaluate verified corrections, unresolved outcomes, boundary enforcement, approval behavior, retries, and failure handling within a defined scope. Review the evidence for individual decisions before expanding authority or applying the workflow more widely.