A Practical Answer: Govern Actions, Not Just Models
Banks can adopt agentic AI without slowing innovation by matching each system’s authority to a specific, continuously enforced control framework. The key distinction is that an ordinary generative AI tool usually produces information, while an agent can interpret that information, call banking systems, select a payment method, and initiate a transaction. Consequently, reviewing the model and its prompt once a year is not enough. The bank must also control identities, data access, permissions, workflows, transactions, exceptions, and evidence throughout the agent’s operating life.
Also worth reading: How Can You Use AI for Safe Budgeting Without Sacrificing Financial Control? · What Is Agentic Payment Governance and How Should Financial Advisors Control AI Spending in 2026? · What Are Agentic Banking Controls and How Should Banks Implement Them in 2026?
A useful starting rule is that no agent receives unrestricted access merely because it performed well in testing. Read-only access should be the default. Permissions should then expand by use case: an account-analysis agent may read balances, a service agent may draft a payment, and only a carefully tested treasury agent may release funds. Controls should be expressed as enforceable limits—such as a maximum transaction amount, an approved beneficiary list, permitted payment rails, operating hours, or a required second approver—rather than vague instructions such as “be accurate.”
This approach does not mean suppressing experimentation. Banks can create a controlled path from research to production, allow sandbox experimentation, and progressively increase authority as evidence accumulates. Innovation comes from testing many use cases cheaply, learning which ones provide value, and retiring weak or risky ones early. The aim is not to make every agent behave like a human employee. It is to ensure that the bank always knows what the system can do, under what conditions it can do it, who is accountable, and how execution can be stopped.
Why Conventional AI Governance Is Not Enough
Traditional AI governance usually focuses on model validation, bias testing, data privacy, output quality, and periodic model review. Those controls remain necessary, but they do not address the operational risks created when software acts through a bank’s identity. A model may produce a plausible and policy-consistent instruction while relying on stale account data, misidentifying a beneficiary, selecting an inappropriate payment rail, or exploiting an overly broad application-programming interface. The resulting issue is no longer only an inaccurate answer; it may be a validly authenticated transaction initiated by flawed logic.
An agent also creates a chain of delegated authority. The user tells the agent what to do, the orchestration layer selects tools, the model chooses parameters, an API executes the instruction, and a payment or core-banking system commits the action. Each component can introduce failure even when the underlying language model behaves as expected. For example, the agent may intend to pay invoice 1048 but may submit invoice 1048 in the wrong currency. Prompt controls will not necessarily catch that type of error if the instruction and tool description both appear reasonable.
Banks therefore need continuous authorization rather than one-time approval. Before every consequential call, the platform should verify the user, the agent’s current mandate, the requested action, relevant data, and any required human approval. After the call, it should compare the executed action with the approved mandate. A transaction outside policy should be blocked, queued for review, or escalated according to a predefined rule. This “policy-as-code” approach turns written policy into an operational control that can be tested, logged, and updated across every agent deployment instead of being left to the discretion of a prompt designer.
A Layered Control Model for Banking Agents
Agentic controls should operate in layers so that a failure of one safeguard does not become an uncontrolled event. Identity and authorization come first because an agent should never act through a shared service account that obscures the initiating customer, employee, and software identity. Transaction controls come next, including amount limits, beneficiary restrictions, payment-rail rules, duplicate detection, and segregation of duties. Behavioral monitoring then identifies unusual sequences, such as an agent accessing 30 accounts immediately before proposing payments to a newly created beneficiary.
Human oversight should be designed around specific risk points rather than added as a ceremonial confirmation at the end. Humans cannot supervise every low-risk action indefinitely, and an approval dialog that merely asks someone to click “accept” may create responsibility without meaningful review. Approval is more valuable when it presents the intended action, source of funds, beneficiary, fees, relevant policy deviations, and evidence supporting the recommendation. For high-value payments, new beneficiaries, urgent instructions, or policy exceptions, the approver should receive enough information to make an independent decision.
| Control Layer | Main Question | Example Banking Control | Expected Evidence |
|---|---|---|---|
| Identity | Who or what initiated the action? | Separate user, agent, workload, and service identities | Authentication record and signed delegation |
| Authorization | Is the agent permitted to do this now? | Treasury agent limited to approved accounts and rails | Policy-decision log |
| Transaction | Is this particular action acceptable? | $25,000 per-payment cap and verified beneficiary | Approval record and limit result |
| Monitoring | Is behavior normal and internally consistent? | Alert on rapid account switching or new payees | Event timeline and anomaly record |
| Human Oversight | Where is human judgment required? | Dual approval for payments above $100,000 | Named approver, timestamp, and rationale |
| Recovery | What happens when the system fails? | Revoke tokens, halt agent, and freeze affected workflow | Incident report and remediation evidence |
Designing Permissions That Enable Speed Rather Than Block It
The most common governance mistake is treating every agent use case as either fully autonomous or entirely manual. That binary approach creates unnecessary bottlenecks for simple tasks and exposes the bank to unacceptable risk for consequential ones. A better method classifies agents by the authority they exercise. A reporting agent may summarize cash positions; a recommendation agent may propose investments; a workflow agent may prepare beneficiary records; and an execution agent may initiate payments. Each level deserves different authentication, approval, monitoring, and recovery standards.
Permissions should be narrow, temporary where possible, and purpose-specific. Instead of giving an invoice-processing agent broad access to a corporate banking platform, the bank can expose a tool that accepts only an approved invoice identifier, verified beneficiary, permitted currency, and amount within a defined range. The tool—not the model—should enforce those constraints. This reduces both malicious prompt injection and ordinary reasoning error because the model has fewer ways to request an unintended action.
Temporary access is especially useful during controlled pilots. A bank can authorize an agent to operate against synthetic data for 30 days, then against read-only production data for another 30 days, and finally permit a small number of low-value transactions for a 60-day limited release. During each stage, the bank can establish whether users need the functionality, whether exceptions are manageable, and whether the agent improves cycle time or processing accuracy. Authority should expand only when measured performance and control effectiveness justify it.
The 30-, 60-, and 90-day periods are illustrative rather than regulatory requirements. Banks should adapt them to the use case and risk profile. Their value is that they make deployment evidence-based and reversible. A pilot that performs well should not receive unlimited production authority automatically, just as a weak pilot should not be allowed to continue merely because senior leaders have announced an agentic AI strategy.
Human Approval Must Be Meaningful and Proportionate
Human involvement is often presented as the universal answer to agentic risk, but approval at every step can make the technology slow, expensive, and less useful. If employees must approve routine reconciliations that the agent performs accurately, the bank has introduced automation in name only. Conversely, allowing payments, customer communications, credit decisions, or account closures without review may create unacceptable exposure.
Approval thresholds should reflect potential harm, reversibility, novelty, and uncertainty—not just transaction value. A $5 million payment to an established supplier may follow well-tested rules, while a $2,000 transfer to a new beneficiary may deserve additional verification. Similarly, an agent recommending a readily reversible balance transfer differs from one closing a customer’s account or changing a credit limit. Banks should include data confidence, policy deviations, unusual timing, account takeover indicators, and model uncertainty in the routing decision.
Approvers also need time and authority to challenge the recommendation. Showing a green status indicator beside a vague summary encourages rubber-stamping. A better interface might display the exact payment instruction, a plain-language explanation of why the agent proposed it, the beneficiary verification state, the expected account balance after settlement, any deviation from treasury policy, and a clear route for correcting the instruction. The system should preserve the approver’s identity, the version of the information displayed, and the final decision.
Segregation of duties remains important. The employee who configures an agent should not also be the sole approver for every payment that agent produces. Developers should not silently raise transaction limits in production, and business owners should not override model or risk alerts without recording a reason. This does not require inflexible two-person control for every low-value action. It requires that sensitive authority cannot be concentrated in one person and exercised invisibly.
Continuous Monitoring, Evidence, and Emergency Shutdown
Banks need evidence not only to investigate misconduct but also to demonstrate that controls operated as intended. For each consequential agent action, the record should identify the initiating user, model and prompt version, retrieved data sources, tool selected, authorization decision, approval, transaction response, and subsequent outcome. Logs should be tamper-evident, access-controlled, time-synchronized, and retained according to the bank’s legal and regulatory obligations. They should be capable of reconstructing what the system knew at the moment it acted.
Real-time monitoring should distinguish operational exceptions from cybersecurity threats. A failed API call, stale exchange rate, duplicate invoice, and incorrect payment rail may require a different response from an attempt to manipulate the agent through prompt injection. Banks should monitor tool calls, privilege changes, data access patterns, beneficiary creation, transaction frequency, unusual beneficiaries, repeated approvals, and differences between the user’s instruction and the executed action.
A tested shutdown capability is essential. Banks should be able to revoke an agent’s credentials, disable a particular tool, block a payment path, pause an orchestration workflow, and notify responsible teams without waiting for a model provider or software vendor. A kill switch that exists only in documentation is not an operational control. It should be exercised through scheduled simulations at least as often as other critical continuity controls, with results recorded and defects remediated.
Recovery plans should also address actions already taken. Stopping an agent prevents further activity but does not reverse a completed payment, disclosure, or account change. Procedures should identify which transactions can be recalled, which require customer notification, how access is preserved for investigation, and when legal, compliance, cybersecurity, and public-relations teams must become involved. The response should not automatically restore service simply because the immediate defect has been patched.
Common Mistakes That Turn Innovation Into an Incident
A first mistake is allowing the model to bypass specialist systems by giving it direct access to powerful banking interfaces. It is faster to build this way, but the model becomes responsible for enforcing limits that are better embedded in software. Another common error is treating the agent’s identity as a convenience. Shared credentials and broad API keys make attribution difficult and increase the impact of token theft, configuration errors, and misleading instructions.
Banks also err by measuring only model accuracy. Useful evaluation includes task completion, false-action rate, exception rate, approval burden, cycle-time savings, customer outcomes, control overrides, and losses from near misses. An agent with 99.5% accuracy across 20,000 proposed actions would still produce about 100 questionable actions, so production volume matters. Accuracy must be assessed by action type and amount, not reported as one impressive aggregate percentage.
Another failure is treating compliance approval as deployment approval. Compliance may confirm that a use case is permissible, while technology, cybersecurity, operations, risk, and the business owner must still determine whether identity, integration, monitoring, staffing, and recovery work. The opposite mistake is giving every new agent a full year-long validation process regardless of risk. That encourages teams to hide low-value pilots in informal tools or launch them before scrutiny.
Finally, banks should not set a universal percentage for agents that require human approval. A fixed target such as “80% autonomous” can reward unsafe behavior by rewarding removal of controls. Autonomy should be earned per action and can be withdrawn when circumstances change. Published commentary from EY, Bain, Boston Consulting Group, J.P. Morgan, and banking leaders between 2024 and 2025 broadly supports strong controls, authentication, and governance, but operational design remains a bank-specific responsibility. Industry attention is not evidence that the bank has implemented an adequate control environment.
When a Bank Should Pause, Limit, or Stop an Agent
Banks should pause deployment when an agent cannot reliably identify its initiator, when its access exceeds the stated purpose, or when developers cannot explain why a consequential action was taken. They should limit deployment when performance evidence is incomplete but the use case has clear value and can be confined to read-only or recommendation-only work. A useful interim measure is to narrow the account scope, remove execution tools, cap values, or require dual approval until the evidence improves.
Near misses should trigger review even when no customer loss occurs. Suppose an agent proposes a payment to a beneficiary whose name differs by one character from an approved supplier. If the beneficiary control catches it, the event is reassuring, but it also reveals a predictable failure mode. The bank should determine whether the agent’s data preparation, matching logic, and interface can be improved rather than treating the saved transaction as proof that the system is safe.
Material changes also require renewed review. Replacing the underlying model, connecting a new payment provider, giving access to a new account system, altering retrieval logic, or changing a transaction limit can change risk even if the model version remains the same. Banks should conduct targeted validation before the change and use a staged release where feasible. Emergency patches deserve urgency, but not the abandonment of logging, approval, and post-incident testing.
Stopping an agent entirely is appropriate when its business benefit cannot be demonstrated, when its error pattern is not controllable, when recovery repeatedly fails, or when regulatory and customer protections cannot be maintained. An agent that only produces 4% time savings after six months but requires constant intervention may not merit continued production expense. Conversely, an agent that shortens cash forecasting from two hours to 20 minutes while improving data accuracy may justify investment even if it never initiates a payment. The relevant question is not whether the technology is autonomous; it is whether the bank has made the delegated authority safe, measurable, and economically worthwhile.
A Governance Path That Rewards Safer Experimentation
Banks that slow agentic innovation usually make every deployment depend on the same heavyweight approval process. More effective organizations create tiers that match oversight to authority. Research environments use synthetic or anonymized data and cannot affect customers. Advisory environments can use live read-only data but cannot execute transactions. Workflow environments may create drafts or cases. Execution environments are isolated, tightly restricted, and monitored, with authority granted only after control testing.
Each tier should have entry criteria, named owners, performance thresholds, expiry dates, and automatic downgrade rules. For example, an execution agent might begin with a $1,000 daily cap for one corporate account. Expansion to $25,000 might require 60 days without a material control breach, at least 99.9% successful reconciliation of approved actions, and verification that every emergency shutdown test succeeded. No expansion should occur simply because calendar time has passed; the bank must examine what happened during the period.
This model lets the bank learn at different speeds. Low-risk experiments need limited technical review, while high-impact systems receive deeper analysis. Teams can also differentiate between shadow mode, in which the agent observes and recommends without acting, and production mode. Shadow deployment allows banks to compare agent decisions with employee decisions before granting authority. Later, the agent can prepare actions for human execution before finally receiving narrow execution rights.
For cash and treasury operations, this staged progression can produce an attractive balance between control and speed. Forecasting, reconciliation, liquidity alerts, payment preparation, and routine supplier payments can be introduced separately rather than bundled into one broad “agentic treasury” program. Each capability can prove its value and controls before the next is connected. The bank then treats faster deployment as the result of clear boundaries and reusable infrastructure, not as permission to bypass accountability.