How to choose an AI agent vendor for your bank

Choosing the wrong AI agent vendor costs months, not just budget. With a generalist, taking a banking use case to production takes 9 to 18 months; with a specialist that already brings banking skills and integrations, the realistic range is 4 to 8 weeks. This checklist covers the compliance, experience and verifiable-timeline criteria a large bank should demand before signing.

Every technology team at a large Latin American bank is having the same conversation at some point in 2026: we are taking conversational AI to production, but with whom? The choice matters. Picking the wrong AI agent vendor costs more than budget. It can mean months lost, a project frozen by Risk or Compliance halfway through, or one more chatbot that resolves nothing. At Delto we support banks including Ficohsa (Honduras), Banreservas (Dominican Republic) and Banco Patagonia (Argentina) through that process, and we see the same evaluation mistakes repeat. This guide gathers the criteria a large bank, with high volume, regulatory exposure and a brand to protect, should demand before signing with any vendor. Are you comparing a chatbot or an AI agent? Many evaluations start badly because they compare different things under the same name. A traditional chatbot answers inside predefined decision trees. An AI agent reasons over customer context, bank policy and current regulation, and executes real operations in the core: it collects, sells, resolves, transacts. Dimension Traditional chatbot AI agent Logic Rigid decision tree Reasoning over context and policy Scope Answers frequent questions Executes operations in the core banking system Maintenance Manual rules for every new case Reusable skills trained by domain Continuity Loses context across channels Keeps the thread across WhatsApp, web and voice If the vendor cannot explain which channel it executes real actions in, rather than just informing, it is probably selling a chatbot under a new name. It is the same reason so many "smart chatbots" still frustrate banking customers . The criteria that separate a banking vendor from a generalist Compliance by design, not compliance bolted on top. Ask the vendor to show, not describe: masking of personal data before it reaches the model, immutable traceability with at least 7 years of retention, per-country data residency, and authentication controls (MFA, SSO, granular RBAC). If this gets solved "later", the project was born with risk. Worth remembering that traceability and data quality are not a new demand created by AI: they have been supervised territory since the Basel Committee published its principles on risk data aggregation and reporting . An agent acting on the core falls squarely inside that perimeter. Real banking experience, not generic experience. A vendor coming from e-commerce or retail will have to learn, on your bank's budget, what a card dispute, a delinquency reason or a KYC policy actually means. Ask how many banks it serves today, in how many countries, and at what end-user volume. A proven skill library, not a build from scratch. Every banking flow has rules and exceptions already solved at some other bank: resolving a dispute, offering a product, running early-stage collections. A vendor with hundreds of audited banking skills already built saves you months. One starting from zero bills you for them. Verifiable production timelines. With general vendors, taking a banking use case to production takes 9 to 18 months. With a specialist that brings skills and integrations already built, the realistic range is 4 to 8 weeks. Ask for banking references with a real go-live date, not a contract signature date. And pay attention to what happens after launch: that is where most projects stall between pilot and production . Continuity across channels. The customer opens the dispute on WhatsApp and finishes it by voice without repeating anything. That requires the agent, not the channel, to be the unit of design. If every channel has its own disconnected bot, there is no real conversational banking: there are several chatbots sharing a logo. Frictionless human handoff. Sometimes AI agents will not resolve a case, and in banking that is fine. What is not fine is the customer losing context when they move to a human. Require the handoff to transfer the full conversation, not an empty ticket. Portability and a clean exit. Asking what happens the day the bank decides to leave is not distrust, it is diligence. Conversational flows, interaction history and configured business rules are the bank's assets. If the vendor cannot explain how they are exported, the cost of switching later will far exceed any price difference today. What should your bank ask any vendor? What happens to my customers' data before it reaches the language model? How many banking skills are production-ready, and how many have to be built from scratch? What is the real go-live date of your last banking client, verifiably? How are the actions the agent executes in the core audited? What happens to the bank's data and flows if we decide to switch vendors? Does the same agent keep context when the customer changes channel? A useful test: ask the vendor to answer these six questions in writing, as an annex to the proposal. What gets committed on paper is what actually gets delivered, and it gives Risk and Compliance something concrete to review before the contract stage. If the vendor answers with generalities where it should answer with numbers, that is a bad sign. A specialist has those figures at hand because it measures them every month. Four red flags in an evaluation process Beyond the checklist, some signals reliably predict a project in trouble. They show up early, almost always during the demo stage: The demo always runs on fictional data. A vendor with banking experience can show the agent operating against a sandbox that mirrors a real core data structure, even anonymized. It cannot explain what happens to data at the model provider. If the answer is "we use the best available model" and does not cover what gets sent, what gets masked and what gets retained, the part Risk will ask about is missing. The proposal does not separate what already exists from what must be built. When everything "is configurable", the timeline is an optimistic estimate dressed up as a catalog. The team that sells is not the team that implements. Ask to meet whoever will be in the sprints. The gap between a polished demo and a sustained go-live usually lives in that team. None of these disqualifies a vendor on its own. Two or more together usually explain why a pilot never reaches production. What makes Delto different in this process? Delto was built for banking. It is not generic AI with compliance added later. The suite runs on BLAM, a Banking Large Action Model with more than 340 proven banking skills, and includes five specialized agents across WhatsApp, web and voice: Collections , Advisor , Customer Support , Retail Banking and Corporate Banking . The first agent can be in production in 4 to 8 weeks, with data masking, 7-year audit retention and per-country data residency solved from day one. Today we serve more than 3 million end users across more than 15 countries in LATAM and the Caribbean. If you want the detail behind that compliance posture, it is on the security page , and if you are running a broader generative AI evaluation, we wrote earlier about how to choose a company to implement generative AI at your bank . Choosing an AI vendor is not a technology decision. It is a decision about risk, brand and time. If your bank is in that evaluation process, we can show you how other banks in the region solved it, with names, dates and numbers. Let's talk .

What is the difference between a banking chatbot and an AI agent? A chatbot answers within predefined rules and cannot step outside its script. An AI agent reasons over customer context, bank policy and regulation, and executes real operations in the core: it collects, sells, resolves and transacts, rather than just informing.

How long does it take a bank to get its first AI agent into production? With a banking specialist that arrives with skills and integrations already built, 4 to 8 weeks for a scoped use case. With a generalist starting from scratch, the usual range is 9 to 18 months. The difference is not the language model, it is how much has to be built before the first real conversation.

What must an AI vendor guarantee to be safe inside a regulated bank? Masking of personal data before it reaches the model, immutable traceability with at least 7 years of retention, per-country data residency, and access controls such as MFA, SSO and granular RBAC. These must be built into the product, not bolted on mid-project.

What should a bank ask before signing with a conversational AI vendor? What happens to customer data before it reaches the model, how many banking skills are production-ready, the real verifiable go-live date at another bank, how the agent's actions in the core are audited, and what happens to the data and flows if the bank switches vendors.

Is a generalist AI vendor or a banking specialist the better choice? A generalist coming from e-commerce or retail will learn what a card dispute, a delinquency reason or a KYC policy means on the bank's budget. A specialist has already solved those flows at other institutions and arrives with the exceptions mapped. In banking, where the cost of an error is regulatory and reputational, the bank pays for the vendor's learning curve.