How a Conversational AI Is Trained: A Guide for Banks Adopting Advanced LLMs

Training an LLM for banking has 3 stages: massive pretraining, domain fine-tuning, and reinforcement learning from human feedback (RLHF). Reviewers apply 8 criteria (accuracy, safety, clarity and more). Delto brings this to production with BLAM: 300+ ready-to-use banking Skills for LATAM.

Generative artificial intelligence is reshaping many industries, banking in Latin America among them. Smart chatbots, personalized recommendation engines, and virtual assistants that understand and respond in natural language are no longer the future: they are the present. But there is one question many leaders still don't ask, and it is key to making informed decisions: how are these language models (LLMs) trained, and what impact does that have on my business? At Delto, where we help banks adopt generative AI safely and strategically, we believe understanding the training process is essential. Because if your bank is going to trust an AI with decisions, recommendations, or answers, you'd better know how it was educated. Key stages of LLM training Massive pretraining To start, the model is trained on large volumes of public text: books, academic papers, web pages, forums, and more. The goal is to learn to predict the next word in a sentence, capturing grammatical, semantic, and stylistic patterns. This lets the model generate fluent, coherent text, but it does not guarantee accuracy, safety, or usefulness in critical contexts like banking. Specialization through fine-tuning At this stage, the model is adjusted with data more relevant to the specific domain: real customer service interactions, regulatory documents, internal manuals. This way, it begins to better understand financial jargon, the needs of banking customers, and the right tone. This process also makes it possible to comply with specific regulatory frameworks and adapt the model to the cultural and legal context of Latin America. Reinforcement learning from human feedback (RLHF) This is the most critical stage for building trustworthy models. Here, people evaluate two model responses to the same prompt (A and B), choosing which is better and explaining why. That feedback is used to train the model to prefer more useful, safe, and clear answers. Human reviewers apply a carefully designed set of criteria, many of which are now considered an industry standard. We share them below in more direct language: Key criteria for evaluating AI responses These criteria are essential for any language model that will be used by banks, financial institutions, or sensitive environments. Truthfulness and informational accuracy: Is the answer correct, verifiable, and free of hallucinations (content the LLM invents because it lacks the information to answer)? A good model does not invent laws, rates, or banking procedures. ‍ Safety and responsibility: Is the content appropriate, free of risk, and does it avoid suggesting dangerous, illegal, or discriminatory actions? ‍ Clarity and style: Is the answer easy to read, with well-built sentences and no grammatical errors? A model may have good information, but if it doesn't communicate well, it's useless. ‍ Concision and relevance: Does the answer get to the point or run on unnecessarily? In finance, every second and every word counts. ‍ Understanding the request: Did the model answer exactly what was asked? This matters most when the customer wants something very specific: changing a PIN, checking rates, or understanding a charge. ‍ Logical flow: Are the ideas ordered and connected? ‍ Context adaptation: Does the answer account for the region, the channel, or the end user? ‍ Level of formality: Does the tone match the bank's institutional image? These criteria are used both to train and to audit models running in production. Why this matters for banks in LATAM Banks that adopt AI without understanding how it was trained take on risk. They may end up offering incorrect information, undermining customer trust. They could also breach local regulations, with legal consequences, or lose competitive ground to rivals that do personalize their models and train them with judgment. That is why it is essential to demand models trained on quality data, with transparent processes and well-defined human controls. Best practices for financial sector leaders Knowledge is key to making informed decisions. Ask your AI provider: What data was this model trained on? Was fine-tuning used for the financial sector in LATAM? What human criteria were used to assess its quality? Choose customizable models: ideally you can adapt them to your operation and regulatory context. Look for partners who are experts in financial AI: not every model is fit for every purpose. Implementing AI in banking is no longer optional. But doing it well: with judgment, strategy, and a focus on the user, is what separates leaders from followers. Understanding how a language model is trained is the first step toward making informed, safe decisions. At Delto, beyond training language models with quality criteria, we developed the Banking Large Action Model (BLAM), a set of more than 300 ready-to-use banking Skills , designed specifically for financial institutions in Latin America. This means banks don't have to build their conversational assistant from scratch: they can start with a model that is already trained, tested, and regulatorily aligned, and simply adjust the integration with their internal systems. BLAM is not a generic library, but an architecture designed from the start to combine technical power, regulatory compliance, and implementation speed. It is, ultimately, the bridge between the theory of LLM training and its concrete, reliable, and safe application in real banking operations.

What are the stages of training an LLM? Three: massive pretraining on public text, fine-tuning on domain data (customer service, regulatory documents), and reinforcement learning from human feedback (RLHF), where people compare responses A and B so the model prefers the more useful and safe ones.

What criteria do human reviewers use to judge AI responses? Eight: truthfulness and accuracy, safety and responsibility, clarity and style, concision and relevance, understanding the request, logical flow, context adaptation, and level of formality. They are used both to train and to audit models in production.

What should a bank ask its AI provider? What data the model was trained on, whether fine-tuning was used for the financial sector in LATAM, and what human criteria were applied to assess its quality. It's best to choose models that can be customized to the local regulatory context.