Billing support at a Tier-1 bank does not scale by adding headcount. Every ticket starts the same way: a confused line item on a statement, a charge nobody recognizes, a fee that looks wrong. The answer is almost always sitting in the platform already, but finding it means reading logs and configuration most support engineers were never trained to read.
The problem that would not go away
L1 support at Zafin spent most of its time on the same handful of question shapes, asked about different accounts. Average resolution time sat stubbornly around 8 to 12 minutes a ticket. Almost all of it was spent translating a billing record into something a non-engineer could explain to a customer. Adding headcount would not fix that. The platform had to start explaining itself.
Why a separate Python service, not more Java
The pricing and billing platform itself is Java and Spring Boot, and it stays that way. The copilot is a standalone Python service built on FastAPI, talking to the platform over REST rather than living inside it. That separation was deliberate: an LLM integration brings its own dependency surface, OpenAI's SDK, tokenizer libraries, a PyTorch training loop, that has no business inside a pricing engine processing two million transactions a day. If the copilot service falls over, pricing and billing keep running.
Streaming, not spinning
The first version returned a single JSON blob once the model finished generating, and it felt slow even when it was not. A multi-paragraph explanation of a charging flow can take several seconds to generate end to end. Switching the FastAPI endpoint to Server-Sent Events let the UI render tokens as they arrive, the same pattern a modern chat interface uses. Perceived latency dropped even though the total generation time did not change at all. Perceived speed and actual speed are not the same problem, and SSE is the cheap fix for the first one.
Self-training on its own history
The interesting part is not the chat interface, it is what happens after each ticket closes. Every resolved query, the context pulled to answer it, and whether the agent accepted the answer becomes a training example. A PyTorch model trained on that growing history reranks which billing records and charging-flow explanations get surfaced for a new question that looks similar to one already resolved. The model does not write the final answer, it decides what evidence the LLM gets to see. That is a deliberately small, supervised job, not an open-ended one, and it is why it was safe to let the model retrain on live data without a human reviewing every update.
What it actually moved
L1 resolution time dropped by roughly half. The honest reason is not that the LLM is smarter than a support engineer, it is that the copilot always reads the full charging flow and never gets tired of the fortieth near-identical ticket of the day. The bar it has to clear is low: do better than a human skimming logs under time pressure, consistently.
What I would tell someone building this today
- Keep the LLM service out of the core platform's process and deployment lifecycle entirely. It should be able to fail without taking anything else down with it.
- Stream the response. The generation time is the same either way, only the wait feels different.
- Let the self-training loop rerank retrieval, not author the answer. That keeps a bad training update from turning into a confidently wrong response.
- Measure resolution time, not model accuracy. The metric that matters is how fast a human closes the ticket, not how the model scores on a held-out set.