Get the app
Back
Building Cleo • 05/29/25

Let's have an Espresso: MLOps at Cleo

Data engineer Fabio Badali talks about why we introduced Espresso, our MLOps framework.

Machine learning is a critical component of Cleo’s technology landscape. We are currently maintaining machine learning models to help with the most disparate parts of our pipelines and services, ranging from transaction enrichment to scoring.

In the fast-paced world of data science, there’s a natural temptation to cut corners when developing machine learning solutions given the complexity involved in these models. However, at Cleo, we think compromising on our engineering principles could not only harm your project’s quality but also introduce unnecessary headaches down the line. We think we can level up our machine learning solutions by adopting software development best practices, such as thorough testing, continuous integration/continuous delivery (CI/CD), and monitoring, but we are not going to do that at the cost of our values.

That's why we introduced Espresso, our MLOps framework. We wanted to provide an easy and standardized way to serve, deploy, and monitor our machine learning models. With Espresso, we shifted the paradigm from machine learning models to machine learning services.

Our coffee beans: Before Espresso

Every good espresso needs good coffee beans, and our Espresso makes no exception.

Having high-quality machine learning solutions is a game-changer when it comes to building a MLOps framework. For us, this meant we were able to trust the data science team’s implementation of the core logic of our services, enabling more flexibility during development for data scientists and saving a ton of time to invest in perfecting Espresso.

However, before Espresso, we had basically no standard solutions to build and serve our great services, while monitoring was almost nonexistent.

Roasting the coffee beans

We believe every taste is unique, so just as a coffee connoisseur roasts their own beans to achieve the perfect cup, we crafted our own MLOps framework to deliver the best possible outcomes for our needs. While there are plenty of off-the-shelf options available, we wanted to ensure that our framework was tailored for us, rather than trying to fit us into a predefined box. We also realized that, just as coffee blends can provide richer and more complex flavors, mixing different ideas from different solutions could result in better outcomes for Espresso.

We needed a production-grade solution, not just a fancy playground, so we made the conscious decision to build Espresso prioritizing compatibility and ease of migration from day zero. That's how the first iterations of this framework (before it got this wonderful name!) were already solving most of the needs we explained above. However, they were still relying on Amazon Sagemaker as a hosting solution so that we could iterate faster, learn at speed, and make it happen.

Eventually, we realized that Sagemaker was not the right serving tool for us (although we still use it for training our models) for several reasons:

  • We thought we could do better with costs and deployment times. Spoiler alert: We were right!

  • We had relatively unique needs that were incompatible with some Sagemaker features. As a result, we quickly found ourselves using Sagemaker solely as a hosting solution.

  • Sagemaker’s requests throttling is very aggressive and depends on the number of cores. We have both I/O and CPU-bound services, and we quickly ended up with a lot of replicas with very small CPU utilization, wasting resources and money.

  • HTTP headers are stripped out, and there is no support for W3C tracing headers. This means we were not able to follow a trace entirely from our tools without implementing custom code.

That’s why Espresso started to take shape as a full-fledged serving solution based on Kubernetes and Istio.

How we brewed our Espresso

Of course, we put a lot of love into crafting our Espresso, but we also relied on a variety of enabling technologies to handle the nitty-gritty work behind the scenes, including:

  • Deployment: CircleCI and Terraform to build an infrastructure-as-code pipeline to seamlessly deploy our ML services in our Kubernetes cluster.

  • Serving: s2i, FastAPI, Istio and Docker to abstract complexities for the data science team, offering a solution that “just works” while also introducing advanced capabilities such as server-side batching.

  • Monitoring: Integrations with Rollbar, New Relic, Prometheus, Cloudwatch, Slack, and our own Espresso Data Capture for model monitoring.

Tasting the Espresso 

Just like a well-crafted hot espresso shot, reviewing the results of your work can be invigorating, and evaluating the numbers is a fundamental step in every Cleo project.

  1. Time to create a new service dramatically dropped, from days to potentially minutes, thanks to our Espresso CLI which enables you to generate a skeleton project in a few seconds.

  2. Monitoring capabilities were very basic before Espresso. Now, being able to automatically generate dashboards and analyze different spans of one request, we found important bottlenecks in our recurring transaction matcher service in just 10 minutes. This analysis brought a 50x speedup to the service.

  3. After migrating out of Sagemaker, full CI/CD runtime went from ~25 minutes to ~8 minutes.

  4. The adoption of Espresso server-side batching feature guaranteed a 3x average speedup to anti-fraud services.

  5. We achieved cost savings of ~40% while keeping the same performance.

  6. Internal surveys revealed that developer experience improved in basically all phases (development, deployment, and monitoring), with Espresso consistently cited in the good quadrants of the data science squad retrospective.


Bottom line: When you’ve got all the right ingredients in place and everything is executed just right, you can craft something truly delicious, and the end result is gonna be pretty darn satisfying.

Let’s raise a cup of espresso to a successful MLOps framework. Cheers to great taste and even better results!

Interested in working on our data or ML team? Check out Cleo's open roles here.