Paco Orozco: Growing at the Edge: Doubling Traffic While Changing the API Gateway
About this talk
In this talk, Paco Rosco discusses the challenges and solutions associated with replacing API gateways while doubling traffic for Adyina's services. He details the transition from an obsolete application gateway to KrakenD, an open-source gateway that proved to be faster and more efficient. The speaker describes how they conducted thorough assessments and proof of concept evaluations before deciding on the new API gateway. He emphasizes the importance of testing, communication with users, and addressing operational challenges during the migration process. Despite facing multiple deployment issues, the team successfully onboarded new marketplaces and improved performance metrics significantly. The talk highlights key lessons learned about managing transitions in technology infrastructures and maintaining service quality.
Full transcript
[Music] ladies and Gentlemen please welcome our next speaker Paco Rosco presenting the topic growing at the edge doubling traffic while changing the API Gateway hello everyone uh you look very well from here a bit you know a bit uh aggressive let's say but I'm very happy to be here in front of all of you uh before I start my talk I would like to know how many
of you are thinking to visit Barcelona or have been in recently okay that's good that's good I'm uh from Barcelona I was born in Barcelona uh Barcelona is the capital of Catalonia and Catalonia what that we call Catalonia is uh officially an autonomous Community within Spain uh but we have our own um Parliament we have our own President we have our own police force and we have
our own language that is the Catalan the Catalan is a language that was V for more than 30 years and uh it H become largely underground so the old Generations like me that were raised between the 70s and the 90s prefer to speak Catalan in order to recover the language that was kind of uh disappeared before the new generations are not using it at all that's why
uh since some years ago I started to collaborate with an initiative to extend the use of the Catalan and I thought what better than to do it today with you that are in front of me right so what I'm going to suggest you is to learn some wordss in that's good right so you will come to Barcelona and you can speak with the people the natives there
if you find any uh but in order to do that I think that we can start from the basics so I'm going to ask you to follow me and join me please so the idea is that we are going to know how are the names of the phas an starting point so the phase in Catalan is called K car well then we have uh the the are
very well the nose is the nas very well that's a the next one is very important lips are Javis Javis well done and uh the last one for today it's ears that are ores okay let's try to recap again because I've seen that some people are a bit lost so we we have the face that is car cool Ace Nas yavis andas awesome I think that I
have the cleverest audience ever now I'm going to try something different I'm going to indicate one part of my face and you need to say the name don't be shy I want to feel that you know how this part of the face it's name it okay o well perfect Javis uras and big Applause for all of you and for your first wordss in Catalan so you are
prepared to visit Barcelona so let's come back to the speech so uh as I said I'm PCO I've engineering manager at ad as you can see uh I joined it seven years ago I'm a product manager and engineering manager of an oo team that is called DH I have two friends here so I hope that they have not brought any Rotten Tomato to throw at me but
just in case I will keep some distance uh on the nonprofessional uh side uh I love hiking uh this picture is from the Pines where I spend long weekends hiking and suffering a bit but you know then it's enjoyable and uh please raise your hand the knows the ones that knows what ad vinta is and what ad vinta does D I need to speak to our marketing
and hiring departments okay I would uh so adaina is a leading online classified group who operates more than marketplaces uh in 10 different countries H we want to help to find a new know we want to help to find everything and everyone and new purpose like we are in the secondhand uh world where we are trying to that people can easily sell things that people can easily
find a new job new houses new cars Etc okay H We Want To Be A Champion for sustainable uh Commerce making a positive impact in the environment in thec Commerce and the society and we do that mainly in Europe as you can see where our main Market places are operating last year our European platform received two and a half billion visits monthly that was a lot only
five of Our Brands were generated ating 70% of that visits so we are have we are we have some very strong Brands around Europe we employ uh more than 5,000 uh people which me which means that we are a very big Tech employer in Europe and H it's a very challenging but a very lovely place to work and to grow and in our case the edge please
raise your hand I have two friends here from The Edge uh The Edge is a team composed by four Engineers I have two here the other two are the ones that are working and in our case we have two services that we are in charge one is the gdpr middleware or privacy handers that is here and the other is jams the Privacy handers is the more boring
one but for uh your understanding is the one that is leading with gdpr request that if you know is a low European law that is trying to protect the rights of the users so uh when a user in one of our Market places wants to delete their data or request all the data that we have from there our service is the one that orchestrates that gdpr request
so we receive the request and we send that request to all the data stores that are the ones that are responsible of deleting the data or taking out the data very boring we are going to put side but the other one is the one that we are going to talk today it's called jams and jams is the acronym for your at the vinda media service we are
not very good at naming this service it's h used by our marketplaces every time that someone wants to publish an act let's imagine that all of us are thinking to sell something I'm going to sell my laptop where is the company laptop but well anyway so I will take some pictures of my laptop I will publish in one of the marketplaces I will upload these pictures that
are taken with my brand new phone which means that these pictures have a very good resolution very good quality but it means as well that they are very big very heavy okay so then in the marketplaces or front ends wants to show that pictures with lower latency okay with low latency and that means with maybe doing some transformation or service does exactly that we offer the possibility
to our front ends to transform these images in images that are more optimal to be published on the websites easy which kind of Transformations you we have a lot this is only some of them but usually the ones that are more used are resizing the image cropping the image adding some Watermark maybe file format change in order to to have a format that is more suitable for
websites Etc H lately we have added some transformation that are powered by uh AI like the capacity to remove the background of an image or to hide in some part of the image automatically like the car plate uh in case of of cars or the possibility to identify the objects inside the image so we can use that information in order to uh improve the search ad parameters
or to improve the descriptions of the ad we work as well with videos and documents and uh this service it's probably the highest adopted service in the company just for some numbers we are uploading or they are uploading around uh 700 million images every month and we have around 800 th000 requests per day delivering images okay that's quite big it's quite critical because if you are going
to sell or to buy something and you don't have images probably you are not going it and this is how we do it this is the jams architecture so uh there are a bunch of microservices that we deploy in a us and kubernetes and when a Marketplace uploads let me think oh it works perfect I'm going to put here sorry if I'm F you so when a
Marketplace uh uploads the images they use the this AP Gateway that is here which H is in charge of the authorization authentication and request preparation to call other services that will be the ones that uh will store at the end the image in AWS S3 packets okay that's how the images are reaching or Services we have uh here at the upper right some Services these are the
services that are ER in charge of uh this powered uh AI features and the ones that are uh up right are the ones that are satellit services to head metrics locks for work to send locks to our users so they can be autonomous uh managing that services and then when the front end or the or the sites wants to show these images they call it through fetch
API that we are offering as well in front of that fetch API Gateway there is always a CDN because as you can imagine each ad is going to be seen several times if you remember we have like two and a half billion visits monthly so we don't want to do two and a half billion Transformations more monthly so we cach that Transformations inside the CDN and inside
the service as well so then we have a bunch of services that are the ones that are doing these Transformations depending of the object type there are ones for animated GS another one for images for documents for videos Etc okay once this transformation has been done we cash the result so the next time that someone is going to call the same object with the same transformation we
already have it and we save some Computing there all clear till now and uh today what we are going to talk is about this ping AP gateways that were the ones that we replac it and it's what uh this talk it's about uh all the services are developed using Java and Goan we use a spring boot in Java and and that's important for the continuing with that
presentation so these AP gateways were replaced are you wondering why of course they were failing that was the reason and the reason is that uh these Services uh were implemented using zul zul is a liar 7 uh application Gateway that is Netflix and it's currently used by Netflix is the service that uh receives all the mobile and website calls in Netflix and then calls the rest of
the services the backend services in Netflix we were using the same that was back in 2021 and H we were having a lot of problems these two IP gateways were with us since the beginning of the service more than seven years ago and we already detected some problems that we were like working around we were like trying to avoid we were adding technical dep to our backlog
okay and we were not um able to replace at all the service or fixed it at all okay back in 2021 this service was a bit discontinued and Netflix was using an internal service that they were not H open sourcing with the rest of the community okay so we started with them at some point they diverge we continue with the open source version and is where all
this things happened in h q1 2021 we had a lot sorry q1 2022 we had uh three incidents that were very very big that impacted a lot our users and it's you know the time that we decided hey folks we need to do something with that and we need to to replace these AP gateways replacing uh this critical part was not uh you know an easy thing
we need to think thoroughly how to do it that's why we started to do an assessment of the AP gateways implementations that they were available and the the teammates did an assessment of nine AP gateways this is a result of that uh assessment where we were trying to to uh map the characteristics that we need the features that we need with the state ofth art on the
ipwi at the end we had like two finalists Kraken d and a spring Cloud Gateway remember that we were already using a spring Boot and spring Cloud was a good candidate because it was very aligned with our Java application that we were deploying uh what we decided is to take this two finalist and do a POC with each of one PC is a proof of concept that
it's implementing part of the AP Gateway with both implementations and a test uh their performance to see um the stability of these two services and something important the ease of development how we can develop using these two AP gateways to ensure that later we will not have any problem after these PCS uh the chosen one was Kraken D Kraken D who knows what is D Co Kraken
D it's a an open source Gateway which is which has been proven to be the fastest one at least this is what they say in their website and I truly believe and use a declarative language to create the end points so it's jaml files Json files that you can configure all the RightWay and uh it allows a lot of extensions like plugins that you can add to
this I gateway to customize it at your own taste uh it's written in go that is one of the H preferred coding language in our team so everything looks good and uh for that time Kraken D joined to the Linux Foundation which was ensuring the continuity of the project over the time so at the beginning they were a small community then once they joined to the Linux
Foundation you know the community growth a lot so for us was like Hey this service is not going to end tomorrow so we have a finalist was Kraken D that would solve all the problems that we had in our services but or bistan uh implementation or replacement of this AP goway was kind hijacket by some company updates and it was that adaina was acquiring eay classified Group
which was a very good news okay some of us had opt have a stock options of adinda so that was very good H we became the world largest uh Marketplace and it means that we were the big player in in not only in Europe we were present in um South uh South America North America uh Europe of course Africa and Australia so we were everywhere okay that
was very good news for the whole company we did uh this adquisition and everybody was very happy but that update is the one that made me be today because that means as well that all the marketplaces that we B with that adquisition shoot on board in jumps they were using a media service that was coming from eBay and they should deprecate that service and use the one
that we had in adita that was yam and we had only nine months to do that because after after 9 months both companies evay and evay classifier group should not have any relation so we had this hard deadline remember we have two epic gateways that are failing that has been involved in three big incidents and now we are having more marketplaces ding or we were estimating that
we will double the traffic on these who is feeling fear okay we were in that time but we say hey challenge accepted we are going to do it so we stop the immigration for a while and what we started to do was preparing our marketplaces to on board in our service okay uh jams it's a service that offers a lot of authonomy to the users okay uh
we provide a lot of sales serving tooling which which translates into very low operational F food uh footprint okay how we can do that with only four teammates if it was not in the same way so it means that most of the work most of the configuration of the deployment is done by all users we offer some cell serve tooling and they can configure everything they are
responsible for changing blah and blah to deploy new things etc etc etc but this autonomy works as a charm when you know about the service when you have been using it for a while since all of these marketplaces were newcommerce we needed to prepare several workshops first to ensure if the features that they were using in the old service were covered by yams that's important thing and
secondly to train to use jums as it is expected so instead of asking can you change that configuration is here you have the tool you can do it by yourself so first we did an assessment about the uh features Gap luckily we didn't have had and later uh we organized several handson to uh explain how the service works how the self serf tooling can be used and
how the migration process to adopt jams will be and this required a gillion of documents a gilon of documentation we revamp or service portal explain and everything that is our service PL portal the one that has all the conf all the um documentation has use cases has the reportings etc etc uh that uh work implied more than 60 people to think about the immigration because of course
each Marketplace has their own it team so we need to on board and engage them ET okay and after preparing all these steps uh is where our new colleagues can start thinking how to do the immigration how to prepare themselves to do this and we like continue it with the AP Gateway replacement because again remember we were failing with our current traffic but we wanted to add
the traffic again so two times the traffic that we had so that was the recipe for failure for sure so uh what we decided is like we implemented h both H AP gateways using Kraken D and then we decided to start the migration of one of the AP gateways and we will use the Gateway can anyone know why the Fetch and not the management no so we
started with the fetch because if you remember the the the map that I show at the beginning the fetch has a CDN in front of it so any D time D time in the fetch could be somehow uh fixed by the CDN so we thought that that was the best and less risky approach so we plan a various straightforward Implement and replacement that was uh using the
new Gateway with Canary releases okay do you know what is a canary release okay I think that is important to stop here and uh to show you how we release code in production so this is our releasing pipeline more or less okay it's a simplification but we have our code in GitHub every developer when they need to do a change make a branch work on the branch
at some point they send our pool request that this pool request uh automatically will execute some unit test some integration test and we ask for approval of other members in the in the team okay to ensure that the viables names are correct once everybody it's uh has approved well not everybody but at least two people has approved this pull request we merch to the main branch the
main branch will um trigger the deployment and we will run again unit and uh integration test and then we will buil the docker images the imis everything that is needed for the deployment we will deploy it in dep and in Pre and in Pre we will run again some testings okay and when we are super sure that everything is working we will go to Pro and in
Pro we will use Canary release Canary release it's a the topic to to um name when you have a lot of instances in production and you replace only few of them with the new code and you start watching how they behave and if they are like increasing increasing the latency or they are producing a lot of errors it's a would say safe way to put something in
production and that's what we thought to do so a we are going to deploy or new fetch Gateway using the brand new Kraken D that is going to fix our problems while everybody it's on boarding and easy I have one of my teammates that the word easy is the default word that he can throw when you put a challenge in front of him but uh that straightforward
plan became a BM Kurby mtin road with a lot of problems okay so basically we need we were forced to do seven de deployments which means six roll boxs and this is what this stock is about so of course we didn't success at beginning we encountered a lot of problems that we had to fix so uh we ended deploying seven times and we had to deal with
very high CPUs where we were deploying because the way that we were logging the in in our service was not the right one with the new AP Gateway uh there were other errors that were produced because some heers that we were using in the service were not there were not implemented we forgot and the one that for me was probably one of the big biggest one and
is the one that I'm going to highlight is when H we were we deployed the instances and we started to see that there was a very used HTTP method that was missing and it was the head method we use the head method a lot service and we were expected that all modern AP gateways will Implement all the http methods but that was not true we H reviewed
the documentation we contacted the maintainers of the Kraken the uh software and they confirmed us that they were not implementing the head say come on we have invested more than five months and now we discover that one of the most used head most one of the most used HTTP methods is not there that was thought but uh what we did is we implemented that method we make
it available for the for the community now it's an extension in Lura it's worth to mention that Lura and Kraken D are the same are the same base code but Kraken D is the commercial version so you can use Kraken D for free but if you want they can offer you commercial support and some extensions paid so uh we ended having a AP Gateway that was not
Lura was not Kraken D was Gateway uh this Gateway is the one that we are using it's uh fully compatible with Lura so every time that the community fix something or launch a new feature we can adopt it because the differences are very very few one so once we fix it that big problem we did it so we were super happy you could be super happy with
us because we had a fetch Gateway that was working that new market places were on boarding and we saw the metrics and everything looks like hey we nail it so it was time to go to the next one the management gway we have deployed seven times the fet gway we knew a lot about this Gateway now we uh improve or through shooting skills with a new gway
so deploying the management gway should be as easy as you know whatever you want to put at the end of the of the sentence so we plan it again a canary release replacing only some instances to have only percentage of the traffic and uh we did the same that we were expecting in the fed and of course we failed it Mr rle so new five deployments new
four roll backs new four incidents some of the incidents were you know as weird as putting a wrong timeout configuration that was making that the calls could not be completed before the timeout that we set which was like hey why we have set this low number instead of one bigger well you know then H cm it's one of our authorization services so we put a gr calls
cannot be authorized so all the C were failing basically the third one is that um the new Gateway was U dealing with the trailing slashes in a different way than the previous one and in some cases that was failing and the fourth one was that uh we realize that in some cases when when we were using multipar request the authorization was not working at all so again
that but we did it it took us a lot of roll backs we included a lot of incidents but we were happy the marketplaces were on boarding then I will show you how fast they were on boarding and we were seeing the metrics and everything was going I I need to fill you know that your engagement with that so probably you have been in projects that has
been successful without being without cing 10 incidents we were okay and this is what we want after replacing both IP gateways the new code was using less CPU half of the CPU to be honest uh this means fewer instances and with the same traffic which means that we were allowed to reduce the number of instances which reduce the cost remember that we were that we are ping
AWS so that's how it works uh the Lura Kraken d uh or now jumps Gateway was able to reach more request in parallel so we were we were having like 17 uh th000 second while the old schol Gateway could handle only 2K 2,000 requests okay last week we reached 65 th000 request per second with that Gateway uh the logs H we had if you remember one of
the problems that we had at the beginning is that the request IDs were lost during the the old SCH gateways so now the logs were H consistent we have all the information that we were able we can trace all the request uh using that request IDs and uh having consistent locks makes us feel more confident uh after 10 incidents as I said before we were better through
well shooting than the Gateway and uh given that the whole team was working in a new AP Gateway at the same time we broke we broke some silos that we have before because of course before they were like the expert in zul and the rest were like trying to catch him up with him but at that time we were all of them all of us in the
same level okay that that was a very good benefit uh uh at all and we would we did that while all the marketplaces that you saw at the beginning were on boarding we double the traffic we had three times the traffic that before we had double the number of marketplaces on boarded inor service while we replaced that while we were creating new features that were demanded by
your users so that was a very very very complete year here you can see see for example how the fetch Gateway was uh behaving before and after okay that's the difference between using uh the Kraken D or not using it as you can see the I don't know if you can read it from there but the blue line on top is half so it's double that the
the one that are below this is how the or kpis were increases from since the beginning to the end so we multiply by four number of request that we were receiving by the number of new uploaded images monthly that we were receiving and as I said before how this on board was done it was you know like an exponential growth so here more or less in December
is where we uh started to deploy the F the first uh fetch Gateway we failed it six times I remember then we started the the management and the number of request was like escalating uh exponentially thanks to the adoption of the new marketplaces and to finish this uh this speech I would like to share with you some of the takeaways of that uh complete year that we
had replacing the AP gateways and doubling the trffic at the same time first of all don't be afraid of spending time doing the assessment and any time invested in the assessment should probably uh have help us to identify some of the problems that we had like for example the lack of the head method how hell we forgot to test that or some of the configurations and that
will uh you know reduce the number of incidents that we had which at some point having more incidents makes you less confident about the work that you are doing uh test are super important I'm not going to insist on that because we are in a deos company everybody knows that this are super important with Su we were it was very difficult to test the way that zul
works was very very difficult and we were we had like a test coverage around 60% of the use cases now we are 96% which means that every time that we need to do a refactor a deployment we are more confident we are more relaxed we can do it every time every hour during the day the next one is about logs and metrics for for me that's very
important logs and metrics are vital for your work why because if not you are blind and if you don't are blind you don't know what is happening what is going on what is what is work so we treat uh logs and Metric as a first class citizens it means that in every change that we do we think about what how the logs and the metrics the next
one is a another one that is usually save it's uh communicate the changes with your users even that you think that they are going to be transparent communicated so over communicate that's not a problem if you do that I can uh tell you that they will be more engaged on that every time that you fail go there and say hey we have failed we forgot to implement
that head method hey about how you we forgot sorry we are humans and we forgot that uh make us uh a very close Community between the users and the team and we were very engaged on that and that brought us the possibility to reach the last one that is win and lose together we were losing together when we were failing a deployment when we were causing an
incident that that means money in our users but at the same time when we finish that when we we complete the replacement the whole migration we celebrate it as one team as well fact this is one of the the picture that was taken in one of these celebrations as you can see the faces of all of them are like yeah yeah we have failed a lot we
are suffering a lot well it was not that way okay and this is the end of my speech and now uh now what I would like to to know is what you have been would you done different friendly or if there is any question any comment something that was not clear I will be more than happy we have actually we have quite a lot of questions to
you thank Applause uh I really like this first one what it's coming here it's really interesting one have you fired any your bad Engineers come on six roll out okay just count the number of people that are here five math doesn't lie no we have not fighted anyone uh I truly believe that any failure is an opportunity and that was always an opportunity if we were not
failed so much we for sure have not been engaged our customers because that was key to our customers to be engaged with us and to see that we were failing and they were new in the company they don't knew anything in the company so I can imagine that they were feeling like with imposto syndrome they were coming to a new company everything was running perfectly but they
were they needed to leave a service that they knew to to adop a a new service that they don't know anything about it so this feeling I think that was shed when we were failing we were all together in that so I don't think that uh that was a reason to fail someone of course I ALS believe about it okay how did you get to as far
as to deployment in BR without finding out uh H HTP get method was missing in the crack in D La well it was not the get method it was the head method he okay but yeah that's a question that I sometimes uh wake up during the night thinking how was possible okay it was written in Thea was tested yeah yeah yeah it was written and you know
I can say that uh we were very confident that or test coverage was covering this case that we did not review with it that's why it's so important to have a very good test coverage we had 60 now we had next6 probably we are missing something for sure but the head method it's fully tested okay clear uh did you do low testing too and if yes could
you elaborate on tool you used and in which Step you would do the test we do we did uh stress testing when we had the two finalist you remember we did we did a PC we are using uh can you help me with the tool I know that it's a stress test tool but I don't remember now the name sorry I I will come later it will
go later if the one that has ask that is here please reach me later because I will okay it's it's a very wellknown uh testing framework that now I cannot remember and we do it in the in the when we were doing the the two finalist and now every time that we adop Sorry J meter yeah this is the the tool J meter is the tool that
it's an open source tool as well where you can send use cases and you can grow as much as you want and every time that we are implementing a new major version we are doing some stress test in order to ensure that the performance is the same sorry U again this soft questions as a Spanish why did you go for the spring Cloud ah why because it
was not the best so we we thought that uh that would be the finalist because um we were using a spring both and it was the natural way that and that's that was the thing if you saw in the in the Matrix was not General available it was still in Veta release candidates and so on so we thought okay and once uh we tested it it does
not behave and the other one okay again maybe the that's data from 2022 now Everything Has Changed now Netflix zul is very good because they have reached a new version and it's probably another time to to review with it okay many of some of the questions was already asked answered here why you have two AP gateways so yeah we have two AP gateways basically in order to
be able to deploy one or the other because they are they are doing two totally different things even that they are the same AP but ones have authentication the other the fetch Gateway has no authentication so they are very different even that they call AP Gateway and they use the same uh framework to be implemented they are so different that was one problem that we had before
that the Z Gateway was deployed twice which means that every time that we wanted to change something in the fetch we were deploying the management as well which makes us to run a risk that we don't want to run again then why we have to separate a with one more lovely question is Kraken is as stable as you expected before the migration what the challenges related to
the new Gateway are you trying to solve now kaken is super stable it's rock solid you can send a lot of request and it will stay there okay what are the challenges So lately we have been uh migrating the IP to graviton instances graviton armm 64 okay that we have found a lot of uh problems there and um I think and and more now what we are
trying is to reduce the cost to make them more optimal etc etc I think that this is end of our time but Pasco you really popular today so don't don't be shy please connect to Pasco he has really nice blue t-shirt it's really easy to recognize him and I would like to say thank you and warm applauds to Mr thank you very much yes thank you very
much thank you
More from this event
See all 58 talks →
Halil Ibrahim Kalkan: Building a Kubernetes Integrated Local Development Environment
45:20
Viktor Vedmich: Ideal Blueprint Versus Reality for CI/CD Pipelines
46:03
Koray Oksay: Continuous Deployment: The GitOps, The Pipelines, and The Ugly
43:03
Marek Grzenkowicz: How to Automate Dependency Updates with the Renovate@Roche Bot
44:47