About this talk
In this talk, the speaker, Yamaha, head of data platforms at Sligro Food Group, discusses the company's transition to a data-driven organization. He highlights the challenges faced during this transformation, including the need to consolidate multiple data sources and improve data governance. The speaker outlines the strategic approach taken to create a centralized data platform using Google Cloud's ecosystem, incorporating tools like BigQuery and Soda for data quality monitoring. He explains how they built a data lakehouse architecture and implemented data hubs to facilitate flexible data access and sharing within the organization. Additionally, he emphasizes the importance of business alignment and ownership of data across departments to ensure success.
Full transcript
Good afternoon, ladies and gentlemen. Can you hear me correctly? Because I have no idea because I can't hear myself. Uh well, thank you so much for joining me actually in this adventure because most of the people are actually drinking beer, but um it's uh it's nice that you're actually here uh to listen to our story. Just a quick introduction, who am I? Uh I'm Yamaha. I'm the
head of data platforms at Sleer Food Group. Um as you can see in our company, we all have a passion for food. So besides that, I do something with technology. I also sometimes work in our stores actually as a show cook because that's what we have in our DNA. So that's what you actually need to have to work with our company. So what am I going to
tell you today? Um I'm going to introduce our company to you and I'm not going to do that myself, but I'm just going to use a little video explaining what we do. Um and then I'm going to talk about our data transition that we've actually started. Um and I'm first going to talk about how to sell that actually to your business. So the first part is going
to be slightly businessoriented. Uh the second part is is what did we build technology-wise? Uh and the third part is how did this concept actually prove itself when disruption of course as always happened and we needed to respond very quickly. Everybody okay with the audio because I see people switching? Yes. Great. So, let me first introduce Sliggo Food Group to you for the people that don't know
us or that the people that have an idea of us that we just have some stores, but we have a little bit more than just stores. [Music] Yeah. [Applause] Just a minute. Nothing. Heat. Heat. Heat. So, as you can see, a very diverse company. And just to mention that none of the people that you saw in this video were actors. They were all our own people, our
own people starring in what So now about our data transition because I probably most of you have some affinity with data. So I expect a little bit of background there. With Slero, we had a long history of all sorts of data initiatives. We have a formal data warehouse with terod data. We have some illegal data warehouses built on all sorts of tools which were not meant to
be data warehouses but turned out to be data warehouses. We had a digital department that wanted it their way and started their own data warehouse. somewhere on Google and we went to SAP and also we started using SAP uh uh analytical cloud. So a whole mixed bag but eventually that's not the way to go if you want to be truly datadriven. Um and it's been since actually
corona that we found out that we need to be much faster and more agile in our ways of working with data. Um we have an international ambition which has a lot of impact on how we use data. We faced corona meaning we need to scale down our entire organization and rebuild it. Scale down again and rebuild it. Uh we have a new generation with total new demands,
total new wishes. Uh we have the climate is that's changing and then we faced a recession. So when I first started the company in in uh 2017 uh everything was extremely predictable. Uh and right now we have no more predictability in the way our customer orders in the way that the trends are. So everything is new. So we really need to steer on data. And when we
said to ourselves we want to be datadriven, we needed to ask ourselves a couple of questions in order to actually start. What is our maturity level? Is it really our ambition to be datadriven? Uh what organization fits that purpose? What technology do we need to use? And how can we actually make that successful? So what do you do? You ask gardener and you ask them okay so
what is a maturity level for a datadriven um and this is a very common model that you actually see you can be aware you can be reflective or you can be truly AIdriven which is well world league uh uh soccer so we all saw ourselves to be well where do some predictions and we are very much there but in the reality if you look at all the
boxes that you need to tick to be there this is where you are we had no good data quality We held multiple data warehouses. Uh we had no data governance and so on. So we really had a whole way to go to get to the point to where we wanted to be. And this is why we created a program, a strategical program to actually make that change
within Slero. Um and you can run this program from it or from data department. And you can think, okay, we're going to change Sle and we're going to change data governance and we're going to change the way people work with data. But it's not going to happen unless you have full commitment from your board and unless the board is actually in on the program. So that's why
we put this in our digital ambition which is part of the corporate roadmap and we made the sponsorship of the program not it or data but we actually made it our business. Our COO himself is the main sponsor of this program. That is how we got started. So then you choose your right organization and again there are multiple models. You can have a decentralized organization, you can
have a federated organization or you can have a central organization where you say well we have our analytics central. We can have some federated users. Uh that's basically the models that you have and of course we we were doing everything in all the departments. Everybody was just getting data doing the PowerBI thing locally and thinking they were doing analytics. But what fits our maturity level was actually
to centralize everything to have a central organization and then eventually to start federating that giving people the opportunity to build some dashboards themselves based on control data sets. Moving forwards now we know what type of organization we need to choose towards the technology that uh we wanted to use. So we need to choose the right fundament. So you can go for a out of the box tool
uh like terod data that we have now uh combined it with Informatica or combine it with other big tools. Um but we actually chose to go for a ecosystem. So we chose a hyperscaler and use all the tools of that hyperscaler the out of the box tools of the hyperscaler to actually build our own data ecosystem just the way we like it. And by using a hyperscaler
you can be extremely quick in implementing and it won't give you that 99% fit that you might want that the big tool like Informatica terod data could give you but on the other end it gives you the speed and agility to very quickly deliver business value uh and to very quickly respond to changes and that's why we actually choose to build our ecosystem on top of a
hyperscaler in this case Google. Um and of course you don't you're not doing that yourself you're doing that together with a partner. So we chose Sabia uh to be our partner in that to learn us how to deal with this technology and to start building our own team that can eventually take over the technology. And what did we want to build? Well, of course, we want to
build a data platform with a data lakehouse with advanced analytics with a uh regular analytics, some data governance tools on top and we could make all sorts of nice architectural drawings for it. But how do you sell that to the business? Well, within Slegro, you use food and you actually tell the people, "Okay, so we have an analogy here. We're going to talk about potatoes coming from
the field being worked up until you actually get a plate of steak and flies because this is the language our business could understand. So we basically explain them the way data travels from a source system towards a data lake towards a structured data warehouse towards data products towards analytical products that eventually can be consumed. And by explaining it this way we could get the business along in
our journey that we want to make because without the business you cannot make this journey. Now executing this mission so building this platform and we build a green field platform. So that was a luxury that we could start just from scratch. First thing to do, get your GCP connected, get a landing zone, get connected towards your regular Siggo background. But we wanted to be independent. So we
wanted to be fully devops and fully independent from our more traditional technology side of Slegro. So we set up our own landing zone. uh connected that to the major tools that SLGO was offering but kept the uh um kept kept being solely on that specific platform on the landing zone. So we were fully in control there. Next we chose a few data operators to get data out.
We chose five trend to be specifically with our AS400 system. We use dbt which is very common and of course GitHub uh for all your uh git actions and the CI/CD. And then we came to our data warehouse and we came to a very specific question that we needed to ask ourself. Do we want to have Kimble models or do you want to go for something which
is more resilient? Uh so data vault we at sle are in the midst of actually a transition from a AS400 legacy platform towards a fully new SAP platform. So we will be running a hybrid business the next two or three years. Further on we do a lot of takeovers. So we take over a party, merge it into our own organization and what is more helpful than a
very resilient way of modeling where you can actually combine multiple systems and that's why we chose data vault as to be the technology or the modeling technology for our data So that's it. You start building your data warehouse. Uh we use BigQuery as a regular database. Um we have a stage a business vault and then eventually we have Kimal Marts ready to be used and consumed by
our analytics department and then came the next challenge. Everybody wanted to use PowerBI because then they were extremely flexible and they could hand out the data to everybody. They could make their own insights but they could also change the data. But we didn't want that. We wanted to have control over the definitions. We wanted to be very clear in what data they could be transporting to others
because we were at maturity level zero and everybody was having a discussion about what a specific KPI meant or what wasn't a specific KPI. So we went very specifically for Lucer because Lucer allows you to be very much in control in what people can do with the data and who can share it and who can't share it. So you can give people a clean shape on a
data model. they can build their own insights, but it's still going to be your data with your definitions and it's going to be very much in control. So that's why we specifically made that choice to on top of this stack add our Lucer fundamental uh for our analytics. Of course, moving forward, you also have some advanced analytics. We're now running a PSC with Vert.x AI um which
makes this whole thing of course complete. As to data governance, we have a lot of data in the system. I already told you we were very immature which means that our data quality is very poor. So you need to add some data cataloging tool, you need to add uh a data quality tool and we chose something lightweight, something that wasn't very difficult to implement but still gave
us the value that in our maturity level made sense. So we added soda which is a very lightweight data quality tool. Uh but it's very handy to just start using data quality. And we added data plex from Google as our data catalog where we also just store the definitions of data classification that we can then later use in all our data pipelines for data security. And then
of course we have disruption because you're in the midst of building this project and everything goes well and you're building you're building building but then all of a sudden the business says okay we just thought of that we wanted to change the whole Belgium infrastructure. We wanted to put it on the AS400s and we want to harmonize it um and help we need to translate all the
article information from A to B and it needs to be a live translation for the next couple of So you need a sort of data hub for your master data. That's what at least the the the requirements were from the program. And they started to build that with the people managing the master data. That didn't work out. So eventually they turned us to us and said okay
guys you are thinking composable. Um can you help us to do this on a very flexible manner? Yes we could because what we want to have with all our master data is that it's really loosely coupled within our environment. uh we want to have data hubs with generic APIs everywhere in our environment basically to be filled with the right master sources and consumed by subscribed data consumers.
Uh so we can take away master sources at any given moment and add consumers at any given moment and all the data in this hub is not in a specific model to an application but it's in a generic model. So we ask the business tell us if you look at a customer or if you look at an article or a supplier if you look at this in
business terms how would that look to you? Customer has addresses it has locations it had delivery locations and so on. An article with us in in retail we have about 900 attributes to an article. So it's a pretty complex thing. uh but tell us in business terms how that looks and we'll build a generic model not looking at your data sources not looking at the applications that
need to receive it just generic. So what we do we consume all this data from those data sources put it into a very generic model and then expose that to the consumers which eventually became our data hubs and we already now have two of those data hubs standing by and this is where the Google ecosystem really shows this flexibility because we had everything set up as we
had we had our ingestion pattern set up um we could very easily spin up an infrastructure just to consume the data in a very similar fashion we already did the modeling exercise to get everything in this general generic model because we also use it in our vault. So this is where you could see everything was adding up and you were very able within two months to actually
deliver the first prototype or working prototype for this data hub. And just for the technology-minded people here in the room, how does this then look? Um eventually of course it's pretty complicated. Um we use multiple ingestion patterns to actually get the data. It's fully event driven. So you get an event from the front that something has changed. The event is triggered throughout the entire chain uh and
delivered to your consumers. We use uh uh cloud SQL databases in multiple zones because this is used in our operational systems as well. So we need to be high available and very resilient here. Um and the data quality again of these data hubs is monitored in soda as well. So all your master data is now centralized in a very specific place which is called the data hub
uh and is then monitored in terms of quality and continuously fed back to the business that are trying to improve your data. how are we looking in 2025? Basically we have our system set up. We added our data hubs. We have our uh analytics platform standby. We are doing our first proof of concepts with Vert.exi for the advanced analytics and we're still filling up our generic model
with new entities uh and new data that come from the various sources. So this is just to give you an insight in what we do. Uh and my question to you, do you have any questions about what you just saw? Everybody's very silent. Yes, >> sorry. Ah, there's the microphone. >> Thanks. Um, my question was around the business value that uh this setup delivers to you and
the business. >> Yeah. Yeah. The one of the things of the business value uh the question was what's the business value? um when not having your data centralized, when everybody just has their own definitions on the KPIs, when nobody can actually tell you the truth of what our customer satisfaction is, uh there already lies a very big uh uh business value. So we have one centralized version
of truth, one centralized place where you can actually get your data. And if you then can use the same data in your operational process as well um then you're feeding multiple processes in your business at the same time with the same data with the same definitions and that for us was a very big uh value and also the speed in which we can now do things for
Sle. So that was worth the investment and also getting rid of all the old technology is also a pretty big benefit actually. Other questions? >> Hi. Uh, one question I would have is how did this like technological migration also look like on the let's say human side like did it change how you are organized? Did it create new teams? Oh yeah, it it had a big big
effect of course on how we were organized because prior our data warehouse team was just a mixed bag of people doing both the ETLing, both the front-end reporting. Uh people could just walk in and tell them oh I want to have this KPI change in XY Z. So we created a whole new department. Uh we split the department in a technology side, so a platform side which
I'm heading and a business side which is the BI specialist and the data scientists. So this organization already grew from six people towards roughly 40 in that very short period. So that's just the the part on the people that are building this platform. Now on the people in the business, there's also a big effect. Nobody up until the point that we started our program took any responsibility
for the data that we had. Data governance or particularly data ownership was non-existing. And it took us uh um a lot of missionary work actually to really tell people need to own the data. A director needs to own the data not it a specific commercial director needs to own their own data and they are responsible for the state of the data. They are responsible to provide data
stewards. They are prov responsible to havememes that actually know their data. And this has been a very big turnaround uh for people to understand they are responsible. they're responsible for their KPIs uh to understand what happened. So this is a continuous struggle that we had and it's it's improving improving improving up to the point where we can now actually say okay so one of the 50 entities
that we defined that we have everyone has an owner and every owner knows that he's owner of that specific data entity. So it's it's a it's a really change in the way people think about data. more questions. >> Uh the big question is what's next? >> What's next? Well, we still have a lot of entities that we need to process. So, we first started out with our
old entities. So, the entities that we already had in our data warehouse. Um we're halfway there modeling all those and getting those all in in ingested actually in the system. But of course there's a huge bunch of use cases with entities that we simply do not have yet that people want to have delivered. But what's next for us is basically our migration to SAP which means that
we really now need to facilitate that hybrid environment that's going to exist over the next year. So this year we have five data hubs scheduled to be delivered and we have a full generic model that needs to be fed from both the AS400 side and the SAP side. So we have still one glass plane of reporting uh that we can run our company for. So that's going
to take up all our year uh this year. But there's already a big wish list of course of all cool analytical use cases or predictive use cases that people want to start. Uh so there's not a a lack of ideas from the business that we would what we want to have. If no further questions then I thank you for your attention. If you find this very interesting
or if you think I would like to work with Siggo because it's a great organization, go to our website Pentanel. Uh we still have a lot of analytical engineering jobs. We have some data engineering jobs, some cloud engineering jobs. Uh and it's super fun. And if you're interested, that guy in the white, he is actually the guy to talk to. Thank you guys for listening.