DevDays Europe 2025

Bobby Bahov: AI Image Generation: Visual Artistry with Midjourney, Adobe Firefly, and DALL-E3

46:30 · 20 May 2025 – 23 May 2025 · YouTube

About this talk

This talk explores the evolution of AI image generation tools, focusing on MidJourney and Adobe Firefly, with a brief mention of DALL·E. The speaker, Bobby, a senior business analyst and AI product manager, provides insights into how AI has transformed image rendering over the past two years. He explains the differentiation between various models, including autoencoders, convolutional neural networks, generative adversarial networks, and diffusion models, highlighting their roles in creating photorealistic images. The session also showcases live demonstrations of these tools, offering real-time image generation and comparisons of outputs across different versions of MidJourney and Adobe Firefly. Bobby emphasizes advancements in diversity representation within generated images, and discusses potential ethical implications concerning the use of AI in art.

Full transcript

[Music] Hello everybody welcome to the last talk before the lunch break and it looks like a very interesting talk which will be from Bobby who is a senior business analyst and consultant in Mentor at Mentor mate and his talk will be AI image generation visual artist with mid Journey adob Firefly and doll E3 so Bobby if you ready I'll leave the sage to you thank you so

much hi everyone I'm very glad to be here and to present today the topic of visual Artistry with mid journey and Adobe Firefly mainly we'll also explore a bit about um DOL um but before we start um I will introduce myself in more details um I'm a senior business analyst at Mentor mate but I'm uh also a AI product manager I've been working in the AI field

since 2018 so about six years more than six years now uh my background is in software development I have a few years of experience as a project and product manager and then I enter the AI field as a consultant and a product manager nowadays I'm also doing a PhD on uh artificial intelligence specifically on synthetic data simulation models and digital Twins and the goal I have for

this session today is I hope to give you a good understanding of how AI image generation models work what their evolution is so we'll explore some um we we are going to see how in the past two years or so things changed and evolved and it's very interesting to uh see the differences in the models from two years till now from two years ago and we also

uh uh do a comparison on the um different tools as I said mainly mid Journey which is my tool of choice but we also see some um comparison with Adobe Firefly and uh at the end during the live demo we can also check um DL if we have time in the meantime if you have any questions feel free to um put them in the chat and towards

the end of the presentation I'll try to answer as many as I can as as many as time allows of course um the best thing is in addition to questions I would like to invite you throughout the presentation to give me ideas on specific images that you want to see generate life during the demo at the end I will show you some images that I have generated

I will generate some images live in front of you but I'm also inviting you to um give me your ideas so if there's time I will choose a few ideas from the chat and uh we'll see them generate life okay as a starting point I want to make the disclaimer that this is not a technical Workshop what I mean by this is that I'm not going to

be explaining the technical details of how image generation models are trained or how they actually work but I will be explaining more um the UI of the tools that uh that that that I already mentioned so you can as quickly as possible start generating images yourself if you haven't done so already and throughout the presentation you'll see some images and you will see um yeah there will

be a lot of visual stuff obviously if you see anything bothersome don't blame it on me the models are not mine and I um it's not I'm just showing you here and it's all purely for educational purposes um okay let's start this first we're going to explore very little of the theoretical background so I hope I don't lose you here um it's a just a bit of

theory just to show you how um this field has evolved in the past decade or two so even though image generation is something that we have seen a boom of uh in the past couple of really in 2022 there was so much discussion around um possibilities and the ethical complications and considerations of uh generating images and the discussion keeps ongoing till today uh but this is nothing

new we just reached a point where the images look photo realistic and we're going to see this in a bit the evolution but it's it's nothing new because researchers and experts have been trying to get here for quite some time and here to the left we see an image uh from a scientific paper from 20 uh from 2006 describing AO encoders and um Auto encoders were one

of the first algorithms that were designed towards and used towards images um so let me explain in a simple way I'm going to try to oversimplify this so uh even folks who are not with who don't have technical background can hopefully understand imagine the auto encoders as a magical organizer what does this mean you have a closet for for example full withd clothes old things you know

it's a mess it's it's like this and even though it's a big mess and chaotic you have a magical organizer who can who um every time you put something in the closet they know what exactly that thing is and how it looks like and where its exact position is and it can immediately even if you reassemble the uh the closet the magical organizer will be able to

assemble it again in its original um in its original state so even though it was a mess it knows how to put it back together so think so this metaphor now try to think about it within an image an image can be a mess of pixels a mess of colors shapes shadows and all to encoder all to encoders try to um uh encode and then decode there's

a encoder and a decoder they try to to uh encode and decode the image trying to recreate a fuzzy image into its original state so hope this makes sense if it doesn't it's it doesn't matter that much for the rest of the presentation but just keep in mind the important thing is it's nothing that new then CNN's came around um a bit later I think uh still

around 2010 we had some CNN 2012 even um CNN's um the abbreviation stands for convolutional neural networks and this is an image of a paper from 2015 and those of you who are who have been following the field might remember that in 2015 um there was this think um a big discussion on the so-called deep dreams AI dreams um it's because scientists from uh um I think

MIT yeah they used the convolutional neural network to generate completely new images from scratch so Auto en colders were trying to recreate an existing image here with CNN scientists tried to create something totally new and the results were horrific to say to say it mildly um as you can see the the created images weren't just an a proper photorealistic image not even an artistic rendition they were

images consisting of images so if we look a bit uh closer in the details here we see that uh um they were there are images constructing the bigger image so here for example we have this small doors that build up the ones so yeah it's as I said it's very abstract very horrific we don't need to dwell too much into it too much on it but yeah

it was weird but this was the first time that we saw something generated completely from scratch okay and then we have guns guns were a big thing they they still are guns stand for generative um adversarial Network and the gun to oh uh wait I forgot to give you the um wait I forgot to give you the Met metaphor here how to remember CNN sorry about that

just forget about guns for a second CNN remember them as an assembly line approach so think of a factory and you have an assembly line and the CNN basically works by identifying small parts of an image so um or even generating just very small parts so it can uh one one station on the assembly line can specialize in colors another can specialize in small shapes or shadows

and a third one can um specialize in the bigger silhouettes of the object that we are trying to illustrate so think of it like this it it works um on small chunks of the image and each station um expands on the on the work of the previous ones okay now I'm going to talk about guns and why this is important um guns were even though as uh

machine learning algorithms they're not new they were one of the first uh to be used to train models that can get to something useful so the images generated by guns weren't that photo realistic but guns are very good at um expanding upon existing images so here in this example from uh from this paper we can see how a gun is able to paint a drawing so here

we have a sketch of a cat and it paints it to look a bit more real here we have a sketch of a woman's face and the gun is able to turn it into a paint looking a painting or at least something resembling one but the the biggest um the most popular applications of guns were recoloring so look at this person and uh the background behind him

or here we have I think this is a Dell it cuts away the background right so um this was a very useful application a few years ago and to to understand how guns work and try to remember them think of an Art Forger and an art detective so guns consist of two networks and in the metaphor I'm going you imagine The Art Forger tries to come up

with the best fake possible while the art Detective is um trying to detect all the fake images generated by The Art Forger so um they are in a continuous competition with each other so The Art Forger is trying to fool the detective and the detective is trying to catch the forger um in their steps and this competing works or the forger and the detective compete until they

reach a state of perfection where most of the time the forger is able to consistently fool the detective what's real and what's fake uh or it could also be the other way around when the detective is always consistently uh figuring out when an image is fake even if it looks real so that's how we have some applications but they don't work that well some applications for detecting

uh problems in anomalies okay and finally we get to diffusion models and that's that's what allowed us to have this explosion of phot realistic generated images the fusion models are they are also not that new here here here's a diagram explaining how it how portion of it works from uh paper dating in from 2020 so diffusion models what they do is uh remember all how oops how

how they have encoder and decoder so they first um blur the image and then they try to get it back to uh something resembling the first one as much as possible the initial one well the diffusion model does basically a very similar job but much much better and instead of just blurring the image a bit they completely destroy it so they uh there's um um there's a

process of introducing white noise to an image and then the diffusion model tries to recreate a new image by gradually slowly removing the noise so here we have um this the noising process where we start with a white noise blurred image and we slowly start removing the white noise and trying to shape an image underneath that's how the diffusion works okay this was all the theory I

promise from now on I'm going to only show you cool generated images um so I hope I didn't lose you yet so recent developer developments let's see how how things have changed recently so there was a reason why I went through this um journey of showing you the various models that have emerged in the past couple of decades because most of the image generation applications that we

see nowadays are based on diffusion models and here's a quick example of how a diffusion model Works in real life right this is this is mid Journey so we start with the Blurred image here it's not even a um from the beginning but this is around 10% of the work of the mod on the model then uh if we want to generate this if we want to

get to this final result these are the steps that it goes through and in mid Journey you can uh see the different steps of the model by using the command dash dash stop and then giving it a percentage number so these are the um these are 10 stages of that process and as you can see if it's initially just removing the blurriness and uh shaping the overall

silhouette and then it starts adding more and more details okay now let's look at the uh evolution of the model this is ugly I know but I decided to show you three examples so we're going to see three main examples of how the different algorithms evolved in the past couple of years so the first example is an image of a person I wanted to generate an elderly

French woman with deep wrinkles and a warm smile sitting in a Charming Soho Cafe filled with plants looking out the window wearing a bright pastel linen laser and floral silk blouse so mid Journey version one dating March 2022 so more than two years now was able to generate this ugly thing doesn't even look like a proper human silhouette mid Journey version 2 just a month later was

able to more accurately produce a human silhouette but the details are horrific I'm going to change it it's ugly version three in July 2022 was a bit better it managed to actually add some more details uh in the background here well that's not really the background but and the plant and her um clothing but still the face is terrible and then this is when the boom started

when all the news CH Channel started uh showing AI generated images and talking about um how things are going either really good or really bad depending on channel but this is when things started really drawing people's attention and this is when I also started using mid journey in November 202 to when things went from this to this and you can see the difference it's not perfect but

it's really good especially compared to the old version this looks like an oil painting kind of uh there's a bit of a problem here with her glasses maybe a bit in on the eyes but it's it's good and it's for the first time it's actually able to follow the prompt so she does have bright pastel linen Blazer and floral blouse um and she's looking out the window

right there are plant around her for the first time we have something that actually follows it so this was November 22 big jump right get ready for jump this is mid Journey version 5 dating one year ago March 2023 more than a year ago and here we are now starting to talk about actual photo realism uh again it's not perfect the anatomy is a bit weird in

some places but honestly if you don't know that this is a AI generated image and you see it especially in small scale like a thumbnail most people won't even think twice that that's a real image big jump right in the quality and at the in the same at the same time March 2023 Adobe Firefly released its first version of um of its model for generating images now

I need to talk just for a bit on the differences between these tools mid journey is a a proprietary software and you need a subscription to use it Firefly is Adobe property uh and it's been trained completely on anobi on Adobe um Adobe stock images so mid Journey's data set is for training is secret we don't know what images they've been using they say it's been only

public they've only used publicly available images there were some scandals there but um Adobe at least we know that they're using only the images that they own so they own Adobe stock images me Journey requires a subscription to use it's not super expensive but um for $10 a month I think you can uh you can use it a lot while Adobe Firefly is still free in the

web version they initially said it's going to be free just for the beta but it's still free I I I believe yeah in October 2023 we got Adobe Firefly version two which finally got to the level that we wanted here in the first version it wasn't that good the face looks weird the blouse and the Blazer are reversed in um style of how I describe them but

here they're in the right style following the prompt so much better also I want to point out the difference in the diversity here Adobe Firefly from the get go pushed a lot for using or generating people who are not just plain um just normal common white images of white people which is the majority of what you will find on the internet most probably so Adobe Firefly tried

to push for higher level of diversity and we'll see some examples of that in a bit mid Journey version six this is the newest uh officially released in February 2024 got to a whole new level the level of details in the background is exquisite and honestly at this point we can barely find a reason to think this is not um this is not a real image the

newest version of Adobe version three dating April 24 April 2024 get a bit more artist art a more artistic feel and style it's not that F realistic I would say but it's good the level of details amazing okay let's observe the same journey of evolution on on the second example which is a more simple prompt a lake in the mountain beautiful landscape bright sky mid Journey version

one try to get there the water looks weird the sky looks okayish version two the sky looks better the trees look okayish the water doesn't look that good version three they try to create to recreate reflection in the water but it doesn't look real and then version four this big jump remember version 4 from November 2022 um very artisty uh artisty looks like an oil painting and

then version five looks like an mobile phone camera image very good level of detail on in the sky the Reflection In The Water the trees uh Adobe Firefly the the first version the the sun looks a bit weird not that good version two much better and the grass is good the the water still not perfect though mid Journey version six the latest one looks like a DSLR

camera image at this point the lightning the the Shadows it's it's just great and the latest version of of adobe Firefly again as I said a bit more a bit more artisty style more Arty it's not as if taken by a real camera it looks edited the colors are a bit too too yeah too contrast contrasty and for the third example I decided to show you uh

an image of an object and I went for a burger so for those of you who don't eat meat or are against it please bear with me so here I wanted to have an image of a burger with egg tomatoes green salad ketchup pickles um proper bread and juicy meat Etc and there's a bowl of french fries next to the Burger that's the that's the prom that

I gave it mid Journey version one and two that don't look like anything that you would want to eat looks like plastic honestly version three looks again like something you don't want to eat but at least at least it starts to look like spoiled food or something version four the big jumping quality but very much like an oil painting starts to look okayish version five it's like

it came out of a burger commercial very good and it's properly the the the prompt mid Journey version one didn't look that good version two better but still not great latest mid Journey version very photo realistic and it kind of makes me hungry I don't know about you and the latest uh Adobe Firefly looks quite good again the lightning is not as dramatic as with mid Journey

but looks good okay now let's very quickly look at some of the more Niche model evolution that we have observed probably you have all heard about um the problem of image Generation Um one of the big problems and that's generating hands so AI struggles a lot with hands so here I Tred to illustrate the evolution mid Journey version one and two I'm not going to even give

you an example because it's pointless as you can see version three look doesn't look like anything it looks like a blop of human flash it's terrible version four adds a bunch of fingers to each hand so we know that occasionally AI can generate um images of hands with additional fingers well this is why we know that like this this is when the memes started it just adds

like two three four fingers to each hand fingers version five not great we have some weird stuff happen here but much better the first version of Adobe Firefly uh I want to again draw your attention to the diversity so when I am asking mid Journey it's all hands of male white man here we have a woman's hand or a female looking hand and a hand of a

person of a color a person of color so big push for diversity by by Firefly Firefly version two went even further where we have a male looking hand with polish uh but still extra fingers are present not great and finally mid mid one doesn't add extra fingers occasionally you can still see an image with extra fingers but it's uh mostly intact with five fingers the late T

Adobe also has the right amount of fingers and again we see a lot of a lot more diverse um level a higher level of diversity between the people that we see because I haven't specified I just said two Business Leaders right mid Journey just assumes that two Business Leaders means white men Adobe Firefly pushes for this diversity so um an interesting take on this we're going to

see more in a second faces faces was are also something that uh the the models struggled a lot and this is why look at Mid Journey version three when I ask it to give me a closeup of a group of adults with clearly visible faces and eyes it's not even humans I don't know what that is it's some weird wooden thing I guess I don't know version

four very artisty looks a bit weird but it's it's good at least it human version five the the second big jumping quality it gives us actual human faces um Adobe Firefly again look at the diversity even though here we have um again more diversity with mid Journey version five but here it's um whole new level and I love it Adobe Firefly version two generated this uh music

band kind of looking people and also by the way not exaggerating with the exact same prompt in mid Journey version two I got this image so a group of adults with clearly visible faces and eyes sometimes produced dogs so yeah funny version six finally we get to a very diverse level uh a very good level of diversity not only on the color um the gender but also

finally in age so they're finally not the same age we see some more elderly people here and Firefly uh this is version three this is a mistake um yeah that looks creepy as I said the latest Firefly version is more um Artistry so very quickly because we are running out of time and I want to show you real life uh demo we have just 10 minutes left

I'm going to quickly go through some some additional features of uh these tools so we can zoom out if this is the originally generated image we can say expand upon it or we can even give it an a real image not AI generated and just ask it to expand it so this is like zooming out of the image and it can zoom out a lot so here

zoom out one 1.5 and two so it adds extra details extra people in this case even and it can also describe images so we can go the other way around instead of generating images we can ask it to describe the images for us uh let's see this is the image and this is the original prompt and when I ask me journey to describe it this is what

it me so um yeah this is the description of image it's kind of okay uh it adds a bit of extra description that I wouldn't think of putting myself there like politics and yeah specific uh style of clothing and when you ask it to generate an image with the description it provided it gives me this so not too far it's kind of the same style so it's

very good with the description and what we can use this process uh for is have an image describe it with the model and then use the same description to generate more similar images okay we can also generate images in different as aspect ratios or we can say uh we can tell the algorithm to exclude specific uh specific parts and not generate something uh we can also create

images that can be repeatable in vertical or horizontal uh line and um this is great for tapestry for example so next time you're read redecorating your home you can use this for your wall decoration and we can also see the process of generating it remember the yeah we can also in paint what does this mean well we have this image and we can say change the path

to a river or we can select this part here and say give me a hot air balloon and it generates this it doesn't touch the rest of the image just the selected area or a castle and it gives you this we can also specify specific things how they should be separate from each other within the image so if we say a space ship it will generate a

Sci-Fi ship if we say space colon colon ship it will generate a wooden medieval ship in space so it looks at the two things separately now industry applications this is an example of a goblin Warrior that I decided to generate for game or concept design because I've spoken with some game game development Studios or designer um Studios and they all say that this saves them a lot

of time those that are using such models one example that I heard was a client comes and asks for 100 goblins two years ago we would have to that that's what they say we would have to draw the Goblins by hand a 100 of them now it takes an hour or two to generate them with with the AI and just fix them a bit and that's just

for con concept art right for prototyping we can use the same for um for example Hightech City or even for a logo so next time you're developing a new mobile application for your Prototype at least before you get to a point where you can afford hiring a real designer a human designer you can use this and the logo looks great I asked you to give me a

logo of a mobile fantasy RPG game or we can even generate characters for a commercial like this one for Coca-Cola that I generated we can even use it for um ideas on decoration interior design or fashion design this dress totally generated can give us ideas or even product design look at this case for an iPhone or Furniture Design just look at the level of details on this

wardrobe amazing we can even generate complex images for I don't know for books for covers for um ideas for even movies or stories look at this it's a very complicated prompt but the image is just perfect it's just it's as if it came out of an Japanese anime very often I get the the question when I talk about image generation whether it can be used for web

design it can ugly the the layout is okay but it's not like a designer would say that it can be better aligned probably the Ping here and there could be better and if we ignore the images that they just look disjointed and ugly yeah it can be but if only if you're stuck and now the more interesting things of late is the character reference until recently we

weren't able to reuse the same character and generate multiple images of the same character why is this important so think of generating a concept of a comic book we couldn't have the same characters appear now we can and here's an here's an example I generated these images of of a robot original prompt is this and then I said give me the same robot or similar robot painting

or on the International Space Station so it's possible now very quickly I'm going to go through the ethical concerns wow time flies really quickly Um this can be used for fake news as you can imagine and it is already being used for fake news there was a scandal uh last year where images of I think last year where images of Donald Trump were circling around where fake

images that is that depicted him being arrested and it was all fake so in a world with such photorealistic um images be aware that a lot of what we see might not be true and finally demo time we have just a couple of minutes I'm going to very quickly show you something cool and uh yeah I don't see any questions but if you have questions now is

the time to shoot them this is the this is the how mid Journey UI works this is basically Discord for those of you who are who have used Discord this will be familiar Discord is a separate program in which you use the mid Journey bot and to use it we just type IM imagine and here we use a prompt I'm going to just copy paste an existing

prompt that I have just so you can see how it works in the last couple of minutes um so I'm asking it to give me an IM of of Batman and the comic style and we can see how it works and uh why is this important it can generate things that are completely new here I generated an image of a dock as you can see it's uh

a Golden Retriever with Majestic fur and it looks cute great Dogo but we can also generate things uh popular pop culture so like this comic book style Batman and it even gives us a female Batman why not I haven't specified that it has to be male so yeah I'm going to generate something else now just for for this remember the cool anime image let's see if it's

going to I think this is the same prompt as uh this uh and um I'm going to also put another image to be generated in the last one minute earlier this year I organized an event where we with my friends we went to a dog shelter to walk docks and for that event I use this prompt to generate an image so why not you can you can

use this in your daily life for anything uh I didn't specify the aspect ratio here but uh we can see it's going to be pretty here it is it's like an for the fans of anime in the audience I think you appreciate this uh and the image of docs finally just as a final prompt while we wait for the dog image going to use this prompt to

generate a really cool image of a model you're going to see it in a second I see a question what about ethical stuff concerning AI stealing artists art it this a concern yeah well my take very quickly in the last 30 seconds is that AI is just you are using AI to teach the model uh to to um to use images if I'm looking at an artist

and I learn from the artist style and learn to do some something on my own is this stealing yeah there's no definition thank you connect with me on LinkedIn this is the QR code thank you so much Bobby I let you use your time until the very end because we had only one question but now we are really at the end of our time but it was

really very interesting it was exciting to see from starting from the early versions of all these tools and at that point we have arrived right now thanks a lot for this interesting talk and enjoy the rest of the conference thank you see you see you

From event

DevDays Europe 2025

20 May 2025 – 23 May 2025

All event videos
Back to Watch