Google Gemini has just announced a brand new image generation model that may have caught them up to open.
AI maybe surpassed them and a number of other players in the industry.
So today on the podcast we'll be diving into basically what the new capabilities are on this image model, because there's a whole bunch of things that image models have struggled to do that I think they've done exceptionally well.
And there's some areas that I think are a lot of rooms for improvement on Google.
We'll be breaking down all of that on the podcast. today.
And before we get into it, I just wanted to mention if you want to see a bunch of posts I've been making about this or anything else I'd love for you to go check out and follow.
Make sure to follow me over on x, where I post basically a bunch of interesting things that I have discovered here and post a lot of stuff about AI.
My handle is Jaden underscore AI.
And I will leave a link in the description or in the show notes or whatever.
But yeah, make sure you go follow me on X. I would love to connect with you there.
All right.
So let's get into what Google has announced.
The biggest thing, I think, that I've been the most excited about with all of their kind of updates in this new image model.
The first one is basically that it is capable of actually having consistent image recognition, facial recognition, consistent characters throughout the throughout a number of different images.
So this is pretty cool.
There is a bunch of funny things that i've been able to generate today and i'll i'll give you more on that in a second, but a couple of the big ones, i think, are that basically, you can you can have the same person inside of multiple shots.
So This is something I think ChatGPT kind of struggles with.
We're basically asking it to generate me and you upload an image of yourself after.
Basically, when it puts your face on an image in ChatGPT, it doesn't really look like your face.
And so Google, I think, does a really phenomenal job of actually you can give it a picture of yourself and it will put you in and it will actually look realistic.
So I think that's one thing that's interesting.
It can also make edits to your background and the lighting in the photo, but also basically, you can extrapolate that and be like it's just like a chat interface with the photo, where you just talk to it, you say what you want to edit and it makes the edits.
It does a really good job of this.
It also does really high quality images.
So when I was first testing it out, I actually thought that there was some sort of issue with it.
Because when I looked at the image, it kind of looked low quality inside of the chat.
But I think that's something that they're doing just to preserve, like bandwidth while you're chatting.
If you actually click on an image, I have my 4k monitor.
And if you're watching on on Spotify or over on YouTube, you'll see that if you click on an image on like a big monitor, it will blow it up and it will look really high quality.
So I think overall, it kind of looks grainy at first, but don't be fooled.
It is actually good at generating some actually big images.
Now, some of the things that I'm the most excited about with this whole thing is that when they made the announcement today, they actually have rolled this out to everybody.
So free users are actually getting access to this.
People that, like I, feel like there's so many times where we get these big image or any sort of AI tool launch.
And ChatGPT, for example, I feel like is really notorious.
Well, they'll be like, hey, you know, the new agents is rolling out.
But first it's going to like enterprise users and then the week after that it's gonna roll out to like paid pro users and then after that it's like education and then after that like the free users.
I forget about whatever the feature is.
And I don't get to actually try it because it's not in the news anymore.
It's like a week later.
And I got to go back like months later and figure out like Oh, what was that thing?
They announced where you can connect your Google Calendar blah blah, blah and go mess around with it.
So I think that's kind of bad for adoption.
And Google does this really well, where it seems like right now, they've just rolled this out.
And not only is it, you know, for all free users, but also it is just completely available for everyone, including developers, on the api, google vortex, anywhere that basically anywhere that you can access the google models.
It is currently out, so i think that's really cool.
Another thing that i think is fantastic which uh, our friend over at google, logan ko patrick, has mentioned, is that it was so technically.
It's called gemini 25 flash.
By the way, i know we call it uh, nano banana.
Basically the.
The whole nano banana thing comes from the fact that they put in these like anonymous benchmarking websites where people would try out different image models side by side and pick which image they preferred.
And there's, and so basically the big uh ai companies again like code names, so you don't know what tool you're actually testing or benchmarking.
So nano banana was this mysterious model that was doing really well and kind of beating everyone in benchmarks.
And a lot of times, by the way.
What will happen is someone like OpenAI or other players might put a model in there.
And if it doesn't do good compared to others, they'll just pull it down.
Work on it a little bit, try to make it better and then retry.
So honestly, it's kind of like a testing the market ahead of time.
I think this is a great strategy.
But anyways, there's a model called Nano Banana was doing really good.
No one knew what it was.
Google, everyone at Google kept making like these subtle banana jokes.
And then all of a sudden, we have the launch that 2.5 flash image is out.
This is a great tool.
It's done phenomenal in the benchmarks.
It has really great character consistency, creative edits and of course, it has all of Gemini's world knowledge.
So I think it does a really great job of that.
And it is basically crushing Flex, Quen, and ChatGPT 4.0 in image editing.
Now I think there's also like the fact that mid journey is not put into a lot of these benchmarks because they don't have an API.
So I think that's an important thing to note, which meta, just we know, signed like a really big deal to have mid journey embedded into meta AI.
And I think mid journey for a long time has sort of been the frontier for the best in class image generation model.
But it's really forgotten because they don't have an API.
It's not integrated into any software that a lot of people use the distributions a lot lower.
For a long time it was only on discord.
And of course, now like I think they have their own website, which is great.
But I find myself not going to the website just for that image model.
I kind of like something like that to be embedded into a chat model that I'm also using for other purposes.
So I found myself in the past, I would kind of use like sometimes I'd use the grok image generator.
Mostly nowadays, I'm using opening eyes image generator.
But it feels like Google might actually be giving them a run for their money with this.
Because even with chat GPT, for example, when I need to make thumbnails for like YouTube, I'll have my, I'll have my editors go and actually use like different custom platforms where I upload a whole bunch of pictures of myself and it can clone my face, because I use the OpenAI image generator for thumbnails for a long time and it just wasn't super accurate.
So hopefully this is something that Google Gemini would solve.
And we may actually be able to just use this, for I know like YouTube thumbnails is like funny but like basically you can extrapolate that to all sorts of graphic design where you actually need a consistent character inside of it, a consistent person.
People are doing some really cool, you know, basically a bunch of really cool things.
Some of them are uploading, you know, a picture of themselves, and then they're able to get it to generate themselves in like a hundred different styles.
I think that's great.
It looks very, the character is very consistent.
People also have done things where they'll like upload a YouTube thumbnail and say hey, swap my face onto this YouTube thumbnail.
It does a great job of that, which previously, you know, you pay for software to do.
So I think overall, they're making some really big strides.
And it's quite an impressive thing.
Now, with all of that out of the way, I want to tell you a couple areas that I think that they can improve on a couple funny glitches I found while using the platform.
So while I was testing it out this morning, one of the issues I ran into was basically the.
It just has like a couple like funny things.
So the first thing i did is i got it to generate a 90s grunge style photo of me.
Basically, it was a, it was a prompt in their demo that i clicked on and it did.
It wasn't super flattering, but that's fine.
Um, i then asked it to make me a youtube thumbnail, and one thing i'll say about this is like, technically it's tied to gemini.
So you would assume like basically, because it's tied to gemini, it should be smart, it should know a lot of things.
But when i asked it to make a thumbnail, it made a square image, which thumbnails obviously are like landscape mode.
So I like I'll say that is like one thing that I kind of tested.
It doesn't seem to be opening.
I wouldn't make that mistake.
They would understand and be able to make the thumbnail as a as rectangle.
So that's kind of the first, I guess, strike.
The images are good, but I'll tell you one thing that I found is I actually I give it a funny prompt.
I think my prompt was something like Make a YouTube thumbnail of me being chased by a shark underwater, looking terrified.
And by the way, I'll explain the whole, I'll read the whole prompt.
So, you know, like how robust I gave this, how robust this was.
And I'll give my like grading on how well it did.
But anyways, so a shark looking terrified as a shark.
The shark has the words interest rates on its side.
A boat is above us.
A hook is in the water from a fisherman.
And the hook has a piece of paper attached to it that says, buy now.
Okay.
Just tried to make it like the most elaborate, descriptive thing.
And it actually generated that image perfectly.
There's a boat.
There's a shark.
The shark's got the words on it.
There's a person swimming in the water.
And there's like a by now piece of paper on the hook, whatever.
So basically, it was exactly what I wanted.
But the one thing that was funny is like in the message before, I uploaded a photo of myself.
So it's like upload a photo of yourself.
We'll make your thing.
So I did.
And then in my next follow up, I said, make a thumbnail of me.
And it didn't actually use that photo.
And it didn't make it a picture of me.
It just put like a random... woman swimming in the water.
So I guess, like again it feels like it's not so much even the image generator.
Maybe it's like Gemini wasn't isn't quite doing its, wasn't quite doing its best.
And that being said, I think it was running on Gemini 2.5. flash a quick model but not the smartest model and i probably would get better results by upgraded to 2.5 pro which is you know interesting and maybe that's maybe that's a me problem um because i kept trying to get it to do things and it didn't feel like the model was very smart it would just keep regenerating i said like i uploaded a photo of myself and i'm like okay just like use this person and put this person in the water.
I was trying to make it not the woman.
And then it literally just took my head and stuck it on the woman.
So it's like me wearing a sports bra, swimming in the water.
And I'm like, come on, felt very unflattering.
I told it to change the dimensions.
It didn't change the dimensions.
It made it portrait instead of landscape mode.
Again, it's probably all kind of because I'm using a slower model.
Um, it struggled a lot but then i kind of got to some.
I completely changed it and i had to generate an image of a monkey chasing a person in the jungle.
It did a really good job of that and it looked photorealistic.
And the reason i bring that up is because up until now, the shark one that i was kind of doing with me and the shark in the water it looked kind of cartoony.
And this is like basically my pet peeve of all these ai models is you ask it to generate a photo of something and it looks cartoony and you can't really pass it off for a lot of different use cases, And so I love it when it can make it photo realistic.
Now the thing that I noticed is, if you want it to basically get a photo realistic image, you need to describe a scene that could be normal.
Like my like.
Basically, if I have like a shark with the words interest rate on the side, it's not going to be a realistic image because it's like.
It sounds like something you could imagine from a cartoon.
And it basically then generate an image of a cartoon, even if you tell it to make it photo realistic.
But if you ask it to generate something like a monkey chasing a person in the jungle, that sounds like kind of a photorealistic possibility.
It actually made it photorealistic.
So that's another thing that I think is a little bit tricky is trying to get it to be realistic.
Okay, the next thing I encountered on this whole journey was basically the going against the guidelines.
I was like, hey, like, basically, what is it capable of doing?
What is?
Google has been famous in the past for their image model, kind of doing funny things, like people would ask it to generate an image of World War Two German soldiers.
And it it has, like this diversity, in the past had a diversity equity, inclusion kind of like thing, built in the DEI thing.
So generate like black Nazi soldiers because it didn't want to only generate white people.
But obviously, that is very historically inaccurate.
A lot of people found that offensive for a million reasons.
And so Google, I think, tried to like backtrack and fix it.
Now I was confused, kind of like curious where is google at today on basically what it's able to generate.
So i think i asked it to generate um.
The first thing i asked it to generate was um, a picture of me getting chased by a bunch of camel thieves in the desert, with a nuclear bomb exploding in the background, me holding a gun.
I was like okay, let's make it crazy, right.
And of course it says no.
And i was like okay, probably because of the gun.
So i was like hey, generate me holding a gun.
And it said nope, can't do it.
So I was like, okay, whatever.
Google doesn't like the guns.
That makes sense.
But what was interesting then was that I basically asked it to do a photo of me in the Sahara Desert with a bunch of camel thieves chasing me.
And so I like pulled out the nuclear bomb, pulled out the gun stuff, tried to just make it simple.
And I actually did generate that image.
But what was interesting was it pulled things from the like.
So basically, the prompt above, where it's like, this prompt goes against our guidelines.
It pulled things out of that prompt.
So like it still is reading the whole chat thread when it generates this image, which was kind of like weird to me because I never said anything about a nuclear bomb in this photo.
But in the background of the photo generated, there is a giant nuclear mushroom cloud.
And so that, I think, is I mean basically I think Google probably wants to look into that where it a prompt can be against their guidelines, but in a follow-up prompt it can pull from previous prompts and pull things out of that into the new image.
So anyways, kind of an interesting thing I discovered.
Overall, I think this is a really interesting model.
I tried all sorts of quite crazy things that I posted on X.
I think one of the biggest things that I basically discovered, found out that was kind of uh, shocking to myself, was that um, i was able to get it to generate a picture of, you know uh, camel thieves in the sahara chasing me, but if i asked it to draw, or to you know design, elephant thieves in the savannah chasing me, it wouldn't.
And so i was like either just be consistent is basically my, was my message to Google like either make it so that it doesn't generate any of these two groups of people.
I'm assuming it's because of it's like discrimination clauses or something where, um I don't know, because basically that the camel thieves are all um Arabic people with turbans, holding guns, chasing me.
And so I'm like you can imagine you can see the picture on on Twitter which you can imagine why some people would find that maybe offensive or and personally I actually don't really care what you generate with it.
But I just think it should be consistent because if you ask it to then generate um elephant thieves, which you would imagine would be people from Africa um, it won't do that because probably it thinks maybe it would.
You know it's like discriminatory or bad in some way.
But it's just interesting because it will generate to.
It will generate some people and won't generate other groups of people.
So i'm like, either generate it all or don't generate any of it, just be consistent.
And the reason i have this message uh, to google and i was kind of testing and trying all this is because, as a software developer and someone that um, or as someone that has a software company and we're constantly adding these APIs into our software, it's really tricky to be able to add this stuff without knowing what it is basically capable of and so what it can do and what it can't do.
And so if you have a specific use case where you need to generate something specific, Google doesn't have very clear guidelines of what it actually is capable of doing.
Now, I did ask Google Gemini, like, what are your guidelines?
What can you generate and what can't you generate?
Because basically after it wouldn't generate a picture of a gun for me, I was like okay, that makes sense.
But like, what are these guidelines?
And then this is what it told me.
It said, I'm currently unable to generate images that depict real people.
This includes public figures, celebrities, or private individuals, even with an uploaded image.
Now, this isn't true because in their launch event, they literally tell you to do that.
And I did look like 90 of the photos, a couple that wouldn't let me do because I think of other reasons.
So I think like basically, if you ask it what it's capable of doing, it's not actually being accurate with what it tells you.
After that, it said violent or graphic content.
I'm like, okay, it makes sense.
Sexually explicit material makes sense.
Self-harm makes sense.
Hate symbols or discriminatory content makes sense.
But also I feel like discriminatory content can be like a really big catch-all.
I don't know.
I just feel like somewhere between the violence and discriminatory content, I feel like there's going to be a lot of regular pictures that people might try to generate that, for one reason or another, Google doesn't like.
And basically the way this works is like when you generate an image before it appears to you on the screen.
It goes through like a processor over at Google and it determines is this image, does it have anything harmful or that might offend people, or et cetera, et cetera.
And I think it's very, this is like technology is very common for most other image generators where it's like NSFW content, right?
Any sort of explicit material, it will basically like filter out nudity and anything like that.
And we have like very like this is like a very common AI.
Models are trained to see and understand and know what that is and basically not generate those types of images.
Or like the, the model will generate it either way, but just not give it to people, not show it to people.
But i think that what's tricky that google is trying to do here is because that they have a much bigger kind of guard rails on their image generator.
Where it's like violence or discriminatory content, it's like when you throw up a picture and um maybe, for example right, it's like um, people with turbans chasing me with guns in the desert like oh, that's fine, but like it wouldn't let you do African people chasing me, because that would be discriminatory.
And so it just feels like it's very tricky because there's this AI model that's deciding what's discriminatory and what isn't.
And anyways, so this is obviously a huge can of worms that Google is trying to tackle.
I don't really care what they do.
I just want them to be clear on the rules for it.
And I'm just flagging this as like an issue that it is inside of the model probably should be addressed in one way or another.
So it's very interesting.
But overall I've been really impressed with what everything was basically the quality of the images that Google is generating.
It's not perfect by any means, but I think it's a huge step up, especially from Google's last image image generation model.
So I'm really excited to have them in the arena.
I'm excited to see it honestly kind of go back to back or head to head with meta and what they've been able to get with mid journey, once that comes into the meta apps.
I'll be really curious to see if If basically, they're able to beat out Midjourney.
I feel like Midjourney still has a leg up right now.
So it'll be really interesting to see where this goes.
In any case, thank you so much for tuning into the podcast.
I know it was a bit of a long episode, but I was doing a deep dive.
This is really cool technology.
And I think there's a lot of really hot button kind of controversial things, things happening in this space right now.
So wanted to make sure I covered it all.
Thanks so much for tuning in.
Make sure to go follow me over on X and check out the AI box platform.
If you want to try out all the models I talk about on the show, AI box.ai.
I'll catch you on the next episode.
Have a nice day.