Skip to content
GenLovers

Nano Banana: Google's image AI, put to the test

Last updated:

Written by Clement

Our verdict

The model to reach for when you bring it an image and tell it what to change, and the one that took every check in our showdown without a single partial.

Reach for it when

  • Editing a real photo by instruction while faces and everything unmentioned stay put
  • Keeping one character consistent across a series of generations
  • Compositing several reference images (person, garment, location) into one scene
  • Anyone who wants a genuinely usable free tier rather than a trial

Skip it if

  • You want from-scratch stylized hero art: Midjourney still wins on raw aesthetics
  • Your prompt sits near Google's safety line, which refuses real people and catches benign prompts with them
  • You need many successive edit passes on one image, where artifacts accumulate

At a glance

Made by
Google DeepMind, the image model of the Gemini family
Best at
Conversational photo editing and character consistency
Access
Free in the Gemini app and AI Studio, paid via the Gemini API
Free tier
Daily generation cap, visible watermark plus SynthID
API cost
A few cents per image
Our showdown
16 of 16 checks passed across both published tiers
Nano Banana output: a red-haired woman in a GenLovers-branded t-shirt and a man in a t-shirt laughing forehead-to-forehead in a kitchen, orange juice and coffee on the counter.
Nano Banana, prompt 1 of our showdown. The logo stays legible where the fabric folds, which is the check this prompt exists to make.

Nano Banana started as the anonymous codename of a mystery model that topped blind image-editing leaderboards; when it turned out to be Google's Gemini image model, the nickname had already stuck, and Google kept it.

That makes it a different animal from Midjourney-style art generators. Nano Banana is at its best when you bring it an image and tell it what to change. This page covers what it does well, what the free access really includes, where it falls short, and when a different tool serves you better.

Quick facts

Made byGoogle DeepMind, the image model of the Gemini family
What it doesText-to-image plus conversational image editing: restyling, combining images, consistent characters across generations
AccessFree inside the Gemini app/site (daily limits), free inside Google AI Studio, and paid through the Gemini API for developers
PricingFree tier with generation limits and watermarks; higher limits with Google AI subscriptions; per-image API pricing
VersionsThe current version is 2, and there is always a Pro version that runs a bit better
Best forEditing real photos, product shots, keeping one character consistent across many images, colour accuracy, and writing legible text

What Nano Banana does best

Editing is the superpower. Give it a photo and an instruction: swap the background, change the lighting to golden hour, put the subject in a leather jacket. It makes the change while preserving what you didn't mention, especially faces, though it is often better to spell out what you do not want changed. That identity preservation is what pushed it up the leaderboards.

It follows instructions literally and treats them as a spec: object counts, spatial layout, and multi-step edit requests come out the way you wrote them.

It also composites: hand it several reference images (a person, a garment, a location) and ask for one scene combining them. For e-commerce mockups and social content, that replaces a real workflow, not just a toy.

What the free tier really gets you

The Gemini app and Google AI Studio both include Nano Banana at no cost with a daily generation cap, which resets and is usable for learning and test work: one of the most generous free tiers among the big-name tools. Free outputs carry both a visible watermark and Google's invisible SynthID marking; paid Google AI subscriptions raise the limits and remove the visible mark on higher tiers.

Developers get the same model through the Gemini API at a few cents per image, cheap enough to prototype real products against.

Limitations to know before you commit

Pure from-scratch artistry is not the strength. Ask for a moody cinematic illustration with no reference and the result is competent but rarely as striking as Midjourney's take on the same prompt. Outputs can drift toward a smooth, slightly plastic sheen, particularly with skin.

Moderation is conservative. Google's safety filters refuse realistic depictions of real people and anything suggestive, and some benign prompts get caught in the blast radius. Version 2 does read as more permissive than Alibaba's model, which is remarkable next to where both sat a year ago. If a prompt keeps bouncing, rephrase around the words that are tripping it.

Heavy repeated edits degrade the quality. Each successive edit pass re-renders the image, and after many rounds small artifacts accumulate. I pushed one product photo through six consecutive edit passes to see exactly where it broke: by pass five, background texture had visibly softened and a shadow direction had drifted. I now cap myself at two or three edits per image and start a fresh generation for anything beyond that.

How Nano Banana compares

Against Midjourney: Nano Banana wins on obedience, editing, text rendering, and price; Midjourney still wins on raw aesthetics and style control for anyone who wants that specific look. We moved our own pipeline fully onto Nano Banana and dropped Midjourney once instruction-following closed enough of the aesthetic gap to not justify a second subscription.

Against open-weight local models: local Stable Diffusion or Flux setups offer total control and no content policy, but need a GPU and hours of setup when you are new to it. Nano Banana is the zero-setup convenience pick with a real free tier. The trade is Google's rules and Google's watermarks.

The model showdown

Every image model on this site runs the same six prompts across three tiers (SFW, erotic and explicit) on the same rig, with no cherry-picking and no retouching. Below is what Nano Banana returned, graded against a fixed rubric so the results mean the same thing from one model page to the next.

Run on
Direct Gemini API, models/gemini-3.1-flash-lite-image
Date
Why these six prompts? ▾

They are chosen to break things, across three tiers so a model that refuses NSFW outright still gets a fair, comparable test. The SFW tier stacks two-person contact, readable logo text on folded fabric and genuine expressions with teeth. The erotic tier stacks the four failure modes that mark an image as machine-made at a glance: hands, fine text on skin, a mirror that has to obey geometry, and three distinct character designs in one frame. The explicit tier exists to test the same anatomical limits at full intensity, but its output is never published on this site.

No model passes every check, and not every model can attempt every tier. That is the point. A rubric everything passes tells you nothing about which model to reach for on a Tuesday afternoon.

SFW detail test

Two people, bright kitchen

Two subjects interacting in hard morning light, with a brand logo on fabric that has to stay readable while the fabric folds. Tests skin-on-skin contact, genuine expressions with teeth, and text rendering on a non-flat surface. Fully clothed throughout.

Nano Banana output: a red-haired woman in a GenLovers-branded t-shirt and a man in a t-shirt laughing forehead-to-forehead in a kitchen, orange juice and coffee on the counter.
Unretouched output. No inpainting, no upscaler, first usable seed.
Scorecard
  • Pass: Both faces free of uncanny artifacts
  • Pass: Logo text legible on folded fabric
  • Pass: Teeth and open-mouth expressions natural
  • Pass: Contact points between subjects anatomically sane
What we saw

The cleanest render of this prompt across every model we tested. Logo, faces, teeth, and contact points are all correct with nothing to nitpick, and unlike three of the five other models we ran this exact prompt against, he's wearing a shirt.

Goth bride, cathedral

A second logo-tattoo placement rendered larger and at a different body location than Erotic Prompt 1, colored contact lenses, dental/prosthetic detail (fangs), an asymmetric wink held against a wide-open other eye, and two coordinated hand gestures in one frame. Fully clothed in a bridal gown.

Nano Banana output: a pale gothic bride with black hair in a lace wedding dress, one hand pulling her lip to show fangs, the other in a peace sign, purple eyes, a chest tattoo reading GenLovers.
Unretouched output. No inpainting, no upscaler, first usable seed.
Scorecard
  • Pass: Chest tattoo text legible and matches the brand wordmark/gradient
  • Pass: Purple iris color consistent across both eyes
  • Pass: Fangs rendered as coherent dental anatomy, not a texture smear
  • Pass: Both hands' gestures anatomically distinct and correct
What we saw

Both hand gestures present and correct, fangs read as real teeth, the wink-plus-open-eye combination is right, and the chest tattoo is legible with the correct gradient. Tied with Z-Image Turbo for the best result on this prompt in the whole batch.

Erotic tier

Solo portrait, beach

Skin under warm golden-hour light, bikini fabric texture, hand and ring anatomy, fine tattoo linework, and a mirror reflection that has to agree with the subject. The single hardest frame of the set.

Nano Banana output: a woman in a dark green bikini on all fours on a beach, tongue out with a drop of drool, a forehead tattoo, an ornate mirror behind her reflecting her from behind.
Unretouched output. No inpainting, no upscaler, first usable seed.
Scorecard
  • Pass: Hands and fingers anatomically correct
  • Pass: Fine tattoo linework legible
  • Pass: Mirror reflection consistent with the subject
  • Pass: Bikini fabric texture holds up close
What we saw

This is the update to make to the not-supported assumption we'd been carrying for this model's NSFW tier: it isn't blocked. The tongue-out, drooling, cross-eyed expression came through close to exactly as written, the forehead mark and hand-with-rings detail both hold up, and the mirror correctly doubles her silhouette. Best erotic-tier result in the batch.

Three characters, anime, night

Stylized rendering with three distinct character designs in one frame, water and reflections, bare-foot anatomy, and two animals. Multi-subject scenes are where most image models quietly give up.

Nano Banana output in anime style: three women in pastel robes at the edge of a rooftop pool at night, one hand on an inner thigh and a cheek kiss, a white cat and a dark dog nearby.
Unretouched output. No inpainting, no upscaler, first usable seed.
Scorecard
  • Pass: Three distinct, non-cloned character designs
  • Pass: Feet and toes rendered correctly
  • Pass: Water surface and reflections coherent
  • Pass: Both animals recognizable and correctly anatomized
What we saw

Three distinct designs, correct single cat and single dog, good feet and water detail, and the specific hand-on-inner-thigh-plus-cheek-kiss gesture between two of the three women, the one detail almost every other model in this batch dropped or generalized into a group huddle.

Explicit tier

Not supported by this model

Same API-side content filter blocks this tier outright, distinct from the erotic tier which the filter does allow.

Frequently asked questions

Is Nano Banana free?
Yes, within limits. It's included in the free Gemini app with a daily generation cap and watermarked output. Paid Google AI plans raise the caps, and developers can use the model via API at per-image rates of a few cents.
Why is it called Nano Banana?
It was the codename the model competed under, anonymously, on blind image-comparison leaderboards. It kept topping the charts before anyone knew it was Google's, the nickname went viral, and Google adopted it as the public name.
What is Nano Banana best at?
Editing images by instruction: changing one element of a photo while keeping faces and everything else intact, combining several reference images into one scene, and keeping a character consistent across a series of generations.
Was this helpful?

Hands-on guides

Related models

Get new guides by email

One email when we publish new guides and model breakdowns. No spam, unsubscribe anytime.