Google has launched Nano Banana 2.1, its latest image generation and editing model, bringing improvements to visual design, text rendering, mask-based editing and subject consistency. The model is positioned as a more efficient successor to Nano Banana 2 and is based on Google’s Gemini 3.6 Flash architecture.
Updated 7 October 2026: Google formally launched Nano Banana 2.1 on 6 October, following earlier reports of a model name appearing in Flow. The specifications and paid API prices below come from Google’s model card, developer documentation and live pricing table. Rollout is described by Google as gradual, so availability may vary by product and account.
The new model is being integrated across Google’s AI ecosystem, including the Gemini app, Google Search’s AI Mode, Google AI Studio, Google Flow, Google Stitch, Google Ads and the Gemini Enterprise Agent Platform. For developers, Nano Banana 2.1 is also available through the Gemini API, with support for 1K, 2K and 4K image generation.
Key takeaways
- Google has launched Nano Banana 2.1 as an update to Nano Banana 2.
- The model is based on Gemini 3.6 Flash.
- Google says it improves visual design, image editing and subject consistency.
- It can use up to 14 reference images in a single prompt.
- Google says users can maintain consistency for up to four characters and 10 objects.
- The model supports 1K, 2K and 4K image generation.
- It adds stronger text rendering and infographic capabilities.
- Search grounding allows the model to use Google Web and Image Search.
- Google has introduced configurable Thinking levels.
- API image-generation costs are lower than Nano Banana 2 at comparable resolutions.
What is Nano Banana 2.1?
Nano Banana 2.1 is the latest member of Google’s Gemini image-generation family. While the Nano Banana name became popular as a shorthand for Google’s image-generation technology, the underlying system is part of the Gemini model family.
Google describes Nano Banana 2.1 as a high-efficiency image generation and conversational editing model. Unlike a traditional image generator that primarily turns a text prompt into a new picture, the model is designed for iterative workflows in which users can provide images, request changes and continue refining the result.
The model is based on Gemini 3.6 Flash. Google’s model documentation lists text and images as core inputs and image and text as outputs, while the Gemini API documentation supports text, image, video and PDF inputs for the model. This gives Nano Banana 2.1 a broader working context for image-editing tasks than a simple text-to-image system.
The important change is therefore not simply that Google has released another image generator. Nano Banana 2.1 is aimed at making AI-generated visuals more controllable and useful for repeated production work.
Better editing is the main upgrade
One of the biggest improvements is image editing.
AI image generators have traditionally faced a difficult problem: changing one part of an image without unintentionally changing everything else. A request such as changing a person’s clothes, replacing a product’s background or modifying a particular object can cause the model to alter facial features, proportions, lighting or unrelated elements.
Nano Banana 2.1 is designed to improve this type of workflow through stronger mask-based and multi-turn editing.
Mask-based editing allows users to identify a particular area of an image and instruct the model to modify that region. This can be useful for replacing backgrounds, changing individual objects or making localized design adjustments without rebuilding the entire image.
Google’s evaluation data also shows an improvement in its mask and ink-based editing benchmark compared with earlier Nano Banana models.
For professional users, this matters because image generation increasingly resembles an editing workflow rather than a one-shot generation process.
Subject consistency gets a major focus
Another major upgrade is subject consistency.
Keeping the same person, character or product recognizable across multiple generated images is one of the hardest problems in generative image technology. A model may produce a good first image but change facial features, clothing details, product shapes or other characteristics when asked to create another version.
Nano Banana 2.1 is designed to reduce that problem.
Google says the model can process up to 14 reference images in a single prompt. Its documentation also highlights consistency for up to four characters and object fidelity for up to 10 objects.
This opens the door to more complex workflows involving multiple people and products.
For example, a marketing team could provide several reference images of a product and ask the model to place that product in different environments. A creative team could provide several character references and produce multiple scenes while attempting to preserve the identities and visual characteristics of those characters.
The model is not perfect, however. Google’s own model card acknowledges that character consistency can still fail in some cases, meaning users should treat the capability as an improvement rather than a guarantee.
Better text and infographic generation
Text rendering has become another important battleground for AI image generators.
Older image models often struggled to place readable words inside generated images. Letters could be misspelled, typography could become distorted, and longer text could turn into visual noise.
Nano Banana 2.1 is designed to improve this area, particularly for posters, diagrams, infographics and other graphics where the generated image itself contains information.
Google’s model evaluation includes infographic design and infographic factuality tests. The company reports substantially higher scores for Nano Banana 2.1 than the previous Nano Banana generation on these internal evaluations.
This could make the model more useful for marketing materials, presentations, social-media graphics, product concepts and explanatory visuals.
However, Google’s own model card says small text can still become blurry and that long paragraphs remain a limitation. The improvement therefore should not be interpreted as meaning AI-generated typography is now universally error-free.
4K image generation and wider formats
Nano Banana 2.1 supports image generation at 1K, 2K and 4K resolutions.
The model also supports unusually wide and tall aspect ratios, including formats such as 1:4, 4:1, 1:8 and 8:1. Google says it has addressed tiling artifacts that could occur with wide and panoramic images at higher resolutions.
That is particularly relevant for banners, website layouts, advertising creatives and other applications that do not use conventional square or portrait formats.
For creators, the combination of higher resolution and improved consistency could reduce the amount of manual post-processing required after generation.
Google Search grounding adds another layer
Nano Banana 2.1 can also use Google Web and Image Search for grounding.
Grounding allows the model to retrieve information from Google’s search systems instead of relying entirely on information contained within its underlying model.
For image generation, this can be useful when a prompt refers to real-world subjects, locations, products or other information that benefits from current or externally retrieved context.
For example, a user could ask for an infographic about a current subject and have the model use search information while producing the visual.
This also reflects a broader direction in Google’s AI strategy: connecting generative models with the company’s search infrastructure rather than treating image generation as an isolated capability.
Thinking levels give users more control
Nano Banana 2.1 also introduces configurable Thinking levels.
The Gemini API documentation lists minimal, medium and high Thinking settings, with medium as the default.
The idea is to give users a choice between faster responses and more reasoning-intensive generation. A simple image may not require extensive reasoning, while a complicated infographic, multi-reference edit or detailed composition can benefit from additional processing.
This is particularly relevant for developers building production applications because different requests can have different latency and quality requirements.
Lower image-generation costs
Price is another important part of the Nano Banana 2.1 launch.
Google’s current Gemini API pricing lists standard image output at $30 per million image tokens. Based on Google’s token consumption figures, that works out to approximately $0.0336 for a 1K image, $0.0504 for a 2K image and $0.113 for a 4K image.
The previous Nano Banana 2 model is listed at $0.067 for 1K, $0.101 for 2K and $0.151 for 4K images.
That means Nano Banana 2.1 significantly reduces the image-output cost at 1K and 2K resolutions, while the reduction is smaller at 4K.
Developers can also use Google’s Batch API for lower-cost bulk processing. The Batch pricing for Nano Banana 2.1 is half the standard image-token rate.
The lower output cost could be significant for companies generating large numbers of images automatically. E-commerce platforms, advertising companies, design tools and content-generation services can potentially run substantially more image-generation requests before reaching the same output budget.
Nano Banana 2.1 vs Nano Banana 2
| Feature | Nano Banana 2 | Nano Banana 2.1 |
|---|---|---|
| Underlying model | Gemini 3.1 Flash Image | Gemini 3.6 Flash |
| Image generation | Yes | Yes |
| 1K output | Yes | Yes |
| 2K output | Yes | Yes |
| 4K output | Yes | Yes |
| Multi-image references | Supported | Up to 14 reference images |
| Subject consistency | Supported | Improved, up to 4 characters and 10 objects |
| Mask-based editing | Supported | Improved |
| Text rendering | Supported | Improved |
| Search grounding | Supported | Google Web and Image Search |
| Thinking controls | More limited | Minimal, medium and high |
| Standard 1K image output | About $0.067 | About $0.0336 |
The comparison shows why Google is positioning the new model as an efficiency-focused upgrade rather than simply another experimental image generator.
Google’s own benchmarks show a sizable improvement
Google’s published model card reports higher evaluation scores for Nano Banana 2.1 across several image-generation and editing categories.
In text-to-image testing, the Thinking version recorded an overall preference score of 1,050, compared with 990 for Gemini 3.1 Flash Image, the model behind Nano Banana 2.
The difference becomes more pronounced in some editing tests. Nano Banana 2.1’s multi-character consistency score was 1,106 in Google’s evaluation, compared with 978 for the earlier model. Its mask/ink-based editing score was 1,049 versus 965.
These numbers are useful for understanding Google’s internal testing, but they should not be treated as universal industry rankings. The evaluations were conducted using Google’s methodology and benchmark sets.
In practical use, image quality can vary significantly depending on the prompt, reference images and editing task.
Why this matters for creators and businesses
The significance of Nano Banana 2.1 extends beyond consumer image generation.
For creators, better subject consistency means fewer attempts may be required to build a series of images around the same person, character or object.
For e-commerce businesses, maintaining product appearance across different backgrounds and marketing environments is particularly valuable.
For advertising agencies, multi-turn editing can make it easier to adapt a campaign visual into different formats without recreating the entire concept.
For developers, lower image-output pricing can make automated image-generation workflows more economically viable.
The combination of editing, reference-image support, search grounding and lower output costs is therefore arguably more important than any single improvement in visual quality.
Google is building a broader AI creation stack
Nano Banana 2.1 also demonstrates how Google is integrating image generation across its broader product ecosystem.
The model is distributed through the Gemini app, AI Mode in Search, Google AI Studio, Google Flow, Google Stitch, Google Ads and Google’s enterprise AI platform.
This creates multiple entry points for the same underlying image technology.
A consumer can encounter it through Gemini or Search. A designer can use it through creative tools such as Flow or Stitch. A developer can access it through the Gemini API. A business can integrate it into enterprise workflows.
That distribution strategy could become an important competitive advantage because image-generation models are increasingly competing not only on raw quality but also on accessibility and integration.
The limitations are still important
Despite the improvements, Nano Banana 2.1 is not a perfect image editor.
Google’s own model card lists several limitations, including blurry small text, difficulty with long paragraphs, imperfect character consistency and occasional failures to follow editing instructions precisely.
The model can also retain elements of an original subject’s pose when an edit is expected to change its structure. Spatial relationships such as left and right can sometimes be confused.
These limitations matter because highly controlled professional workflows require predictable output. A model can produce impressive images while still requiring human review.
The result is that Nano Banana 2.1 should be viewed as a stronger creative-production tool rather than a replacement for professional designers or image editors.
The Bigger Picture
The launch reflects a shift in the AI image market from simple text-to-image generation toward controllable visual production.
Early generative-image systems competed largely on whether they could produce realistic pictures from prompts. The next stage is about whether they can maintain identity, preserve products, understand complex references, edit only the requested region, render useful text and repeat those operations reliably.
Nano Banana 2.1 is built around those requirements.
Google is also making the economics of that workflow more attractive by reducing image-output costs for developers. If quality improvements continue while inference becomes cheaper, automated image generation could move from an occasional creative experiment toward a routine production component for software companies and businesses.
Looking Ahead
The immediate test for Nano Banana 2.1 will be how it performs outside Google’s controlled evaluations. Users will determine whether the improvements in consistency, editing and typography translate into fewer failed generations and less manual correction in everyday workflows.
Google’s broader challenge will be maintaining that quality advantage while keeping image generation fast and inexpensive. With AI image models increasingly embedded into search, productivity, design, advertising and developer platforms, the competition is shifting from individual model launches toward complete AI creation ecosystems.
FAQs
What is Google Nano Banana 2.1?
Nano Banana 2.1 is Google’s latest Gemini-based image generation and editing model. It is based on Gemini 3.6 Flash and is designed to improve visual quality, editing, text rendering and subject consistency.
Can Nano Banana 2.1 generate 4K images?
Yes. The model supports 1K, 2K and 4K image generation.
How many reference images can Nano Banana 2.1 use?
Google’s API documentation says Nano Banana 2.1 can use up to 14 reference images in a prompt. Google also highlights consistency for up to four characters and 10 objects.
Is Nano Banana 2.1 cheaper than Nano Banana 2?
Yes, based on Google’s current API pricing for image output. A standard 1K image is listed at approximately $0.0336 compared with about $0.067 for Nano Banana 2, while 2K output is approximately $0.0504 versus $0.101.
Sources and previous coverage
First-party evidence: Google DeepMind model card (6 October), Gemini API model documentation and Google API pricing. Independent original reports from Android Authority, Decrypt and ITmedia NEWS corroborate the launch, rollout and pricing direction. Google’s performance results are its own tests; independent benchmarking had not been published at launch.
Lapaas Voice earlier covered the pre-launch Flow sighting; that report should be read as dated evidence, superseded by Google’s launch. For context on creative-AI competition, see our report on Microsoft’s MAI-Image-2.6 and on Google’s AI media watermark policy.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



