In short
The comparison of Nano Banana 2 and 2.1 tests only three tasks and relies on one-off generations. We examine why such a test helps identify potential differences but does not confirm the new version’s overall superiority.
The promise to improve quality, text in images, and character consistency sounds broad. But comparing Nano Banana 2 and 2.1 on three prompts does not make it possible to determine which version is better overall.
The author ran the same requests through both models and evaluated the results using a checklist. This is a useful way to see possible differences on specific tasks, but the available text does not include the actual results, so it is impossible to verify the conclusion about where the new version performed better.
The limitations here are significant: one run per model, just three tasks, and a single evaluator. Such a test does not measure result consistency and is no substitute for an independent benchmark. According to the author, as of the date specified in the article, there were no independent measurements, and publications about the release referred to a Google spreadsheet.
When choosing a model for your own tasks, what matters more: the developer’s stated tests or your own runs on familiar prompts?
Source: All Articles in a Row / Artificial Intelligence / Habr