GPT Image 1.5 Lags Behind Banana2 in Comprehensive AI Image Generation Test

Open Source Talent Scout
Comparative display of GPT Image 1.5 and Banana2 interfaces, showing performance difference.

OpenAI's updated GPT Image 1.5, featuring enhanced instruction following, more precise image editing, and a fourfold increase in generation speed, has been rolled out to all users with a new user interface. Despite these updates, a comparative analysis against Banana2 across twelve scenarios indicates that GPT Image 1.5 currently underperforms its competitor, particularly in areas like multi-text generation, text-based posters, and world knowledge integration.

Abstract digital grid with fragmented and clear text, symbolizing multi-text generation challenges.

Abstract digital grid with fragmented and clear text, symbolizing multi-text generation challenges.

Multi-Element and Multi-Text Generation Challenges

The evaluation began with a complex "hell case" requiring the generation of a 6x6 grid containing 36 distinct elements. Banana2's output, while exhibiting some element repetition and exceeding the specified column count, generally produced more aesthetically pleasing individual elements than GPT Image 1.5.

In a subsequent test involving the generation of a 3:4 image with a complete Chinese poem, "Song of My Thatched Hut Ruined by Autumn Wind," in calligraphy with Hanyu Pinyin annotations, GPT Image 1.5 struggled with Chinese character accuracy. Although its Pinyin annotations were surprisingly more precise, numerous characters were incorrect. Banana2 demonstrated superior performance in this multi-text generation task.

Stylized world map with accurate and inaccurate digital data points, representing AI knowledge integration.

Stylized world map with accurate and inaccurate digital data points, representing AI knowledge integration.

World Knowledge and Infographic Accuracy

A third round focused on world knowledge, specifically image-to-image generation using a photograph of China's Huajiang Grand Canyon Bridge. The prompt requested basic information about the bridge, including dimensions, width, height, and completion date, to be added as handwritten annotations. It also required schematic diagrams of the main cable's cross-section and a suspension bridge. While GPT Image 1.5 produced an impressive initial result, several data points, apart from the bridge deck's height above water and maximum span, were inaccurate. Banana2 exhibited significantly higher data accuracy in this test.

Further testing involved creating historical posters for specific locations using latitude and longitude coordinates. The task was to depict events from May to October 260 BC at a given geographical range, providing a detailed infographic in Chinese. Banana2 again outperformed GPT Image 1.5, suggesting a consistent gap in knowledge integration and infographic generation capabilities.

Hand interacting with a distorted digital image, symbolizing AI image modification limitations.

Hand interacting with a distorted digital image, symbolizing AI image modification limitations.

Image Modification and Stylization Limitations

OpenAI had noted several quality updates for GPT Image 1.5, including improved handling of numerous small faces. However, when tasked with generating an image of "thousands of people gathering in front of the Oriental Pearl Tower in Shanghai, with every person's face clearly visible," GPT Image 1.5 produced highly distorted and "uncanny valley" faces upon zooming, whereas Banana2's faces began to distort only from the fourth column.

In multi-image blending, both models showed limitations. When asked to blend ten furry characters into a single scene, both GPT Image 1.5 and Banana2 either missed characters or generated duplicates. GPT Image 1.5 generated one fewer character, while Banana2 generated one more than requested.

For precise image modification, GPT Image 1.5 was tested against Banana2's updated feature allowing users to specify modifications using circles, arrows, and text. While GPT Image 1.5 successfully erased lines, circles, and text, changes to light and shadow were less discernible. OpenAI has acknowledged that GPT Image 1.5 is less effective in stylization compared to its predecessor, offering only 13 filter types.

Nine-panel storyboard showing logical and illogical sequences, illustrating AI storyboarding challenges.

Nine-panel storyboard showing logical and illogical sequences, illustrating AI storyboarding challenges.

Redrawing and Sequential Storyboarding

A "re-drawing" test, converting an image of Conan into a real-person street snap with 2D illustration elements in a specific style, showed GPT Image 1.5 performing surprisingly well, nearly matching Banana2.

However, in a nine-panel grid generation task, where the goal was to infer a timeline of events from an image and create a chronological storyboard, GPT Image 1.5's output lacked logical sequence and exhibited significant style deviation. Banana2's result demonstrated much stronger logical progression.

The analysis, summarized by Grok, concluded that GPT Image 1.5 currently lacks competitiveness against Banana2, with its primary advantage appearing to be generation speed. Even cases previously endorsed by Greg for GPT Image 1.5 were surpassed by Banana2. An additional test involving coloring a comic page and translating text into Chinese within the image further highlighted Banana2's superior performance.

The evaluation also included a complex request to professionally explain AI video generation training principles and create an English PowerPoint presentation in the style of Crayon Shin-chan's hand-drawn illustrations. This test, conducted over an extended period, underscored Banana2's more robust understanding of complex instructions and Chinese language nuances, which GPT Image 1.5 struggled with.

ToolMesh
ToolMesh Weekly

Stay Ahead of the AI Curve

Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.

No spam, unsubscribe at any time.