Search
Menu

OpenAI has broken a clear boundary. 4o image generation is now directly in GPT-4o. You generate images directly in chat. No additional tools. No DALL-E. No hassle. You describe what you want and Chat delivers it in real time. With exceptionally good results. This is not just a standalone feature. It's a new era.

Finally more usable AI images

AI images looked impressive for a long time, but were often not employable. Text was distorted, hands didn't match, and details often felt cluttered. GPT-4o is making great strides. Text in images is more readable. Proportions are more realistic. Reflections and shadows make more sense. Chat uses the context of your conversation and follows instructions more accurately. Images are not yet perfect, but they are much more consistent. For the first time, AI image generation feels usable. Not just for testing, but for immediate deployment.

What's behind 4o Image Generation?

OpenAI's image generation now works based on the Sora model. Sora was first developed for video, but is now also the basis for the new image function in GPT-4o. So instead of DALL-E, ChatGPT uses an "omnimodal" model that combines text, image, audio and speech. As a result, it understands context better and the image more closely matches what you mean.

This is all you need

A computer. Internet, and a ChatGPT account. That was it. It's that simple. Those who work with the free version can already do a lot, but with a Plus subscription you get access to GPT-4o and the new image generation. You don't have to install anything. And you don't have to learn complicated prompts. Just type what you see in front of you, and Chat creates the image. In plain language. From one place. Image creation has never been so accessible.

What makes 4o image generation so much better?

So GPT-4o finally understands and shows better what you mean. Here's what specifically works better about this brand new model:

More visual understanding through better training

GPT-4o is trained on a huge amount of images and text simultaneously. As a result, the model learns not only how words and images belong together, but also how images relate to each other. That makes for visual output that feels much more logical. The composition fits better, the style remains more consistent and there is more attention to the right details. That makes the images usable rather than just beautiful.

Text in picture now works well

One of the biggest improvements is how GPT-4o handles text in visuals. Think of menus, street signs or posters with headlines. Where previous models struggled with letters and legibility, GPT-4o places text in the right place and in the right style. Symbols, words and images work together. This makes image generation useful for communication, not just atmosphere.

From street signs to wedding cards

The AI now understands much better how image and context work together. Two witches studying a crowded street sign? A rustically designed menu for a Korean restaurant? Or a wedding card where image and typography blend seamlessly? GPT-4o can visually build these types of scenes based on context and instruction. Including details such as hair color, sign texts, formatting or style references. This creates images that tell a story, not just set a mood.

Conversations now form the image processing

Image generation is right in the conversation. You say what you want to see, GPT creates it. Then you give an additional clue. The model adapts the existing image without starting over. That makes GPT-4o suitable for concept development. Think character design for games, multi-step infographics or fine-tuning content visuals. Everything stays within the same style and context.

More control over complex images

Where other models struggle to handle many elements at once, GPT-4o maintains overview. Images with ten, fifteen or even twenty separate objects remain correct. A map with icons, an infographic with several parts or a sticker with multiple visual layers? GPT-4o keeps the proportions and positions right. This keeps the result usable even when the image becomes more complex.

In-context learning with your own images

You can also now upload your own image and let GPT-4o think along. Upload a rough sketch, photo or example. The model picks out style, composition or details from there and uses those to create a new image. That "in-context learning" makes it possible to generate AI images that match your brand or design - without having to explain anything technical.

Tests & applications

Since its launch on March 25, we have been testing GPT-4o image generation quite a bit. Not just to play with it, but to discover what you can really do with it in practice. We create visuals for socials, infographics and ads. Sometimes immediately usable. Often 75% good. Perfect as a semi-finished product, a source of inspiration or a starting point that you can easily fine tune. Below you can see some of our favorite outcomes.

Supplementing product photos

For us, this is the most concrete and immediately applicable point. Need a cool setting for your product? It's now possible with two clicks. Yes, really. Boring product photos are now no longer an excuse, but something of the past. Below you can see an example: a simple product photo, the corresponding prompt, and the stunning result. This is beyond belief, isn't it?

Example 4o Image Generation - product photo addition

Example 4o Image Generation - product photo addition agency 

Clothing change

This is a super practical application, especially for web shops. Expensive photoshoots for all products are a thing of the past. With AI, you simply let a model adjust clothes in a few clicks (also fun to prank your colleagues).

Example 4o Image Generation - clothing change

Example 4o Image Generation - clothing change 

Person change

You can now easily put someone in a different setting. In the example below, you can see how Gido is put in the place of me (Igor). The result is not perfect, but it is frighteningly good!

Example 4o Image Generation - person change

Campaign materials in a snap

In the example below, you can see how we turn a coffee machine into an image for social ads. Whether it's a poster, banner or ad, it can now all be done quickly and instantly ready for publication. Bizarre!

Example 4o Image Generation - social ad

Example 4o Image Generation - social ad 

The limitations

GPT-4o image generation is impressive, but not yet finished. OpenAI is aware of several limitations that will be further improved after launch. For now, it is important to know where the model falls short. So that you take that into account in your workflow.

Cropped images and fabricated details

The model sometimes mis-cuts longer images. Consider posters or visuals with a lot of vertical content. The bottom then falls away, losing important elements. In addition, with little context, elements that you didn't ask for may pop up. These kinds of "hallucinations" occur especially with short or vague prompts.

Constraints on many objects

GPT-4o can handle up to about 10 to 20 individual objects in a single image. After that, errors in placement, scale or relationship between elements occur. This is called a binding problem. Consider complex views such as a complete periodic table, a full menu or a technical infographic. The more loose elements, the greater the likelihood of confusion in the image.

Problems with small text and details

Text on a small scale remains a challenge. In images with a lot of information in limited space, such as disclaimers, labels or info cards, text often becomes illegible or distorted. Handwriting or decorative letters also remain difficult to process. For sharp typography at small size, the model is not yet accurate enough.

Editing parts is not precise enough

You can give instructions to modify a specific part of the image, but that doesn't always work as intended. Sometimes the model changes other parts as well, or new errors arise. So correcting a typo, replacing an object or adjusting a color doesn't always work without side effects. OpenAI is working on more precision, but it is not there yet.

Problems with multilingual text

GPT-4o has difficulty with non-Latin characters, such as Arabic, Korean or Chinese. These are sometimes rendered incorrectly or replaced with random symbols. The more complex the text, the more likely characters will be rendered incorrectly. For visuals with multiple languages, this is an obvious limitation.

Graphs and data display

The model is not currently suitable for accurate graphs or precise data representation. Bar graphs, axes or numerical visualizations are often constructed inaccurately. GPT-4o understands global form, but lacks precision.

Shift in the content landscape

For years, producing visuals was a cumbersome process. With briefings, correction rounds, waiting times and coordination. But those barriers are largely gone. Images can now be generated simply in ChatGPT. For social posts, infographics, campaigns and presentations. Just to name a few. The quality is often good enough for immediate use. The speed is many times faster. This causes a major shift in the content landscape. Less emphasis on production. More on creation, brand and direction.

Branding becomes more important than ever

When anyone can create images, recognition makes all the difference. The pace is picking up, the amount of content is increasing. But only brands with a clear style and story will stick around. AI generates the image, but you set the direction. Color usage, tone of voice and visual line are no longer a detail. They are the foundation. Branding is no longer something for later in the process. It is the foundation by which everything begins.

Endless possibilities

With AI, you can now get started in ways you didn't think possible before. Whether you want to create social ads, posters, infographics, comics or even cartoons, it can all be done in an instant. But that's just the beginning. Also consider enhancing your website with new icons, fresh illustrations or visual upgrades that instantly enhance your design.

Example 4o Image Generation - OMA

Example 4o Image Generation - OMA

Leave a Reply

Your email address will not be published. Required fields are marked *

Most frequently asked questions about this blog