- Genial
- KI-Webinare für Unternehmen
- AI for Image Generation Webinar April 23rd
AI for Image Generation Webinar April 23rd
This week's session was entirely about image generation, comparing Google's Nano Banana Pro and Nano Banana 2 with ChatGPT's brand new GPT Image Gen 2, released just two days before the webinar. It is aimed at business users, especially the brick-and-mortar businesses that keep saying AI will not affect them: landscaping, housing, renovation, interior design. No live build this time, since the model was too fresh to build a full project around, so instead we ran use cases side by side.
For the people joining for the first time: I used to work at Schneider Electric as the Generative AI Change Manager, where I helped upskill about 100,000 people by deploying upskilling programs across the company, and I run this AI weekly webinar every Thursday at 5 PM Paris time with practical use cases and live builds (1:48).
Why image generation suddenly matters (4:31)
I never did an image generation session before because I considered the technology to be mainly about creating cute images. That stopped being true when Nano Banana Pro and Nano Banana 2 came out. These models do not just generate pretty pictures, they do visual reasoning: the model actually understands the subject of the image it is producing and uses the correct terminology inside it. That is the step change, and it is the reason this whole session exists.
Nano Banana Pro and 2 in practice (6:06)
We started with the baseline: a cute cat drinking a beer, comic style. You can swap styles freely, use a reference image, and get very consistent characters across iterations. The result had a couple of small artifacts, but nothing out of this world, and far beyond anything we had before (9:00). One thing to know: images generated inside Gemini always carry the watermark. If you want Nano Banana output without it, you need a service like Higgsfield on top (9:21).
Then we kicked it up a notch with the demo that shows what visual reasoning means: a step-by-step infographic of how gold is extracted, all the way to the finished gold bar (9:47). This requires the model to actually know the process. The result was very detailed, and the words on it are the exact terms used inside the mining and refining industries. That is what separates Nano Banana Pro from creativity-first tools like Midjourney, which aim mainly at communication assets for creative roles (10:41).
It also understands audiences. I took the same gold infographic and asked for the wording to match an audience of kids five years old or less. Terms like electrolytic refining, electrowinning, and carbon absorption disappeared, the wording changed completely, and the meaning stayed the same (11:54). A friend of mine built a business on exactly this: a platform where parents narrate the stories their kids dream up, with the kid's favorite plushie as a consistent character and wording that fits the age (13:06).
Two more examples from my own testing. I fed Nano Banana Pro problems from a first year master's physics exam, and it wrote out detailed, correct solutions on the sheet itself (13:56). And you can hand it a history book as context and get an infographic-style timeline that matches all the knowledge you gave it, completely accurate (15:14).
Higgsfield, the dedicated creative platform (7:51)
While images generated, I made a case for people in marketing or heavily creative roles: paying licenses left and right only for image generation is not efficient. If you have a lot of image and video generation use cases, get one dedicated tool built around them. The best one I have found is Higgsfield. It gives you Nano Banana, image generation by Kling, and GPT image generation in one place, with a user experience that actually circles around generation: extra settings and freedoms you simply do not get inside ChatGPT or Gemini. Practical tip from the end of the session: they run 30 percent discounts all the time, and if you wait a couple of weeks you can usually catch 50 percent off, at which point it is very good value (49:43).
GPT Image Gen 2 and thinking mode (16:01)
Now the new player. GPT Image Gen 2 is live in ChatGPT, and when you click Create an Image you get two types of cards: normal ones, the classic make-me-a-cool-image mode, and cards with a Thinking label. Thinking mode is where the true power is: the model reasons for a long time before producing the image and fetches information from the web while doing it (18:08).
The infographic I generated with thinking mode the night before was very, very good: all the information pulled from the web by GPT, extreme levels of detail in the composition, and a map that held up under zoom, Japan fine, the Koreas fine, maybe some issues in China, all from a single prompt (16:26).
Then the business demos, aimed squarely at the real world:
- Cracked wall repair (18:33). I generated a synthetic cracked wall, and that is a trick in itself: if you do not have a photo of something, generate it and test your use case on top. Then I asked for three ways to fix the wall at three budget levels, with before and after results for each in one image. The model produced exactly that, and on the version I made earlier it added its own tips: basic patch and paint, spackle, reinforced repair, professional repair, maintenance (20:31).
- Your own pricing, applied (19:42). If your company has a pricing guideline, say 1,000 euros per square meter, you give ChatGPT the document, then photos of the zone to fix, and it prices the repair proposals following your own rules.
- DIY step-by-step guides (20:56). I asked for the how-to to run the fixes myself, and got a detailed visual tutorial: clean the area, apply spackle, hand smooth, prime, paint, then prepare the crack, mesh tape, first coat, second coat, bonding agent. All accurate, all from a single prompt. This is lead magnet material for a lot of businesses that used to make these guides manually (23:46).
- Landscaping from satellite view (24:09). Starting from a satellite image of a house, you can modify it well enough to show a client what their property will look like after the work.
- Interior staging by persona (24:52). I generated an empty hyperrealistic 20 square meter apartment, then asked for three furnished versions side by side: an 18 year old student, a young couple around 25 looking for their first apartment, and a 60 year old lady wanting a calm place. Same apartment, same radiator, same AC position, consistent floors, different lives. I know someone whose actual job is staging furniture in apartments for photos, and I do not know where that job ends up in the near future (21:40).
- Angles (28:40). These models now understand perspective, so you can ask for an angle you never showed. My live attempt at re-shooting the student apartment from the entrance failed, to be fair, but my prompt was vague, and it shows the current limits.
Worth noting on limits and settings: I am on the 20 dollar ChatGPT plan and generated around 20 images during the session before finally hitting a cap near the end, which honestly surprised me, since Gemini used to limit image generation much more aggressively (26:53). You can also manually pick the resolution of every image instead of relying on auto mode (28:19).
Claude versus ChatGPT, and the Codex question (30:23)
A lot of people have been migrating to Claude lately, in my opinion for the wrong reasons, and then complaining about the experience. My honest take: if you have exactly 20 dollars to spend on AI, you are better off with ChatGPT. People say "but I need Claude Code". Codex is literally the same thing as Claude Code, OpenAI is just bad at marketing it. Learn to use one and you know how to use the other. The real difference is budget: Claude Code technically starts at 20 dollars but you will get almost nothing out of it there, and realistically you want around 100 dollars per month. Codex at 20 dollars gets you quite far, and at the 100 dollar mark the two become quite similar (31:44).
A first pass at Claude Design (32:36)
Since we were on the topic of visuals, I did a live aparté on Claude Design, the recently launched tool people call, quote-unquote, "the Figma killer". Honest verdict: not quite there yet, but a very good user experience, and user experience is what will distinguish these tools going forward, because underneath they are all becoming very similar.
What it does: wireframes and high-fidelity mockups of anything, plus a design system setup where you can import Figma files, logos, assets, and your communication tone. That design system step matters, because good design systems are exactly what AI cannot do well on its own yet (33:49). One key thing to understand: Claude does not generate images, it generates code, and everything you see, the mockups, the slides, the animations, is code under the hood (34:55).
I demoed a landing page mockup for an imaginary cookie brand, and while it generated, showed a lyric video I had made just before the session: I gave Claude the lyrics of Franz Ferdinand's Take Me Out and asked for a lyric video in that style. It picked up the theme of the song and animated the catchphrases accordingly (35:58). You can also upload a video of yourself and have it add motion design keyed to what you say: mention that 50 percent of people never go to Asia and it puts a big 50 percent on screen, because it understands that is the key number (37:13).
Two caveats: it is slow, painfully so for a live demo, and right now you cannot export the video as an MP4, so the workaround is recording your screen (39:35). But the strategic point stands: Claude Design closes the gap between designers and developers, because a design made there exports as code straight into Claude Code for a developer to work on, and you can select individual items and have it modify precisely those (38:11).
Where this is heading (40:22)
My three closing takeaways. First, image generation with visual reasoning is going to be big within the next six months. Second, the adopters will come from the physical world: architects, landscaping, renovation, and I am genuinely curious about plastic and reconstructive surgery, where you could show before and after estimates, whatever the legal side turns out to be. Third, hairdressing, and I speak as a bald man.
So we tested that live: generate a bald guy, front and top view, then create a timeline of how he would look after a hair transplant, in jumps of one to two weeks over two months (41:30). After a couple of interface fumbles on my side, and briefly hitting my generation limit, the result was genuinely good: a reasonable timeline, the same face frame after frame, realistic red dots on the scalp early on, and visible scarring in the right places (48:37). You could plausibly do the same for reconstruction cases.
I left everyone with one cultural data point: a 100 percent AI-generated film on YouTube, 21 minutes long, steampunk theme cranked to the maximum, with real human voices behind the characters but every visual made with AI. It pulled 5.7 million views in 11 days, and it simply would not have been possible to make one year ago (50:51).
Q&A highlights
- Which speech-to-text am I using? Wispr Flow. The killer feature is snippets: I say "my permanent webinar links" and the full text drops in, same for emails or referral codes (7:27).
- Is it really possible to generate this many pictures on a regular subscription? I was on the 20 dollar ChatGPT plan and got roughly 20 images before hitting a limit near the end of the session, so yes, with a cap that I have not fully mapped yet (26:53).
- Will Anthropic develop design or image capabilities to keep up with OpenAI? I do not think they should. Image generation is a completely different discipline from what they build, and Claude Code can already connect to state-of-the-art image models via API. That is what I used to do: plug an OpenAI API key into Claude Code or Codex and industrialize the production of visual assets. Anthropic's edge is a model that is very good at using tools, while other players build the tools it uses (43:32).
The AI weekly webinar happens every Thursday. All previous sessions are available on the webinars page.


