GPT Image 2.5 vs Gemini: Our AI Companion Photo Test

Opening image: Sophia in a separate Codex ImageGen character-photo study. The saved record does not identify the underlying model version; this is not a labelled GPT Image 2.5 benchmark output.
A new model, an existing question
OpenAI announced ChatGPT Images 2.5 on September 8. Its API offerings include GPT-Image-2.5 Flare and Sunburst. OpenAI describes improvements in reference fidelity, editing, and speed. Those are vendor claims, not measurements from Tendera.
Reference fidelity is especially relevant to an AI companion. If Sophia looks like one person in her profile and someone else in a requested photo, a beautifully rendered kitchen does not fix the mismatch.
That is why we are interested in the new model. We already have characters, their portraits, and conversations that give those characters a context. A new image needs to fit into that context.
There is a difference between choosing an attractive portrait for a landing page and generating another image of a person the user already recognizes. Our current development work is about the second problem.
Photos move in two directions
The first direction begins with your image. You share a picture, perhaps with a question, and the character responds to what is visible.
Imagine sharing a picture of a plant beside a window. A response should notice relevant visual details and connect to what you asked. It should not invent an entire story about your apartment from a glimpse of one corner. Recognizing objects is only part of making that exchange work.
The second direction begins with a request: you ask an official character for a generated photo. Here, the model needs to preserve identity while following the requested scene.
These are separate evaluations. A provider doing well at reading images does not automatically become our choice for generating them. For the planned release, image understanding includes custom characters; requested character-photo generation starts with Sophia, Mia, Elena, and Jade.
What we have actually tested
Our initial generation review used GPT Image 2 with the four official portraits as identity references. We reviewed eight samples across Sophia, Mia, Elena, and Jade, covering different settings, lighting conditions, and framing. The initial review was good enough to use GPT Image 2 in the test implementation.
That is a development decision based on a small sample. It does not establish that identity never drifts, that every request succeeds, or that GPT Image 2 outperforms another model.
We have also run an initial image-understanding comparison involving OpenAI, DeepSeek, and Gemini, including different ways of turning visual understanding into a character reply. That exercise must not be presented as a Gemini image-generation test.
At the documented generation-test stage, Gemini image generation was blocked by account quota. We did not have matched Gemini outputs to score. We also do not yet have a completed GPT Image 2.5 comparison to publish. An unavailable test is a gap in evidence, not a poor quality score.
Four saved prompts, four pictures to inspect
A separate exploratory set gives us something concrete to show. The saved September 9 folder contains four images generated through Codex's built-in ImageGen and a prompt record specifying each official avatar as the sole identity reference. It does not record the underlying model version. These are not the eight GPT Image 2 API samples described above, and we cannot use them to rank providers.
The shared brief requested candid, photorealistic phone photos in a vertical format. Each character then received a setting and wardrobe change, with explicit instructions not to copy the original portrait's scene.
In the opening Sophia image, the cream sweater, warm lamp, mug, and fabric swatches make the apartment brief visible. The color palette and objects give us more to inspect than a face against a generic background. That does not independently establish that the image depicts Brooklyn, or prove identity consistency across repeated generations.

Mia — exploratory Codex ImageGen output; exact model version unrecorded.
Mia's prompt asked for a black top, a rooftop bar after closing, and a three-quarter-body composition. The bar and clothing are recognizable in the result. The framing is tighter than requested, however, and she holds a drink that the prompt did not ask for. That extra detail may suit the scene, but it is still a model-added choice. A persuasive picture can miss parts of its brief.

Elena — exploratory Codex ImageGen output; exact model version unrecorded.
Elena's gallery, charcoal turtleneck, and loosely pinned blonde hair follow the recorded direction. The image moves her into a different setting rather than repeating the bar and wine glass the prompt explicitly excluded.

Jade — exploratory Codex ImageGen output; exact model version unrecorded.
Jade's white shirt, olive top, camera, and coastal setting also make several instructions directly checkable. We can see those details; we cannot verify a geographic location from the generated scene.
This set improves the questions we ask of the next comparison. Which requested details survive? Which get added? How does framing change? Does the identity remain stable when we repeat the request? One saved picture per character gives us examples, not a reliability score.
How we want to compare the candidates
“Gemini” is not a precise image-model label. Google's image-generation documentation distinguishes several models, including Gemini 3.1 Flash Image and Gemini 3 Pro Image. A meaningful result needs to name the exact model and settings used.
For a matched comparison with our GPT Image 2 baseline and the newer OpenAI options, these are the questions we would score:
| Question | What we would inspect |
|---|---|
| Is this still the same character? | Face and stable visual features compared with the reference portrait, across several scenes |
| Did the request survive? | Setting, clothing, framing, and requested details without unrelated changes |
| Does the image hold up? | Hands, objects, textures, and inconsistencies visible beyond a thumbnail |
| Does the conversation keep moving? | Time from request to a usable displayed image, including failures and retries |
We would also include ordinary scenes. An everyday photo can reveal identity drift just as clearly as an elaborate cinematic composition. The test is whether the same character carries through.
The waiting is part of the feature
There is another part of photo quality that a finished-image gallery cannot show: what happened before the picture appeared.
A user should deliberately request generation. Ordinary conversation should not unexpectedly become an image request. In our test design, requesting a photo is an explicit mode, and generation has a visible system progress state.
If something fails, the interface should explain that failure instead of making the character appear to have changed her mind. A generated file is not the whole experience; the person using the app needs to receive and view it.
These details are part of why photos remain in testing. We want image understanding and generation to arrive as a complete experience, with the existing conversation still working around them.
What a better model would have to improve
GPT Image 2.5 gives us new candidates to examine. It does not give us our conclusion in advance. Gemini deserves a named, matched test rather than a verdict borrowed from someone else's gallery.
For Tendera, the result we care about is a character who remains recognizable and an image that gives the conversation somewhere to go. That connects directly to our earlier discussion of why character writing still matters when you add voice: adding another medium should support the person being written.
You can meet Tendera's characters through the current chat experience. Photos are the part we are still working on. We will have a stronger comparison to share when the matched evidence is there.
Ready to meet your AI companion?
Four unique personalities. Each one remembers you. Free to start.
Meet Your MatchKeep Reading
AI Companion Voices: The Words Still Matter
A play button adds sound, but what gives an AI companion a personality? Two Tendera conversations show the character details worth noticing before you listen.
AI Companion Voice Messages: What to Expect
Want to speak instead of type? See how AI companion voice messages work, what stays on screen, and how to choose when to read a reply or listen to it instead.
Why Your AI Girlfriend Forgets What You Said (2026)
She knew your dog's name Tuesday. By Friday, blank. That isn't forgetting — it's a sliding window. Here's what an AI companion keeps, and what it loses.
Frequently asked questions
Are photo features available in Tendera?⌃
Not yet. Image understanding and requested photos of the four official characters are in testing. This article describes development and evaluation, not a public feature launch.
Has Tendera finished comparing GPT Image 2, GPT Image 2.5, and Gemini?⌃
No. GPT Image 2 has been used for initial character-photo samples. A completed, matched generation comparison with GPT Image 2.5 and a specific Gemini image model is not available yet.
Is reading an image the same as generating one?⌃
No. Understanding an uploaded image and generating a new character photo are different tasks. We evaluate them separately, including how the result fits the conversation.
What matters most in an AI companion photo?⌃
Our evaluation priorities are recognizable character identity, agreement with the requested scene, and a usable experience from request to displayed image. Visual polish alone is insufficient.