← Back to Blog
AI CompanionGPT Image 2.5GeminiAI ImagesCharacter Consistency

GPT Image 2.5 vs Gemini: Our AI Companion Photo Test

Tendera Team6 min read
GPT Image 2.5 vs Gemini: Our AI Companion Photo Test

Opening image: Sophia in a separate Codex ImageGen character-photo study. The saved record does not identify the underlying model version; this is not a labelled GPT Image 2.5 benchmark output.

In brief
We are testing image understanding and character-photo generation for Tendera. GPT Image 2 is our initial generation baseline; GPT Image 2.5 and Gemini are part of the broader comparison we want to resolve. We do not have a completed three-way ranking. The question guiding the work is specific: can a photo belong to the character you have been talking to, and can the character respond meaningfully when you share an image? These features are still in testing.

A new model, an existing question

OpenAI announced ChatGPT Images 2.5 on September 8. Its API offerings include GPT-Image-2.5 Flare and Sunburst. OpenAI describes improvements in reference fidelity, editing, and speed. Those are vendor claims, not measurements from Tendera.

Reference fidelity is especially relevant to an AI companion. If Sophia looks like one person in her profile and someone else in a requested photo, a beautifully rendered kitchen does not fix the mismatch.

That is why we are interested in the new model. We already have characters, their portraits, and conversations that give those characters a context. A new image needs to fit into that context.

There is a difference between choosing an attractive portrait for a landing page and generating another image of a person the user already recognizes. Our current development work is about the second problem.

Photos move in two directions

The first direction begins with your image. You share a picture, perhaps with a question, and the character responds to what is visible.

Imagine sharing a picture of a plant beside a window. A response should notice relevant visual details and connect to what you asked. It should not invent an entire story about your apartment from a glimpse of one corner. Recognizing objects is only part of making that exchange work.

The second direction begins with a request: you ask an official character for a generated photo. Here, the model needs to preserve identity while following the requested scene.

These are separate evaluations. A provider doing well at reading images does not automatically become our choice for generating them. For the planned release, image understanding includes custom characters; requested character-photo generation starts with Sophia, Mia, Elena, and Jade.

What we have actually tested

Our initial generation review used GPT Image 2 with the four official portraits as identity references. We reviewed eight samples across Sophia, Mia, Elena, and Jade, covering different settings, lighting conditions, and framing. The initial review was good enough to use GPT Image 2 in the test implementation.

That is a development decision based on a small sample. It does not establish that identity never drifts, that every request succeeds, or that GPT Image 2 outperforms another model.

We have also run an initial image-understanding comparison involving OpenAI, DeepSeek, and Gemini, including different ways of turning visual understanding into a character reply. That exercise must not be presented as a Gemini image-generation test.

At the documented generation-test stage, Gemini image generation was blocked by account quota. We did not have matched Gemini outputs to score. We also do not yet have a completed GPT Image 2.5 comparison to publish. An unavailable test is a gap in evidence, not a poor quality score.

Four saved prompts, four pictures to inspect

A separate exploratory set gives us something concrete to show. The saved September 9 folder contains four images generated through Codex's built-in ImageGen and a prompt record specifying each official avatar as the sole identity reference. It does not record the underlying model version. These are not the eight GPT Image 2 API samples described above, and we cannot use them to rank providers.

The shared brief requested candid, photorealistic phone photos in a vertical format. Each character then received a setting and wardrobe change, with explicit instructions not to copy the original portrait's scene.

In the opening Sophia image, the cream sweater, warm lamp, mug, and fabric swatches make the apartment brief visible. The color palette and objects give us more to inspect than a face against a generic background. That does not independently establish that the image depicts Brooklyn, or prove identity consistency across repeated generations.

Mia in a black top at a rooftop bar, holding a drink

Mia — exploratory Codex ImageGen output; exact model version unrecorded.

Mia's prompt asked for a black top, a rooftop bar after closing, and a three-quarter-body composition. The bar and clothing are recognizable in the result. The framing is tighter than requested, however, and she holds a drink that the prompt did not ask for. That extra detail may suit the scene, but it is still a model-added choice. A persuasive picture can miss parts of its brief.

Elena in a charcoal turtleneck inside an art gallery

Elena — exploratory Codex ImageGen output; exact model version unrecorded.

Elena's gallery, charcoal turtleneck, and loosely pinned blonde hair follow the recorded direction. The image moves her into a different setting rather than repeating the bar and wine glass the prompt explicitly excluded.

Jade wearing a white shirt over an olive top and holding a camera beside the sea

Jade — exploratory Codex ImageGen output; exact model version unrecorded.

Jade's white shirt, olive top, camera, and coastal setting also make several instructions directly checkable. We can see those details; we cannot verify a geographic location from the generated scene.

This set improves the questions we ask of the next comparison. Which requested details survive? Which get added? How does framing change? Does the identity remain stable when we repeat the request? One saved picture per character gives us examples, not a reliability score.

How we want to compare the candidates

“Gemini” is not a precise image-model label. Google's image-generation documentation distinguishes several models, including Gemini 3.1 Flash Image and Gemini 3 Pro Image. A meaningful result needs to name the exact model and settings used.

For a matched comparison with our GPT Image 2 baseline and the newer OpenAI options, these are the questions we would score:

QuestionWhat we would inspect
Is this still the same character?Face and stable visual features compared with the reference portrait, across several scenes
Did the request survive?Setting, clothing, framing, and requested details without unrelated changes
Does the image hold up?Hands, objects, textures, and inconsistencies visible beyond a thumbnail
Does the conversation keep moving?Time from request to a usable displayed image, including failures and retries
The reference portraits and scene requests should stay the same across candidates. We should review repeated outputs, not choose the best single image from each provider. Provider-specific settings need to be recorded rather than treated as interchangeable.

We would also include ordinary scenes. An everyday photo can reveal identity drift just as clearly as an elaborate cinematic composition. The test is whether the same character carries through.

The waiting is part of the feature

There is another part of photo quality that a finished-image gallery cannot show: what happened before the picture appeared.

A user should deliberately request generation. Ordinary conversation should not unexpectedly become an image request. In our test design, requesting a photo is an explicit mode, and generation has a visible system progress state.

If something fails, the interface should explain that failure instead of making the character appear to have changed her mind. A generated file is not the whole experience; the person using the app needs to receive and view it.

These details are part of why photos remain in testing. We want image understanding and generation to arrive as a complete experience, with the existing conversation still working around them.

What a better model would have to improve

GPT Image 2.5 gives us new candidates to examine. It does not give us our conclusion in advance. Gemini deserves a named, matched test rather than a verdict borrowed from someone else's gallery.

For Tendera, the result we care about is a character who remains recognizable and an image that gives the conversation somewhere to go. That connects directly to our earlier discussion of why character writing still matters when you add voice: adding another medium should support the person being written.

You can meet Tendera's characters through the current chat experience. Photos are the part we are still working on. We will have a stronger comparison to share when the matched evidence is there.

Ready to meet your AI companion?

Four unique personalities. Each one remembers you. Free to start.

Meet Your Match

Frequently asked questions

Are photo features available in Tendera?

Not yet. Image understanding and requested photos of the four official characters are in testing. This article describes development and evaluation, not a public feature launch.

Has Tendera finished comparing GPT Image 2, GPT Image 2.5, and Gemini?

No. GPT Image 2 has been used for initial character-photo samples. A completed, matched generation comparison with GPT Image 2.5 and a specific Gemini image model is not available yet.

Is reading an image the same as generating one?

No. Understanding an uploaded image and generating a new character photo are different tasks. We evaluate them separately, including how the result fits the conversation.

What matters most in an AI companion photo?

Our evaluation priorities are recognizable character identity, agreement with the requested scene, and a usable experience from request to displayed image. Visual polish alone is insufficient.