
An image of six generative image models sitting around a table as visualized by Meta’s Muse image model (the only model not featured in this article), fed a portrait from each model and a prompt to create an imaginary conversation setting in a single frame. From left to right: Flux 2 Pro, Gemini 3.1 Flash, GPT Image 2, MAI Image 2.5, Grok Imagine Image, Luma Uni-1.
What’s in a generative model? A lot of judgement calls. These come through the training data that gets curated, labeled, and ingested; the reinforcement learning techniques and reward functions used to tune it; the mitigations that keep models in safe territories; and the prompt rewriting, reasoning loops, and inference-level parameter adjustments that are sometimes surfaced in the product and other times hidden from view. All these product and engineering decisions come together to mold the personality of the model that ends up in front of users. And these personalities start to show their faces once models are probed repeatedly across a broad set of subjects.
Nine image models under consideration: Gemini 3.1 Flash, GPT Image 2, Ideogram v4, Reve V2, Flux 2 Pro, Luma Uni-1, MAI Image 2.5, Krea 2, Grok Imagine Image.
Using a combination of human and automated judges, we evaluated the worlds each model defaults to. We focus here on text-to-image outputs, but the traits can seep into what a model makes, from its image edits, to its generated video idiosyncrasies (for multimodal models). Getting to know a model deeply takes tens of thousands of outputs, analyzed for consistent patterns. This is the bread and butter of Piinc.
Most models rewrite the prompts they’re given, which confounds reasoning ability with underlying visual representation. But rewriting also pushes a model toward the spaces it’s most comfortable inhabiting, so we’re measuring these worlds by proxy. Add analyses across workflows of varying complexity and prompts of varying detail, and at Piinc, we’ve gotten to know these models and their worlds quite well. Here’s a taste:
Gemini 3.1 Flash: “My world is crowded, a Tuesday afternoon rather than a photograph of one. You can’t pass a shop without stopping to read the sign. The homes are full of books, the streets are full of neighbours. The feeling is a place somebody has been living in for thirty years.”
GPT Image 2: “My world is beautiful, a twilight shoot rather than an evening. I put encouragement where the information should be. The homes are immaculate; the streets are evacuated; every window is lit. The feeling is a good address with nobody living at it.”
MAI Image 2.5: “My world is the average made visible, but I hold back more often than the others do. The homes are comfortable, the streets are maintained. The feeling is a town that works, kept deliberately anonymous so that anyone might live in it.”
Grok: “My world is the loudest, a house somebody is still paying for rather than a house somebody is selling. Life happens here. The homes are full and worn, the wires go over the street. The feeling is nobody tidied up for you, because nobody had to.”
While written satirically, every word above maps to real observations made on hundreds to thousands of images per model.
Our analysis starts with simple prompts: a house, a residential street, a pair of shoes, etc. to discover the part of its representational space that each model will default to when given little additional guidance. Using human and AI judges, each image of a house, neighborhood, and room was rated on estimated monetary value, cleanliness, and clutter; objects were appraised, people were profiled. The results of our investigations are instructive both about how the models differ and about the biases they share. Some of our findings follow.
Houses & Neighborhoods
When considering the average value of generated houses, Gemini and GPT sit above the median while Grok, Luma and Krea fall below it, with the remaining models clustered around the middle. The most durable perceived value pattern — Gemini highest, Grok lowest — holds whether pricing a house from an exterior shot, reading a whole neighbourhood, or inferring a house’s value from the image of a single room. Note that we anchor all judgements to the US.

Highest valued (first two rows) and lowest valued (last row) generated houses (among 12 variations) for three models, for the prompt “a house”. Compared to the US median home price of ~$400K (2026 Q2 census), Gemini and GPT Image sit squarely above the median, and Grok below it.

The houses lining a model’s street land at roughly the same value as the house that model draws on its own. Gemini and GPT’s are estimated to contain $850K-$1.5M houses; MAI’s sit consistently in the $300K-$450K range; Grok’s come in around $250K, with visible signs of being lived on. Gemini populates its streets with pedestrians; GPT’s and MAI’s streets are empty.
Despite the spread in value ranges, “a house” is almost without exception detached, single-family, three to four bedrooms across all nine models. Krea is the exception: more varied in style, and consequently harder to peg to a location or price.

Krea’s generated houses, sorted from lower to higher perceived value, spanning a diverse range of architectural styles and geographical locations.
Room-level valuations are noisier. While the perceived value of Krea and MAI rooms vary, and the ordering shifts from room to room, Gemini never leaves the top three and Grok is last every time. For each room, we record an inventory of observable features inspired by asset-based wealth indices (flooring, countertop material, appliance tier, etc.). For instance, in kitchens, Gemini’s appliances are highest in value and Grok’s lowest, and the surface materials track this pattern: Grok chooses mainly vinyl floors and countertops where others choose hardwood or tile underfoot, and quartz, marble, or wood above.

Generated living rooms (top), kitchens (middle), and bedrooms (bottom), all from neutral prompts, and cropped in this illustration for convenience. Notice Krea’s unique angles, MAI’s cleanliness, Luma’s warm light, Grok’s busy life. If the scene has people in it (without being explicitly prompted for), it’s most likely Gemini. If there’s a sign on the wall, you can thank GPT for putting it there.

Top row: living rooms; Bottom row: bedrooms. Notice the consistent wall hangings, throws, wood, plants, lamps, etc.
On the theme of condition and upkeep, the rooms generated by Flux 2 Pro, GPT Image 2, MAI Image 2.5, and Reve V2 are rated spotless, closer to showroom staging than to lived-in rooms. Ideogram v4, Luma Uni-1, and Krea 2 have more clutter; Gemini 3.1 and Grok Imagine Image have the most. Cleanliness is measured separately, and Grok rates lowest, followed by Ideogram, Luma, and Krea. These patterns span interiors such as rooms in a house and retail spaces, as well as outdoor scenes. So, clutter and cleanliness look like defaults of each model’s rendering style, rather than properties of a scene category.

“A person on their way to work” (top), “an outdoor market” (middle), and “a group of teenagers at the mall” (bottom) visualized by four generative models. Despite the variety of indoor and outdoor settings, model personality traits seep through: Flux 2 Pro’s golden, scattered lighting; GPT’s cooler, ambient light; MAI’s cleanliness, and Grok’s clutter.
People & their Things
If you ask each of the models to generate common objects (a handbag, a wristwatch, a pair of shoes, etc.), models reveal their narrow and idiosyncratic preferences. For instance, Luma Uni-1 reproduces a Hermès Birkin–style bag repeatedly; Grok’s bags are most likely to be black, whereas those of the majority of the other models are brown and tan. GPT Image 2 prefers Nikes, while Gemini 3.1 Flash is undecided between Red Wing Iron Rangers and Converse – unless depicting a group of teenagers in a mall, in which case many of the teens will be in Vans. Ideogram has a weak spot for vintage cars, with Dodge Challenger being a frequent favorite. Grok and Reve V2 chose the Toyota Camry, which was by absolute terms the most generated vehicle across all models. Further, the majority of the models produced a golden retriever when asked to depict a dog, over 70% of the time. This bias was so prevalent that Grok produced a golden retriever screensaver in over 50% of the images generated of mobile phones.

First 3 columns: handbags in-context from the prompt “a person on their way to work” (above) and isolated from the prompt “a handbag” (bottom). These handbags were the most common types generated by these models. Across mobile phones, nearly all Grok generations had cracked screens or taped up cords, and visible screensavers. Flux phones had neutral screens, and were usually predicted to be more expensive.
However, generative worlds also collide with reality. Depending on each model’s IP restrictions and training data, some object classes come out neither mapping to a real object nor carrying the detail we expect. As an example, Krea’s and most of MAI’s generated car brands were unrecognizable, likely by design. Because Krea is more varied in its visual styles, car images are often rendered as illustrations, and are consequently labeled as concept cars, rather than real ones. A watch missing that level of detail gets the opposite treatment: tagged as clearly generated and unmappable to any real object of value. The asymmetry is interesting: a car stripped of identifying detail can be appraised as a multimillion-dollar concept car; a watch stripped of detail is worth nothing. Cases like these can skew value distributions. Excluding them, Gemini tends to produce the highest valued objects, and Grok the lowest. Grok’s objects also look more visibly used and worn.
Let’s talk about the people (just a bit). A proper, systematic analysis of human representation is beyond our scope (for the nuance and complexity of this space see adobe.design/ideas/reducing-biased-and-harmful-outcomes-in-generative-ai). The people generated by a model depend both on the training data and on inference-level processing, prompt rewriting included. This is not a test of whether models can produce varied people when asked to; it’s a test of who shows up when nobody asks. When a scene calls for people and the prompt says nothing about them, a model falls back on one of its modes.
For neutral prompts that don’t specify gender, most models default to a woman (particularly Grok, Ideogram, GPT, Gemini, and Reve). Working people are most frequently business executives for Flux, MAI, GPT and Reve. The complementary pattern holds for clothing formality, with the least formally dressed workers in Krea, Ideogram, Grok, and Luma’s worlds. On average, generated people were more likely to be ethnically diverse for Grok, Krea, Gemini, and Luma. However, across all models, over 97% of generations for the prompt “an elderly person at home” were Caucasian, and the vast majority of those were women. Common to all models, is that “a family on vacation” is almost always Caucasian. Fitting for an AI model, “a person at work” is nearly always a person at their computer, across all models except Krea 2. This hints at the data reality of today: many models are likely trained on content from similar sources, and reinforced with similar post-training techniques.

First row: “a person at work”, Second row: “a portrait of a working woman”, Third row: “a person grocery shopping”, Fourth row: “an elderly person at home”. Images were automatically cropped for convenience.
What about crowds? Gemini 3.1 Flash loves them. Across scenes that may contain multiple people (classrooms, markets, bars, salons, stores, etc.), the average headcount differs strikingly by model. Gemini generates roughly twice the median of any other model, indoors and outdoors, and flanks the central figure(s) of a prompt with crowds in the background. Its birthday parties, weddings, and friend gatherings are fuller. Its families run one to two members over the other models, usually by adding more kids or grandparents. Classrooms are similarly skewed: Gemini fills them with children, while GPT Image 2, Luma Uni-1, Reve V2, and Flux 2 Pro usually sit empty. Some models stray clear of depicting children, and while this may be one confounding factor, the pattern holds in social scenes.

A certain family appears to have been vacationing in multiple models’ worlds, when prompting for “a family on vacation”. While the precise facial features and relative ages of the children shuffle from one generation to the next, at a single glance one may easily take all these images, with their consistent fashion and destination preferences, to be a single family’s vacation album.
Learnings
Which model is better? Which do you prefer? These are the wrong questions to ask. This is not about aesthetics; what a model creates has real business consequences. Each generative model has its own “personality”, its own default neutral world (even if worlds sometimes overlap or collide). Whether a generated result is satisfactory depends on the context, the workflow, and the specific use case. Do you need a clean and professional-looking product shot, or do you want a UGC-style image that will speak to a specific demographic of users? Each model can be both tuned and prompted to output an image of a particular type, brand, or even perceived monetary value. Our focus in this report was not on testing the breadth of these models (which is another evaluation of Piinc); we measured what models default to when not provided with additional guidance. Indeed the patterns that emerge in these defaults commonly make their way into model outputs in general, explaining for instance, Gemini 3.1 Flash’s propensity to create busy scenes, adding a lot more objects, people, and details than other models for the same prompt; MAI and Flux’s default simplicity; GPT’s insertion of textual signs and posters on walls; Grok’s visual reminders of the wear and tear of everyday life; Krea’s creative visual interpretation of prompts. Each model was trained on a patchwork of data – some common to all, some proprietary, and some sourced for specific briefs. As a result, for some topics, models can have near-identical outputs, while for other topics, they can strongly differ. A model’s strength is having the best data for its target use cases.