Early reports about Mona Lisa 1 are mixed. A small number of same-prompt comparisons suggest that Mona Lisa 1 may produce more natural photorealism and may handle busy prompts better in some cases. Other users still report noise, text artifacts, and only modest differences from GPT Image 2. No controlled benchmark establishes an overall winner.
That distinction matters. Mona Lisa Img has not run a controlled test of the two models, and the community examples linked below are not our results. This review separates official documentation from original user reports so that an interesting early signal does not become an unsupported product claim.

Editorial illustration. It does not show model output or benchmark results.
Short Answer
GPT Image 2 is an officially documented OpenAI model with a published API model page and image-generation guide. Mona Lisa 1 is different: users report encountering that name after blind comparisons in Arena, but no official developer, model card, API, price, or release plan has been published for it.
The available comparisons support continued testing, not a recommendation to switch models. The strongest early signals concern natural-looking texture and one complex prompt-following example. The evidence does not currently support claims that Mona Lisa 1 is clearly better at illustration or anime, wins across all categories, or is confirmed to be an OpenAI model.
What Is Confirmed
GPT Image 2 is a documented production model
OpenAI documents gpt-image-2 as an image generation model and lists a dated snapshot, gpt-image-2-2026-04-21. Its official guide covers generation and editing through the Image API and the Responses API. These are product facts, not inferences from social posts.
GPT Image 2 also appears at the top of Arena's public text-to-image leaderboard at the time of this review. Leaderboard positions can change, so that placement should be treated as a current snapshot rather than a permanent ranking.
Arena comparisons are blind before voting
Arena explains that users enter a prompt, receive outputs from two anonymous models, vote for the better response, and then see the model identities. This makes a revealed model label useful evidence of what the interface displayed in that individual battle.
It does not, by itself, provide a model card, identify the developer, or guarantee that the same experimental model will remain available. We found original posts reporting a mona-lisa-1 reveal in Arena, but no first-party Arena announcement or documentation for that model name. Its appearance should therefore be described as reported, not officially announced.
Why People Suspect an OpenAI Connection
One early report says that an output associated with Mona Lisa 1 was recognized by OpenAI's Verify tool. That is a meaningful provenance clue, but it has a narrower meaning than many summaries give it.
OpenAI's provenance documentation says images generated by supported OpenAI products may contain C2PA metadata and an invisible SynthID watermark. It also warns that metadata can be removed and that provenance checks have limitations. Most importantly for this comparison, a successful provenance check does not disclose the exact model that generated an image.
The defensible conclusion is therefore:
- Supported: at least one reported sample produced an OpenAI-related provenance signal.
- Not established: the developer, architecture, product family, or final identity of Mona Lisa 1.
Naming patterns, perceived visual similarities, and release speculation add context, but none can close that attribution gap.
What Original Early Tests Report
More natural photorealism in some matched prompts
Several original X posts show prompts run against Mona Lisa 1 and GPT Image 2 and describe Mona Lisa 1 as more realistic or less synthetic. The examples point to details such as skin texture, lighting, material response, and the overall feel of a photograph.
These are useful observations because they include visible results and, in some cases, a stated same-prompt comparison. They are still small, author-selected samples. We do not know whether all attempts were shown, whether settings were equivalent, or how often the preference repeats across prompts.
The appropriate claim is: some early matched examples suggest a possible photorealism improvement. They do not demonstrate a general photorealism win.
A possible advantage on a busy prompt
One original comparison uses a complex scene with multiple requested elements and reports that Mona Lisa 1 follows the prompt more successfully. This is a good candidate for a future benchmark category because object count, relationships, and scene constraints can be scored more consistently than overall aesthetics.
For now it remains one comparison. It supports a hypothesis that Mona Lisa 1 may handle crowded instructions better; it does not establish instruction-following superiority.
Incremental rather than dramatic differences
Not all early impressions describe a large change. One tester characterized the result as a slight improvement rather than a major jump. A Reddit discussion similarly includes mixed reactions: some users prefer the new output, while others see familiar weaknesses or only a small difference from GPT Image 2.
This mixed response is important because it limits the strongest possible conclusion. The evidence is not converging on an obvious winner across use cases.
Text and dense-detail problems remain visible
The Reddit examples and comments include reports of fuzzy lettering, visual noise, and inconsistent layout in detail-heavy images. These weaknesses matter for posters, interfaces, packaging, diagrams, and any image where readable text or precise spatial relationships are central.
We found no broad, repeatable evidence that Mona Lisa 1 has solved these problems. Any future test should score text transcription separately from aesthetic quality so that an attractive image cannot hide incorrect words or structure.
Claim-by-Claim Comparison
| Question | What the current evidence supports | Confidence |
|---|---|---|
| Has Mona Lisa 1 appeared in Arena? | Original users report seeing the label after blind battles; no first-party model announcement was found. | Medium |
| Is Mona Lisa 1 made by OpenAI? | A reported Verify result is a strong provenance clue, but it does not identify the exact model or establish ownership. | Low |
| Is photorealism better than GPT Image 2? | Some same-prompt examples suggest more natural texture and rendering. The sample is too small for a general conclusion. | Low |
| Is complex prompt following better? | One useful original comparison suggests a possible advantage. Repeated tests are still needed. | Low |
| Is text rendering better? | Current reports still show text and dense-detail failures. No clear advantage is established. | Low |
| Is illustration or anime clearly better? | No reliable original evidence found that supports a broad superiority claim. | Unknown |
| Is Mona Lisa 1 the overall winner? | No controlled benchmark supports that conclusion. | Unknown |
| Are its API, price, release date, and specifications known? | No official documentation was found. | Unknown |
What the Evidence Does Not Support
The current source set is too limited to claim any of the following:
- Mona Lisa 1 is definitively developed by OpenAI;
- it belongs to a specific existing model family;
- it is a final product name rather than an experimental Arena label;
- it consistently beats GPT Image 2 in photorealism, illustration, anime, typography, or editing;
- its knowledge cutoff, training data, safety system, supported resolutions, or latency are known;
- a public launch, API, price, or release date is imminent;
- any future API would be compatible with GPT Image 2.
These are not minor caveats. They determine whether the article is reporting observable evidence or turning speculation into product information.
How a Fair Benchmark Should Be Run
A useful comparison should begin with a fixed prompt set and a scoring rubric published before results are selected. At minimum, it should cover:
- natural portraits, interiors, outdoor scenes, and product-style photography;
- illustration and anime prompts with repeated characters and specified visual traits;
- exact text transcription, typography, and structured layouts;
- object counts, left/right relationships, reflections, hands, and spatial composition;
- long prompts containing both important and distracting details;
- editing and reference-image tasks, if both interfaces offer comparable controls.
Each model should receive the same prompts and the closest equivalent settings. Testers should disclose the interface, date, output size, number of attempts, selection procedure, and failures. Multiple outputs per prompt should be retained, and reviewers should score them blind before model identities are revealed.
Most importantly, every evaluated result should be available. A collection of hand-picked successes can illustrate capability, but it cannot measure reliability.
Current Verdict
Mona Lisa 1 is worth watching, but the comparison is not settled. Original early tests provide a credible signal that some outputs may look more natural and that busy prompts may improve in certain cases. They also show remaining noise, lettering, and layout problems, while at least one tester describes the change as modest.
For now, GPT Image 2 is the known quantity: it has official documentation, API guidance, and a public leaderboard position. Mona Lisa 1 remains an anonymously reported Arena model with interesting samples and major unanswered questions. The evidence supports a future benchmark, not an overall winner.
Official Sources
- Arena FAQ - how blind battles, voting, and model reveals work.
- Arena Image Battle guide - how to use the image-generation comparison interface.
- Arena text-to-image leaderboard - current public ranking, which may change over time.
- OpenAI: C2PA and SynthID in generated images - what provenance signals can and cannot establish.
- OpenAI GPT Image 2 model page - official model and snapshot information.
- OpenAI image generation guide - official generation and editing documentation.