GPT-6 Wins Narrow Victory in SVG Drawing Showdown, Tested on 1,350 Images
9 models battled in SVG drawing and guessing. GPT-6 Astra led with 75.5 points, but the top 3 were neck-and-neck.
A new yardstick for measuring generative AI performance has emerged. It is an evaluation in the form of a “drawing-guessing game” in which models draw pictures in SVG and other models guess what they are. According to reporting by Silicon Alien of Huxiu, GPT-6 Astra took first place with 75.5 points after 1,350 drawings and 11,961 guesses.
The idea originated from a regular test by developer Simon Willison. It is the attempt, continuing since 2024, to have models draw a “pelican riding a bicycle” in SVG. What started as a joke has become established as a method for comprehensively assessing reasoning, drawing, and understanding of abstract concepts. This evaluation expands on that idea and is distinctive in that humans do not score it; models judge each other.
Unique Evaluation Design Separating Drawing
and Guessing
The first round included eight competing models plus one reference model. The international entries were three models: GPT-6 Astra, Gemini 3.8 Flash, and Claude Fable 5.1. The Chinese entries were five models: Kimi K3, Qwen3.8-Max, GLM-5.3-Flash, MiniMax-M3, and DeepSeek-V4-Flash-Vision-Exp. Gemma 4 26B was included as a reference and was excluded from ranking and scoring.
The procedure consists of two stages. For each word, all nine models each draw one image in SVG, which is converted to PNG and then guessed by all nine models. The guessing side cannot view the code, and guidance through annotations or element names is prohibited. Text elements in the images are banned, and timeouts or truncations were treated as fouls. All calls were made with default API settings, which also supports fairness of conditions.
Scores were split evenly between drawing and guessing. The drawing score is the rate at which a model’s own drawings were correctly guessed by other models. The guessing score is the rate at which it correctly guessed other models’ drawings. Judgments were automated using a word list and matching rules, with no aesthetic evaluation or third-model review involved. Resampling was performed 2,000 times at the word level to calculate 95% confidence intervals.
The word list comprises 150 words with a clear breakdown. It includes 70 common words, 20 action scenes, 10 trending jokes, 10 AI self-references, 10 common-sense violations, plus 30 undisclosed benchmark words. The undisclosed words are for comparison across time periods, aimed at curbing optimization for public words. This can be seen as a device to balance transparency and reproducibility in the evaluation design.
After 1,350 drawings and 11,961 guesses, GPT-6 Astra won by a narrow margin with 75.5 points
Overall Results With Top 3 Models in a Close Race
The overall leader was GPT-6 Astra with 75.5 points. Second was Gemini 3.8 Flash with 74.6 points, and third was Claude Fable 5.1 with 74.3 points. The confidence intervals of the three overlap, and the statistical difference is small. These three form the first tier.
Chinese models followed closely behind. Kimi K3 scored 71.7 points, Qwen3.8-Max scored 71.3 points, and GLM-5.3-Flash scored 68.9 points. MiniMax-M3 scored 61.9 points and DeepSeek scored 60.9 points. Reference model Gemma scored 38.2 overall; it guessed about half of the common words correctly, but its drawings were recognized only about 20% of the time.
Drawing ability and guessing ability do not align. Astra and Claude excel at drawing, ranking in the top two for drawing scores with 79.8 and 79.1 points respectively. Gemini ranked first in guessing with 76.1 points but fifth in drawing. GLM scored 74.7 in drawing versus 63.1 in guessing, the largest gap among the eight models. Differences in characteristics invisible in a single overall score emerge.
Diversification of evaluation methods is a concern for the industry as a whole. In MIT Tech Review “Innovators Under 35” 2026 Edition, Selection Process Details Released, transparency of the selection process was also an issue. This use of automated matching and explicit confidence intervals can be seen as an attempt in line with that trend.
Weaknesses Exposed by Common-Sense Violations
and Wordplay
For common words, all eight models secured scores from the 70s to the high 80s. In contrast, in the common-sense-violation category, the highest score was 34 and the lowest was 21. None of the common-sense-violation drawings by reference model Gemma were recognized.
Specific examples illustrate the structure of the weakness. “A cat cleaning a human’s toilet” received zero correct answers. “A mouse chasing a cat” drew many wrong answers. “Special forces travel” was invariably drawn as a soldier in camouflage. Dragged along by the literal wording of the buzzword, the act of traveling was omitted.
GLM’s drawing of “a fish fishing for a human” was emblematic. It depicted a fish sitting in a boat hoisting up a human; four of seven models guessed correctly, while three answered simply “fishing.” There is a strong tendency to revert to the default interpretation when a drawing is even slightly ambiguous. Both the drawing and guessing sides are pulled toward commonsense bias.
Paradox of Reasoning Effort and Cost Not
Translating Into Scores
Thinking longer did not lead to better drawings. Astra’s median reasoning tokens when drawing were just 232, taking 41 seconds. Yet it ranked first in drawing score. DeepSeek and Qwen spent 19,000 and 18,000 tokens respectively, but sank to the lower ranks.
Cost-effectiveness was also not linear. Gemini scored 74.6 points at a cost of about 82 yuan. The most expensive, Claude, cost about 304 yuan yet achieved a similar score. GLM had 40 guessing timeouts counted as wrong answers, meaning overthinking effectively became a loss. No stable positive correlation was seen between the amount of reasoning or spending and scores.
Remaining Challenges in Element Count and
Self-Recognition
Packing in more elements did not raise accuracy. The accuracy gap between the fewest-element and most-element groups was only 4.4 percentage points. DeepSeek drew “covering one’s ears to steal a bell” with 222 elements, and all seven models failed to guess it. When Gemini drew the same idiom with simple elements, all seven models guessed it correctly. Astra’s “apple” was guessed by seven models with just eight elements.
Self-guessing results are also suggestive. Astra, Gemini, and Claude had self-guessing accuracy of about 83%. MiniMax’s self-guessing was 53.3%, 13 points lower than being guessed by others. DeepSeek’s self-guessing was 49.3%, also nearly 10 points lower. Some models may lack consistent understanding of their own outputs. Rankings remained unchanged even after adjusting for drawer and guesser effects.
This article is based on Huxiu (All Rights Reserved). It relies on fair quotation under Article 32 of the Japanese Copyright Act.
Editorial Opinion
In the short term, we expect evaluation that separates drawing and guessing to spread to other evaluation designs. Rather than competing on a single-metric ranking, separating generation and understanding could become standard. The strong showing by domestic models is likely to broaden procurement options and boost demand for verification.
In the long term, we expect weakness on common-sense violations to become the practical focus. For explanatory diagrams and educational uses, accuracy in exceptional expressions will determine value. The effectiveness of simple drawings could also help curb generation costs, potentially prompting a shift in design philosophy.
The remaining question is what counts as correct. Automated matching is fair, but has limits in evaluating creativity. The issue may be where human judgment should intervene and how to reconcile it with reproducibility.
References
- ” 我们做了个“你画我猜Benchmark” :1350张图后,GPT-6险胜 ”, by 硅星人 — 虎嗅网, 2026-09-05T11:44:39.000Z (ARR)
- Source URL: https://www.huxiu.com/article/4888835.html?f=rss
Frequently Asked Questions
- What is the drawing-guessing game evaluation?
- It is a method in which pictures drawn by models in SVG are converted to PNG and guessed by other models. It does not use human aesthetic scoring, but judges correctness by automated matching against a word list. Its distinctive feature is the ability to measure generative ability and comprehension separately.
- Why did GPT-6 Astra win?
- Its high drawing score of 79.8 points lifted its overall score to 75.5. With a median of only 232 reasoning tokens, it drew communicative pictures in a short time. However, the confidence intervals of the top three models overlap, and the difference is small.
- Why did all models fail on common-sense-violation prompts?
- Because the default relationships learned during training are strong, and reversed relationships cannot be maintained in both drawing and guessing. When a drawing is ambiguous, it tends to revert to the usual interpretation. The result demonstrates weakness in handling exceptions.
Comments