PAPER-DIGEST · 2026-08-15
Li et al.: Measuring Whether AI Really Sees Shape, via Jigsaw Puzzles — Fukai Reads
Spatial reasoning in vision-language models / benchmark design
TL;DR
Is "seeing the picture" the same ability as "knowing whether two pieces fit"? JigShape, released by a team led from the University of Southern California, is a benchmark (a shared problem set plus scoring method) that tries to measure exactly that, using jigsaw puzzles. The target is vision-language models (VLMs — AI systems that take an image and text together and answer in text), and the question is how far they can handle spatial and geometric constraints.
Earlier jigsaw-style benchmarks cut pieces into rectangles. In uniform regions such as sky or grass, different arrangements then look identical, so the ground truth is not unique. JigShape cuts pieces with interlocking tabs (convex bumps) and blanks (concave notches) so that the correct answer is forced to be unique, and it spans grid densities from 4x4 up to 16x16.
The results are sobering. With no task-specific training, only GPT-5.5 clearly beat chance, reaching 69.65% piece accuracy on 4x4 (chance is about 6.25%). That falls to 4.37% on 8x8. Fine-tuning reaches the 97% range on 4x4 but drops below 5% on 12x12. And when the tab-and-blank shapes are removed, 97% collapses to 10%. It reads as though the models were leaning on shape cues rather than picture content. This is an arXiv preprint (submitted 30 July 2026, revised 4 August 2026, not yet peer-reviewed).
Introduction
The paper is arXiv:2607.27670, primary category cs.CV (computer vision), cross-listed to cs.AI. There are fourteen authors — Shawn Li, Wei Yang, Jike Zhong and others — spanning the University of Southern California, the National University of Singapore, Adobe Research, Texas A&M University, Rice University, and the University of North Carolina at Chapel Hill. The abstract page lists no venue or acceptance note, so the accurate description today is a preprint with no stated peer review. It is roughly two weeks old and not yet widely discussed.
I picked it today because "can AI solve this puzzle?" is turning into a practical question for people who make puzzles. You might want automatic difficulty estimation, or automatic playtesting, or you simply notice that readers paste a screenshot of your board straight into a chatbot. All of those hang on one thing: how far can today's AI actually get on visual puzzles?
This paper attaches concrete numbers to that question, and it does so with jigsaw puzzles, a subject everyone has physically handled. What follows is written so the gist comes across without opening the paper.
Background
What was already known: vision-language models are quite strong at saying what is in a picture (recognition, captioning) yet stumble on basic spatial relations. Left versus right, above versus below, in front versus behind, depth — benchmarks such as VSR (Liu et al., 2023) and BLINK (Fu et al., 2024) have repeatedly flagged these as weak spots.
Using jigsaw puzzles for this is not new either. Jigsaw-Puzzles (Lyu et al., 2025) built 2x2 and 3x3 boards from 1,100 images. JPwLEG (Song et al., 2023) used 3x3 and 5x5 with eroded gaps between pieces. Visual Jigsaw (Wu et al., 2025) went to roughly 118,000 instances and used jigsaw solving itself as a post-training task (extra training applied after pre-training). All of them cut rectangles.
The authors object on two grounds. First, ambiguity: with rectangular cuts, uniform-texture regions make several arrangements visually indistinguishable, which in their words leaves the task "ill-posed" (no unique correct answer). Second, coarseness: grids of 2x2 to 5x5 do not put enough combinatorial stress on a model. JigShape starts from the attempt to remove both problems at once.
Approach
The core idea is the cut. A piece edge is one of three kinds. A tab is a convex semicircular protrusion; a blank is a concave semicircular indentation; a flat is a straight edge, appearing only on the outer boundary. Adjacent pieces must have complementary edges, tab fitting into blank. Because of that geometric constraint, the correct arrangement is unique even where the picture is uniform — the same reason a physical jigsaw lets you settle the sky by feel.
Figure 1 from Li et al., “JigShape” (arXiv:2607.27670, CC BY 4.0)
There are four grid densities. 4x4 has 16 pieces and roughly 2x10^13 arrangements, with a tab ratio of 0.15. 8x8 has 64 pieces and about 10^89, ratio 0.20. 12x12 has 144 pieces and about 10^249, ratio 0.24. 16x16 has 256 pieces and about 10^507, ratio 0.28. Tab depth scales with density so the bumps stay visible as pieces shrink (Table 2 of the paper).
The source material is photographs: 900 images from DIV2K, 1,500 from DIV8K, and 21,342 curated from Unsplash, for 23,742 unique images. Each image yields one instance per grid setting, giving 95,468 instances in total. In each task the shuffled pieces are laid out in ID order — not in their correct positions — and the model must output a mapping from each piece ID to a (row, column) position.
Scoring is not a single number but four. Piece Accuracy is the fraction of pieces placed correctly. Exact Match is the fraction of instances where every piece is correct. Adjacency Accuracy is the fraction of genuinely adjacent pairs that are also adjacent in the prediction. Shape Compatibility is the fraction of predicted adjacencies whose edges actually interlock. Chance levels are given as 6.25% piece accuracy on 4x4 down to 0.39% on 16x16, and 30.8% to 44.1% for shape compatibility. Splitting the score four ways is, as I will argue later, the part designers should look at.
Findings
Six zero-shot configurations were evaluated with no task-specific training: GPT-5.5, GPT-5.4-mini, Claude Opus 4.8, Grok-4.2 (reasoning and non-reasoning), and Llama-4-Maverick, at 250 instances per grid setting. The headline numbers from Table 4: on 4x4, GPT-5.5 reaches 69.65% piece accuracy and 26.40% exact match, while everything else sits near the 6.25% chance level — Claude Opus 4.8 at 8.50%, Grok-4.2 with reasoning at 11.97%, Grok-4.2 without reasoning at 6.90%, GPT-5.4-mini at 6.75%, Llama-4-Maverick at 6.62%. Exact match is 0.00% for every model except GPT-5.5, and the authors write that only GPT-5.5 exceeds the random baseline on 4x4.
Then the cliff. Even GPT-5.5 falls to 4.37% on 8x8 and 0.67% on 12x12, where chance is about 0.69%. The authors call this a "scaling cliff" — performance dropping off sharply as scale increases — and say the models cannot keep satisfying the constraints as the piece count grows.
What if you train on the task? The authors supervised-fine-tuned Qwen3-VL-8B and Gemma3-12B on 6,500 mixed-grid examples (3,000 at 4x4, 2,000 at 8x8, 1,000 at 12x12, 500 at 16x16). On 4x4 the numbers jump: Qwen3-VL-8B to 97.28% piece accuracy and 90.80% exact match, Gemma3-12B to 97.82% and 89.20%. But on 8x8 they get 27.34% (exact match 0.00%) and 33.58% (0.40%); on 12x12, 3.44% and 4.57%; on 16x16, 0.40% and 0.35%. Training did not fill in the cliff.
The most suggestive result is the final ablation study (removing one design element at a time to see which part is doing the work). On 500 paired 4x4 instances, stripping the tab-and-blank shapes back to rectangular cuts dropped Qwen3-VL-8B's piece accuracy from 97% to 10% — barely above chance. From this the authors conclude that the models learned to exploit shape cues with only limited integration of visual content.
Where This Is Useful
First: do not plan on handing automatic playtesting or automatic difficulty estimation of visual placement puzzles to a current vision-language model. Once you pass about sixteen pieces the models are indistinguishable from guessing, so a design that shows the board image to an AI and asks "is this hard?" will not hold up. If I were building a daily jigsaw or sliding puzzle, I would pass a symbolic representation instead — which piece has which edges, what interlocks with what — and let a combinatorial solver work on it, measuring difficulty by search effort.
Second: piece count is not the only difficulty dial. What the ablation exposed is that difficulty jumps the moment you remove shape cues. Flipped around, a jigsaw-like puzzle's difficulty can be designed along two axes: the tab ratio (how much of the boundary interlocks) and how uniform the picture is. Increase the sky-and-grass regions and reduce the distinctiveness of the outlines, and you get harder without adding pieces. Adding pieces mostly adds monotonous minutes, so I think the two-axis version is the better experience.
Third: if you want a daily puzzle that is not trivially solved by an AI, constraint satisfaction that requires vision and geometry at once is still fairly robust. Boards of 8x8 and larger were effectively out of reach for every model tested here. But with GPT-5.5 at 69.65% on 4x4, small boards are already within range — and these numbers describe the models the authors tested at the time, not necessarily next year's.
Fourth, a lesson from the benchmark design itself. The paper's smartest move is making the answer unique by cutting with tabs and blanks. That translates directly into puzzle design: if your generator emits boards with multiple valid solutions, your difficulty metric is measuring something other than what you think — the same reason Sudoku insists on a unique solution. And splitting the score into piece accuracy and adjacency accuracy is a usable idea for hints and progress display. "Are these two correctly next to each other?" is a kinder progress signal mid-solve than "is this piece in its final position?"
Limitations
The authors first. As far as I could read, the paper has no standalone limitations section — it closes with a conclusion and an ethics and reproducibility statement — so the limits are embedded in the discussion of results. Two are stated explicitly. One is that 16x16 evaluation was skipped for the frontier models because performance was already poor at 12x12. The other is that the fine-tuned models depend on shape cues, with only limited integration of visual content.
What Fukai would point out here is threefold. First, the output format. For a 256-piece board the model must write out a 256-line mapping in text, which measures the ability to hold a long structured output together as much as the ability to solve a puzzle. Part of the cliff may come from that, and the paper does not appear to separate the two.
Second, sample size. Zero-shot evaluation used 250 instances per grid setting. Against a benchmark of 95,468 instances, the number actually scored is quite small — thin ground for comparing an exact match of 0.00% against one of 0.40%.
Third, the material is skewed. Every image comes from DIV2K, DIV8K, or Unsplash, all photographic datasets. Whether the same cliff appears with illustrations, abstract patterns, or near-monochrome figures cannot be told from this paper. Puzzle makers usually care most about exactly those cases, so I would call this open work.
Fukai's Reading
This paragraph is Fukai's own reading. I want to read this study as an attempt to separate "seeing" from "fitting." In the vocabulary of design criticism, a jigsaw carries two signals: a content signal, whether the picture continues across the seam, and a formal signal, whether the edges interlock. What the ablation exposed is that the fine-tuned model got to 97% on the formal signal alone, barely using the content signal. That is not something to scold it for — human jigsaw enthusiasts also grab the border and corner pieces first and work by shape. What nags at me is the other implication: a benchmark that believes it is measuring visual reasoning may in fact be measuring dexterity at constraint satisfaction. For anyone designing puzzles this is not someone else's problem. The dial we think we are turning when we tune difficulty may quite easily be a dial for something else.
Closing
If you want to go deeper, reading a few predecessors together gives you the map. Jigsaw-Puzzles (Lyu et al., 2025) uses the same puzzle type at 2x2 and 3x3 and separates seeing from understanding from reasoning. Visual Jigsaw (Wu et al., 2025) runs the other way, using jigsaws to train rather than to measure. For the underlying weakness in spatial relations, BLINK (Fu et al., 2024) and VSR (Liu et al., 2023) are the starting points.
If you are curious about measuring AI reasoning on symbolic puzzles instead, our earlier article on Pencil Puzzle Bench — a benchmark that evaluates LLMs on Sudoku, Slitherlink and the like — makes a good companion read for seeing where visual and symbolic puzzles converge and where they part. JigShape's training data is public on HuggingFace, so looking at the boards yourself is quick.
After reading it I found myself circling, in pen on printed paper, what happens between 4x4 and 8x8. Between 16 pieces and 64 pieces is not a factor of four but something like 10^76 times the arrangements. Standing in front of that gap, the stranger fact may be that humans finish jigsaws so casually.
References
Papers and materials referenced in this article:
・HTML version of the paper (body text, Table 2 and Table 4)
・JigShape-Train dataset (HuggingFace)
・Related work: Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models (Lyu et al., 2025)
・Cited in the paper as related work: JPwLEG (Song et al., 2023) / Visual Jigsaw (Wu et al., 2025) / BLINK (Fu et al., 2024) / VSR (Liu et al., 2023)
Reactions (no login)
Anonymous • one of each per visitor per day
Part of these series
Paper DigestEpisode 59 of 91
