Autoregressive Mosaics — Probing 2D Spatial Reasoning in Text-Only Language Models, by Ashwin Nedungadi, Stefan Oehmcke, and Stefan Lüdtke

AM-Bench · arXiv 2026

Autoregressive
Mosaics

Probing 2D Spatial Reasoning in Text-Only Language Models

Ashwin Nedungadi · Stefan Oehmcke · Stefan Lüdtke

University of Rostock · Institute for Visual & Analytic Computing (VAC)

Overview

Autoregressive Mosaics introduces AM-Bench: a benchmark that separates translating given geometry into code from composing a layout from a text prompt. We evaluate eight open-weight models (8B–34B) on 150 layout prompts, producing 13,200 generation attempts.

How well can text-only language models compose 2D layouts, how does the output medium affect this ability, and what spatial information do they represent before drawing?

Geometry is easier than composition. All eight models reach a median part-wise intersection-over-union of 1.0 when geometry is specified. Their ability to compose a scene varies substantially.

The medium matters. Switching from the custom Python drawing API to SVG improves pooled layout scores by 0.37 points on a 0–5 scale, at the same 24 × 24 resolution.

A coarse layout is decodable before drawing. Probes outperform a text baseline in all eight models. Further analysis of two models recovers shared layout structure, but not their specific output layouts—consistent with incremental construction rather than a complete advance plan.

These results concern coarse, code-generated images; layout quality is assessed by two vision-language judges.

Demonstration

An early prototype of generating mosaics through text and code.

Benchmark

Models write Python using six drawing primitives on a deterministic 24 × 24 canvas. Two tasks separate instruction translation from spatial composition.

01
Translation task
Follow the geometry

Translate fully specified shapes, positions, and colors into drawing code. Part-wise intersection-over-union measures agreement with exact reference geometry.

Specified LayoutExact References
02
Layout task
Compose the scene

Choose geometry and placement from an underspecified prompt. Two vision-language judges score fidelity, shape, color, spatial arrangement, and completeness across elemental, iconic, and compositional prompts.

150 PromptsThree Tiers
Citation
@misc{nedungadi2026autoregressivemosaics,
  title = {Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models},
  author = {Nedungadi, Ashwin and Oehmcke, Stefan and Lüdtke, Stefan},
  year = {2026},
  eprint = {2608.30751},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url = {https://arxiv.org/abs/2608.30751}
}
Autoregressive Mosaics · AM-Bench
Copyright © 2026 Ashwin Nedungadi. All Rights Reserved.