CHROMA · A controlled benchmark for machine map understanding

Toward AI-Friendly Cartography:Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps

CHROMA measures how hue, color ordering, and lightness contrast shape foundation models' spatial reasoning on sequential choropleth maps.

Yonghe Sun · Zhenjia Liu · Hua Liao · Wenjia Xu · Nai Yang · Weihua Dong · Zhiwei Wei

Randomized encodingOrdinal structure broken
ClassesMixed
Sequential encodingMachine-readable
LowHigh
+9.2accuracy points
from ordered color
5,760controlled maps
28,800reasoning tasks
21multimodal FMs
5task dimensions

The question

Cartography was designed for people.
Do its visual rules transfer to machines?

CHROMA turns three classical color-design principles into controlled experiments, holding geography, values, questions, and rendering constant while changing only the visual encoding.

H1

Hue palette

Do different sequential hue families carry stable semantic advantages for foundation models?

Limited, inconsistent effect
H2

Ordinal organization

Does preserving the familiar light-to-dark sequence help models compare and rank regions?

Large, systematic effect
H3

Lightness contrast

How much class separation is enough for reliable machine-readable thematic values?

Asymmetric threshold effect
CHROMA framework from controlled map construction to benchmark evaluation
The complete CHROMA pipeline: controlled hypotheses, map construction, multi-level tasks, human validation, and evaluation of 21 models.

The benchmark

Five levels of map reasoning

Every controlled map is paired with questions that move from reading a single region to recognizing higher-order spatial patterns.

United StatesFranceGermanySwitzerland
01

Attribute identification

Match regions to thematic classes and classes to regions.

D1
02

Spatial recognition

Resolve direction and adjacency between geographic regions.

D2
03

Comparison

Compare thematic magnitudes between specified regions.

D3
04

Ranking

Find local or global extrema and rank regions by value.

D4
05

Pattern delineation

Recognize clusters, trends, and spatial structures.

D5

Core results

Structure beats decoration.

Models care far more about whether colors preserve magnitude and remain separable than about which hue family a map uses.

H2 · Color ordering−9.2 pts

Breaking the sequence breaks reasoning.

Mean accuracy falls from 53.7% to 44.5% when class colors are randomized. Ranking suffers most: D4 drops by 24.6 points.

Sequential53.7
Randomized44.5
H1 · Hue4.3 pt range

No universally best hue.

Across 18 sequential palettes, mean accuracy ranges from 51.2% to 55.5%. The winner changes across models and task types.

51.255.5
H3 · Contrast−2.0 pts

Avoid compressed lightness.

Average accuracy changes from 55.3% at standard contrast to 55.8% at high contrast and 53.3% at low contrast.

55.8High55.3Standard53.3Low

Main experiment · Accuracy (%)

Performance by model and color condition

H2 compares sequential and randomized ordering. H3 compares standard, increased, and reduced lightness contrast.

ModelH2 · Color orderingH3 · Lightness contrast
SequentialRandomizedStandardHighLow
Qwen3.5-9B68.653.670.870.967.6
Qwen3.5-4B62.649.163.663.660.2
Qwen3.5-2B49.541.451.252.049.1
Qwen3-VL-8B-Instruct62.148.063.364.160.3
Qwen3-VL-4B-Instruct58.845.859.760.256.7
Qwen3-VL-2B-Instruct43.436.344.544.742.6
Qwen2.5-VL-7B-Instruct51.041.652.253.150.1
Qwen2.5-VL-3B-Instruct44.238.046.247.445.1
GLM-4.6V-Flash64.351.966.566.662.7
Gemma3-12B-IT35.735.336.836.736.2
Gemma3-4B-IT29.629.729.629.429.8
InternVL3.5-8B59.146.362.264.260.4
InternVL3.5-4B58.647.059.960.957.0
InternVL3.5-2B41.235.441.041.940.2
Kimi-K2.678.965.180.981.879.4
Qwen3.6-plus44.341.146.947.744.7
Doubao-Seed-2.0-lite47.844.949.550.149.0
MiMo-V2.549.438.350.751.349.4
ERNIE 5.040.437.243.343.143.2
Gemini-3.5-Flash77.958.980.480.576.7
GPT-5.559.749.261.461.658.6
Average53.744.555.355.853.3

For H1, averages across the 18 hue palettes range from 51.2% (Purples) to 55.5% (YlGn); the complete palette-by-model table is available in the paper.

Ordering effect across models

The effect is broad, not model-specific.

Most tested systems lose accuracy when the same palette is detached from its natural ordinal progression. Human readers remain much more robust.

Contrast effect across models

Losses are asymmetric.

Reducing contrast hurts consistently, but extra contrast beyond an already separable standard condition adds comparatively little.

Design implications

Three practical rules for AI-readable choropleth maps

01

Preserve darker-means-more

Use the familiar directional convention, not merely any monotonic ordering. Reversing the convention can be especially harmful for comparison and ranking.

02

Protect class separability

Avoid compressed lightness steps. Once adjacent classes are clearly distinguishable, further expansion offers diminishing returns.

03

Choose hue for the audience

Because no hue family is consistently best for machines, hue can prioritize human accessibility, semantics, and context without sacrificing a universal model preference.

Matched 2×2 ablation

Where do the errors actually come from?

The same maps, questions, values, and answer choices are tested while attribute values and spatial relations are independently supplied as clean text. This separates color-and-legend decoding, spatial reasoning, and their integration without introducing OCR confounds.

Neither supplied64.0%Full visual task
Region values67.6%Spatial relations inferred
Spatial relations57.9%Values decoded from map
Both supplied68.9%Text-supported ceiling

What it shows

Color is useful, but the bottleneck is compositional.

  • Supplying values mainly improves attribute identification, comparison, and ranking.
  • Supplying spatial relations mainly helps topology-dependent questions.
  • Supplying both performs best, revealing an additional integration cost between the two information streams.

Can models learn around it?

Fine-tuning helps—design still matters.

LoRA adaptation substantially raises overall accuracy, yet the relative sensitivity to randomized ordering and insufficient contrast remains. Better models and better maps are complementary, not interchangeable.

93.2%Qwen3.5-9B after LoRA
on sequential maps

Open resources

Read, reproduce, and build on CHROMA.

Citation

Cite CHROMA

@article{sun2026chroma,
  title   = {Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps},
  author  = {Sun, Yonghe and Liu, Zhenjia and Liao, Hua and Xu, Wenjia and Yang, Nai and Dong, Weihua and Wei, Zhiwei},
  journal = {arXiv preprint arXiv:2608.15736},
  year    = {2026},
  doi     = {10.48550/arXiv.2608.15736}
}
Expanded research figure