COLM 2026

The Percept-V Challenge:
Can Multimodal LLMs Crack Simple Perception Problems?

1Computer Science and Engineering 2Yardi School of Artificial Intelligence Indian Institute of Technology, Delhi

Modern MLLMs solve olympiad maths from an image. They cannot reliably count ten circles. Percept-V isolates basic visual perception from reasoning and knowledge β€” and models fail.

6,000
program-generated images
30
domains × 20 problem sizes
7
TVPS-4 perception skills
95.4%
human accuracy
66.9%
best model accuracy

One task per skill

Seven instances from Percept-V, one for each TVPS-4 skill. Each card shows the summarised prompt, the gold answer, and GPT-4o's response. Every image is drawn by a Python program, so the gold answer is known exactly and the image cannot have been seen during pre-training.

Figure 1. Sample questions from Percept-V illustrating tasks related to different visual perception skills as defined by the TVPS-4 framework.

Abstract

Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large Language Models (MLLMs) match up to humans in basic perception? Even though many benchmarks evaluate MLLMs on advanced reasoning and knowledge skills, there is limited research that focuses evaluation on simple perception. In response, we introduce Percept-V, a dataset containing 6000 program-generated uncontaminated images divided into 30 domains, where each domain tests one or more TVPS-4 skills. Our focus is on perception, so we make our domains quite simple and the reasoning and knowledge required for solving them are minimal. Since modern-day MLLMs can solve much more complex tasks, our a-priori expectation is that they will solve these domains very easily. Contrary to our belief, our experiments show a weak performance of SoTA proprietary and open-source MLLMs compared to very high human performance on Percept-V. We find that as the number of objects in the image increases, performance goes down rather fast. Our experiments also identify the perception skills that are considerably harder for all models. Fine-tuning an open-source MLLM shows considerable gains in performance, though the gains only marginally carry over to other related datasets, pointing to limitation in generalization abilities of the learned representations. Code and data are available at github.com/dair-iitd/Percept-V.

Key findings

The dataset

Every instance is a pair (Q, I): a domain-specific question prompt and a Python-generated image of simple geometric objects. Ground truth is computed during generation, so there is no annotation noise and no contamination. Each domain spans 20 problem sizes with 10 instances each β€” 200 per domain, 6,000 in total.

Grounded in TVPS-4

The Test of Visual Perceptual Skills is a standard clinical instrument for human perception. Its seven skills give the benchmark a construct that predates and is independent of LLMs.

Problem size is a dial

Size 1–20 operationalises perceptual complexity: number of objects, grid dimension, distractors, or processing steps. This is what makes the scaling curves possible.

Plug-and-play generation

One script per domain, driven by --num_image and --num_size. Any size of dataset can be regenerated on demand with fresh, unseen images.

The seven skills and the domains that use them

A domain may appear in several rows: solving its instances can require more than one skill.

Table 1. Domains using each skill. 17 domains test a single skill, 7 test two, and 6 test three jointly.

Answer formats

Each prompt specifies its output format explicitly so that scoring is automatic. Proprietary models comply almost perfectly β€” for them, errors come from perception rather than formatting. Open-weight models are noticeably worse: DeepSeek fails the format on 61.6% of fixed-length answers, and even Gemma-4 and Qwen3-VL slip on 4–15%.

TypeDescriptionDomains
Format error rate (%)
TypeGPT-5-miniGPT-4oo4-miniGeminiGemma-4Qwen3-VLQwen2.5-VLDeepSeek

Explore all 30 domains

Filter by skill or answer type, then open a domain to see its exact prompt and a sample image at problem size 9.

Skill
Answer type

Results

Four proprietary models (GPT-5-mini, GPT-4o, o4-mini, Gemini 2.5 Flash) and four open-weight models (Gemma-4-31B-it, Qwen3-VL-8B-Thinking, Qwen2.5-VL-7B-Instruct, DeepSeek-VL2-Tiny), zero-shot, temperature 0. All except GPT-4o, Qwen2.5-VL and DeepSeek support β€œthink with image”. Cells are shaded by accuracy; click a column header to sort.

Table 2. Comparison of MLLM performance across skills and domains (%), zero-shot.

Nothing is solved

Gemma-4, the strongest model, is still ≀50% on 9 of 30 domains. GPT-4o is ≀50% on 24, o4-mini on 17, and DeepSeek-VL2-Tiny scores exactly 0.0% on 16.

The floor: colours_present

2.62% average β€” the hardest domain for every model. It only asks which of 20 named colours appear, with swatches and a legend provided. Humans get 90%; the best model gets 6.67%.

The ceiling: inside_circles

91.0% average. A single red dot, some circles, one yes/no answer. When the task reduces to one object relation, models are fine β€” the difficulty is in holding many objects at once.

Thinking beats provenance

Open-weight Gemma-4-31B tops the table at 66.9%, and Qwen3-VL-8B edges past GPT-4o. Every thinking model outranks non-thinking GPT-4o; the two non-thinking open models trail far behind.

Accuracy against problem size

One panel per skill plus an average. Every model degrades monotonically as instances grow β€” the single most consistent pattern in the paper.

Line plots of accuracy versus problem size for each of the seven TVPS-4 skills plus an average panel, showing a consistent downward trend for all eight models.
Figure 2. Overall accuracy of all models across problem sizes, per skill.

Counting, OCR and grid understanding

Three skills recur across the dataset outside the TVPS-4 taxonomy: counting (13 domains), OCR (9) and grid understanding (7). Counting and OCR degrade sharply with size; grid tasks stay nearly flat for reasoning models, plausibly a training-data effect. Counting is slightly harder for most reasoning models (GPT-5-mini: 47.7 counting vs 60.9 non-counting), but Gemma-4 is the exception β€” it is better on counting (70.5) than on non-counting (64.2). Either way, weak performance on non-counting domains shows that counting alone does not explain the failures.

Accuracy versus problem size for counting, grid understanding and OCR skills.
Figure 3. Accuracy across problem sizes for counting, grid understanding and OCR.

Single-skill domains

Restricting to the 17 domains that require exactly one skill isolates skill difficulty. Visual Sequential Memory is the weakest skill for 4 of 8 models; Visual Spatial Relationship is the strongest for all of them. Only five skills appear β€” Visual Figure Ground and Visual Closure have no single-skill domains, and the number of domains per skill varies sharply (vsr: 6, vm: 5, vfc: 3, vd: 2, vsm: 1), so the columns are not equally reliable.

Bar chart of model accuracy on single-skill domains grouped by skill, one panel per model.
Figure 4. Average accuracy on single-skill domains, grouped by skill.

Human study

16 college students answered randomly selected questions spanning all 30 domains and all problem sizes. The models were evaluated on exactly the same samples under exactly the same protocol.

Table 3. Overall accuracy on the human-study samples. The gap to the closest MLLM, Gemma-4-31B, is 26.7 points. On colours_present it is 90% versus 6.67%.

Radar chart comparing human accuracy against seven models across the seven TVPS-4 skills; the human polygon encloses every model polygon.
Figure 5. Model accuracy versus human accuracy across the seven TVPS-4 skills.

Fine-tuning

We build Percept-Vtrain β€” 60K samples, 2K per domain β€” and fine-tune Qwen2.5-VL-7B, varying which modules are trainable: the language model (L), the vision encoder (V), and the MLP projector (P) between them. Out-of-domain transfer is measured on CV-Bench and Color-Bench.

LLMVision
encoder
MLP
projector
Percept-V
overall
size
1–10
size
11–20
CV-Bench Color-Bench

Table 4. LoRA (rank 8) and full fine-tuning of Qwen2.5-VL-7B. LoRA ordering on Percept-V: (L+V) > (L+V+P) > (L) > (L+P) > (V+P) > (V) > (P).

+66 points in-domain

18.1% β†’ 84.3%. LoRA gets within 1 point of full fine-tuning, so the capability is a small, learnable delta rather than something the model fundamentally lacks capacity for.

Almost no transfer

CV-Bench moves 73.9 β†’ 77.6 and Color-Bench 46.3 β†’ 46.0. The model learns the tasks, not the perception. No catastrophic forgetting either.

The projector is inert

Training it alone gives 49.2. Adding it to the best pair lowers accuracy (83.4 β†’ 80.3). One MLP cannot absorb the mismatch between the two modalities.

That the projector is inert may imply models struggle to query, align and transform visual information, and motivates stronger interfaces between the vision and language modalities.

Ruling out the easy explanations

If models fail at counting circles, the natural objections are that the images are too small, too flat, or the prompts too odd. None of them hold.

In-context examples make it worse51.3% vs 54.8%

One-shot prompting degrades every proprietary model. A short few-shot run with three examples on Gemini dropped average accuracy on Percept-V to 21.5%. Inspecting responses, models confuse the details of the in-context example with the problem to be solved, and hallucinate wrong theories after mis-perceiving the exemplar itself. Weak performance is therefore unlikely to be a task-understanding problem.

SkillGPT-5-miniGPT-4oo4-miniGemini

Table 5. One-shot accuracy (%). Zero-shot averages of skills for reference: 54.77 / 29.26 / 53.63 / 52.27.

Accuracy versus problem size in the one-shot setting, showing the same downward trends as zero-shot.
Figure 6. Accuracy across problem sizes in the one-shot setting.
Resolution is not the bottleneck+0.9 at 2×

Percept-V images average 670×650 pixels. Upscaling 2× gives no benefit (GPT-5-mini 54.77 β†’ 55.62; Qwen 14.64 β†’ 15.01), so the information needed was already present β€” these are not failures of acuity. Downscaling 0.5× does hurt (βˆ’13.6 and βˆ’2.1), confirming the models genuinely read the images rather than guessing. The skills that suffer most under downscaling are exactly those requiring many small objects to be individuated at once.

SkillGPT-5-miniQwen2.5-VL
Orig.2×0.5×Orig.2×0.5×

Table 6. Accuracy (%) at original, 2× and 0.5× resolution.

3D rendering does not rescue it73.5% → 55.3%

A potential criticism of Percept-V is that its images are auto-generated and may not occur naturally in a model's training data. Would models do better on more realistic images? Six domains were rebuilt as 3D Blender scenes in the style of CLEVR β€” solid forms, naturalistic lighting, cast shadows β€” with 100 images each over sizes 1–10. Even at these small problem sizes, where 2D accuracy is comparatively high, 3D performance stays sub-par: GPT-5-mini gets worse on all six domains, falling 18 points overall. Qwen, already near its floor, gains slightly on the two counting domains, plausibly because shading separates adjacent objects that flat silhouettes leave ambiguous. Neither model improves where it would matter most: colours_present stays at 9.0 and 0.0. Photorealism redistributes difficulty rather than removing it.

Grid of example images from 3D-Percept-V rendered in Blender, showing solid shapes with realistic lighting and shadows.
Figure 7. Examples from 3D-Percept-V, created via Blender for a subset of domains.
DomainGPT-5-miniQwen2.5-VL
2D3D2D3D

Table 7. Accuracy (%) on six domains rendered in 2D and 3D, matched on problem sizes 1–10.

Prompt wording barely matters±1 point

The original prompts were paraphrased into concise and verbose styles. Average of skills moves by about one point for both models, and the original chain-of-thought style prompting is often the best of the three. Concise beats verbose for both.

SkillGPT-5-miniQwen2.5-VL
ConciseVerboseConciseVerbose

Table 8. Accuracy (%) across prompting styles.

Ordered answers are harder than unordered oneslist vs set

Breaking accuracy down by answer type across problem sizes shows that thinking models (GPT-5-mini, Gemini, o4-mini) are relatively robust to size for the unordered set type, but not for the ordered list type. Producing an answer in the right order is a distinct challenge from producing the right elements.

Accuracy versus problem size split by answer type: single, fixed, list and set.
Figure 8. Accuracy across problem sizes by answer type.

BibTeX

@article{ghosh2025percept,
        title={The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?},
        author={Ghosh, Samrajnee and Agarwal, Naman and Garg, Hemanshu and Mittal, Chinmay and Singla, Parag and others},
        journal={arXiv preprint arXiv:2508.21143},
        year={2025}
      }