Benchmarking Commonsense Visual Reasoning for Vision-Language Models
Co-author with Prof. Ernest Davis (NYU) · arXiv cs.AI · Submitted Mar 2026
Visibility-reasoning benchmark for vision-language models with a 2×2 XOR design across 100 families and 300 headline cells, with automated graders and scoring.
My role
- Evaluated 9 VLMs across three tiers (flagship, prior-gen, open-source)
- Surfaced a 26% abstention rate in GPT-5 as a key finding
- Top scores: GPT-4o (0.728) and Gemini 3.1 Pro (0.727) effectively tied