Papers & Research

Work that leans more into research, evaluation, and writing than product UI, but still grounded in concrete systems and data.

Benchmarking Commonsense Visual Reasoning for Vision-Language Models

Co-author with Prof. Ernest Davis (NYU) · arXiv cs.AI · Submitted Mar 2026

Visibility-reasoning benchmark for vision-language models with a 2×2 XOR design across 100 families and 300 headline cells, with automated graders and scoring.

My role

  • Evaluated 9 VLMs across three tiers (flagship, prior-gen, open-source)
  • Surfaced a 26% abstention rate in GPT-5 as a key finding
  • Top scores: GPT-4o (0.728) and Gemini 3.1 Pro (0.727) effectively tied

Adversarial Robustness of 3D U-Net Radiation Therapy Dose Prediction

Co-author with Prof. Birjoo Vaishnav (U. Maryland) & Prof. Weixian Liao (Towson) · Accepted at SERA 2026 (IEEE/ACIS); submitted to ASTRO 2026 · 2025 – 2026

Evaluated adversarial (FGSM/PGD) and clinically plausible CT perturbations on the OpenKBP dataset using clinical DVH metrics.

My role

  • Identified resolution downsampling as the only perturbation with immediate clinical impact
  • Bone shift late takeoff at level 4-5; noise, bias field, and dental flat through extreme amplitudes
  • Junior Investigator Travel Award application for ASTRO 2026

Legal Text Classification with BERT and LegalBERT

Co-author · NYU Natural Language Processing · Course project · 2024

Fine-tuned BERT-family models on legal corpora; compared against classical baselines with reproducible training and reporting.

My role

  • Compared BERT-Double, Legal-BERT, and Custom-Legal BERT models
  • Implemented end-to-end NLP pipeline over legal case corpora
  • Proposed and implemented a composite evaluation metric for legal QA