Ben Vardi

My research focuses on improving the robustness and reliability of vision-only and vision-language models. Although state-of-the-art models achieve impressive results on standard benchmarks, they frequently exhibit brittle behavior in real-world settings. I develop efficient interventions to make large pretrained foundation models more robust and trustworthy.
One example of this problem comes from vision-language models. These models demonstrate remarkable visual understanding capabilities and can answer complex questions. However, they can make unnatural errors, such as providing answers to unanswerable image-related questions, for example, questions asking about objects that do not appear in the image. In our work CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering, we introduce a novel lightweight method that substantially reduces such errors.
Another direction of my research focuses on digital pathology. Foundation models can now analyze histopathology images and perform complex diagnostic tasks, yet they remain sensitive when tested on images originating from different hospitals. In collaboration with The Computational Pathology Laboratory at Sheba Medical Center, we are exploring ways to alleviate this issue.
I am also interested in the reliability of text-to-image models, particularly their ability to follow compositional prompts, involving multiple subjects and attributes.
Before starting my PhD, I completed a BSc in Biology and Psychology with an emphasis on Neuroscience at Tel Aviv University. I then earned an MSc in Computer Science from Ben-Gurion University, where I worked on puzzle-solving algorithms. I also worked as a software engineer at Waves Audio and as a computer vision engineer at Snap Inc.