• Event Date: April 4, 2011
  • Event Start Time: 12:00 PM
  • Event End Time: 7:00 PM
  • Event Location: Stony Brook University, Department of Computer Science
  • Event Type: Human and Computer Vision Series
  • Event Semester: Spring 2011
  • Event Contact: Dr. Tamara Berg
  • Event Extra info: <a href="http://tamaraberg.com/">Dr. Tamara Berg</a>


People communicate using language, whether spoken, written, or typed.  A significant amount of this language describes the world around us, especially the visual world in an environment, or depicted in images or video.  Such visually descriptive language is potentially a rich source of 1) information about the world, especially the visual world, and 2) training data for how people construct natural language to describe imagery. In addition there exist billions of photographs with associated text available on the web; examples include web pages, captioned photographs, and video with speech or closed captioning.  In this talk I will describe several projects related to images and depiction, including: automatically labeling faces in news photographs, discovering visual attribute terms from noisy web collections, and generating simple natural language descriptions for images.  All papers, created data sets, and demos are available on my web page at: http://tamaraberg.com/

Background Readings:

Reading 1: http://tamaraberg.com/papers/berg_tnips.pdf
Reading 2: http://tamaraberg.com/papers/attributediscovery.pdf