Yusuf Aytar
Massachusetts Institute of Technology
66 Papers
1K Citations
Yusuf Aytar is an academic researcher from Massachusetts Institute of Technology. The author has contributed to research in topics: Computer science & Reinforcement learning. The author has an hindex of 23, co-authored 54 publications. Previous affiliations of Yusuf Aytar include University of Oxford & University of Central Florida.
Chat about Author
Papers
•Posted Content
SoundNet: Learning Sound Representations from Unlabeled Video
TL;DR: In this article, the authors leverage the natural synchronization between vision and sound to learn an acoustic representation using two-million unlabeled videos and propose a student-teacher training procedure which transfers discriminative visual knowledge from well established visual recognition models into the sound modality using unlabelled video as a bridge.
910
Learning Cross-Modal Embeddings for Cooking Recipes and Food Images
Amaia Salvador,Nicholas Hynes,Yusuf Aytar,Javier Marin,Ferda Ofli,Ingmar Weber,Antonio Torralba +6 more
- 21 Jul 2017
TL;DR: This paper introduces Recipe1M, a new large-scale, structured corpus of over 1m cooking recipes and 800k food images, and demonstrates that regularization via the addition of a high-level classification objective both improves retrieval performance to rival that of humans and enables semantic vector arithmetic.
•Proceedings Article
SoundNet: Learning Sound Representations from Unlabeled Video
Yusuf Aytar,Carl Vondrick,Antonio Torralba +2 more
- 01 Jan 2016
TL;DR: This work proposes a student-teacher training procedure which transfers discriminative visual knowledge from well established visual recognition models into the sound modality using unlabeled video as a bridge, and suggests some high-level semantics automatically emerge in the sound network, even though it is trained without ground truth labels.
Temporal Cycle-Consistency Learning
Debidatta Dwibedi,Yusuf Aytar,Jonathan Tompson,Pierre Sermanet,Andrew Zisserman +4 more
- 15 Jun 2019
TL;DR: It is shown that the learned embeddings enable few-shot classification of these action phases, significantly reducing the supervised training requirements; and TCC is complementary to other methods of self-supervised learning in videos, such as Shuffle and Learn and Time-Contrastive Networks.
•Posted Content
With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations
TL;DR: Nearest-Neighbor Contrastive Learning of visual representations (NNCLR) as mentioned in this paper samples the nearest neighbors from the dataset in the latent space, and treats them as positives, which provides more semantic variations than pre-defined transformations.
264