English 箭头
Podcast Cover

[Human-Centric Vision: Advancing AI Through Few-Shot Learning and Self-Supervised Research]-[NVIDIA’s Shalini De Mello Talks Self-Supervised AI, NeurIPS Successes - Ep. 140]

NVIDIA AI Podcast · B2 · 2021-04-07

Technology
Or study on the web version

📋 Summary

Introduction to Human-Centric Computer Vision

In this episode of the NVIDIA AI Podcast, host Noah Kravitz interviews Principal Research Scientist Shalini DeMello, a specialist in computer vision and machine learning. DeMello’s work is fundamentally driven by a desire to solve "real-world problems" that impact human lives. Her research focus, which she terms "human-centric vision," involves analyzing human behavior through technology—specifically by tracking head orientation, hand gestures, and gaze estimation. A major catalyst for this research has been the automotive industry, where DeMello has worked on NVIDIA’s Drive IX platform to address the critical issue of distracted driving, noting that 80-90% of traffic accidents involve some form of human distraction.

The Challenges of Generalization and Few-Shot Learning

One of the primary hurdles in human-centric vision is that "one size doesn't fit all." DeMello explains that because every individual has unique anatomical features—such as interpupillary distance or corneal refractive indices—standard deep learning models often struggle to achieve high precision. To address this, she employs "few-shot learning," a technique that enables neural networks to train on a very limited set of examples (often fewer than 10).

DeMello highlights the risk of "overfitting" in these scenarios, where a model simply memorizes the small sample set rather than learning generalizable concepts. To combat this, she utilizes "meta-learning" or "learning to learn," which involves distilling information into a "lower dimensional space" to help the model learn more efficiently without requiring massive datasets.

NeurIPS Contributions: Mesh Reconstruction and Gaze Redirection

DeMello discusses two papers presented at the NeurIPS 2020 conference, both of which emphasize reducing the need for supervision or labels:

  1. Consistent Mesh Reconstruction in the Wild: This research explores how to reconstruct the 3D shape of non-rigid objects (like birds) from 2D video. By exploiting the "temporal signal" of video, the model enforces consistency across frames, allowing for 3D shape estimation even without 3D ground-truth models.
  2. Self-Learning Transformations for Gaze and Head Redirection: This project focuses on the controllable manipulation of faces. DeMello and her team developed a way to "disentangle" extraneous factors like lighting and background from the subject's head pose and gaze. This allows for precise, controllable editing, such as gaze correction in group photos, which has significant implications for both data synthesis in training and content editing.

The Evolution of AI and Research Philosophy

Reflecting on her career, DeMello notes that the industry has shifted significantly since she began. Early in her studies, neural networks were often dismissed as "black boxes." However, the breakthrough of convolutional neural networks around 2012 proved that these models could outperform "hand-designed features."

Looking ahead, DeMello is focused on making AI more "ubiquitous and available" by moving away from brittle systems that rely on massive, biased datasets. She emphasizes that understanding the "underlying compositional concepts" of data is essential for fairness and for enabling researchers in regions with limited data access to benefit from AI advancements.

Conclusion and Advice for Future Researchers

DeMello concludes by offering advice to the next generation of scientists: "Follow your passion" and "always be curious." She stresses that deep, genuine curiosity—questioning concepts rather than accepting superficial understandings—is what separates a great researcher from a good one. Her journey from physics and signal processing to cutting-edge AI serves as a testament to the power of aligning one’s core values with complex, human-focused problem solving.

🎯Key Sentences

1
I'm sure I've left a few things out.
2
I'll start by asking you to tell the audience.
3
What do you work on?
4
That's really what inspires me at the end of the day.
5
I won't throw stones in a glass house.
Expand All

📝Key Phrases

1
go hand in hand
2
put it more concretely
3
at the end of the day
4
not an apples to apples comparison
5
one size doesn't fit all
Expand All

📖 Transcript

Hello and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. Today we're looking internally, if you will, speaking with one of NVIDIA's own research scientists about her latest work and her career in computer vision and machine learning in general.
Shalini Demillo is a Principal Research Scientist at NVIDIA with a focus on self-supervised and few-shot learning, 3D reconstruction, viewpoint estimation, and human-computer interaction.
During her time at NVIDIA to date, Shalini has invented technologies for viewpoint estimation learned propagation, gazed estimation, 2D and 3D head pose estimation, hand gesture recognition, face detection, and video stabilization. and I'm sure I've left a few things out.
We're going to speak today about a recent poster Shalini created for last year's NeurIPS conference, that outlines a way to create animated avatars based on video.
And I'm sure we'll get to some other things as well.

ListenLeap Brings You Into Real Context Learning

🎨 Interesting Content
🌍 Real Materials
📱 Listen Anytime
Or study on the web version