
Vasu Sharma
Head of AI
PocketFM
About Vasu
I am presently the Head of AI at PocketFM working on long form video generation, content creation and understanding, building foundational models for expressive TTS, creative writing models and foundational research on AI x Narratology. PocketFM is a late stage, 500M$ ARR company which is the leader in long form AI native media entertainment platform. Previously, I worked as a Senior Staff Research Scientist at Tesla Optimus, working on building Causal, Dynamic, Real time World Models to enable closed loop RL training for humanoid robotics. I am also working on building Multimodal foundation models, efficient video generation models, Robotics Foundation models and high fidelity autolabelling pipelines working at massive scale to enable reliable training for Robotic action models and to develop true generalization capabilities in a diverse set of real world usecases. I have also worked as an Applied Research Scientist Lead at Facebook AI research, working on building Multimodal foundational generative AI models. I am also interested in the domain of self supervised learning. I have published 100+ papers across top AI conferences like NeurIPS, CVPR, ACL, EMNLP, TMLR, ICLR, NAACL, COLM, EACL, WACV, Interspeech among others garnering over 16k+ citations. I routinely work with multi-billion scale datasets to train these massive multimodal models. In the past, I have also worked as Quantitative Researcher at Citadel where my work involved leveraging the power of Machine Learning and Statistical methods in an attempt to fathom the enigmatic world that is the financial markets. I have also worked at Amazon Alexa AI on large scale multimodal models and Embodied AI applications to bring smart robot intelligence to Alexa devices. I actively advise several early stage startups and often guest lecture at Stanford, CMU, MIT, Oxford, Oreilly among others.
I graduated from Indian Institute of Technology, Kanpur completing my Bachelors in Computer Science and Engineering and then completed my graduate school in Machine Learning and Artificial Intelligence at the Language Technologies institute at Carnegie Mellon University.
- Track
- Vision, Spatial, & World Models [Technical]
- Industry
- Media & Entertainment
- Job Function
- AI/ML
- Company Size
- 501-1,000 employees