Nayan Saxena
Teaching machines to think, express and create.
Nayan Saxena is an AI research scientist training frontier models at scale. His work studies how machines learn, reason, perceive and represent knowledge.
Work
Liquid Foundation Models 2 and 2.5Liquid AI × MIT CSAIL · Research Scientist · 2025 to 2026Open-weight multimodal models built with researchers from MIT CSAIL, shown in AMD's CES keynote, past a million downloads on HuggingFace. Joint vision-language architecture, mid- and post-training under tight parameter and latency budgets, and the video and grounding work for the VL models. Co-author on the technical report. Runs in a browser tab.
Phoenix, and the Canva Foundation ModelLeonardo AI, acquired by Canva · Founding researcher · 2024 to 2025Australia's first text-to-image foundation model, trained from scratch in Sydney with DeepSpeed and FSDP on billion-scale curated data. Canva bought Leonardo AI in August 2024 to get it, and it ships as DreamLab inside Magic Studio for around 200 million people. TIME Best Inventions 2024.
The AI data commonsMIT Media Lab, Data Provenance Initiative · 2024 to 2025The first systematic audit of robots.txt, terms of service and licence drift across 4,000 datasets in 67 countries, then the same audit for speech and video. NeurIPS 2024, ICLR 2025, then the New York Times, Nature and MIT Technology Review.
Value-based biddingRoyal Bank of Canada × Google · Data Scientist, Chief Data Office · 2022 to 2023A small real-time model behind every Canadian search for mortgage and investment products, answering in under 500 ms. Conversions rose 27 percent and conversion value 110 percent, worth about $100 million over the life of the model.
Wombo DreamWombo AI · Machine learning engineer, generative AI · 2022The first image-generation app most people ever touched, and the first to commercialise text-to-image. VQGAN+CLIP, then diffusion, then image-to-video, at 2 million daily users. Inference work on 800-million-parameter latent diffusion models, with techniques later adopted by the Diffusers library. Google's App of the Year.
Research
What intelligence is made of, studied in machines and in people.
How a system learns, reasons, perceives and knows. Every paper here answers one of the four.
When Does Continual Learning Require Learning?
Continual learning is not one capability. New domains, drifting facts and accumulating state each call for a different update, and the paper says which must be learned inside the weights and which can live in scaffolding. Project page.
LFM2 Technical Report
How a family of small, open-weight foundation models was built to run fast on phones and laptops: a hybrid architecture found by search under latency limits, and a training recipe that reaches the quality of much larger models.
Consent in Crisis: The Rapid Decline of the AI Data Commons
The open web is closing to training faster than any dataset can be rebuilt, and nobody had counted. 4,000 datasets, 67 countries.
Don't Think of the White Bear
Tell a language model not to think about something and, under load, it thinks about it more. Ironic rebound survives the move from people to transformers.
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
Instead of generating every candidate reasoning chain in full, prune the uninformative branches early using a training-free signal from the model itself, keeping the accuracy at a fraction of the cost.
ToDo: Token Downsampling for High-Resolution Diffusion
High-resolution diffusion carries redundant tokens in attention. Dropping them speeds generation with no visible cost. Shared by Gradio.
- 2026When Does Continual Learning Require Learning?Harrington A.*, Saxena N., Murphy M., Borovykh A., Yun Z., Kamath S., Kyi A. E., Darrell T., Malik J., Bai Y.*Preprint · UC Berkeley · project page
- 2026WelfareQA: Welfare and Social Protection Questions for Nigeria, India and KenyaSaxena N.First place, Uncharted Data Challenge, Adaption Labs
- 2025LFM2 Technical Report33 authors in equal contribution, alphabetical, Amini A.* to Tumma N.*, including Saxena N.*Liquid AI × MIT CSAIL
- 2025Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive LoadMann L.*, Saxena N.*, Tandon S.*, Sun C.*, Toteja S.*, Zhu K.CogSci · Workshop on Interpreting Cognition in Deep Learning Models, NeurIPS
- 2025Inference-Time Chain-of-Thought Pruning with Latent Informativeness SignalsLi S.*, Huang N.*, Saxena N.*, Luo N., Lin V., Zhu K., Dev S.Workshop on Efficient Reasoning, NeurIPS
- 2025Bridging the Data Provenance Gap Across Text, Speech and VideoLongpre S., Singh N., Cherep M., Tiwary K., Materzynska J., Brannon W., Mahari R., Dey M., Hamdy M., Saxena N., et al.ICLR
- 2024Consent in Crisis: The Rapid Decline of the AI Data CommonsLongpre S., Mahari R., Lee A. N., Lund C. S., Oderinwale H., Brannon W., Saxena N., et al.NeurIPS
- 2024ToDo: Token Downsampling for Efficient Generation of High-Resolution ImagesSmith E., Saxena N., Saha A.IJCAI
- 2022Sign-to-Speech Model for Sign Language Understanding: A Case Study of Nigerian Sign LanguageKolawole S., Osakuade O., Saxena N., Olorisade B.IJCAI, spotlight · Workshop on ML for the Developing World, NeurIPS, oral
- 2022Dynamic Strategy Selection in Active Function LearningSaxena N., Gelpi R., Buchsbaum D., Lucas C.CogSci · thesis
- 2022Towards One Shot Search Space Poisoning in Neural Architecture Search (Student Abstract)Saxena N., Wu R., Jain R.AAAI
- 2022NeuralArTS: Structuring Neural Architecture Search with Type Theory (Student Abstract)Wu R., Saxena N., Jain R.AAAI, oral
- 2021Sampling Heuristics for Active Function LearningGelpi R., Saxena N., Lifchits G., Buchsbaum D., Lucas C. G.ICCM · CogSci
- 2021Poisoning the Search Space in Neural Architecture SearchWu R.*, Saxena N.*, Jain R.*Workshop on Adversarial Machine Learning, ICML
- 2021Statistical Consequences of Dueling BanditsSaxena N., Chen P., Liu E.Workshop on Reinforcement Learning for Education, EDM, spotlight
- 2021IRT++: Improving Student Response Prediction With Gaussian Initialisation and Other ModificationsSaxena N., Lodaya V., Thakur T.ICALT
- 2020Political and Socioeconomic Influences on Social Distancing Behaviour in the United StatesKeating L., Saxena N., Cooper E., Tirico J., Khain D., Imahori D.preprint
- 2020The Nearest Neighbour Problem: An Algorithmic ApproachSaxena N., Shum A.technical report
* equal contribution
Teaching
Online courses, corporate training, and university teaching.
More than a hundred thousand people have learned from this work, online, in companies and in classrooms.
- Mastering Reasoning Models: Algorithms, Optimization, and Applications
- Advanced Quantization Techniques for Large Language Models
- Building Secure and Trustworthy LLMs Using NVIDIA Guardrails
- Mastering Large Language Models with the Cohere API
- AI Orchestration: Developing and Testing Your AI Prototype
- Hands-On Generative AI with Multi-Agent LangChain
- Hands-On AI: OpenAI Realtime API for Voice Conversations
- Hands-On Generative AI with Diffusion Models
- Jul, Aug and Sep 2026AI-Assisted Software Development with GitHub Copilot
- Jun 2026Researching with AI Assistants
- Jun 2026Data Quality in Python
- Nov 2025Multi-Agent Systems in LangGraph
- Sep 2025Generative AI Applications with LangChain
- Jul 2025Prototyping AI Agents with Langflow
- Jul 2026 and Sep 2026Foundations of AI Agents, AI Workflow Automation and Langflow
- Sep 2024Everything You Need to Know About LLMs in 2024
- Apr to Oct 2023Instructional specialist, AI and Data Analytics, School of Continuing Studies and edX
- Fall 2020 to Winter 2022Teaching assistant: Introduction to Statistical Reasoning and Data Science; Probability with Computer Applications; (Graduate) Methods of Data Analysis I
- Winter 2023Instructor, Diploma in Data Science, School of Continuing Education
News
The research and the models, in the press.