Our Podcast
Conversations with leading women in AI research from around the globe.

August 19, 2026
An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics the field relies on, and why CLIPScore, one of the most widely used measures for scoring image descriptions, stops correlating with human judgment the moment you introduce context. We also get into why longer descriptions aren't necessarily more informative, what happens when you just ask a model to "be concise," and whether AI models trip over charts and graphs the same way humans do.Elisa's research sits at the intersection of linguistics, accessibility, and multimodal AI, and this conversation covers the full arc of her work, from the theoretical question of why humans never describe images the same way twice, to the practical question of what that means for building systems that actually work for blind and low-vision users.In this episode:Why "context matters" is more radical than it sounds for image description evaluationThe hidden reason CLIPScore breaks down once context enters the pictureWhy length is a bad proxy for information density - and what to use insteadWhat happens when you prompt a model to just "be concise"Why charts and photos need completely different evaluation approachesWhether AI models make the same mistakes as humans when reading data visualizationsWhat NeurIPS's Top Reviewer Award taught her about writing a genuinely useful peer reviewResources & Links:Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsWhen More Words Say Less: Decoupling Length and Specificity in Image Description EvaluationCHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Modelsđ§ Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:â â LinkedInâ â â â Blueskyâ â â â X (Twitter)â â â â â WiAIR websiteâ

July 22, 2026
Only 19% of Americans say AI has actually improved their productivity - so why the gap between the hype and reality? In this episode of Women in AI Research, Dr. Malihe Alikhani (Northeastern University, Contextual AI Lab) unpacks the hidden failures in how we build and deploy AI: why sycophancy is really a collapse of alignment, why bigger models aren't better aligned, and why "thin" alignment breaks down in the real world.Key topicsThe impact of moving across different AI contexts on system designThe role of language as performative and active in shaping realityInteractive inference and uncertainty in AI systemsThe importance of context in meaning and system designAI policy, transparency, and societal impactSycophantic behavior in large language modelsMeasuring AI alignment: thin vs. thickAI adoption across sectors and demographic groupsThe role of policy in AI development and safetyEthical considerations in AI research and deploymentResources & Links:Breaking the AI Mirror: Sycophancy, productivity, and the future of collaborationHype and harm: Why we must ask harder questions about AI and its alignment with human valuesHow are Americans using AI? Evidence from a nationwide surveyConnect with Dr. Malihe Alikhani:https://x.com/malihealikhaniđ§ Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:â LinkedInâ â Blueskyâ â X (Twitter)â â â WiAIR websiteâ

June 17, 2026
Do large language models truly understand languageâor are they sophisticated pattern matchers?In this conversation, Dr. Anna Ivanova (Asst. Prof. at Georgia Tech) explores one of the important questions in AI: the relationship between language, thought, and intelligence. Drawing from neuroscience, cognitive science, and AI research, Anna explains why language understanding is harder to define than most people realize, why reasoning and language are not the same thing, and what today's LLMs can and cannot tell us about human cognition.Key Topics:Do LLMs understand language or merely generate convincing text?The difference between formal and functional linguistic competenceWhat LLMs can learn from language aloneâand what they cannotWhy human cognition and AI cognition may be fundamentally differentTheory of mind, reasoning, and common misconceptions about AI capabilitiesHow cognitive scientists evaluate the "thinking" abilities of LLMsWhat neuroscience can teach AI researchers about interpretabilityWhy understanding AI requires studying both behavior and internal representationsThe future of multimodal models and AI cognitionResources & Links:What does it mean to understand language?Dissociating language and thought in large language modelsHow to evaluate the cognitive abilities of LLMsHow Do LLMs Use Their Depth?True LensConnect with Dr. Anna Ivanova:https://bsky.app/profile/neuranna.bsky.socialhttps://x.com/neurannađ§ Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:LinkedInBlueskyX (Twitter)â WiAIR websiteâ

April 17, 2026
What actually happens when AI systems fail in the real world?In this final part of our conversation with Saadia Gabriel (UCLA), we unpack one of the most urgent challenges in modern AI: why even the most advanced models remain vulnerable to manipulation - and what that means for safety, fairness, and society.From multi-turn jailbreaking attacks with near 100% success rates to misinformation shaping human beliefs, this conversation goes beyond surface-level concerns and dives into how harms actually emerge in deployed systems.We explore:Why current guardrails are not enoughHow realistic attack scenarios differ from academic benchmarksThe connection between model vulnerabilities and societal harmWhat AI can (and cannot) do about misinformation and persuasionThe open research problems that still donât have solutionsResources & Links:Generative AI in the Era of 'Alternative Facts'ModelCitizens: Representing Community Voices in Online SafetyTranslation as a Scalable Proxy for Multilingual EvaluationConnect with Dr. Saadia Gabriel:https://x.com/GabrielSaadiahttps://bsky.app/profile/skgabrie.bsky.social

April 15, 2026
What does it mean to build AI systems we can actually trust?In this first part of our conversation with Saadia Gabriel (UCLA), we explore the deeply personal and technical journey behind her work on AI safety, misuse, and responsible NLP.From experiencing targeted hate speech firsthand to receiving a best paper nomination, Saadia shares how her lived experience shaped her research â and why language models must be designed with both capability and risk in mind.đ§ In this episode, we cover:How personal experiences influence AI research directionsThe intersection of NLP, security, and privacyWhy LLMs can be both powerful and dangerousWhat it means to build trustworthy AI systemsLessons from working across multiple research paradigmsHow to pursue high-impact research as a PhD or early-career scientistResources & Links:X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-AgentsConnect with Dr. Saadia Gabriel:https://x.com/GabrielSaadiahttps://bsky.app/profile/skgabrie.bsky.social

April 13, 2026
What does it actually mean for a model to understand audioPaper: https://arxiv.org/abs/2601.19673In this episode, I talk with Iwona Christop, a PhD student at Adam Mickiewicz University, about her recent EACL paper introducing ART (Audio Reasoning Tasks) â a new benchmark designed to evaluate whether multimodal LLMs can truly reason over audio, not just transcribe or classify it.Most existing benchmarks test audio skills in isolation (like ASR or classification). But real-world intelligence requires something deeper: combining signals, comparing sounds, tracking context, and making decisions.This work takes a different approach:No text-only shortcuts â tasks canât be solved via transcription aloneReasoning-first design â models must combine multiple audio cuesNo expert knowledge required â anyone can verify correctnessWe also dive into the diverse task design, including:Audio arithmetic (counting and comparing sounds)Cross-recording speaker & language identificationSound-based reasoning (e.g., inferring properties from audio)Speech feature comparison (accents, variations)Multimodal reasoning across text and soundThe dataset includes 9 tasks, 9,000 samples, and 30+ hours of audio â all generated in a scalable way using templates and TTS.đ If you care about multimodal reasoning, evaluation, or the limits of current LLM capabilities, this conversation is for you.Iwona Christop:https://www.linkedin.com/in/iwona-christop/đ Like & subscribe for more deep dives into cutting-edge AI researchđ New episodes from EACL 2026 coming soon#WiAIR #EACL2026