The Missing Piece in AI: Why Understanding Audio and Video Matters More Than Raw Computing Power
A startup called Mundo AI just raised $20 million because it believes the next big leap in artificial intelligence won't come from bigger models or faster computers, but from teaching AI systems to actually understand what they see and hear. The company is building datasets and evaluation tools designed to help AI models grasp audio, video, and other sensory information in ways that current systems largely miss.
Why Can't AI Understand a Conversation the Way Humans Do?
Here's the gap Mundo is trying to close: when two people have a conversation, humans pick up on far more than just the words. We notice tone of voice, facial expressions, timing, gestures, background sounds, and social context. All of these details change what a conversation actually means. An AI trained primarily on text transcripts misses almost all of that.
Mundo describes this as the difference between recognizing speech and genuinely understanding a conversation. A system might transcribe every word perfectly but completely miss whether someone is being sarcastic, frustrated, or joking. On the video side, the company focuses on areas where vision systems can miss context, sequencing, and physical intent. These gaps matter because they represent real limitations in how AI systems interact with the physical and social world.
What Problem Is Mundo Actually Solving?
Mundo AI emerged from Y Combinator's Winter 2025 batch with an initial focus on multilingual training data. The founders noticed that AI models performed significantly better in English than in many other languages, largely because there wasn't enough high-quality native-language training data available. They built a platform to collect, generate, and annotate datasets by working directly with native speakers in countries where target languages are spoken.
But the company has since expanded that mission. Its current platform tackles the broader challenge of teaching models to understand real-world sensory experiences, including how events unfold over time and how audio, visual information, and context interact with one another. For emerging AI modalities where established training approaches don't yet exist, Mundo is designing datasets around the specific capabilities researchers are trying to develop.
How to Build Better AI Perception Systems
- Collect Multimodal Data: Gather datasets that include speech, video, gestures, facial expressions, and environmental context simultaneously, rather than treating each modality separately.
- Develop Evaluation Metrics: Create benchmarks that measure how accurately AI systems perceive real-world situations, not just how well they recognize individual elements like words or objects.
- Work with Domain Experts: Partner with native speakers, video annotators, and subject matter experts to ensure datasets capture nuance and context that generic data collection misses.
The Series A funding, led by GreatPoint Ventures and including participation from Y Combinator, E12 Ventures, and Next Frontier Capital, brings Mundo's total funding to $24 million. The company plans to use the new capital to expand its research, engineering, and operations teams as it develops additional data and evaluation infrastructure for what it calls "perceptual AI" systems.
Why Does This Matter for the Future of AI?
Mundo's thesis reflects a broader shift in AI development toward multimodal systems that need to understand more than written language. The company argues that the next major improvements in model capabilities will require not only larger models and additional compute, but new forms of training data and better methods for measuring how accurately AI systems perceive real-world situations.
This perspective aligns with emerging research in multimodal learning analytics, which examines how AI can better understand complex, real-world learning and interaction scenarios. Researchers have noted that while much AI research happens in controlled laboratory settings, deploying these systems in authentic environments reveals significant challenges. More than half of multimodal learning analytics studies have been conducted in lab settings rather than real-world contexts, highlighting a gap between what works in research and what works in practice.
Mundo's founders bring diverse backgrounds spanning machine learning research, quantitative finance, Amazon Web Services, Binance.US, Cohere, and Hugging Face. Their combined experience suggests the company is positioned to tackle both the technical and operational challenges of building perception-focused AI infrastructure at scale.
The company's datasets and evaluations are already being used by leading AI labs developing multimodal models, indicating that the market recognizes the value of this work. As AI systems become more integrated into real-world applications, from customer service to healthcare to education, the ability to understand context, tone, and visual information becomes increasingly critical. Mundo AI's focus on building the foundational data layer for perceptual intelligence suggests that the next generation of AI breakthroughs may depend less on raw computational power and more on how well we teach systems to perceive the world as humans do.