Audio AI startup Noiz has closed a seed round worth tens of millions of yuan, with GSR Ventures as the latest backer. But unlike most audio models, it isn't chasing music generation, speech recognition, or text-to-speech. Instead, Noiz wants AI to understand the physical world behind the sound.
Hear a door slam, and the model shouldn't just recognize 'door closed.' It should know what material the door is made of, how much force was used, how far away the sound originated, and how the room's structure shapes the acoustics. That's why Noiz is collecting data that carries spatial, material, distance, and directional information.
The bigger ambition is something like a 'sound-based world model': given the current environment and what's happening in it, the model predicts what should be heard at that exact moment. That could push audio models beyond dubbing and sound-effect tools into real-time games, virtual worlds, and robotics — letting machines use sound to understand what's going on around them.