San Francisco-based Sonilo has secured $11 million in funding to expand its generative audio technology. The round was led by B Capital, with participation from Redpoint Ventures. The company plans to use the capital to scale its computing infrastructure and broaden its reach among creators and developer platforms.
Sonilo was established by former TikTok employees who worked in the social media giant’s AI and music divisions. The leadership team includes CEO Shawn Song, who previously led multimodal AI efforts at TikTok. CTO Alex Yin holds a doctorate in computer music from the University of York. COO Keli Li managed TikTok’s global music operations, and CMO Trista Taylor has experience launching consumer AI applications.
Sonilo AI audio technology
The company’s core product is the Sound World Model, a multimodal system designed to generate audio that matches visual content. Unlike tools that create music or sound effects in isolation, this model analyzes video footage to understand pacing, emotion, and narrative shifts. It then produces synchronized audio tracks tailored to the specific scene.
Users can generate audio by uploading a video clip, providing a text description, or combining both inputs. The system is built to handle commercial use cases, targeting short-form video creators, short drama producers, and enterprise production teams. It also offers AI dubbing with lip-sync capabilities, allowing studios to localize content into different languages without manual post-production editing.
“We built Sonilo to understand the visual story,” said Shawn Song. “You give us the footage, and Sonilo understands what’s happening frame by frame, including the pacing, emotion, and story, and generates music and sound effects that naturally fit the scene.”
Partnerships and licensing
Sonilo has signed early distribution agreements with TapNow, an AI creator aggregator, and ComfyUI, a platform for assembling AI video tools. On the rights side, Shutterstock serves as an early licensing partner. The company states that its models are trained on licensed music, ensuring that artists and rights holders receive licensing fees and a share of revenue.
Daisy Cai, a general partner at B Capital, highlighted the strategic value of the technology. “We believe generative audio will become an essential part of the AI-native video stack,” Cai said. “Sonilo stands out because its multimodal technology gives creators two ways to generate music and sound effects: describe what they want in text, or provide a video and let the model create audio that fits the scene.”
The company is currently finalizing publishing and master recording agreements with major rights organizations. These deals are expected to close within weeks, solidifying the legal framework for its generative audio services.
Sonilo intends to launch its consumer application in the fourth quarter. This release will make the video-to-music tool directly available to individual creators, expanding its presence beyond enterprise and developer channels.
Source: Variety

