David AI, founded in 2024 and headquartered in San Francisco, develops high-quality audio training datasets to support AI research and applications. The company builds the data layer for audio AI, a critical infrastructure component for advancing speech recognition, machine translation, speech synthesis, and conversational AI systems. It has raised $80M in total funding, including a $50M Series B round, and counts several leading AI labs and major technology companies among its partners.
The company's product portfolio addresses distinct audio data needs. Converse provides natural two-speaker English conversations, while Chorus extends this to multi-speaker scenarios. Atlas is a dataset spanning over 15 languages with dialect metadata, and Dialog focuses on expert domain conversations. Together, these products serve researchers and engineers working across a range of audio AI applications.
David AI operates with a structured R&D methodology that moves from hypothesis through design, experimentation, evaluation, and iteration to production and release. This approach reflects the company's emphasis on data quality and rigorous curation in a field where training data is a key bottleneck for model performance.






