
28 Sep 2026
Modulate raises $25M to scale audio-native AI models
Boston voice-intelligence company Modulate said Monday it raised $25 million led by Future Ventures, with Hyperplane and Lakestar participating, bringing total funding to $60 million as it expands audio models used for deepfake detection, voice-agent supervision, and conversation understanding.
A Boston audio company just took another $25 million, led by Steve Jurvetson’s Future Ventures, to expand models that understand voice beyond the transcript. Voice agents and deepfake risk are both getting larger, and listening, on this bet, has to mean more than turning speech into text.
On Monday, 28 September 2026, Modulate said it raised $25 million led by Future Ventures, with Hyperplane and Lakestar participating. The company is based in Boston. It calls itself a frontier audio AI company: software built for sound, especially the human voice, rather than only for text. Modulate said the round brings its funding to $60 million. The release does not give the round a letter, such as Series A or Series B, and it does not print a valuation, the price this announcement would put on the whole company.
Modulate said the money will support AI and machine-learning research, product and engineering, developer relations, and partnerships. It also said the company will keep widening the APIs, models, and deployment options developers can use. An API is a door other software uses to call these models. The same page says Modulate is building new industry models, adding software kits and APIs, adding people who work with developers, creating partner integrations, and supporting new places customers can run the software.
Modulate said its models now analyze more than 10 million hours of audio each month, and that they have processed more than 600 million hours in total. An hour of audio is one hour of recorded or live sound. Those counts are the company’s. They mean the models are already running on a large amount of real voice, not only on a short demo.
Modulate said transcription and deepfake detection both ranked first on public Hugging Face benchmarks. Hugging Face is a public site where labs post models and comparison scores. Transcription means turning speech into text. A deepfake, here, is a voice made or altered by software so it sounds like a real person. The release names the boards: first on Hugging Face’s Open ASR Leaderboard for transcription, and first on Hugging Face’s deepfake speech benchmark. ASR is the short name for automatic speech recognition, the same job as transcription. Modulate also said deepfake detection reaches 98.9 percent accuracy on public benchmark data, and that the transcription API costs $0.03 an hour when the audio is processed in a batch. Those scores and that price are the company’s.
The flagship product is Velma, powered by what Modulate calls an Ensemble Listening Model, or ELM. An ensemble is a group of smaller models working together, instead of one giant model doing every job. Modulate said the ELM uses more than 100 specialized audio models, picked and combined for each task. The company says they read emotion, tone, intent, emphasis, synthetic speech, and how a conversation is behaving. Synthetic speech is a voice a machine generated. Intent is what the speaker is trying to do, not only the words. Modulate said Velma is twice as accurate as a traditional large language model at catching a real hit, and that it produces seven times fewer false alarms. A large language model is the kind of system behind a text chatbot. A false alarm is a flag on something that was fine. The company also said Velma can be up to 1,000 times more efficient than one large model, using less money, energy, and memory. Those comparisons are Modulate’s. It also said Velma can run while a call is still going, so an application can step in before the call ends.
Carter Huffman, chief executive and co-founder, said: “Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript.” A transcript is the text of what was said. He said the company already uses this work to protect organizations from deepfake attacks, help voice agents read emotion and answer with more empathy, spot dangerous behavior in online conversations, and check whether a voice agent is doing the job it was given. He said more than a hundred specialized models sit under that work. Those sentences are his, in the company announcement.
Steve Jurvetson, co-founder of Future Ventures, said Modulate has “gained a significant technical lead in audio-native AI” and that the market is expanding quickly. He said the models have been used in demanding voice settings, and that the need now reaches AI agents, security, and customer experience. He said the investment is meant to help the company move faster, grow the team, and put the models in front of more developers and partners. Those sentences are his, on the same release. The page also lists him as a board member of SpaceXAI.
In plain terms, Modulate sells software that listens to voice calls and to AI voice agents. It looks for deepfakes, scams, emotion, and breaks of a company’s rules. It is not only a tool that turns speech into text. The company says the same models are used to watch how a voice agent performs, to flag harassment, and to notice when a caller’s voice was made by software. Examples it prints include protecting hospitals from deepfake callers, watching voice agents, and spotting harmful talk on social and gaming platforms.
TechCrunch reported the same $25 million round and the same three investors on Monday morning. It cited PitchBook for a separate figure: that Modulate had raised $41 million, at a $170 million valuation, before this round. That earlier total and that valuation are PitchBook’s, as TechCrunch reports them. Monday’s Modulate release does not repeat them, and it does not put a new price on the company. TechCrunch also wrote that Mike Pappas and Carter Huffman founded the company in 2017 after meeting as MIT physics students, that it has about 40 to 45 employees, and that it plans to add about 10. The company page says the founders are MIT alumni. The year, the second founder’s name, and the headcount are TechCrunch’s.
The picture is a screenshot of Modulate’s press page for Monday’s announcement. The “Raises $25M” headline and the Future Ventures opening are in the frame. An indigo band runs down the left side. It is the press page. It is not a photograph of the founders.