Skip to main content

Hume

Freemium

The Emotional Intelligence Lab for Voice AI.

A research lab and technology company building emotionally intelligent voice AI. Hume develops speech-language models that interpret and generate expressive speech, and expression measurement models that analyze emotional signals across voice, face, and language.

emotion recognitionvoice synthesisTTSspeech-to-speechempathic AIprosodymultimodalSentiment AnalysisSpeech RecognitionSpeech-to-TextData AnalysisInformation ExtractionUser Behavior AnalysisanalysisrecognitionAudio & VoiceData Analysis
Visit Website
Hume preview

What is it

Hume AI is a research lab developing models that embed emotional intelligence into voice AI. Its work spans two categories: speech-language models for interpreting and generating expressive speech, and expression measurement models for analyzing vocal, facial, and verbal emotion. The platform is built on over a decade of affective science research and semantic space theory.

What it can do

Engage users through the Empathic Voice Interface (EVI), a real-time speech-to-speech AI that detects vocal modulations and responds with emotionally aligned tone. Generate expressive speech from text via Octave, an LLM-powered TTS system. Analyze emotional expression in audio, video, images, and text through Expression Measurement. Access open-source TADA on Hugging Face and curated speech datasets for training voice models.

Who is it for

AI developers building voice applications, UX researchers studying emotional engagement, healthcare platforms monitoring patient tone, customer experience teams optimizing call interactions, and creators producing narrated audio content.

Key Features

Empathic Voice Interface (EVI) – Real-time emotionally intelligent speech-to-speech AI

EVI measures nuanced vocal modulations and responds using a speech-language model trained on millions of human interactions. It unites language modeling and text-to-speech with superior EQ, prosody, end-of-turn detection, interruptibility, and alignment. Supports external LLM compatibility, back channeling, and expressive instruction following for natural, human-like voice conversations.

Octave Text-to-Speech – LLM-powered expressive voice synthesis

Unlike conventional TTS systems that read words phonetically, Octave is a speech-language model that understands contextual meaning, unlocking nuanced expressiveness in tone, pacing, and emotional intensity. Supports voice design, modulation, cloning, and conversion from both text and natural language descriptions. Available in Octave 1 and Octave 2 preview with expanded language support and lower latency.

Expression Measurement – Multimodal emotion analysis across voice, face, and text

Built on 10+ years of research in semantic space theory, these models capture hundreds of dimensions of human expression. Analyzes speech prosody, vocal bursts, emotional language, facial expressions, and facemesh data across audio, video, images, and text. Provides granular, research-grade emotional insights for analytics, healthcare, and user experience applications.

TADA – Open-source LLM text-to-speech system

Streams text and audio together to reduce hallucinations and latency compared to conventional pipeline TTS architectures. Available as an open-source model on Hugging Face, enabling researchers and developers to build upon, fine-tune, and customize the architecture for proprietary voice AI applications without vendor lock-in.

Use Cases

Building empathic customer service voice agents

Deploy EVI-powered voice assistants that detect caller frustration or distress through vocal modulation and respond with appropriately adjusted tone and pacing. The system reduces escalation rates and improves satisfaction scores by handling emotionally charged interactions with calibrated empathy rather than rigid scripts.

Creating emotionally aware digital companions

Build voice companions for seniors, children, or mental wellness support that recognize emotional states in real time and adjust responses to provide comfort, engagement, or motivation. EVI's interruptibility and back channeling create natural conversational flow without the mechanical turn-taking of traditional voice bots.

Generating expressive audiobook and podcast narration

Use Octave TTS to produce narration that varies tone, pacing, and emotional intensity according to narrative context. The LLM-powered understanding of contextual meaning delivers audiobooks and podcast episodes that sound genuinely performed rather than mechanically read, reducing production costs for independent creators and publishers.

Analyzing user experience research sessions

Process recorded user interviews and usability testing sessions through Expression Measurement to identify sentiment trends, emotional friction points, and engagement patterns. Facial expression analysis combined with vocal prosody detection reveals reactions that quantitative clickstream metrics alone cannot capture, providing actionable UX insights.

Pricing plans

Frequently Asked Questions

Discussion

No comments yet. Be the first to start the thread.