David Senra · Published 2026-09-09

Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs

Open on YouTube ↗

Summary

Overview

  • Speaker: David Senra and Mati Staniszewski
  • Channel: David Senra
  • Main topic: Building an AI company and the future of audio-first communication
  • Purpose: Provide founders, builders, and entrepreneurs with deep insights into starting and scaling an AI-native company from first principles. David Senra interviews Mati Staniszewski, co-founder of ElevenLabs, discussing how they built ElevenLabs as an AI-native audio platform from first principles. They cover the evolution of audio AI models, company building dynamics, small team philosophies, market adoption across fintech, healthcare, and retail, and why focusing on audio provides unique emotional resonance.

Topic Map

Founding ElevenLabs and First Principles Thinking

  • Explanation: How ElevenLabs started in 2022 before the explosion of ChatGPT, focusing on audio and voice generation inspired by cultural observations in Poland.
  • Key claims:
    • Starting before ChatGPT meant the hype was around crypto and metaverse, leaving room to focus purely on AI.
    • The trigger point was noticing how Polish media dubs entire movies with a single narrator voice, losing emotional intonation.
    • Building from first principles instead of formulas yields non-consensus success.
  • Examples:
    • Polish movie dubbing using a single narrator voice.
  • Terminology:
    • ElevenLabs
    • first principles
    • audio-native
  • Why it matters: Demonstrates the importance of tackling overlooked problems in large markets.

Research vs. Product Deployment Under One Roof

  • Explanation: The deliberate organizational strategy of combining cutting-edge AI audio research with practical product delivery.
  • Key claims:
    • Most companies separate research from product, but combining them accelerates real-world impact.
    • Research builds frontier audio models (speech-to-text, text-to-speech, orchestration), while product builds the platform for customers.
    • Small, autonomous teams drive both invention and distribution.
  • Examples:
    • Google's knowledge graph vs. image visual models.
  • Terminology:
    • research and development
    • knowledge graph
    • LLM
  • Why it matters: Unifies technological breakthroughs with rapid market distribution.

The Evolution of Audio Models and Voice Identity

  • Explanation: How voice generation evolved from robotic text-to-speech to emotionally rich, authentic, and customizable AI voices.
  • Key claims:
    • Initial models were robotic and unstable; the breakthrough was achieving human-like emotional intonation.
    • Voice is a profound identity marker; authentic voices drive deep emotional engagement.
    • Marketplace models allow creators and individuals to authenticate, share, and monetize their AI voices.
  • Examples:
    • Recreating voices for individuals who lost them due to ALS or cancer.
  • Terminology:
    • text-to-speech
    • voice cloning
    • AI voice marketplace
  • Why it matters: Shows how generative audio transforms personal and enterprise communication.

Small Teams, No Titles, and Decentralized Execution

  • Explanation: Operating with flat hierarchies, small deployment squads, and intense customer obsession.
  • Key claims:
    • Small teams of under 10 people retain flexibility and autonomy to move fast.
    • No-title culture eliminates bureaucracy and centers focus on solving customer problems.
    • Integrating research, product, go-to-market, and engineering tightly around customer pain points accelerates growth.
  • Examples:
    • Palantir and Honda organizational models.
  • Terminology:
    • flat hierarchy
    • autonomous squads
    • go-to-market
  • Why it matters: Explains how high-growth startups scale output without sacrificing speed or culture.

Key Points

Audio-first communication represents the next frontier in human-computer interaction

  • Explanation: Voice carries emotional nuance, cadence, and empathy that text alone cannot capture.
  • Evidence: Success of AI voice generation across podcasts, customer service, and media creation.
  • Practical implication: Companies should integrate audio interfaces to deepen user engagement.

Customer obsession as a form of R&D

  • Explanation: Using customer feedback loops to continuously refine AI models and product features.
  • Evidence: ElevenLabs scaling across fintech, retail, healthcare, and telecommunications.
  • Practical implication: Treat customer use cases as the primary driver of technological roadmap priorities.

Frameworks, Models & Processes

The Research-Product Integration Loop

  • How it works: Combining foundational AI research with direct product deployment and customer feedback loops.
  • Components:
    • Frontier audio research
    • Product deployment platform
    • Customer feedback and use case iteration
  • When to use: When building deep-tech companies where technical capability must align with market utility.

Examples & Case Studies

A musician who lost their voice used ElevenLabs to recreate their AI voice and perform with their old band.

  • Illustrates: The deep emotional impact and humanitarian value of generative voice technology.
  • Lesson: Technology should amplify human potential and restore capabilities.

Actionable Takeaways

  • Immediate:
    • Adopt first-principles thinking in product design.
    • Eliminate unnecessary organizational layers to maintain execution velocity.
  • Strategic:
    • Combine research and product teams under one roof for faster iteration.
    • Focus on audio as a superior medium for emotional connection.
  • Questions to investigate:
    • How will generative audio change customer service and enterprise communications over the next decade?
    • What safety and ethical guardrails are necessary for voice cloning at scale?

Claims Worth Verifying

  • ElevenLabs has scaled to tens of thousands of business and creator customers globally. (business growth)

Notable Quotes

"The single most powerful pattern I have noticed is that successful people find value in unexpected places." (at 69:13) "We built ElevenLabs to break down language and communication barriers." (at 99:40)

Compressed Summary

  • ElevenLabs builds AI-native audio and voice generation platforms.
  • Combining research and product under one roof accelerates innovation.
  • Small, flat teams maintain speed and customer obsession.
  • Generative voice transforms media, customer service, and accessibility.
  • Keywords: elevenlabs, generative ai, audio, voice cloning, startups
  • Core insight: Generative audio transforms communication by combining technological research with customer-obsessed product execution from first principles.

Core insights

5
Empirical Resultmedium noveltymoderate evidence

The critical quality threshold in speech synthesis is not word-level intelligibility but emotional intonation; an audio-native system must model cadence, emotion, and expressive timing rather than treating text-to-speech as a text-rendering problem.

Why it matters

For teams evaluating or building speech/audio models, this changes evaluation criteria: surface-level naturalness or accuracy metrics can hide the failure mode that matters for real users (emotionless output). It also indicates that future agent voice interfaces will need controllable emotional prosody, not just low error rate.

Generalization

Generative output quality should be measured on the dimensions that carry the intended user effect, not only on surface-level fidelity.

Initial models were robotic and unstable; the breakthrough was achieving human-like emotional intonation.
Open source video
Architecturelow noveltymoderate evidence

ElevenLabs' model of integrating research and product under one roof is not just a hiring tactic; it is an architectural mechanism that shortens the feedback loop between frontier model research, product deployment, and customer pain points.

Why it matters

This is a strong argument against the many AI organizations that separate a research lab from product engineering. When the model itself is the product, model iteration and user-facing runtime need to share the same context and priorities.

Generalization

Deep-tech AI companies should colocate model-building teams and product-delivery teams if the product's core value depends on rapid model improvement.

Most companies separate research from product, but combining them accelerates real-world impact.
Open source video
Small, autonomous teams drive both invention and distribution.
Open source video
Practicelow noveltymoderate evidence

Customer pain points are being used as an R&D input: deployment and customer feedback tell the research organization which model capabilities matter next, making customer obsession a component of the technical roadmap process, not just of go-to-market.

Why it matters

For production AI systems, this means telemetry and support signals should be wired into model/feature prioritization. The highest-value engineering investment may be identifying which customer problem exposes the largest model deficiency.

Generalization

AI product roadmaps can be derived from the gap between what the deployed system does and what repeated customer use cases demand.

Customer obsession as a form of R&D
Open source video
Treat customer use cases as the primary driver of technological roadmap priorities.
Open source video
Mechanismmedium noveltymoderate evidence

Voice is identity-bearing: the same model that creates deep emotional engagement can restore a voice that disease took away, but it also becomes something that must be authenticated, controlled, and monetized consensually. Voice identity therefore becomes a product mechanism with real consent/provenance boundaries.

Why it matters

Voice cloning systems and agents that speak for users need first-class identity infrastructure: consent records, voice provenance, and explicit monetization authorization. Engineering these controls is not peripheral; it is intrinsic to the deployment model.

Generalization

Generative systems that mimic personal identifiers need technical artifacts for consent and ownership, not just policy add-ons.

Voice is a profound identity marker; authentic voices drive deep emotional engagement.
Open source video
A musician who lost their voice used ElevenLabs to recreate their AI voice and perform with their old band.
Open source video
Marketplace models allow creators and individuals to authenticate, share, and monetize their AI voices.
Open source video
Architecturelow noveltymoderate evidence

Small, title-less teams of under ten people are deliberately used as the execution unit because on-the-ground speed of iterating on AI products is considered more valuable than organizational hierarchy.

Why it matters

The insight is that AI-native product evolution is bottlenecked on fast deployment cycles, so organizational abstractions themselves need to be flattened rather than layered. This suggests that architectural governance should be owned inside autonomous squads, not in a central platform bureaucracy.

Generalization

When the product changes weekly, control structures should be distributed to small autonomous deployment teams.

Small teams of under 10 people retain flexibility and autonomy to move fast.
Open source video
No-title culture eliminates bureaucracy and centers focus on solving customer problems.
Open source video

Deep dives

5

Evaluation frameworks for emotional prosody in speech synthesis

Research question

What metrics and test suites can reliably measure whether a voice model preserves emotional intonation, not just word-level intelligibility?

Why

If the breakthrough for ElevenLabs was expressive emotional intonation, then teams building voice agents and TTS need an evaluation method that catches flat, robotic outputs where standard quality scores still look acceptable.

Initial models were robotic and unstable; the breakthrough was achieving human-like emotional intonation.
Open source video
Voice is a profound identity marker; authentic voices drive deep emotional engagement.
Open source video
Source video

Consent and provenance architecture for voice identity platforms

Research question

What concrete technical artifacts—signed consent records, audio watermarks, permission-checking layers—are required to let creators authenticate, share, and monetize their AI voice without enabling impersonation?

Why

Voice is identity-bearing; a marketplace for voices is only deployable if access control and provenance are embedded in the generation pathway.

Voice is a profound identity marker; authentic voices drive deep emotional engagement.
Open source video
Marketplace models allow creators and individuals to authenticate, share, and monetize their AI voices.
Open source video
Source video

Evaluating the coupled research/product organization for AI-native products

Research question

What measured effect does colocating research and product have on model improvement cycle time, and at what organizational scale does the effect dissipate?

Why

The pass-1 causal claim that separated organizations are slower can be tested as an architectural hypothesis. If true, it changes how AI companies should structure their technical teams.

Most companies separate research from product, but combining them accelerates real-world impact.
Open source video
Small, autonomous teams drive both invention and distribution.
Open source video
Source video

Translating customer pain points into model roadmap signals

Research question

How can an AI product systematically identify the model deficiency behind repeated customer complaints and feed it into the training or capability roadmap?

Why

Customer obsession is treated both as a cultural value and an R&D input; engineering this feedback path makes the roadmap more responsive and falsifiable.

Customer obsession as a form of R&D
Open source video
Treat customer use cases as the primary driver of technological roadmap priorities.
Open source video
Source video

Scaling small autonomous research-product squads across verticals

Research question

At what team/company scale do small autonomous squads lose alignment, and what artifacts (shared base model, internal APIs) preserve coordination without reverting to title-based hierarchy?

Why

ElevenLabs uses small teams as its unit of execution while expanding to many verticals; knowing when this model needs new coordination mechanisms is a practical scaling question.

Small teams of under 10 people retain flexibility and autonomy to move fast.
Open source video
No-title culture eliminates bureaucracy and centers focus on solving customer problems.
Open source video
Source video

Article ideas

4

Voice AI is judged by emotion, not accuracy

Speech synthesis products win by modeling emotional intonation—cadence and expressiveness—not by lowering word error rate; benchmarks used by buyers mislead when they omit this dimension.

Angle

Critical product/benchmark take on why TTS evaluation is lagging the capability that matters.

Source video

Research isn't upstream of product at AI-native companies

The fastest AI product companies collapse research and product into one team; this is an architectural decision that keeps model iteration tuned to real customer pain.

Angle

Org design argument based on ElevenLabs' model and its stated acceleration.

Source video

Voice cloning needs an identity layer, not a liability warning

Voice cloning is only safely deployable at scale if consent, provenance and monetization authorization are first-class technical products, not policy add-ons.

Angle

Engineering-centric treatment of voice identity as an authentication system.

Source video

Titles are a tax on AI iteration speed

Small title-less autonomous teams are a deliberate organizational architecture for AI companies because they remove the process overhead that slows weekly product/model iteration.

Angle

Direct challenge to conventional startup scaling: add squads, not hierarchy.

Source video

Project ideas

3

prosody-gap TTS evaluation harness

beyond-evals

Standard auto-MOS/WER metrics cannot detect loss of emotional prosody; an 'emotion intent' listener test will rank TTS models differently on expressive material.

Proof of concept

Build a small set of utterances with marked emotional intent (e.g., grief, joy, urgency) in multiple languages; run several TTS models; produce audio clips; collect listener ratings on whether the emotion came through; compare to standard metric scores.

Measurement

Spearman rank correlation between standard quality scores and emotional-fidelity ratings; responder agreement (Fleiss' kappa) on emotional failure.

Source video

voice-consent ledger sidecar

gatehouse

Synthetic voice generation can be gated by a signed consent artifact for the target voice with lower than 100ms added latency, making identity checks practical at API scale.

Proof of concept

Wrap a TTS endpoint with a consent-checking proxy; issue signed consent tokens for voice profiles; proxy rejects requests without a valid token and tags generations with a watermark or request ID; measure overhead.

Measurement

Added latency P99; invalid-token rejection rate; valid-token acceptance rate; watermark traceability in generated files.

Source video

capability-gap classifier for AI product roadmaps

new

Support conversations about a deployed speech product can be automatically classified into model capability gaps (robotic tone, mispronunciation, lack of emotion) and the resulting gap ranking predicts the model improvements shipped in the next release cycle better than raw feature-request counts.

Proof of concept

Anonymize and label support tickets/customer logs from speech API; use an LLM to tag failure modes; produce a ranked gap list; compare with release notes/changelog over two subsequent quarters.

Measurement

F1 for capability-gap classification; overlap and rank correlation between gap list and actual shipped capabilities.

Source video

Architectural implications

4

The summary explicitly argues that separating research from product slows real-world impact.

Before

Model research, product engineering, and go-to-market are owned by different organizations that hand artifacts across boundaries.

After

Research, product, engineering, and go-to-market stay integrated around customer pain points and deployment feedback.

Consequence

Model iteration speed is bounded by how fast the product team can expose new customer problems to the research team.

Source video

Small autonomous teams with no titles are presented as the unit of execution.

Before

A startup scales by adding management layers, titles, and centralized change-control processes.

After

A startup scales by adding small squads that own invention and distribution end-to-end.

Consequence

Explicit bureaucratic coordination is removed, so the main control mechanism is cultural (customer obsession) rather than organizational process.

Source video

Emotional resonance is the intended product property of the audio platform.

Before

Audio generation systems are optimized as text-to-speech utilities with robotic output.

After

Audio generation systems must be optimized around voice identity, cadence, and emotional impact.

Consequence

Model training, evaluation suites, and interface design all need to be modified to preserve expressive elements rather than merely reproducing words.

Source video

Voice cloning is described as emotionally valuable on an individual level and scalable through a marketplace model.

Before

Voice generation was a centralized capability offered by one provider without a creator identity layer.

After

Voice ownership becomes a marketplace mechanism where creators authenticate, share, and monetize their own AI voices.

Consequence

Products built on voice cloning need authorization and provenance infrastructure that are designed at the platform level from the start.

Source video

Tradeoffs and failure modes

3

Emotionally authentic cloned voices

Benefit

Enables deep emotional engagement and restorative use cases, such as helping a musician who lost their voice perform again.

Cost or risk

The same voice-cloning capability creates a very large requirement for safety and ethical guardrails when used at massive scale.

What safety and ethical guardrails are necessary for voice cloning at scale?
Open source video
Source video

Customer obsession as R&D

Benefit

Keeps technical investment focused on capabilities with immediate marketplace demand and deepens adoption across verticals.

Cost or risk

If customer use cases exclusively drive the roadmap, speculative frontier research that has no current customer may become underfunded or ignored.

Treat customer use cases as the primary driver of technological roadmap priorities.
Open source video
Source video

Small autonomous squads

Benefit

High execution speed and low organizational overhead because teams are under ten people and have no titles.

Cost or risk

The model depends on intense customer obsession as the coordination mechanism; if that cultural constraint weakens, decentralized squads may lack alignment as the company grows.

Small teams of under 10 people retain flexibility and autonomy to move fast.
Open source video
Source video

Open questions

3

What safety and ethical guardrails are necessary for voice cloning at scale?

Why unresolved

The same capability that can restore or extend a person's voice may also be used for impersonation, and the summary does not specify technical controls such as provenance, consent, or abuse detection.

Research direction

Design and evaluate watermarking, voice-identity attestation, and permission-checking layers for agent-generated speech.

Source video

How will generative audio change customer service and enterprise communications over the next decade?

Why unresolved

The summary reports deployment across fintech, healthcare, retail, and telecom but does not give evidence about adoption shape, failure rates, or long-term changes to communication workflows.

Research direction

Study audio-native agents in customer service settings: escalation behavior, emotional-fidelity impact on satisfaction, and latency/cost tradeoffs.

Source video

Can a research/product team stay colocated and small as the company expands into many verticals with very different voice-use cases?

Why unresolved

The summary celebrates small autonomous teams and broad industry scaling, but does not reconcile the tension between vertical specificity and unified research investment.

Research direction

Measure when a single research-product loop needs to split into vertical-specific loops without losing shared base-model improvements.

Source video

Key claims

6
factualVerification needed

ElevenLabs has scaled to tens of thousands of business and creator customers globally.

Evidence

ElevenLabs has scaled to tens of thousands of business and creator customers globally.

Question

What internal or audited growth data supports this customer count, and as of when?

Source video
causalVerification needed

Most companies that separate research from product are slower than combined research-product organizations.

Evidence

Most companies separate research from product, but combining them accelerates real-world impact.

Question

Is there any comparative evidence that colocated research/product teams outperform separate teams across other AI ventures?

Source video
opinionVerification needed

The breakthrough in audio models was human-like emotional intonation rather than only robotic-to-natural speech quality.

Evidence

Initial models were robotic and unstable; the breakthrough was achieving human-like emotional intonation.

Question

What specific model versions or benchmarks establish the transition from unstable robotic output to emotionally expressive output?

Source video
factualVerification not requested

Polish media dubbing entire movies with a single narrator voice loses emotional intonation.

Evidence

The trigger point was noticing how Polish media dubs entire movies with a single narrator voice, losing emotional intonation.

Source video
opinionVerification not requested

Voice is a profound identity marker, and authentic voices drive deep emotional engagement.

Evidence

Voice is a profound identity marker; authentic voices drive deep emotional engagement.

Source video
causalVerification needed

Small autonomous teams can drive both invention and distribution.

Evidence

Small, autonomous teams drive both invention and distribution.

Question

At what team count or company scale does this cease to hold, and what observables would show the failure?

Source video

Connections

4