Listen to this audio :

Voice AI Actor

You just heard the same sentence performed in different ways.
Nothing about the words changed. The meaning didn't come from the sentence.
It came from the delivery.


That's what great actors do.
They don't change the script. They change how the audience feels about it.

Then something changed.
Companies like ElevenLabs started asking,
"Can AI perform this sentence?"


Reading is about saying words correctly.
Performing is about deciding how those words should come alive.


For years, voice AI couldn't do that.
It could pronounce words correctly. It could read sentences clearly.
But it couldn't perform them.

So how does AI know how to perform?


Traditionally all the text was written in markup language, across major providers like Google, Amazon, Microsoft.
Eleven Labs trained their model Eleven v3 on real end to end speech. You can now give the model performance directions directly inside the script.





Before generating any audio, the model reads these directions and plans how the line should be delivered.

It's less like programming.

More like directing an actor.

How can a model infer the speaker's intent and perform the line appropriately every time?


This requires optimizing several things simultaneously:


  1. Planning before delivery:


    A trained actor doesn't discover the emotional arc of a line while speaking it, they plan it.

    Eleven Labs calculates using computation and plan delivery before waveform synthesis.

    It decides:

    • where to pause

    • which words to emphasize

    • when pitch should rise

    • when it should fall

    • speaking rhythm

    • speaking energy


  1. Context awareness:


    Traditional text to speech (TTS) models treats each sentence independently. ElevenLabs' conversational models maintain context across turns, so delivery is influenced by what

    happened earlier in the conversation.


  1. Emotional inference:


Instead of requiring the developer to explicitly say "Speak happily"

the model tries to infer emotion from:

  • the conversation

  • punctuation

  • wording

  • previous dialogue

  • system instructions

    Additionally developers can still guide delivery with prompts or expressive tags like [laughs], [whispers], or [sighs] when they want finer control.

The emotion isn't fixed for an entire conversation.

It changes sentence by sentence.

For example:

"I'm sorry."

"I completely understand."

"Let's fix this together."


Each line has a different emotional intention.

The model performs each one differently instead of applying one emotion to the entire interaction.

What This Optimization Costs?


  1. Speed/Low Latency-

    Highly expressive speech requires more reasoning about context and performance. ElevenLabs explicitly positions v3 as a quality-first model, distinct from its faster sibling models —

    Turbo and Flash prioritize speed and low latency for real-time assistants or call-center bots.


  2. Less predictability-

    An expressive model has more freedom. ElevenLabs encourages developers to constrain emotional behavior through system prompts and brand voice guidelines. A more deterministic model produces more consistent—but often flatter—speech.

What This Means for Designers


For decades, designers have focused on writing the right words.

Voice AI introduces a new design material: performance.


The same sentence can reassure someone, build urgency, express empathy, or create excitement—without changing a single word. The difference lies in how it's delivered.


As expressive voice models become more capable, designers will increasingly make decisions such as:

  • Should this response sound calm or energetic?

  • Where should the assistant pause?

  • Which word deserves emphasis?

  • When should the voice acknowledge emotion, and when should it stay neutral?

  • How expressive is appropriate for this product and this moment?


These aren't technical decisions. They're interaction design decisions.


Made by Tejaswini

VOX. v1

Imagine this.

You tell your friend about your recent holiday adventure only to hear a cold reply.

The first day, you'd smile.

By the tenth day, you'd roll your eyes.

The problem isn't what they say, its the tone and context.

Now imagine that mismatch isn't occasional. Imagine it's every single interaction, because the "person" you're talking to has exactly one personality, permanently, regardless of your mood, the situation, or what you actually need in that moment.

That's what talking to voice assistants has been like for years.


Why Does It Feel Wrong?

Why does it feel so specifically wrong when a voice's tone doesn't match what you need in the moment even when the actual information it gives you is correct?


That's because conversation isn't only about information.

It's also about how that information is delivered.

Think about the range of tone you naturally expect from different people in your life:

Doctor

Confident, reassuring

Close Friend

personal, warm

Colleague

Polite, Supportive

Customer support

Calm, Apologetic

Doctor

Confident, reassuring

Close Friend

personal, warm

Colleague

Polite, Supportive

Customer support

Calm, Apologetic

Voice AI Actor

You just heard the same sentence performed in different ways.
Nothing about the words changed. The meaning didn't come from the sentence.
It came from the delivery.


That's what great actors do.
They don't change the script. They change how the audience feels about it.

For years, voice AI couldn't do that.
It could pronounce words correctly. It could read sentences clearly.
But it couldn't perform them.

Then something changed.
Companies like ElevenLabs started asking,
"Can AI perform this sentence?"


Reading is about saying words correctly.
Performing is about deciding how those words should come alive.

So how does AI know how to perform?


Traditionally all the text was written in markup language, across major providers like Google, Amazon, Microsoft.
Eleven Labs trained their model Eleven v3 on real end to end speech. You can now give the model performance directions directly inside the script




Before generating any audio, the model reads these directions and plans how the line should be delivered.

It's less like programming.

More like directing an actor.

How can a model infer the speaker's intent and perform the line appropriately every time?


This requires optimizing several things simultaneously:


  1. Planning before delivery:

    A trained actor doesn't discover the emotional arc of a line while speaking it, they plan it.

    Eleven Labs calculates using computation and plan delivery before waveform synthesis.

    It decides:

    • where to pause

    • which words to emphasize

    • when pitch should rise

    • when it should fall

    • speaking rhythm

    • speaking energy


  2. Context awareness:

    Traditional text to speech (TTS) models treats each sentence independently. ElevenLabs' conversational models maintain context across turns, so delivery is influenced by what

    happened earlier in the conversation.


  3. Emotional inference:

Instead of requiring the developer to explicitly say "Speak happily"

the model tries to infer emotion from:

  • the conversation

  • punctuation

  • wording

  • previous dialogue

  • system instructions

    Additionally developers can still guide delivery with prompts or expressive tags like [laughs], [whispers], or [sighs] when they want finer control.

The emotion isn't fixed for an entire conversation.

It changes sentence by sentence.

For example:

"I'm sorry."

"I completely understand."

"Let's fix this together."

Each line has a different emotional intention.

The model performs each one differently instead of applying one emotion to the entire interaction.

What This Optimization Costs?

  1. Speed/Low Latency-

    Highly expressive speech requires more reasoning about context and performance. ElevenLabs explicitly positions v3 as a quality-first model, distinct from its faster sibling models —

    Turbo and Flash prioritize speed and low latency for real-time assistants or call-center bots.


  2. Less predictability-

    An expressive model has more freedom. ElevenLabs encourages developers to constrain emotional behavior through system prompts and brand voice guidelines. A more deterministic model produces more consistent—but often flatter—speech.


What This Means for Designers

For decades, designers have focused on writing the right words.

Voice AI introduces a new design material: performance.

The same sentence can reassure someone, build urgency, express empathy, or create excitement—without changing a single word. The difference lies in how it's delivered.

As expressive voice models become more capable, designers will increasingly make decisions such as:

  • Should this response sound calm or energetic?

  • Where should the assistant pause?

  • Which word deserves emphasis?

  • When should the voice acknowledge emotion, and when should it stay neutral?

  • How expressive is appropriate for this product and this moment?

These aren't technical decisions. They're interaction design decisions.

.


  1. Filler Ratio

Take any simple question:

"What's the weather today?"

There are several possible replies:
"24°C."
"It's 24°C.".
"Sure! It's 24°C today."
"It looks like a lovely day today. It's currently 24°C with clear skies."


  1. Lexical density:

Lexical density is simply how much of a sentence is actual information versus conversational language.

High lexical density:

"Timer set. Five minutes."

Low lexical density:

"Sure! I've gone ahead and set a timer for five minutes."

The answer is the same. The experience isn't.


  1. Latency per Persona

Every personality style is implicitly also a latency trade-off. A warmer response takes longer to generate and almost certainly costs real milliseconds in generation and synthesis. Brief mode isn't just a tone choice - it's plausibly the fastest-responding mode by design



What This Means for Designers

2.Pitch variation (Prosody)

This is one of the biggest reasons voices feel alive.

Try saying

"Really."

in three ways.

😐 Really.

😮 Really?!

🙄 Really...

Same word. Three emotions.


Voice personality isn't something you add at the end of the design process. It's about designing the relationship users have with that character.

Good voice design isn't about making AI sound more human.

It's about reducing the gap between what a user expects to hear and how the assistant actually responds.

Because in conversation, people rarely remember the exact words.

They remember how those words made them feel.

Listen to this audio :

Made by Tejaswini

VOX. v1

Create a free website with Framer, the website builder loved by startups, designers and agencies.