Skip to main content

How do I use ElevenLabs voices in Pencil?

How to use the ElevenLabs voice models in Pencil: the models available, where to enable and select them, styling delivery with square-bracket audio tags, and the settings you can adjust.

Written by Michael Whyle

ElevenLabs brings cutting edge text-to-speech voiceover generation into Pencil. You write a script, choose a voice and model, and Pencil generates a lifelike voiceover for your creative. Each voiceover saves to your asset library as a Generation, so you can reuse it across projects like any other asset. It works in the Editor, Canvas and Workflows.

One thing to know upfront: with ElevenLabs v3 you style the delivery line by line, by writing audio tags in square brackets in your script.

This article covers the ElevenLabs models and how to get the best from them. For the full walkthrough of the Voiceover panel and its controls, see the AI Voiceovers article.


Where to generate

As with other AI models, an admin must first enable the ElevenLabs models per workspace under Settings > AI Governance > Voice Generation.

Once enabled they'll appear in the platform anywhere where Voiceover Generation is preset. The two primary areas to work with these models are Workflows and Ads Editor.

  • In Ads Editor: open the Voiceover panel on the left, then choose an ElevenLabs model, language and voice.

  • In Workflows: add the Generate Voiceover node, then set the model and voice in the node settings panel on the right.

In the Ads editor, the length of the audio segment produced is determined by the length of your script in the prompt box, but also constrained by the length in time of your main sequence in the Ads Editor timeline view depends on the length of your ad timeline rather than a setting you choose. Similarly, in Workflows the length of audio generated will depend on the length of the script.


The models

Three ElevenLabs models are available. Choose the one that fits the job.

  • ElevenLabs Eleven V3 is the most expressive model, with rich emotional range and support for audio tags. It suits storytelling and character voices, and covers 70+ languages.

  • ElevenLabs Flash V2.5 is built for ultra-low latency (around 75ms). Use it for quick turnaround and short, conversational pieces.

  • ElevenLabs Multilingual V2 is a high-quality multilingual model recommended for consistent narration across 29 languages. Use it for localisation and longer reads.


Voice and Language

Choose a voice and language before you write. Pencil gives you access to over 15,000 ElevenLabs voices to browse and search. You can star the ones you like to keep them in your Favourites.

Advanced ElevenLabs features such as voice cloning using your own voice are not available in Pencil.


Styling the delivery

The most important think to to know, is that With ElevenLabs latest model, v3, you steer emotion and delivery by writing audio tags in square brackets in your script, such as "[whispers]", "[excited]" or "[angry]".

You can even tag line by line, which gives you granular control over how each part is read:

"[Abrupt] We need to talk about your coffee. [warmly] Because you deserve better. [Conspiratorial, sly] Much better."

Note that results may vary, and it tends to work best with one or two style instructions per generation

This works the same in the Editor and in the Workflows Voiceover Generation Node, using the text node input for your prompt.

​Audio (emotion) tags are an ElevenLabs V3 feature only. For the V2 model family you shape delivery through the script's wording, punctuation and context instead.

A few more tips:

  • Write naturally. Structure and punctuation drive delivery, so use full sentences and clear context.

  • Pacing and emphasis. Ellipses add a pause and weight, and a capitalised word makes it land harder, as in: "It was a VERY long day … nobody listens anymore."

  • Match the voice. Keep tags in character: a calm, professional voice may give unexpected results when prompted to suddenly shout, for example.

  • Iterate. Revise the script or settings and regenerate. Each new version saves to your asset library.

You'll notice that the Voiceover styling instructions box stays greyed out for ElevenLabs models. That field is for Google's TTS 3.1 model, which takes one overall style instruction. The square bracket method gives you per-line control instead, results may vary when running long, complex scripts with many different style instructions.


Settings you can adjust

You choose the model, language and voice at the top of the panel. Everything else sits in Advanced settings, where the options vary a little by model. The same controls appear on the node settings panel in Workflows, so your set-up carries across.

  • Output format. Set the audio quality, for example 128 kbps · 44.1 kHz.

  • Speed. Adjust the read from 0.7x to 1.2x.

  • Stability. Lower for more expression and variation, higher for a steadier, more consistent read.

  • Similarity boost. Control how closely the output tracks the source voice.

  • Style exaggeration. Push the voice's stylistic intensity up or down.

  • Speaker boost. Toggle on to reinforce the character of the chosen voice.

  • Volume gain. Raise or lower loudness to sit over music.

  • Content moderation. Set the moderation option for the generation.


Example use cases

  • Short product videos. Write a short, punchy script for a 10-second product clip, choose Eleven V3 for a lively read, generate the voiceover in the Editor, and place it over your video.

  • Localising at scale. Generate the voiceover in one language, then switch the language and regenerate to build a set of versions. Save each to your asset library and apply them across a template in Workflows to produce localised variants of the same asset quickly.


A note on terms

Use of ElevenLabs models is subject to ElevenLabs' own terms. Please comply with their Terms of Use and Prohibited Use Policy.

For the full list of audio models in Pencil, see the audio models article.

Did this answer your question?