VALL-E's quickie voice deepfakes should worry you, if you weren't worried already

Devin Coldewey

Updated 12 January 2023 at 2:26 pm·6-min read

The emergence in the last week of a particularly effective voice synthesis machine learning model called VALL-E has prompted a new wave of concern over the possibility of deepfake voices made quick and easy — quickfakes, if you will. But VALL-E is more iterative than breakthrough, and the capabilities aren't so new as you might think. Whether that means you should be more or less worried is up to you.

Voice replication has been a subject of intense research for years, and the results have been good enough to power plenty of startups, like WellSaid, Papercup and Respeecher. The latter is even being used to create authorized voice reproductions of actors like James Earl Jones. Yes: from now on Darth Vader will be AI generated.

VALL-E, posted on GitHub by its creators at Microsoft last week, is a "neural codec language model" that uses a different approach to rendering voices than many before it. Its larger training corpus and some new methods allow it to create "high-quality personalized speech" using just three seconds of audio from a target speaker.

That is to say, all you need is an extremely short clip like the following (all clips from Microsoft's paper):

https://techcrunch.com/wp-content/uploads/2023/01/in1.wav

https://techcrunch.com/wp-content/uploads/2023/01/in2.wav

To produce a synthetic voice that sounds remarkably similar:

https://techcrunch.com/wp-content/uploads/2023/01/outcome1.wav

https://techcrunch.com/wp-content/uploads/2023/01/outcome2.wav

As you can hear, it maintains tone, timbre, a semblance of accent and even the "acoustic environment," (for instance, a voice compressed into a cell phone call). I didn't bother labeling them because you can easily tell which of the above is which. It's quite impressive!

So impressive, in fact, that this particular model seems to have pierced the hide of the research community and "gone mainstream." As I got a drink at my local last night, the bartender emphatically described the new AI menace of voice synthesis. That's how I know I misjudged the zeitgeist.

But if you look back a bit, in as early as 2017 all you needed was a minute of voice to produce a fake version convincing enough that it would pass in casual use. And that was far from the only project.

Lyrebird is a voice mimic for the fake news era

The improvement we've seen in image-generating models like DALL-E 2 and Stable Diffusion, or in language ones like ChatGPT, has been a transformative, qualitative one: A year or two ago this level of detailed, convincing AI-generated content was impossible. The worry (and panic) around these models is understandable and justified.

Contrariwise, the improvement offered by VALL-E is quantitative not qualitative. Bad actors interested in proliferating fake voice content could have done so long ago, just at greater computational cost, not something that is particularly difficult to find these days. State-sponsored actors in particular would have plenty of resources at hand to do the kind of compute jobs necessary to, say, create a fake audio clip of the President saying something damaging on a hot mic.

I chatted with James Betker, an engineer who worked for a while on another text-to-speech system, called Tortoise-TTS.

Betker said that VALL-E is indeed iterative and like other popular models these days gets its strength from its size.

The emerging types of language models and why they matter

"It's a large model, like ChatGPT or Stable Diffusion; it has some inherent understanding of how speech is formed by humans. You can then fine-tune Tortoise and other models on specific speakers, and it makes them really, really good. Not 'kind of sounds like'; good," he explained.

When you "fine-tune" Stable Diffusion on a particular artist's work, you're not retraining the whole enormous model (that takes a lot more power), but you can still vastly improve its capability of replicating that content.

But just because it's familiar doesn't mean it should be dismissed, Betker clarified.

"I'm glad it's getting some traction because i really want people to be talking about this. I actually feel that speech is somewhat sacred, the way our culture thinks about it," and he actually stopped working on his own model as a result of these concerns. A fake Dali created by DALL-E 2 doesn't have the same visceral effect for people as hearing something in their own voice, that of a loved one or of someone admired.

VALL-E moves us one step closer to ubiquity, and although it is not the type of model you run on your phone or home computer, that isn't too far off, Betker speculated. A few years, perhaps, to run something like it yourself; as an example, he sent this clip he'd generated on his own PC using Tortoise-TTS of Samuel L. Jackson, based on audiobook readings of his:

https://techcrunch.com/wp-content/uploads/2023/01/samuel_jackson.mp3

Good, right? And a few years ago you might have been able to accomplish something similar, albeit with greater effort.

This is all just to say that while VALL-E and the three-second quickfake are definitely notable, they're a single step on a long road researchers have been walking for over a decade.

The threat has existed for years and if anyone cared to replicate your voice, they could easily have done so long ago. That doesn't make it any less disturbing to think about, and there's nothing wrong with being creeped out by it. I am too!

But the benefits to malicious actors are dubious. Petty scams that use a passable quickfake based on a wrong number call, for instance, are already super easy because security practices at many companies are already lax. Identity theft doesn't need to rely on voice replication because there are so many easier paths to money and access.

Meanwhile the benefits are potentially huge — think about people who lose the ability to speak due to an illness or accident. These things happen quickly enough that they don't have time to record an hour of speech to train a model on (not that this capability is widely available, though it could have been years ago). But with something like VALL-E, all you'd need is a couple clips off someone's phone of them making a toast at dinner or talking with a friend.

There's always opportunity for scams and impersonation and all that — although more people are parted with their money and identities via far more prosaic ways, like a simple phone or phishing scam. The potential for this technology is huge, but we should also listen to our collective gut, saying there's something dangerous here. Just don't panic — yet.

Cosmo
Sabrina Carpenter looks practically naked in completely see-through lace mini dress
Sabrina Carpenter went braless wearing the Mirror Palais Anemone Dress in butter featuring illusion tulle adorned with lace appliqués along the neckline and hem
9 hours ago
Yahoo News Australia
Bizarre Westfield car park scene baffles Aussies: 'Menace to society'
Shoppers got a surprise when arriving at the car park to discover the spaces were taken, but not by cars. Find out what happened.
19 hours ago
Yahoo Finance AU
$500 cost of living payment coming for thousands of Aussie households
Eligible homes can get this lump payment as well as a $180 rebate to go towards their electricity bill.
17 hours ago
Yahoo News Australia
Driver fumes at 'very uncool' find on her car parked on suburban street
A woman is furious about what happened to her car when she left it parked on a street in Bondi. Check out why here.
2 days ago
HuffPost
'How Embarrassing': Trump Mocked For 'Pretending To Be President' In Strange Ceremony
The former president gave a truly bizarre "White House" gift to a visitor.
12 hours ago
Cosmo
Anitta coordinates her teeny bikini with her… fridge?
Anitta shared a series of pics on IG posing in a teeny tiny green string bikini with yellow trim perfectly coordinated to her Smeg fridge and orange juice.
a day ago
Yahoo Lifestyle
Farmer Wants A Wife's Tom reveals surprising behind the scenes fact: 'Uninterested'
EXCLUSIVE: Farmer Wants A Wife star Farmer Tom has told Yahoo Lifestyle a little-known fact from behind the scenes of the show. Read more.
2 days ago
Yahoo Sport Australia
Reece Walsh 'weakness' called out as Broncos move to address ugly truth about NRL star
Reece Walsh is prone to an error or 53. Read more here.
2 days ago
Yahoo Finance AU
Centrelink payment change warning for millions of Aussies: 'May pay early'
Aussies receiving Centrelink payments have been reminded of upcoming closures and payment changes. Here’s what you need to know.
2 days ago
HuffPost
'I Shouldn't Have Said That': Joe Biden Mocks 1 Of Trump's Most Cherished Traits
The president took aim at one of his predecessor's personal trademarks -- and the audience loved it.
16 hours ago
HuffPost
Lara Trump Alarms Critics With 'Frightening' Comment About RNC's Election Plans
"Sounds like a perfect authoritarian election plan to me," fascism expert Ruth Ben-Ghiat commented.
a day ago
Evening Standard
Ukrainian forces 'regain lost positions' in battle for Chasiv Yar against Putin's army
The news from the frontline came as Congress passed a £48 billion military aid package for Kyiv
2 days ago
BBC
Suspended jail terms for pair who left dog home alone
Bentley was reduced to eating rubbish after his owners went on holiday and left him alone at home.
a day ago
The Independent
London horses – live: Runaway horse in serious condition undergoes operation as cavalry inspection takes place
Cavalry horses Vida and Quaker ran loose in the road near Aldwych, central London
35 minutes ago
The Daily Beast
New Complaint Alleges Trump Campaign Hid Millions in Lawyer Payments
John Taggart/Pool via ReutersOn Wednesday morning, The Daily Beast published a report detailing how Donald Trump’s presidential campaign and four associated PACs have been using a GOP compliance firm to pay legal fees, obscuring who is the ultimate recipient of millions of dollars in legal payments.By Wednesday night, the Trump campaign and the four PACs were facing a new ethics complaint over the arrangement.The complaint, which nonprofit watchdog Campaign Legal Center filed with the Federal El
21 hours ago
Yahoo News Australia
RAM driver's 'passive aggressive' note to motorist parked outside home
The motorist hadn't parked over the resident's driveway but returned to find a 'confusing' note on his car. Do you think it was justified?
19 hours ago
Yahoo Lifestyle
MasterChef Australia fans slam judge Andy Allen over peculiar habit: 'Annoying'
MasterChef fans are coming for Andy Allen thanks to a puzzling habit. Find out what it is here.
2 days ago
Yahoo Sport Australia
Cricket fans call out Justin Langer act after Marcus Stoinis' historic century in IPL
Marcus Stoinis has made a massive statement ahead of the T20 World Cup. Read more here.
2 days ago
Yahoo News Australia
Rare Aussie animal you can only see at one zoo
After the species was rediscovered last year, the zoo has been breeding dozens of babies. They're now on display.
2 days ago
Yahoo Sport Australia
Jett Cleary's move to Warriors confirmed as Panthers rocked by sad family development
Jett Cleary wants to forge his own path to the NRL. Read more here.
2 days ago

ALL ORDS

AUD/USD

ASX 200

OIL

GOLD

Bitcoin AUD

CMC Crypto 200

VALL-E's quickie voice deepfakes should worry you, if you weren't worried already

Latest stories

Sabrina Carpenter looks practically naked in completely see-through lace mini dress

Bizarre Westfield car park scene baffles Aussies: 'Menace to society'

$500 cost of living payment coming for thousands of Aussie households

Driver fumes at 'very uncool' find on her car parked on suburban street

'How Embarrassing': Trump Mocked For 'Pretending To Be President' In Strange Ceremony

Anitta coordinates her teeny bikini with her… fridge?

Farmer Wants A Wife's Tom reveals surprising behind the scenes fact: 'Uninterested'

Reece Walsh 'weakness' called out as Broncos move to address ugly truth about NRL star

Centrelink payment change warning for millions of Aussies: 'May pay early'

'I Shouldn't Have Said That': Joe Biden Mocks 1 Of Trump's Most Cherished Traits

Lara Trump Alarms Critics With 'Frightening' Comment About RNC's Election Plans

Ukrainian forces 'regain lost positions' in battle for Chasiv Yar against Putin's army

Suspended jail terms for pair who left dog home alone

London horses – live: Runaway horse in serious condition undergoes operation as cavalry inspection takes place

New Complaint Alleges Trump Campaign Hid Millions in Lawyer Payments

RAM driver's 'passive aggressive' note to motorist parked outside home

MasterChef Australia fans slam judge Andy Allen over peculiar habit: 'Annoying'

Cricket fans call out Justin Langer act after Marcus Stoinis' historic century in IPL

Rare Aussie animal you can only see at one zoo

Jett Cleary's move to Warriors confirmed as Panthers rocked by sad family development