Tech giants sued for using voice samples to train artificial intelligence without consent. Google is facing a new lawsuit under the Biometric Information Privacy Act (BIPA), in which the company is accused of training voice AI models using biometric voice samples from journalists, investigative podcasters, and audiobook narrators.

The lawsuit was filed by seven plaintiffs who allege that Google created its core models based on thousands of hours of recorded speech to extract biometric voice samples. These models were used to power products such as Gemini Live, NotebookLM Audio Overviews, YouTube automatic dubbing, Google Cloud Text-to-Speech, and Google Assistant.

The plaintiffs include award-winning radio journalists Carol Marin and Philip Rogers, investigative podcasters Yohance Lacour, Alison Flowers and Robin Amer, and audiobook narrators Lindsey Dorcus and Victoria Nassif.

Separate but related class action lawsuits filed by the same group of defendants also target Amazon, Apple Inc., Meta Platforms, Microsoft, NVIDIA, ElevenLabs, Adobe, and Samsung Electronics.

According to the allegations, these companies built commercial AI-based voice systems using voice samples collected from the internet and other sources without obtaining written consent, providing notice, or publishing biometric retention policies required by BIPA.

The BIPA deems biometric identifiers to be “biologically unique to an individual,” meaning that once they are disclosed or misused, they cannot be easily replaced or invalidated by an individual.



https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai

Let’s start with what morphing is.

Morphing is an image transformation technique that smoothly changes one image into another, used in film and computer animation.

Voice morphing (or voice conversion) is an advanced digital audio processing technique that seamlessly transforms one person’s voice (the source) into another person’s voice (the target), while preserving the content of the speech. It uses artificial intelligence (AI) algorithms, machine learning, and digital signal processing (DSP). The system analyzes the characteristics of the source voice (timbre, pitch, timbre) and maps them to the characteristics of the target voice.

Researchers analyzing a signal-level approach to voice morphing attacks have revealed vulnerabilities in biometric voice recognition systems. They demonstrated that voice morphing attacks combine identities to bypass voice biometrics.

This is time-domain voice identity morphing (TD-VIM), which allows for the mixing of identities without embedding them in a structure or reference text.

In biometric systems, it’s common practice to associate each sample or template with a specific individual. Advanced voice identity morphing (VIM) allows the generation of a sample that combines the identities of two or more speakers. “The modified voice sample can be used to match all identities whose voice samples were used to generate morphing attacks, which poses a high risk in application scenarios such as banking and finance, where a single identity verification is essential.”

To investigate this issue, the research team created four distinct morphing signals and assessed their effectiveness through a comprehensive vulnerability analysis. The data was compared to the Generalized Morphing Attack Potential (G-MAP) metric, “which measures attack effectiveness in two deep learning-based speaker verification systems (SVS) and one commercial system, Verispeak.”
The results highlight the effectiveness of the TD-VIM method in bypassing advanced verification mechanisms, underscoring the importance of improving SVS security.


The research comes from the Indian Institute of Technology and the Norwegian University of Science and Technology.

more about the voice morphing phenomenon here


Deepfake attacks are AI-based frauds that use a short-form voice sample generated from a source (e.g., social media) to gain unauthorized access to accounts, create additional ads, etc.

Biometric data leak. Unlike a password, a voice cannot be changed. If a voice template leaks from the database, it is irretrievably stolen, creating a long-term risk to the user’s identity. The report “Cyber ​​Threats: What Poles Are Afraid of,” compiled by the Office for Personal Data Protection, among others, shows that one-third of Poles fear information leaks (in general). The most frequently asked questions include where and how the data will be stored, whether it will be adequately secured, and whether it will not be used unlawfully.

Loss of privacy and image. Voice can reveal more than just identity – speech analysis can reveal health, mental characteristics, and emotions, which can also be used against the user.

What is the biggest problem for you?



You can read more about privacy in voice biometrics on our blog
https://biometriq.pl/en/privacy-in-voice-biometrics/

Do you know what watermarking is in voice biometrics? It’s a method of digitally tagging audio. It involves embedding an inaudible marker, called an identifier, into an audio file. The goal is to protect the recording from unauthorized use and verify its authenticity.

Watermarking is a tool that significantly improves the security of voice biometrics systems, mainly by preventing voice-based attacks, so-called deepfakes.

In one of our tools, we developed this proprietary method, a unique technique that protects audio recordings from being used for voice synthesis or access. The method is currently in functional use.

Phase 2 of the Vesper project, a biometrics-based voice communicator, is nearing completion. During this phase, we worked on creating audio stream augmentation technology. We wrote about what this augmentation is here https://biometriq.pl/en/voice-stream-augmentation-what-is-it/

Our proprietary voice stream augmentation engine is currently undergoing perceptual (listening) and blind testing. Their goal is to provide an objective evaluation to confirm proper engine operation in line with the established quality parameters. Furthermore, the built-in voice stream augmentation technology in the voice messenger is designed to aid in detecting unauthorized voice use for further synthesis/conversion without causing degradation of sound to the human ear. This is all to prevent voice theft and ensure the most effective service performance.

It’s worth noting that solutions on the market such as SKYPE, ZOOM, DISCORD, Google Meet, TEAMS, WhatsApp, Signal, Threema, Viber, and Telegra do not support biometric caller authentication.

We are pioneers in this regard.

The comprehensive project completion is scheduled for September 2026.

In the first stage of the project, we mainly tested the far end voice stream source authenticity algorithm, which we informed you about here https://biometriq.pl/en/tests-of-a-voice-communicator-with-a-source-authenticity-detection-module-are-underway/


You can read more about the project on the website https://biometriq.pl/en/vesper-save-voice-communication-platform-with-integration-of-biometric-services/


Project financed by EU funds.

Are you curious about the final solution?

The first international standard for age-assurance technology has been published – ISO/IEC 27566-1:2025. This document establishes a framework for age-assurance systems and describes their core features, including privacy and security, to enable age-based eligibility decisions.

Access permissions refers to the term that authorizes access to applications or services. Definitions of age verification, age estimation, age inference, and subsequent validation are available here.

The standard’s main initiator is Tony Allen, head of the UK Age Check Certification System (ACCS), founder of the Global Age Assurance Standards Summit, and leader of the Australian Age Assurance Technology Research (AATT). He calls the publication of ISO 27566-1:2025 (which he co-authored) “a significant breakthrough in age assurance at the global level.”

A sample of the ISO 27566-1:2025 standard is available free of charge, but access to the full version of the document requires purchase. https://www.iso.org/standard/88143.html

more about the standard https://www.biometricupdate.com/202512/first-international-standard-on-age-assurance-sees-publication

source, photo https://www.biometricupdate.com

  1. The voice biometrics market is relatively young, currently estimated at USD 2-3 billion, USD 2.6 billion according to the Mordor Intelligence report “Voice Biometrics Market Size, Forecast Report, Landscape 2025”.
  2. Depending on the source, forecasts assume growth of approximately $10-15 billion over the next 8-10 years.
  3. The leading region is North America – in the Fortune Business Insights analysis, the share in 2024 was nearly 37%.
  4. Asia-Pacific (APAC) is often cited as the fastest growing region in the coming years.
  5. The “Healthcare and Life Sciences” sector will be the leader in 2025 with a 40% market share.
  6. Growth is driven by: growing security requirements, the need for passwordless authentication, the development of voice and AI technologies, and the digitization of financial and contact services.

sources:

What distinguishes effective voice biometrics systems? The following four indicators determine the advantage of one system over another:
1. Accuracy rate, it means that the effectiveness of biometric systems should be in the range of 95-99%.

2. FAR (False Acceptance Rate), a metric that measures how often a system incorrectly accepts an unauthorized person (e.g., someone impersonating a user) as a valid user. In the most accurate systems, this rate is less than 1%. The lower the rate, the more secure the system and the more difficult it is to impersonate.

3. FRR (False Rejection Rate), a metric that measures false rejections, or the number of times the system rejects a genuine user when it should accept them. Ideally, this figure is below 3%.

4. EER (Equal Error Rate). The point at which the FAR equals the FRR, this metric is often used to compare the quality of biometric systems.

The most effective systems are generally considered to be Phonexia oraz ID R&D systems due to their outstanding performance in comparative tests.

In our research, we primarily use Phonexia engines, but we also utilize others such as Kaldi (X-vector) and ECAPA. The goal is to test our algorithms as extensively as possible in a diverse environment. Security is our top priority.

Phase 1 of the Vesper project is nearing completion. We’ve launched a test version of the messenger with an implemented far-end voice stream authentication module. Tests are being conducted on three different environments: Windows, Android, and iOS. The results are consistent with the project’s KPIs. We’re working to ensure that quality indicators not only meet the design minimums but, where possible, exceed the established goals. Our priority is to develop a product that meets user needs and builds a positive user experience.


We conduct experiments based on 40 speakers, 20-second recordings, testing each recording across 5 channels, and obtaining over 171,500 embeds. This number of recording configurations is designed to help achieve the target parameters, confirming the effectiveness of our messenger.

Vesper Messenger is intended to be a response to the growing problems of cybersecurity and identity theft.

More about the project https://biometriq.pl/en/vesper-save-voice-communication-platform-with-integration-of-biometric-services/

The exhibition was marked by the ubiquitous AI. Many companies presented their latest achievements in constructing systems that communicate autonomously with people. The humanoid robot Ameca (Etisalat) interacting with its interlocutors aroused great interest. The stands with interactive agents (Amdocs) offered an almost unbelievable quality of image and speech generated by the systems.

Google has unveiled Gemini Live, its response to ChatGPT’s voice mode.  Gemini Live has function Share Screen With Live, that allows Gemini to interact with the image displayed on the phone’s screen. Deutsche Telekom has indicated a possible direction for the development of phones by turning the entire phone into a chatbot. The phone has no applications and is a personal assistant that communicates with the user by voice. The basis of the solution is a digital assistant from AI Perplexity, but it is also to be open to, among others, Google Cloud AI, ElevenLabs, and Picsart. South Korean startup Newnal has presented a new operating system for mobile phones that uses historical and current user data to create a personalized AI assistant that is to eventually become an AI avatar behaving just like the user.

All of the above solutions, as well as many others, are connected by the use of voice technologies for two-way communication. The direction indicated at MWC 2025 is clear – our actions will be supported by avatars and bots communicating with us autonomously. The possibility of quick, machine confirmation of who we are talking to is therefore becoming even more important than ever before, because the quality of autonomous voice communication systems does not guarantee correct verification of the speaker by a human.

Photos by Andrzej Tymecki