Tools/Detectors & Checkers

AI Voice Detector

Free AI voice detector: upload an audio clip and check it for the signs that give away AI-generated or cloned speech. Works on voice notes, call recordings and audio pulled from a video, and everything runs in your browser.

Read this before you trust the result

This is a rough screen, not a verdict. It measures pacing, background noise and frequency content in the file you give it. A good 2026-era voice model can pass every one of these checks, and a real human recording that has been compressed, noise-gated or edited can fail several of them. Use the signal list to decide what to check next. Never use it as proof about a person.

Upload an audio clip

Drop an audio file here, or click to pick one

MP3, WAV, M4A, OGG. 15 seconds or more works best.

Your file never leaves this device. It is decoded and measured by your own browser, with no upload and no server.

How to spot an AI voice by ear

Your ears are still better than any free AI voice detector, including this one. Play the clip twice. The first time listen to the words, the second time ignore them and listen only to the sound. Headphones help more than you would think, because a phone speaker throws away the quiet detail that gives a generated voice away.

Breathing

People run out of air. Listen for an intake before a long sentence, and for a breath that lands in a slightly wrong place because the speaker misjudged the line. Many models either skip breaths or paste in the same breath sample over and over. A breath that sounds identical every time is worse than no breath at all.

Mouth noise

Lip smacks, swallows, a dry click at the start of a word, the faint sound of someone shifting in a chair. Real capture is full of small accidents. Generated audio is usually scrubbed clean of them.

Flat feeling

Ask whether the emotion actually tracks the words. AI voices tend to apply one steady mood to a whole passage, so a sentence about something alarming gets the same warm even delivery as the sentence before it. Sarcasm, hesitation and self-correction are still hard to fake.

Odd pacing

Listen for gaps that all feel the same length, stress landing on the wrong word in a sentence, or a comma pause where a person would have pushed straight through. Real speakers also restart sentences, and say "um" in the middle of a thought.

Names, numbers and abbreviations

This is the fastest giveaway. Ask the voice to say an unusual surname, a street name, a long account number or a currency amount. Models mispronounce rare names confidently, read numbers in an unnatural rhythm, and often flatten the digits into one steady stream instead of grouping them the way people do.

Room

Real voices come with a room attached: a bit of echo, a hum, a car outside. If the voice sounds like it is standing in empty space, that is worth noticing, though a synthetic room can be added afterwards.

S sounds and hard consonants

Listen to the S, F and T sounds on their own. Cloned speech often renders them as a thin buzz that starts and stops too neatly, and two different S sounds in the same sentence can come out sounding like copies of each other. A person's sibilance is messier and changes with how hard they are pushing the word.

The same tell twice

If you have more than one clip of the voice, compare them. A clone tends to repeat the exact same quirk across recordings made days apart: the identical breath, the same little rise at the end of a phrase, the same slightly wrong stress on a name. People are not that consistent.

Checking a YouTube video, a voice note or a call recording

This tool reads a file, so the first job is always the same: get the audio onto your device as a file, keep the cleanest copy you can find, and aim for at least fifteen seconds of someone actually talking. A screen recording of a video played out of a laptop speaker has been through two extra rounds of damage, and every one of those rounds makes a real person look more synthetic.

YouTube, TikTok and Reels voiceovers

Platform audio is re-encoded at upload, so treat the high-frequency signal as close to worthless here and weigh pacing, breathing and room tone instead. Narration-over-stock-footage channels are the usual suspects, and the giveaway is often editorial rather than acoustic: perfect diction for ten minutes straight, no stumbles, no restarts, and a script that reads out numbers and foreign names without ever slowing down. Check the channel's older uploads too, because a channel that switched to a generated voice usually did it between one video and the next.

WhatsApp, Telegram and iMessage voice notes

Save the note as a file, then check it. Messenger apps compress hard, which flattens the top end and can gate the background into silence, so a real voice note from a relative can trip two of the checks below on its own. What a voice note does give you is context no detector has: whether it arrived from the number you expect, whether it answers what you actually asked, and whether the person will say the same thing live.

Recorded calls and voicemail

Phone audio is the hardest case for every AI voice detector, free or paid. The network throws away everything above roughly 3.4 kHz on a standard call, which removes the frequency evidence outright and squeezes the rest. If you only have a recording of a suspicious call, read the pause and breath signals, ignore the frequency one, and put most of your weight on verifying the caller another way.

Podcasts, audiobooks and ads

These have usually been noise-gated, de-essed and loudness-normalised by a human editor, which is the single most common cause of a false alarm on this page. A studio-recorded podcast can show near silent gaps, even volume and no breaths, all of which look machine-made. If the show publishes video of the same episode, a few seconds of lip sync settles it faster than any audio check.

Other AI voice detector tools, and what each one is actually for

No single checker covers the ground, so it helps to know what each kind is built to do before you trust a result from it. Vendor accuracy figures are almost always measured by the vendor, on audio the vendor chose. Independent work in the field, such as the ASVspoof challenges that most published research is scored against, has repeatedly found that error rates climb once audio has been compressed, sent through a phone network, or produced by a generator the detector never saw in training.

Generator-made classifiers

ElevenLabs runs an AI Speech Classifier for audio made with its own models, and its documentation is refreshingly blunt about the edges: it reads the first minute of a sample, it does not reliably classify audio from the Eleven v3 model, and it cannot speak for audio made anywhere else. Treat this whole category the same way. A yes is informative, a no only rules out one vendor.

Browser extensions

An AI voice detector Chrome extension scores audio while it plays in a tab, which is handy for YouTube or a video call and saves you the file-wrangling. Hiya and NordVPN both ship one. Before you point any of them at something private, check whether it analyses on your own machine or streams the audio to a server, because the two designs have very different consequences for a recording of your family.

Upload sites and paid APIs

Services like Resemble, Pindrop, TruthScan and the various free web uploaders take your file, run a trained model server-side and hand back a percentage. A trained model genuinely sees things a browser cannot. In exchange you are uploading the audio, and the percentage is a confidence figure from one model on one day, not a measurement of truth. Read the retention policy before sending anything sensitive.

Open source on GitHub

If you can run Python, the research code is public. Architectures such as AASIST, and the ASVspoof datasets they are trained and scored on, are the reference point most papers and most commercial claims trace back to. Results there are published as an equal error rate, and the number you should care about is the one measured on attacks the model was not trained on, which is always the worse one.

Provenance beats detection when you can get it

Guessing from the waveform is the last resort. Some generators sign their output with C2PA Content Credentials, a tamper-evident record of what made a file, and the EU AI Act requires machine-readable disclosure of synthetic media. The catch is that platforms routinely strip metadata on upload, so an absent credential proves nothing at all, while a valid one is worth more than every acoustic signal on this page put together.

Voice tools that answer a different question

AI accent, gender and age detectors

These classify a speaker, not a file. BoldVoice's Accent Oracle, the one people usually mean by an AI accent detector, has you read a short script and then guesses your native language from your English. Gender and age estimators work the same way. None of them tell you whether the voice was generated, and a good clone will be classified just as happily as a person.

Working out which model made a clip

There is no dependable public AI voice name detector that identifies the generator behind an arbitrary file. Vendor classifiers recognise their own output, and occasionally a clip matches a stock voice from a library closely enough to be recognisable by ear. Beyond that, attribution is guesswork, and a tool that names a model with confidence is telling you more than it knows.

AI voice detector versus AI text detector

They catch different halves of the same problem. A text detector looks at word choice and sentence rhythm in a transcript, so it can flag a script a chatbot wrote even when a real person read it out. This tool looks at the sound, so it flags a generated read even when a human wrote every word. If something feels off about a video, running both on the transcript and the audio tells you more than either alone.

Tools that promise undetectable output

Services advertising undetectable voices are selling the same arms race from the other side: add breaths, filler words, a little room noise, vary the pacing, and the acoustic signals that any free checker leans on stop showing up. This is exactly why nothing on this page is presented as proof, and why verifying the person matters more than scoring the file.

Common mistakes when checking a voice clip

Judging a five-second clip

Pacing, breathing and room tone all need a stretch of continuous speech to measure. Under six seconds this tool refuses to guess, and any tool that does return a confident verdict on a short clip is reading noise. Find a longer sample of the same voice before you conclude anything.

Checking a copy of a copy

People often test the version that reached them last: a video re-shared three times, a voice note forwarded through two apps, a clip captured by pointing a phone at a screen. Each hop re-encodes the audio and strips exactly the detail the checks depend on. Go back to the original file whenever one exists.

Treating a leaning as evidence about a person

A result here is a description of a waveform. It cannot tell you who spoke, who recorded it or what anyone intended, and it has no business in an accusation, a dispute or a disciplinary process. If the stakes are real, the answer comes from checking the source, not from scoring the file.

Stopping at one detector

Different checkers disagree constantly, because they were trained on different audio and measure different things. Two tools agreeing is worth something. One tool agreeing with what you already suspected is worth very little, and that is the failure people fall into most often.

Reading a human-leaning result as all clear

The best voice models now add breaths, hesitations and room sound on purpose, which means the signals that catch older output are the first things a current generator fixes. A human-leaning result means nothing tripped. It does not mean nothing is there.

If you think a call is a cloned voice

Voice cloning scams follow one shape: someone you love is in trouble, it is urgent, and you must send money or read out a code right now. The urgency is the attack. It exists to stop you from checking. Cloning a recognisable voice now takes only a few seconds of clear speech, which a scammer can lift from a voicemail greeting, a wedding video or any clip posted publicly, so assume the voice on the line can be faked and verify the person instead.

Hang up

You owe a caller nothing. Stopping the call costs you a few seconds if you are wrong, and saves everything if you are right. Do not stay on to argue or to test them, because every second of your voice is training material.

Call back on a number you already had

Use the number in your own contacts, not one the caller gave you and not the one that appeared on screen. Caller ID is trivially spoofed. If you cannot reach the person, call someone else who would know where they are.

Agree on a family code word

Pick a word or short phrase with the people closest to you, something no one could guess and that has never been posted anywhere. In a real emergency, ask for it. A clone of a voice is not a clone of a memory.

Ask something only they would know

Not a fact from social media. Ask about a small shared moment: what you ate last time you met, what you argued about at the airport. A scammer working from a scraped voice sample will stall, get angry, or push the urgency harder.

Never move money or codes during the call

Gift cards, crypto, wire transfers and one-time passcodes are the whole point of the call. Anything real can wait ten minutes while you verify. Then report it to your bank and to your country's fraud line.

How this AI voice detector works

It all happens on your device

The file is read by your browser, decoded with the Web Audio API and measured in memory. Nothing is sent anywhere, there is no account, and closing the tab wipes it. That matters when the clip is a voicemail from a relative or a recording of a suspicious call.

What gets measured

The audio is cut into 20 millisecond frames. Loud frames count as speech, quiet ones as gaps. From that the tool works out how much the pause lengths vary, how loud the background is between phrases, how many quiet-but-not-silent moments look like breaths, and how much the volume moves across the clip. A 2048-point FFT on speech frames then measures how much the tone-versus-hiss balance shifts and how much energy sits above 7 kHz.

Why there is no percentage

A number like "93% AI" would be made up. These measurements were never calibrated against a labelled dataset of every voice model in circulation, and no browser tool could be. You get a leaning and the reasoning behind it, so you can weigh each signal yourself.

What trips it up

Phone calls, video platform audio and low-bitrate MP3s all strip high frequencies and gate the background, which makes real people look synthetic. A podcast that has been through noise removal and loudness normalisation can trip three checks at once. Meanwhile the newest voice models add breaths, filler words and room sound on purpose, so they sail straight through.
FAQ

Frequently asked questions

Is this AI voice detector free?

Yes. Free, online, no sign-up, no upload limit and no watermark on the result. It runs entirely in your browser, so there is no server cost to pass on to you.

Can an AI voice detector tell me if a voice is fake for certain?

No, and be suspicious of any tool that says it can. Even paid detectors trained on large datasets publish error rates, and they slip further every time a new voice model ships. Treat any detector, this one included, as one input next to context: who sent the file, what they want, and whether the story checks out elsewhere.

How accurate are AI voice detectors?

Accurate enough to be useful, not accurate enough to settle an argument. The high percentages vendors advertise are usually measured on audio the vendor picked, and independent testing in the field has found that error rates rise sharply once a clip has been compressed, sent through a phone network or produced by a generator the detector was never trained on. A real phone recording is close to the worst case for every tool on the market.

Does it detect ElevenLabs, Resemble or PlayHT voices?

It is not tied to any one generator. It looks at properties of the audio itself, so it reacts to whatever a model leaves behind rather than to a vendor fingerprint. In practice older or lower-tier voices from any provider show more of these signals than a current flagship voice does. ElevenLabs also runs its own classifier for audio made with its models, which covers that one generator in more detail than any general checker can.

What file types work?

Whatever your browser can decode: MP3, WAV, M4A, OGG and usually FLAC. If a file will not open, convert the clip to WAV and try again. Longer is better, since pacing and breathing need something to measure. Under six seconds of speech and the tool will say it cannot judge.

Can I check a YouTube or TikTok video with this?

Extract the audio first, then upload the file. Be aware that platform audio has been re-encoded at least once, which eats the high frequencies and can push a real human clip toward an AI-looking reading. Weigh the pacing and breathing signals more heavily than the frequency one, and check whether the channel used a different voice in older uploads.

Is there an AI voice detector app or Chrome extension?

Several browser extensions score audio as it plays in a tab, which suits YouTube and video calls better than a file uploader does. Before installing one, check whether it analyses on your own device or sends the audio to a server, since that decides whether a private recording stays private. This page itself works on a phone browser, so there is nothing to install for a file you already have.

Is there an open-source AI voice detector on GitHub?

Yes, the research code is public. Architectures such as AASIST, trained and scored on the ASVspoof datasets, are what most academic and commercial claims trace back to, and you can run them yourself with Python. Expect research-grade setup rather than a click-and-go tool, and pay attention to results measured on attacks the model never saw in training, because those are the honest numbers.

Can it tell me age, gender, accent or who is speaking?

No. Those are separate problems, and doing them in a browser badly would only produce confident nonsense. Accent and gender classifiers describe a speaker rather than a file, so they will happily classify a cloned voice too. This tool answers one question: does the audio carry the signs of being generated rather than recorded.

How much audio does someone need to clone a voice?

Far less than people expect. A few seconds of clear speech is enough for current cloning tools, which is roughly the length of a voicemail greeting or one sentence from a posted video. That is why a family code word and a callback on a number you already had are worth more than any listening test.

What should I do if a call sounds like a cloned family member?

Hang up, then call the person back on the number already in your contacts, not one the caller gave you. Do not send money, gift cards, crypto or a one-time passcode while the call is live, and do not stay on the line to test them, since every second of your voice is material. If you cannot reach them, call someone else who would know where they are, then report it to your bank and your country’s fraud line.

Is my audio uploaded anywhere?

Never. There is no upload step. Open your browser network tab and watch if you like: the file is read straight into memory on your machine and discarded when you leave.