AI writing is now the baseline, not the exception
The pressure to detect AI-written resumes comes from a real shift in candidate behavior. In a survey of 1,000 US job seekers run by ResumeBuilder.com, 46 percent said they were already using ChatGPT to write their resumes, cover letters, or both, and that figure was measured in early 2023, before assistive writing tools were built into email clients, browsers, and word processors. The same study reported that 78 percent of those candidates landed an interview and 59 percent received a job offer, so the practice is not a fringe shortcut. It works, and it is spreading.
That changes the question a recruiter should be asking. If a large share of your applicant pool used an assistant to phrase their experience, then "did AI touch this document?" stops being a useful filter, because the honest answer for most strong candidates is now yes. The signal that still matters is narrower: is the experience real, is it relevant to the role, and can it survive a conversation? A detector that cannot answer those questions is solving the wrong problem.
Why resumes are hard for AI detectors
Most AI text detectors work by looking for statistical patterns in language, usually some measure of how predictable each word is given the words around it. That approach is difficult on resumes because resumes are already formulaic. They use action verbs, short bullets, repeated structures, and polished business language. A well-written human resume can look machine-generated. A lightly edited AI resume can look human. The very qualities that make a resume readable, conciseness and consistency, are also the qualities that push a detector toward a false alarm.
The reliability problem is not hypothetical, and it is not limited to small vendors. In July 2023 OpenAI quietly retired its own AI Text Classifier, citing a low rate of accuracy. The company had disclosed at launch that the tool correctly identified only about 26 percent of AI-written text as "likely AI," while incorrectly flagging human-written text as AI 9 percent of the time. If the organization that builds the most widely used text generator could not reliably detect its own output in long-form prose, the case for trusting a third-party detector on a 400-word resume is weak.
This means the detector result is rarely strong enough to drive a hiring decision. It can be one signal, but it should not be the signal. The question recruiters need to answer is not "was this written by AI?" The better question is "are the claims accurate and relevant?"
False positives create real hiring risk
False positives matter because they affect real candidates. Non-native English speakers, candidates using resume templates, neurodivergent candidates, and people who use writing assistants may all produce polished, structured resumes that trigger AI-text suspicion. Rejecting them for that reason alone is unfair and weakly evidenced.
The clearest evidence here comes from a peer-reviewed Stanford study published in the Cell Press journal Patterns in 2023. The researchers ran 91 TOEFL essays, all written by humans who were non-native English speakers, through seven widely used GPT detectors. More than 61 percent of those genuine human essays were misclassified as AI-generated, and roughly one in five was flagged by all seven detectors at once. One detector flagged nearly 98 percent of the essays. On essays written by native English-speaking US eighth graders, the same detectors were close to perfect. In other words, the tools were not measuring "machine-written" so much as "written in a constrained, formal style," which is exactly how many non-native speakers and many resume templates read.
The mechanism is worth understanding, because it explains why the bias is structural rather than a tuning bug. Most detectors lean on perplexity, a measure of how predictable the text is. Writing that uses a smaller, safer vocabulary and simpler sentence construction scores as low perplexity, and low perplexity looks like AI to the model. Senior author James Zou put it plainly: current detectors are "clearly unreliable and easily gamed," so teams should be cautious about using them as a solution. A resume, almost by definition, is short, safe, and formulaic, which means it sits squarely in the zone where these tools fail.
The real-world stakes are not abstract. Reporting by The Markup documented a case where Turnitin's detector flagged more than 90 percent of an international student's paper as AI-generated, even though the student could produce drafts and notes proving they wrote it. Transplant that error rate into a hiring funnel and the cost is a qualified candidate silently filtered out, with no signal to the recruiter that anything went wrong. A candidate should not be penalized for using tools to communicate clearly. They should be held accountable for the truthfulness of what they claim. That is a much stronger and more defensible standard.
Multi-signal review is stronger
Modern screening should combine several reviewable signals: CV content, application answers, role requirements, profile links, work samples, interview depth, reference checks, and application metadata. The goal is not to catch every AI-written sentence. The goal is to identify claims that need verification before the candidate advances.
The reason multiple weak signals beat one strong-sounding signal is straightforward. A single AI-text score is a binary verdict dressed up as a probability, and it correlates with writing style rather than with deception. A bundle of independent signals is harder to fool, because a fabricated application that beats one check usually trips another. The recruiter is not hunting for a confession that the document was AI-assisted; they are looking for internal contradictions that genuine experience does not produce.
In practice, the signals worth weighting fall into a few buckets:
- Consistency: do the dates, titles, and seniority on the resume line up with the candidate's public profiles and the answers in the application form?
- Specificity: does the experience include the kind of concrete detail, tools, metrics, constraints, that is hard to invent and easy to verify?
- Provenance: is the contact information a real, reachable identity, or a disposable email and a profile link that resolves to a look-alike domain rather than the real one?
- Evidence: is there a portfolio, a repository, a work sample, or a reference who can confirm ownership of the claimed work?
For example, a resume that mirrors the job description exactly may not be suspicious on its own; tailoring a CV to a posting is good advice that career coaches give every day. But if it also uses a disposable email domain, links to a LinkedIn-like domain that is not LinkedIn, and cannot be supported by a coherent phone-screen answer, the recruiter has a stronger reason to investigate. No single one of those facts is damning. Together they describe a pattern that warrants a closer look.
Structured interviews still matter
The easiest way to test fabricated experience is to ask for detail. Ask what the candidate personally owned, what constraints they faced, what went wrong, which tools they used, what tradeoffs they made, and what evidence exists. Real experience usually has texture. Fabricated experience often stays at the level of polished outcomes.
Structured interviews make this fair. Every candidate is asked comparable questions, and warnings simply tell the recruiter where to probe more carefully. This keeps the process consistent while still addressing risk.
The fairness and compliance angle recruiters cannot ignore
There is a second reason to treat AI-text scores with caution, and it is not about candidate experience. It is about legal exposure. If a screening signal rejects one protected group at a meaningfully higher rate than another, that is the definition of disparate impact, and it does not matter that the tool was never designed to discriminate. The Stanford finding that detectors flag non-native English writers far more often than native speakers is exactly the kind of pattern that can map onto national origin, a protected characteristic in most hiring jurisdictions.
This is no longer a theoretical worry. New York City's Local Law 144, in force since 2023, requires employers using automated employment decision tools to commission an independent bias audit and to publish the results. The EU AI Act classifies systems used to screen and filter job applications as high risk, with conformity and transparency obligations attached. A hidden rule that auto-rejects "AI-looking" resumes is precisely the kind of opaque, unaudited filter these regimes are built to catch. The defensible posture is the opposite of a silent filter: log the signal, explain it, keep a human in the loop, and be able to show why a given application was set aside for review rather than discarded.
Where an ATS should help
An ATS should collect signals, explain them, and keep the review workflow visible. It should not turn AI-text detection into a hidden rejection rule. Treegarden uses advisory warnings for application integrity, separate from AI match scoring, so recruiters can distinguish fit from review risk.
That distinction is important. A candidate can be a strong match and still need verification. A candidate can also have a warning that turns out to be harmless. The software should make both outcomes easy to handle.
Review applications with context
Treegarden helps recruiters manage high-volume pipelines with advisory AI, application integrity warnings, and human review built into the hiring workflow. Book a demo
Frequently Asked Questions
Can AI resume detectors prove a resume was written by AI?
No detector can prove that with enough certainty for a hiring decision. Detectors can provide a signal, but recruiters still need verification and human review.
Should recruiters reject AI-written resumes?
Not automatically. The better standard is whether the claims in the resume are true, relevant, and verifiable.
What should replace AI text detection?
Use multi-signal review: profile consistency, work sample evidence, structured interviews, references, and explainable ATS warnings.
Sources and further reading
- Liang, W. et al. "GPT detectors are biased against non-native English writers." Patterns, Cell Press, 2023 (the 91-essay, seven-detector study; 61 percent false-positive finding).
- Stanford HAI. "AI Detectors Biased Against Non-Native English Writers."
- OpenAI. "New AI classifier for indicating AI-written text" (updated to note the classifier was retired for low accuracy).
- ResumeBuilder.com. Survey: 46 percent of job seekers use ChatGPT to write resumes and cover letters.
- The Markup. "AI Detection Tools Falsely Accuse International Students of Cheating."
- NYC Department of Consumer and Worker Protection. Automated Employment Decision Tools (Local Law 144).
- EU Artificial Intelligence Act. Annex III: High-Risk AI Systems.