Every guide on how to spot AI in interview answers hands you the same checklist. A three-to-five second pause before each reply. Eyes sliding toward a second screen. Answers that read like documentation instead of a person talking. Sudden fluency on one question and not the next. I have read most of those guides, and I ran the checklist myself for a stretch. Then I looked at who it was flagging: nervous people, candidates interviewing in their second language, and the ones who think before they speak. If you are a hiring manager trying to work out who is reading off a model, this is my case for dropping the detection frame, and what I would do instead.
The tells that supposedly detect AI in job interviews are confounded
Take the pause. A candidate goes quiet for four seconds before answering. That is the signature tell in nearly every detection guide, and it is also what a careful person does with a question they have not rehearsed. It is what an anxious person does. It is what someone translating in their head does.
The rest hold up no better. Eyes moving off camera catch anyone with a second monitor, a page of notes, or a habit of looking away while thinking. Polished phrasing catches people who prepared. A jump in fluency between questions catches the very ordinary fact that people know some parts of their own work better than others.
These are not useless observations. They are just not evidence on their own, because for every one of them the innocent explanation is more common in your candidate pool than the guilty one. What feels like a signal is closer to a coin flip with a story attached to it.
One candidate I sat in on hit four items of the checklist at once. Long pause before every answer, eyes off camera while thinking, careful and slightly formal phrasing, and much better fluency on system design than on anything to do with their last team. They were also interviewing late in the evening their time, in their third language, about work they had finished three years earlier. The pauses were someone assembling a sentence, not reading one. On the checklist they looked like the most suspicious person in the pipeline.
The base-rate problem with any AI-cheating tell
Here is the part the checklists skip. Even a genuinely good tell produces mostly false accusations when the thing it detects is rare.
Make up some numbers and watch the shape. These are assumptions chosen to illustrate the arithmetic, not measurements of anything. You interview 100 people and 10 of them are reading off a model. Your tell is good: it catches 8 of those 10, and it misfires on only 1 in 10 honest candidates. That misfire rate sounds tolerable right up until you count. Nine honest candidates get flagged, against eight real ones. More than half of everyone you suspect did nothing wrong.
That is the flattering version, built on a tell far sharper than "they paused." Move the misfire rate to 1 in 5 and you are accusing 18 honest people to catch 8. The arithmetic does not care how confident you felt in the moment.
So detection scales badly in a specific way. The rarer the behavior, the worse your hit rate, and the more your process turns into a machine for quietly insulting good candidates.
What a false accusation actually costs
An interview is not a courtroom, so nobody reads out a verdict. The accusation shows up as a quieter second half of the call, then a "we went with someone else," and a candidate who never finds out why.
The cost lands in two places. The obvious one is the person you lose. The one you do not see is what suspicion does to the interview itself. Once I am watching someone for tells, I am not really listening to their reasoning. I have run interviews in that mode, and my notes afterwards were about where the candidate looked rather than how they had scoped the problem. Neither of us was doing our job.
The longer version of this argument is in why detecting AI cheating is the wrong question, and the same failure shows up on the tooling side in the AI-resistant coding assessment.
Ask questions that ChatGPT answers badly in technical interviews
Say you cannot change your process this quarter. You still have people to interview next week. There is a version of this that works, and it comes down to asking things a model cannot answer on someone's behalf.
A model will happily produce a competent answer about how to structure a service. It cannot tell you what your candidate's team shipped in March, what broke, who argued against it, or what they would do differently now. So go there:
- Ask for a specific instance rather than a practice. "Tell me about a migration you ran" gets you further than "how do you approach migrations." Real specifics have texture that generic competence does not.
- Interrupt. Take the answer and push on its weakest joint. Someone reading off a screen loses the thread. Someone who lived it gets more precise.
- Ask what they got wrong. Real projects have regrets attached to them. Borrowed answers rarely do.
- Follow the consequence. "What did that decision look like six months later?" Models write plans well. People remember aftermaths.
None of that is a lie detector. It is ordinary good interviewing, and it produces better signal whether or not anybody is cheating. That property is the whole point. A detection checklist only pays off in the rare case where someone is reading off a model. Better questions pay off in every interview you run.
Allow the tool and score the AI leverage instead
Quick disclosure. I am building Skillvee, a 60-minute "day at work" simulation that replaces the recruiter phone screen and the technical first round. Candidates solve a real problem, talk to AI peers, and defend their decisions to an AI manager while the screen records. So I have an obvious stake in you giving up on detection. Take the argument on its merits, with the bias in view.
The structural fix is to remove the thing you are trying to detect. If the candidate is allowed to use a model and you can see the session, there is nothing to catch. The question turns into "how well did they use it," which is the one that predicts the job now.
What you watch for is narrow and concrete: what they ask, how they check the answer, which suggestions they throw out, and whether they can explain the finished version with the model closed. The candidates who struggle treat the first output as finished. The ones who are good at this argue with it, and you can hear them doing it. That gap is wide and it needs no inference about eye movement at all. It is the signal our product scores from a recorded session.
Where this gets hard
The friction is real, so here it is. Allowing AI means most of your existing question bank stops working, because a good chunk of it is now one prompt away from being solved. Rewriting an interview loop is genuine work and it competes with everything else on your team's quarter.
It also demands a written rubric. "Score their AI leverage" collapses into vibes unless somebody defines what good looks like before the interviews start. And an interviewer has to actually watch the work rather than skim a summary, which costs more attention per candidate than a pass or fail score does.
That is a fair trade to weigh, and it is why we price by volume instead of per proctored seat. What you get for the effort is a process that does not degrade with the next model release, and does not spend its accuracy budget on accusing the people you were trying to hire.
If you have already made that call, the follow-up question is what you actually write on the scorecard. I worked that through separately in the six behaviors to score once you allow AI in coding interviews.