AI cheating in technical interviews

The AI-resistant coding assessment is a mirage

German Reyes
German Reyes·Jul 20, 2026·5 min read
On this page

Search "AI-resistant coding assessment" right now and you'll find a whole product category selling you the same promise: a test that AI can't solve, so you can go back to trusting the score. Codility announced an AI-resistant task library. CodeSignal shipped agentic assessments. Even Anthropic, whose whole job is building the models doing the "cheating," runs a take-home for its own performance engineers. I run an assessment company, so I watch this category closely, and I think the promise is a mirage. Not because the vendors are lazy, but because "resistant" is a state that expires every time a new model ships. If you're shopping for one this quarter, here's the case for buying something else.

The AI-proof coding test has an expiration date

An AI-resistant assessment rests on one assumption: that you can design a problem hard enough that current models fail it, and that the problem stays hard. The first half is doable. The second half is where it falls apart.

Here's the tell I keep coming back to. Anthropic's performance engineering team has written publicly about their take-home test, where candidates optimize code for a simulated accelerator. Over a thousand people have taken it. And they've had to redesign it after each Claude release, because the newer model outran the old test. Their own account describes Claude Opus 4 outperforming most human applicants on it.

Sit with that. A company that builds frontier models, with every incentive and all the internal knowledge to keep one test ahead, still has to rebuild it every few months. Now picture a vendor selling a static question bank to a few hundred customers. What are the odds their library stays ahead of the next release? The resistance you're buying is real for a while and gone by the next model.

Proctoring can't reliably prevent AI cheating

The other half of the AI-resistant pitch is surveillance. Paste detection, keystroke analysis, webcam proctoring, perplexity scores that claim to know whether a human wrote the code. This is where the category stops being merely futile and starts being harmful.

Paste detection dies the second a candidate retypes instead of pasting. Keystroke analysis flags fast typists and neurodivergent people as suspects. Proctoring quietly pushes your strongest candidates out of the funnel, because anyone with options doesn't want to be filmed through a webcam for an hour to prove they didn't cheat. So you end up with a screen that accuses people you want and waves through the motivated ones who bothered to route around it. That isn't a screen so much as a tax on the candidates you can least afford to lose.

Allow the tool and watch how they use it

Quick disclosure before I go further. I'm building Skillvee, a 60-minute "day at work" simulation that replaces the recruiter phone screen and technical first round. Candidates solve a real challenge, talk to AI peers, and defend their decisions to an AI manager while the screen records. So I have a stake in you rejecting the AI-resistant frame. The argument stands on its own, but you should know the bias is there.

So flip the question. Instead of "can this test keep AI out," ask "how well does this person work with AI." Once AI is allowed and the session is recorded, the arms race just ends. There's nothing to detect, because nothing is forbidden. And what you get instead is the signal that actually predicts the job now: judgment about the machine.

Watching real sessions taught me something I didn't expect. AI doesn't compress the range between candidates. It widens it. Give everyone the same model and the gap between your best and worst applicants gets bigger. A weak candidate takes the model's first answer and ships it. A strong one treats the model like a fast, confident junior who lies sometimes: they prompt with precision, catch the plausible-but-wrong suggestion, and can tell you afterward why the final version looks the way it does. A detector can't see any of that. Worse, it reads the strong candidate as more suspicious, because they used the tool more.

AI-resistant vs AI-leverage, side by side

The two approaches point in opposite directions. One spends its effort keeping AI out. The other spends it watching how AI gets used.

AI-resistant assessmentAI-leverage assessment
Core assumptionYou can design tests AI can't solveAI is part of the job, so make it part of the test
What it measuresWhether the candidate avoided a banned toolHow well the candidate directs the tool
Shelf lifeResets every model releaseGets more predictive as models improve
Effect on candidatesSurveillance pushes strong people outRealistic task, no accusation
Failure modeFalse positives on your best applicantsNeeds a rubric and someone to watch the work

Neither column is free. The right one is just pointed at the durable question instead of the expiring one. I wrote the fuller version of this argument in why "detecting AI cheating" is the wrong question, and the take-home postmortem covers the same failure from the format side.

How to vet an "AI-resistant" claim

If you're still evaluating vendors in this category, don't take the label at face value. Three questions separate real work from marketing:

  1. How often do you refresh the question bank, and what triggers it? If the answer isn't "every major model release, automatically," the resistance is already decaying.
  2. What's your false-positive rate on AI detection, and how do you measure it? A vendor who can't answer this is shipping accusations they haven't validated against your best candidates.
  3. What happens when a candidate retypes AI output instead of pasting? If the whole detection story collapses on that one move, you know how deep the moat is.

The answers usually reveal that you're buying the appearance of resistance, priced like the real thing.

Where this gets hard

I won't pretend the alternative is effortless. Measuring AI leverage means designing a realistic task with enough ambiguity that judgment actually shows up, and a lazy simulation is just a take-home with extra steps. You need a written rubric too, because one manager eyeballing a recording will drift toward vibes. And someone on your side has to actually watch the work or read the report, or you've built signal nobody consumes.

That's a real build-vs-buy decision, and you should make it with clear eyes about your team's bandwidth. What it isn't is a losing race. You're measuring something that stays true across model releases: how this person works when the tool everyone now uses is sitting right there. That's the bet behind how Skillvee scores a session, and it's why our pricing charges by volume instead of per proctored seat. Resistance taxes the wrong thing. Watching people work is the point.

Frequently asked questions

Is there such a thing as an AI-resistant coding assessment?
Only temporarily. A test that stumps today's models can be redesigned to stump them, but the next release usually solves it, so the resistance has a short shelf life. Anthropic has written publicly about redesigning its own take-home test after each Claude release because the newer model outperformed the old test. If a company that builds frontier models can't keep a test ahead of its own releases, a vendor selling you a static one can't either.
How do I make a coding test AI-proof?
You mostly can't, and chasing it costs you strong candidates through proctoring and false accusations. The durable move is the opposite: allow AI, record the session, and score how the candidate directs the tool. What they ask, which output they trust, what they correct, and how they explain it. That signal gets more predictive as models improve, not less.
If I let candidates use AI, how do I still test coding ability?
You still see the code they ship and the reasoning behind it. Watching someone work with AI shows their technical judgment more sharply than a clean-room test, because you see which suggestions they reject and why. Code quality is one dimension to score, alongside communication, collaboration, agency, and AI leverage.
Do AI-resistant assessment vendors work at all?
Their detection features catch the lazy cases and miss the motivated ones, which is the wrong trade for a hiring screen. A vendor investing in harder problems is doing real work, but you're buying a promise with an expiration date tied to the next model release. Ask how often they redesign their question bank and what happens to your data between redesigns.
What should I ask a vendor that claims its test is AI-resistant?
Ask three things. How often do you refresh the question bank, and what triggers a refresh? What's your false-positive rate on AI detection, and how is it measured? And what happens when a candidate simply retypes AI output instead of pasting it? The answers tell you whether you're buying resistance or the appearance of it.
AI-Resistant Coding Assessment: A Buyer's Reality Check