Search "AI-resistant coding assessment" right now and you'll find a whole product category selling you the same promise: a test that AI can't solve, so you can go back to trusting the score. Codility announced an AI-resistant task library. CodeSignal shipped agentic assessments. Even Anthropic, whose whole job is building the models doing the "cheating," runs a take-home for its own performance engineers. I run an assessment company, so I watch this category closely, and I think the promise is a mirage. Not because the vendors are lazy, but because "resistant" is a state that expires every time a new model ships. If you're shopping for one this quarter, here's the case for buying something else.
The AI-proof coding test has an expiration date
An AI-resistant assessment rests on one assumption: that you can design a problem hard enough that current models fail it, and that the problem stays hard. The first half is doable. The second half is where it falls apart.
Here's the tell I keep coming back to. Anthropic's performance engineering team has written publicly about their take-home test, where candidates optimize code for a simulated accelerator. Over a thousand people have taken it. And they've had to redesign it after each Claude release, because the newer model outran the old test. Their own account describes Claude Opus 4 outperforming most human applicants on it.
Sit with that. A company that builds frontier models, with every incentive and all the internal knowledge to keep one test ahead, still has to rebuild it every few months. Now picture a vendor selling a static question bank to a few hundred customers. What are the odds their library stays ahead of the next release? The resistance you're buying is real for a while and gone by the next model.
Proctoring can't reliably prevent AI cheating
The other half of the AI-resistant pitch is surveillance. Paste detection, keystroke analysis, webcam proctoring, perplexity scores that claim to know whether a human wrote the code. This is where the category stops being merely futile and starts being harmful.
Paste detection dies the second a candidate retypes instead of pasting. Keystroke analysis flags fast typists and neurodivergent people as suspects. Proctoring quietly pushes your strongest candidates out of the funnel, because anyone with options doesn't want to be filmed through a webcam for an hour to prove they didn't cheat. So you end up with a screen that accuses people you want and waves through the motivated ones who bothered to route around it. That isn't a screen so much as a tax on the candidates you can least afford to lose.
Allow the tool and watch how they use it
Quick disclosure before I go further. I'm building Skillvee, a 60-minute "day at work" simulation that replaces the recruiter phone screen and technical first round. Candidates solve a real challenge, talk to AI peers, and defend their decisions to an AI manager while the screen records. So I have a stake in you rejecting the AI-resistant frame. The argument stands on its own, but you should know the bias is there.
So flip the question. Instead of "can this test keep AI out," ask "how well does this person work with AI." Once AI is allowed and the session is recorded, the arms race just ends. There's nothing to detect, because nothing is forbidden. And what you get instead is the signal that actually predicts the job now: judgment about the machine.
Watching real sessions taught me something I didn't expect. AI doesn't compress the range between candidates. It widens it. Give everyone the same model and the gap between your best and worst applicants gets bigger. A weak candidate takes the model's first answer and ships it. A strong one treats the model like a fast, confident junior who lies sometimes: they prompt with precision, catch the plausible-but-wrong suggestion, and can tell you afterward why the final version looks the way it does. A detector can't see any of that. Worse, it reads the strong candidate as more suspicious, because they used the tool more.
AI-resistant vs AI-leverage, side by side
The two approaches point in opposite directions. One spends its effort keeping AI out. The other spends it watching how AI gets used.
| AI-resistant assessment | AI-leverage assessment | |
|---|---|---|
| Core assumption | You can design tests AI can't solve | AI is part of the job, so make it part of the test |
| What it measures | Whether the candidate avoided a banned tool | How well the candidate directs the tool |
| Shelf life | Resets every model release | Gets more predictive as models improve |
| Effect on candidates | Surveillance pushes strong people out | Realistic task, no accusation |
| Failure mode | False positives on your best applicants | Needs a rubric and someone to watch the work |
Neither column is free. The right one is just pointed at the durable question instead of the expiring one. I wrote the fuller version of this argument in why "detecting AI cheating" is the wrong question, and the take-home postmortem covers the same failure from the format side.
How to vet an "AI-resistant" claim
If you're still evaluating vendors in this category, don't take the label at face value. Three questions separate real work from marketing:
- How often do you refresh the question bank, and what triggers it? If the answer isn't "every major model release, automatically," the resistance is already decaying.
- What's your false-positive rate on AI detection, and how do you measure it? A vendor who can't answer this is shipping accusations they haven't validated against your best candidates.
- What happens when a candidate retypes AI output instead of pasting? If the whole detection story collapses on that one move, you know how deep the moat is.
The answers usually reveal that you're buying the appearance of resistance, priced like the real thing.
Where this gets hard
I won't pretend the alternative is effortless. Measuring AI leverage means designing a realistic task with enough ambiguity that judgment actually shows up, and a lazy simulation is just a take-home with extra steps. You need a written rubric too, because one manager eyeballing a recording will drift toward vibes. And someone on your side has to actually watch the work or read the report, or you've built signal nobody consumes.
That's a real build-vs-buy decision, and you should make it with clear eyes about your team's bandwidth. What it isn't is a losing race. You're measuring something that stays true across model releases: how this person works when the tool everyone now uses is sitting right there. That's the bet behind how Skillvee scores a session, and it's why our pricing charges by volume instead of per proctored seat. Resistance taxes the wrong thing. Watching people work is the point.