Use case · Assess Engineers for AI
Assess engineering candidates on AI skills before hiring
Traditional coding tests and HackerRank- or CodeSignal-style screens cannot assess these skills. Anthropos runs real builds with Claude Code or Codex enabled. Its AI-skills assessment measures what candidates brief, verify and reject.
“Every candidate says they use Claude Code. Our interview loop cannot tell me which ones can.”
Engineering hiring manager, 200-engineer software company
Three ways to test an engineer
What changes when the tools are allowed and the judgment is scored
Blocking the tools tests the wrong thing. Allowing them without measuring anything rewards whoever types most.
| A coding test | An AI-assisted coding test | Anthropos | |
|---|---|---|---|
| What it measures | Solving a closed problem with AI off | Speed to an answer with the tool on | Judgment around AI-produced work |
| AI tools | Blocked | Allowed, unmeasured | Set per task, stated before the candidate starts |
| Known failure | Penalises engineers who use agents daily | Rewards high token spend, penalises careful work | Volume of tool use is not a criterion |
| The task | A puzzle with a findable answer | The same puzzle with an assistant | A scenario on your stack, with no answer key |
| Scoring | Correct or not | An automated threshold | Criteria set before the session, each cited to evidence |
| Who decides | The threshold | The threshold | Your hiring manager, from the evidence |
What the verified profile is for
One assessment, four decisions
The profile you build at the gate keeps working after the offer is signed.
The session
Candidates build with AI tools
Candidates complete a 15 to 30 minute sandbox build with Claude Code or Codex enabled, or a coding simulation where an AI reviewer challenges their work. You see what they brief, verify and reject, and how they explain each decision.
A coding simulation in progress, with the candidate working against real pushback. Cervato Systems, our demo org.
Candidate experience
Senior engineers complete real builds
Engineers complete replacements and refuse added steps. Most customers replace take-homes or first-round technical screens. Invitations include rubrics and firm time caps, and candidates receive strengths-and-gaps reports. Seniors welcome real, familiar-tool builds over algorithm puzzles they abandoned a decade ago.
The candidate's sandbox: a real repository, a real brief, and Claude Code running in the terminal. Anthropos Workforce demo org.
Scoring
The rubric scores observable decisions
Prompt count, token spend and time to first answer are unscored, protecting careful engineers. Before each session, the rubric specifies problem framing, output judgement, verification and explanation quality. Every score cites the moment that earned it.
The sandbox a candidate works in, with the agent available. Cervato Systems, our demo org.
Anthropos verifies engineering capability before the hire and after it, on one taxonomy of 4,000+ skills and 700+ roles. Datrix trained its engineers on Claude Code with Anthropos AI Academy.





The decision
Hiring managers decide using evidence
AI generates the scenario and plays the characters. Scoring follows a deterministic rubric with criteria fixed before anyone starts. No threshold rejects a candidate automatically. Every candidate receives a report naming strengths and gaps, whether they advance or not.
Eleven criteria, seven from the brief and four from how they worked with the AI, each scored with the reason. Anthropos Workforce demo org.
Defensibility
Testing people is a regulated decision
Recruitment assessment is high-risk under the EU AI Act. We do not argue with the classification, we build for it.
AI generates the scenario and plays the characters. The scoring is deterministic and rubric-based, with the criteria fixed before anyone sits it and applied identically for that role.
A score ranks and sorts. Your hiring manager decides, and no automated step removes a candidate on its own, which is also what GDPR Article 22 requires.
The report cites what the candidate did: what they briefed, what they verified, what they rejected, and which rubric criteria that met or missed. That record is what makes the decision defensible a year later.
Anthropos runs no gaze tracking, flags nobody for looking away, and infers no psychological state. Candidates are told before they start that the session is AI-run and recorded.
Each customer authors its own challenges on its own stack and thresholds, so a candidate rejected by one employer is not rejected everywhere at once.
Candidate data is processed in the EU and is never used to train, retrain or fine-tune any model, ours or a sub-processor's. That is contractual in the DPA. ISO 27001 and GDPR, with the EU AI Act assessment available to read.
Retention is configurable, and the assessment record stays with you.
Rollout
What it takes to start
Pick a role simulation from the catalogue, or an AI Lab challenge with Claude Code and Codex.
Set the tool policy per task, send the rubric with the invitation, and candidates complete asynchronously.
Author your own scenario in Studio from your job description and repository, then delete the round it replaces.
Your IT team connects the ATS through an existing connector and enables SSO. Labs run in sandboxes Anthropos hosts, so there is nothing to install.
Objections
What buyers ask before a pilot
How long does an Anthropos assessment take, and is the time capped?
Anthropos AI Simulations run 5 to 30 minutes and AI Lab challenges run 15 to 30 minutes, in one sitting, from a link the candidate opens when they choose. The cap is real: criteria are calibrated to the stated duration, so nobody is marked down for corner cases the time never allowed.
Which AI tools do candidates use in an Anthropos assessment?
Anthropos AI Labs run Claude Code and Codex in a sandbox Anthropos hosts, and you choose which models are available inside each Lab. If your engineers use a different assistant in their editors day to day, the capability transfers, because the assessment is about briefing an agent, checking its output and correcting it.
Can we build the assessment on our own codebase and stack?
Yes. Anthropos Studio builds a simulation or an AI Lab challenge from your job description, tech stack, repository and evaluation criteria, with no engineers needed to author it. Candidates recognise the work, and a scenario nobody else runs has no findable answer for a candidate looking for one.
Does every candidate get feedback?
Yes. Every candidate who completes an Anthropos assessment receives a report naming what they did well and where they fell short, whether or not they advance. Because Anthropos scores what the candidate produced against criteria published in advance, the feedback is specific to their session rather than a form rejection.
Where is candidate data stored, and is it used to train models?
Anthropos processes candidate data in the EU and never uses it to train, retrain or fine-tune any AI model, its own or a sub-processor's. That is a contractual commitment in the Anthropos Data Processing Agreement. Retention is configured by your team, and candidate deletion requests are supported under GDPR.
Is AI-scored hiring legal in the EU?
Recruitment is a high-risk use under the EU AI Act, which requires transparency, human oversight and documentation. In Anthropos, AI generates scenarios and plays characters, evaluation is deterministic rubric-based scoring, human reviewers hold decision authority, and no emotion inference is performed. Candidate data is processed in the EU.
What if a candidate has never used an agentic coding tool before?
Every Anthropos session opens with a plain-language brief: the scenario, the tools available and what is being assessed. Nobody is scored on interface familiarity. If a role genuinely requires prior agentic experience, set that expectation in the job description rather than letting the assessment format decide it silently.
What happens to the assessment after we hire the candidate?
The verified skill profile carries into Anthropos Workforce, so a new engineer starts with a 0-5 competency level per skill instead of a blank record. Their manager sees the gaps the assessment already named, and development can be assigned in the first week rather than after a review cycle.
Keep reading
Where to go next
Try it on a role you are hiring for now
Bring the job description and your stack. We build the scenario, you take it yourself before a candidate ever does.