Use case · Assess Engineers for AI

Assess engineering candidates on AI skills before hiring

Traditional coding tests and HackerRank- or CodeSignal-style screens cannot assess these skills. Anthropos runs real builds with Claude Code or Codex enabled. Its AI-skills assessment measures what candidates brief, verify and reject.

“Every candidate says they use Claude Code. Our interview loop cannot tell me which ones can.”

Engineering hiring manager, 200-engineer software company

5-30
Minutes per AI Simulation
15-30
Minutes per AI Lab challenge
0-5
Competency level per skill
100+
ATS integrations

Three ways to test an engineer

What changes when the tools are allowed and the judgment is scored

Blocking the tools tests the wrong thing. Allowing them without measuring anything rewards whoever types most.

A coding test An AI-assisted coding test Anthropos
What it measures Solving a closed problem with AI off Speed to an answer with the tool on Judgment around AI-produced work
AI tools Blocked Allowed, unmeasured Set per task, stated before the candidate starts
Known failure Penalises engineers who use agents daily Rewards high token spend, penalises careful work Volume of tool use is not a criterion
The task A puzzle with a findable answer The same puzzle with an assistant A scenario on your stack, with no answer key
Scoring Correct or not An automated threshold Criteria set before the session, each cited to evidence
Who decides The threshold The threshold Your hiring manager, from the evidence

The session

Candidates build with AI tools

Candidates complete a 15 to 30 minute sandbox build with Claude Code or Codex enabled, or a coding simulation where an AI reviewer challenges their work. You see what they brief, verify and reject, and how they explain each decision.

See AI Simulations →

A coding simulation where a candidate works on a real task with an AI reviewer

A coding simulation in progress, with the candidate working against real pushback. Cervato Systems, our demo org.

Candidate experience

Senior engineers complete real builds

Engineers complete replacements and refuse added steps. Most customers replace take-homes or first-round technical screens. Invitations include rubrics and firm time caps, and candidates receive strengths-and-gaps reports. Seniors welcome real, familiar-tool builds over algorithm puzzles they abandoned a decade ago.

Screening for non-technical roles →

An AI Lab workspace with a real repository open and Claude Code running in the terminal

The candidate's sandbox: a real repository, a real brief, and Claude Code running in the terminal. Anthropos Workforce demo org.

Scoring

The rubric scores observable decisions

Prompt count, token spend and time to first answer are unscored, protecting careful engineers. Before each session, the rubric specifies problem framing, output judgement, verification and explanation quality. Every score cites the moment that earned it.

See AI Labs →

The AI Labs sandbox environment available to a candidate during an assessment

The sandbox a candidate works in, with the agent available. Cervato Systems, our demo org.

Anthropos verifies engineering capability before the hire and after it, on one taxonomy of 4,000+ skills and 700+ roles. Datrix trained its engineers on Claude Code with Anthropos AI Academy.

The decision

Hiring managers decide using evidence

AI generates the scenario and plays the characters. Scoring follows a deterministic rubric with criteria fixed before anyone starts. No threshold rejects a candidate automatically. Every candidate receives a report naming strengths and gaps, whether they advance or not.

Screen non-technical roles too →

A scored result showing criteria from the brief and from how the candidate worked with AI

Eleven criteria, seven from the brief and four from how they worked with the AI, each scored with the reason. Anthropos Workforce demo org.

Defensibility

Testing people is a regulated decision

Recruitment assessment is high-risk under the EU AI Act. We do not argue with the classification, we build for it.

The rubric exists before the candidate does

AI generates the scenario and plays the characters. The scoring is deterministic and rubric-based, with the criteria fixed before anyone sits it and applied identically for that role.

No threshold rejects anybody

A score ranks and sorts. Your hiring manager decides, and no automated step removes a candidate on its own, which is also what GDPR Article 22 requires.

Every score opens to the moment behind it

The report cites what the candidate did: what they briefed, what they verified, what they rejected, and which rubric criteria that met or missed. That record is what makes the decision defensible a year later.

No emotion recognition, no proctoring theatre

Anthropos runs no gaze tracking, flags nobody for looking away, and infers no psychological state. Candidates are told before they start that the session is AI-run and recorded.

Your scenario, not a shared test

Each customer authors its own challenges on its own stack and thresholds, so a candidate rejected by one employer is not rejected everywhere at once.

EU-processed, and never training data

Candidate data is processed in the EU and is never used to train, retrain or fine-tune any model, ours or a sub-processor's. That is contractual in the DPA. ISO 27001 and GDPR, with the EU AI Act assessment available to read.

Retention is configurable, and the assessment record stays with you.

Rollout

What it takes to start

Day 1

Pick a role simulation from the catalogue, or an AI Lab challenge with Claude Code and Codex.

Week 1

Set the tool policy per task, send the rubric with the invitation, and candidates complete asynchronously.

Week 3

Author your own scenario in Studio from your job description and repository, then delete the round it replaces.

Your IT team connects the ATS through an existing connector and enables SSO. Labs run in sandboxes Anthropos hosts, so there is nothing to install.

Objections

What buyers ask before a pilot

How long does an Anthropos assessment take, and is the time capped?

Anthropos AI Simulations run 5 to 30 minutes and AI Lab challenges run 15 to 30 minutes, in one sitting, from a link the candidate opens when they choose. The cap is real: criteria are calibrated to the stated duration, so nobody is marked down for corner cases the time never allowed.

Which AI tools do candidates use in an Anthropos assessment?

Anthropos AI Labs run Claude Code and Codex in a sandbox Anthropos hosts, and you choose which models are available inside each Lab. If your engineers use a different assistant in their editors day to day, the capability transfers, because the assessment is about briefing an agent, checking its output and correcting it.

Can we build the assessment on our own codebase and stack?

Yes. Anthropos Studio builds a simulation or an AI Lab challenge from your job description, tech stack, repository and evaluation criteria, with no engineers needed to author it. Candidates recognise the work, and a scenario nobody else runs has no findable answer for a candidate looking for one.

Does every candidate get feedback?

Yes. Every candidate who completes an Anthropos assessment receives a report naming what they did well and where they fell short, whether or not they advance. Because Anthropos scores what the candidate produced against criteria published in advance, the feedback is specific to their session rather than a form rejection.

Where is candidate data stored, and is it used to train models?

Anthropos processes candidate data in the EU and never uses it to train, retrain or fine-tune any AI model, its own or a sub-processor's. That is a contractual commitment in the Anthropos Data Processing Agreement. Retention is configured by your team, and candidate deletion requests are supported under GDPR.

Is AI-scored hiring legal in the EU?

Recruitment is a high-risk use under the EU AI Act, which requires transparency, human oversight and documentation. In Anthropos, AI generates scenarios and plays characters, evaluation is deterministic rubric-based scoring, human reviewers hold decision authority, and no emotion inference is performed. Candidate data is processed in the EU.

What if a candidate has never used an agentic coding tool before?

Every Anthropos session opens with a plain-language brief: the scenario, the tools available and what is being assessed. Nobody is scored on interface familiarity. If a role genuinely requires prior agentic experience, set that expectation in the job description rather than letting the assessment format decide it silently.

What happens to the assessment after we hire the candidate?

The verified skill profile carries into Anthropos Workforce, so a new engineer starts with a 0-5 competency level per skill instead of a blank record. Their manager sees the gaps the assessment already named, and development can be assigned in the first week rather than after a review cycle.

Try it on a role you are hiring for now

Bring the job description and your stack. We build the scenario, you take it yourself before a candidate ever does.