Give them a real sandbox. See exactly how they build with AI.
A Lab is a disposable sandbox. Give it to a candidate, or to one of your own engineers, with a real task: build an application, fix a failing build, configure a server. They work in it with Claude Code or Codex. You get back every command they ran, and a 0 to 5 score for how well they used AI.
Assess candidates and employees on the work itself.
Set a task in a real environment, send a link, and let the person work the way they would on any normal Tuesday. Any language, any stack. Nothing is multiple choice, and nothing is a mock-up of a real environment.
Screening candidates
Checking your own engineers
Any task, any stack
You see everything they did with AI, and how well they did it.
Most technical assessment either bans AI or pretends it is not there. A Lab does neither. The agent is part of the exercise. The whole session is recorded. And the score answers what a code sample never can: does this person get good work out of AI?
The whole session, not the final file
Scored on how they steered it
You can check every score
Let people learn AI tools on real work, with nothing at risk.
Teams avoid practising with coding agents for three reasons: the repository, the keys and the bill. A Lab removes all three. The sandbox never touches your systems, and it is wiped when the session ends. People can try things they would never risk in production. A disposable, isolated sandbox. It boots in about 20 seconds, and people build and test live with the AI tools you approve. Download the work, wipe the sandbox.
Playground
Template
You choose the power, the models and the spending limit.
You decide how powerful each Lab is, which models run inside it, and how much the whole program can cost. IT sets the rules once, and every session inherits them.
As powerful as you need
The models you choose
A ceiling you set
The score joins everything else you know about that person.
AI Labs is part of Anthropos Workforce rather than a separate product. Every session updates the same skills layer that carries the readiness score, the role benchmarks and the development plans for that person.
Feeds AI Readiness
Every session updates the person’s verified AI level in the 0-100 AI Readiness Score, next to simulations and interviews.
Rolls into Workforce Intelligence
Heat maps and gap views by role, team and organization show where AI skill is growing, and where it is not.
Pairs with AI Academy
Each Lab points people to the Academy paths that close their gaps, then verifies the progress with a new challenge in the same area.
AI Academy trains, AI Labs verify on real tools, AI Readiness measures the return.
Built for recruiters, engineering managers and IT.
A Lab produces one thing: the recording and the score. Each of these roles takes something different out of it.
Tech recruiters
Engineering managers
IT and security
Isolated, governed, capped.


What recruiters and engineering managers ask.
Can the candidate just let the AI do the work?
What can someone actually build in a Lab?
Which AI tools do they work with?
How long does a session take?
Is our code or data at risk?
How is a session scored?
Can we cap the cost?
Do Labs work for non-technical roles?
Watch a scored session, end to end.
Book a demo and we will run a real challenge, then read the recording and the score the way a hiring manager would.