AI LabsTechnical assessment

Give them a real sandbox. See exactly how they build with AI.

A Lab is a disposable sandbox. Give it to a candidate, or to one of your own engineers, with a real task: build an application, fix a failing build, configure a server. They work in it with Claude Code or Codex. You get back every command they ran, and a 0 to 5 score for how well they used AI.

An AI Lab sandbox: a live coding environment with Claude Code running inside an isolated Anthropos workspace

~20sec
To boot a disposable sandbox
15-30
Minutes per challenge
0-5
Evidence-backed score
4
Agents: Claude Code, Copilot, Codex, Gemini
Use case 1 · Assessment

Assess candidates and employees on the work itself.

Set a task in a real environment, send a link, and let the person work the way they would on any normal Tuesday. Any language, any stack. Nothing is multiple choice, and nothing is a mock-up of a real environment.

Screening candidates

A candidate opens a link, works for 15 to 30 minutes and finishes. You get the recording and the score, without booking anyone’s calendar.

Checking your own engineers

Run the same challenge across a team you already employ, so you know who is ready for AI work today and who needs training first.

Any task, any stack

Write code. Fix a broken build. Configure a server. Query a live database. If a developer can do it at a terminal, they can do it in the sandbox.

What is different

You see everything they did with AI, and how well they did it.

Most technical assessment either bans AI or pretends it is not there. A Lab does neither. The agent is part of the exercise. The whole session is recorded. And the score answers what a code sample never can: does this person get good work out of AI?

The whole session, not the final file

Every prompt, every command, every tool call, in the order they happened. You can replay it and watch the work take shape.

Scored on how they steered it

How they framed the problem, what they checked, and when they rejected an answer from the model. Anyone can generate code; the score is about whether they got it right.

You can check every score

Every rating links to the moments in the session that produced it. A hiring manager can check the number instead of trusting it.

Use case 2 · Practice

Let people learn AI tools on real work, with nothing at risk.

Teams avoid practising with coding agents for three reasons: the repository, the keys and the bill. A Lab removes all three. The sandbox never touches your systems, and it is wiped when the session ends. People can try things they would never risk in production. A disposable, isolated sandbox. It boots in about 20 seconds, and people build and test live with the AI tools you approve. Download the work, wipe the sandbox.

Playground

An open sandbox. People pick up Claude Code or Codex and build whatever they want: try an idea, break something, find out what the tools can really do.

Template

A sandbox prepared for one task, such as shipping a small application. People build it, run it, test it and download what they made. Then it is wiped.

You’re in control

You choose the power, the models and the spending limit.

You decide how powerful each Lab is, which models run inside it, and how much the whole program can cost. IT sets the rules once, and every session inherits them.

As powerful as you need

Size each Lab to the work: allocate more compute and memory when people run heavier builds or bigger workloads.

The models you choose

Allowlist the models and agents available inside each Lab, so your people work only with the tools you have approved.

A ceiling you set

Cap Labs by credits or by minutes. Give a team $1,000 of Lab time this quarter, and that is the limit.
Claude CodeGitHub CopilotCodex ClaudeGPTGeminiMistral

In the platform

The score joins everything else you know about that person.

AI Labs is part of Anthropos Workforce rather than a separate product. Every session updates the same skills layer that carries the readiness score, the role benchmarks and the development plans for that person.

Feeds AI Readiness

Every session updates the person’s verified AI level in the 0-100 AI Readiness Score, next to simulations and interviews.

Rolls into Workforce Intelligence

Heat maps and gap views by role, team and organization show where AI skill is growing, and where it is not.

Pairs with AI Academy

Each Lab points people to the Academy paths that close their gaps, then verifies the progress with a new challenge in the same area.

AI Academy trains, AI Labs verify on real tools, AI Readiness measures the return.

Who it is for

Built for recruiters, engineering managers and IT.

A Lab produces one thing: the recording and the score. Each of these roles takes something different out of it.

Tech recruiters

Screen for real ability before an engineer spends an hour interviewing. You hand the hiring manager a recording and a score instead of an opinion.

Engineering managers

Find out which of your engineers can genuinely ship with AI, which cannot yet, and which gap to close for each of them.

IT and security

One governed place to use coding agents, with the model list, the compute and the spending limit all set by you.

Enterprise-ready

Isolated, governed, capped.

Isolated per session Your model allowlist Cost caps by credits or minutes GDPR & EU AI Act aligned

Trusted by leading enterprises FS GroupItalgas Together AIDatrix FidesOrbyta
Common questions

What recruiters and engineering managers ask.

Can the candidate just let the AI do the work?
They can, and that is exactly what gets measured. The record shows every prompt and every command. One person takes whatever the model hands them. Another frames the problem, checks the answer and fixes what is wrong. Those two sessions look nothing alike, and the score says so.
What can someone actually build in a Lab?
Anything you would do at a real terminal: an application in any language, a fix on a real repository, a server configuration, a query against a live database. It is a full environment, not an editor in a browser.
Which AI tools do they work with?
Claude Code, GitHub Copilot, Codex and Gemini today, on top of the models you approve: Claude, GPT, Gemini, Mistral and more. The allowlist is yours.
How long does a session take?
Fifteen to thirty minutes for a typical challenge, and the sandbox boots in about 20 seconds. Nothing to install, on your side or the candidate’s.
Is our code or data at risk?
No. Labs are isolated sandboxes fully managed by Anthropos. Nothing runs on your systems, and every sandbox is wiped after the session.
How is a session scored?
By the same engine that scores AI Simulations. The full session is the evidence. The 0 to 5 score is about how the person directed the AI, not just what came out at the end.
Can we cap the cost?
Yes. Every Lab is metered, and you cap spend by credits or minutes, per Lab, per team or per program.
Do Labs work for non-technical roles?
Labs are built for hands-on technical work. For business roles, AI Simulations and the AI Academy cover assessment and training, in the same platform.

Get started

Watch a scored session, end to end.

Book a demo and we will run a real challenge, then read the recording and the score the way a hiring manager would.