Use case · AI Skills for Engineers
Assess and develop your engineering team’s AI capability
Anthropos distinguishes AI awareness from proficiency. Scored Lab challenges verify what engineers can do with Claude Code and Codex. Anthropos then builds an AI upskilling plan for the gaps it finds.

“Every engineer has Copilot. I have no idea if it's making us faster.”
VP Engineering, B2B software company, 400 engineers
Coding test vs Anthropos
What changes when the AI tools stay switched on
A test that blocks the tools measures work your engineers stopped doing years ago.
| A coding test | Anthropos | |
|---|---|---|
| What it measures | Whether the answer is correct | How the engineer briefs, reviews and corrects AI output |
| AI tools | Blocked, or an assistant inside the test | Claude Code and Codex, in a sandbox Anthropos hosts |
| The task | A closed problem with a known answer | A real build, run live and tested before it counts |
| Who takes it | Candidates, once, at the gate | The engineers you already employ, session after session |
| The output | A pass mark | 0-5 per skill, feeding a 0-100 AI Readiness score |
| Integrity | Proctoring against a findable answer | Recording, plus a format with no answer key |
What it is for
Six things you can decide once you know
Every engineer has the licence. These are the calls you cannot make from a seat count.
Assessment
Engineers build in a secure sandbox
Engineers launch a secure sandbox and build with Claude Code or Codex. Anthropos scores their problem framing, requests, checks and discarded work inside the session. Each skill receives a 0–5 score instead of a pass mark.
An engineer mid-Lab: the repository, the brief in the inbox, and the agent working in the terminal. Anthropos Workforce demo org.
Practice
Playground Labs let engineers practise unscored
Playground Labs use the identical sandbox with scoring off. Engineers use Claude Code or Codex for anything, for as long as they want. Nothing is recorded or reported. They practise before measurement and arrive with experience using the tools.
What the engineer gets back: the score, and each AI skill rated out of five with what went well and what to work on next.
Datrix trained its engineers on Claude Code with Anthropos AI Academy. Together AI, Orbyta Tech, Italgas and FS Group run on the same platform.





The durable skills
Labs score the Eight Durable Skills
Automated coding shifts seniority towards the Eight Durable Skills: Problem Framing, Business Judgment, Taste, Output Judgment, Systems Thinking, Spec Writing, Orchestration and Learning Velocity. Labs score these skills. Your CTO, engineers and HR team share vocabulary without translating three ladders.
How the work is scored: what the brief asked for, and separately how they worked with the AI. Anthropos Workforce demo org.
The ladder
Scores support defensible promotions and external benchmarks
Scores map to a 0–5 competency scale: 1 needs supervision, 3 delivers consistently and trains others, 5 sets the standard. Anthropos customers share the scale and rubric structure, making levels comparable. Repeat challenges next quarter for individual and team capability trends beyond licence counts.
Lab activity across the engineering population. Anthropos Workforce demo org.
Code and sandboxes
Where the code runs, and what happens to it
Handing engineers a sandbox with an agent in it raises exactly one question first.
Every Lab runs in an isolated sandbox Anthropos hosts and manages. Nothing installs on your machines, and the environment is metered and destroyed after the session.
A Lab never touches your internal systems unless you deliberately wire it in. You choose which models are available inside each one.
Nothing written in a Lab, and nothing you bring into one, is used to train, retrain or fine-tune any model, ours or a sub-processor's. That is contractual in the DPA.
A Playground Lab is the same environment with scoring off. Nothing is recorded, scored or reported, which is what makes engineers willing to use it.
ISO 27001 and GDPR, hosted in the EU, with the EU AI Act assessment available to read.
Rollout
What it takes to start
SSO goes in, engineers auto-map from CV, HRIS and LinkedIn, and the first coding simulation reaches a pilot squad.
AI Lab challenges run with Claude Code and Codex, scored by the same engine as the simulations.
A baseline per engineer and per team, with Academy paths assigned against the verified gaps.
IT sets up SSO and, if you want role and org data, one HRIS connection. Labs run in sandboxes Anthropos owns and manages, so there is nothing to host.
Objections
What buyers ask before a pilot
What is the difference between AI Labs and an AI Simulation?
AI Simulations assess broad workplace skills in a 5-30 minute scenario with AI actors, including coding scenarios with stakeholders in the loop. AI Labs are the technical, hands-on layer: a live sandbox with Claude Code and Codex, scored on what actually gets built. They complement each other rather than replace each other.
Where do the sandboxes run, and is our code and data safe?
Every Anthropos AI Lab runs in an isolated sandbox that Anthropos hosts and manages. Nothing installs on your machines and it never touches your internal systems unless you wire it in. Each Lab is metered and torn down after the session, so cost stays capped and no data lingers. Region and retention are configurable for enterprise customers.
What does senior engineer mean now that AI writes most of the code?
Anthropos frames it with the Eight Durable Skills from The Engineering Reset: Problem Framing, Business Judgment, Taste, Output Judgment, Systems Thinking, Spec Writing, Orchestration and Learning Velocity. Seniority moves up the stack from producing code to framing the problem and judging output. A Lab session and a coding simulation observe and score exactly those, 0-5.
Does this work for data engineers, DevOps and platform teams?
Yes. Anthropos AI Labs cover technical work generally: building software, running servers, debugging, and using Claude Code to manage and analyse data. Each Lab is configured for the languages and frameworks the team works in. Business roles get prompt-engineering and GenAI workflow Labs; non-technical development runs through AI Academy instead.
Can we build custom AI Labs challenges for our company's specific tools and workflows?
Yes. Anthropos Studio builds custom Labs and coding simulations around your own tools, documents and workflows, with no technical background required, and most are ready to deploy within a day or two. Studio also lets you create your own organization skills and map them to your own role definitions.
How long does implementation take?
Most Anthropos customers are live within 2-4 weeks: account configuration and SSO in week 1, simulation selection and HRIS integration in weeks 1-2, administrator training and pilot launch in weeks 2-3, results review in weeks 3-4. Large deployments with custom Labs and deep integration typically take 6-8 weeks. No systems-integrator engagement is required.
Keep reading
Where to go next
Run a Lab challenge on your own stack
Bring a repository and a squad. We build the challenge, they run it with Claude Code, and you read what came back.