Skills Taxonomy in the AI Era: How to Stay Current Without Losing Rigor

Skills Taxonomy in the AI Era: How to Stay Current Without Losing Rigor

Taylor Sullivan

Vice President of Product and Assessments

Your Skills Taxonomy Can Finally Be Specific. That's the Hard Part.

I trained as an I/O psychologist, which means I spent years learning how to take a job apart. Watch the work, break it into tasks, figure out what a person has to know or be able to do to perform each one, build up from there. It's careful work and it tells you things nothing else does. Do it across an organization and the result is a skills taxonomy: a shared description of what the work requires and what it looks like to do it well.

The work is also slow and expensive, and for most of my career that was the binding constraint. You could describe a job in real detail, but it took months of expert time, and by the time you finished the job had moved. So organizations did the rational thing and described work at a level general enough to stay true. Strategic thinking. Drives results. Fine as names. But for most organizations, that broad name was all that guided talent decisions.

AI is changing that, and it's the development in my field I'm most encouraged by. We can now process work at a granularity that was never affordable before, which means a skills taxonomy can finally be as specific as the work it describes.

Granularity has a shelf life

There's a catch, though, and it's the reason I'm writing this. The more precisely you describe what someone needs to be able to do, the faster that description dates, because you've stopped describing abstractions and started describing practice, and practice moves. "Strategic thinking" will still be a thing in 2035. "Knows when to use extended reasoning versus a single-pass prompt" may not survive the year.

The way out is to separate two layers rather than retreat to the general labels. The name of a broad capability can stay put for a decade. What sits underneath it is what has to move. That's the capability as it actually shows up in practice, described closely enough to observe and assess. Keep the container stable and version the contents.

That's a trade worth making. Precise and perishable beats vague and permanent. But it only works if something keeps the description current, and that's the part most organizations haven't built. The constraint on understanding work has moved from cost to maintenance.

Two ways a skills taxonomy fails

This would be a minor problem if a skills taxonomy were still what it used to be: an HR artifact, consulted occasionally, easy to ignore. It isn't anymore. As organizations have reorganized around skills, the taxonomy has become the thing hiring, deployment, development, and workforce planning all run on. So when it fails, it fails everywhere at once.

And without a method for staying current, it fails in one of two directions. I've watched both happen.

It drifts. Someone builds it carefully, it gets reviewed annually (if at all), and it gradually becomes an accurate description of how the work was done a few years ago. Nothing announces this. Reports still run and dashboard percentages still update, which is what makes drift so durable: a stale taxonomy produces confident numbers right up until somebody digs in and questions their relevance.

Or it bloats. Anything new with a name starts to look like a new skill: a model, a tool, a protocol, a practice, a job title that showed up in enough postings to seem real. Entries accumulate until the taxonomy catalogs everything anyone has touched and signals nothing about capability. This one usually comes from good intentions; adding feels like keeping up.

Both failures land you in the same place: talent data that leaders keep using and, at the same time, can no longer trust.

A durable core and a fast-moving edge

Our answer is to stop asking one taxonomy to be both durable and current, and it has two parts. The first handles bloat: two tracks, so new tools have somewhere to go that isn't the canonical set. The second handles drift: versioning in place, so the canonical set stays current without changing shape.

Our Signature Catalog is our canonical standard. It's a deliberately small set of durable, tool-agnostic capabilities, organized by audience, that we version in place rather than expand by default. Alongside it sits a companion track of tool-specific capabilities: quick-build assessments shipped in days when the field moves, and expected to age out. That companion track is the pressure valve that keeps the canonical set small and credible. The core stays durable. Currency lives at the edge.

What it takes to earn a place in the canonical set

A capability only enters the Signature Catalog if it clears a high bar, enforced at admission rather than cleaned up afterward.

  • Skill, not tool. We measure the underlying capability, not the product executing it. "Design a multi-agent system," not "use this framework."
  • Evergreen over trendy. A slot is earned only if the underlying construct still matters across several model generations.
  • Coverage through structure. Breadth comes from how the catalog is organized by audience, not from sheer count.
  • Each entry earns its keep. If a strong performer on one would pass another, they merge or one steps down. Differentiation is enforced up front.
  • Audience-anchored. Every entry maps to a real population that needs it and would value proving it.
  • The capability is the stable container. The name and scope are the durable layer. The skills, behaviors, and items inside are the versioned contents.

Maintenance runs as an agentic system

Keeping a catalog current can't depend on someone remembering to check. We run it as an agentic system that watches the field continuously and alerts against thresholds we set.

Agents compare each capability against the live state of the art and flag meaningful drift. Review cadence follows how fast each area moves. Leadership and judgment don't change much year to year; prompting and agent design change in months. So the fast-moving areas get looked at monthly and the durable ones far less often. When something crosses a threshold, the agent drafts a specific recommendation, a human expert reviews it, and we refresh what drifted, retire what's stale, and absorb what has proven to last.

The automation finds the drift. It doesn't get to redefine the standard. That division isn't incidental. Deciding what counts as a capability is a judgment call, and I want a person accountable for it.

Net additions to the canonical set are rare and evidence-based. When the instinct is to add an entry for a new tool, that instinct gets routed to the companion track unless the tool exposes a genuinely new, durable, assessable capability nothing else covers.

An example: Prompt Engineering for Developers

Here's what this looks like in practice. Prompt Engineering for Developers has kept its name, scope, and tier since we built it. What's inside it has changed.

Added for 2026: knowing when to ask a model to reason step by step and when that hurts; prompting that reliably calls tools and returns structured output; and treating chain-of-thought as something to apply rather than something to recall.

Kept durable: evaluation discipline, model and cost selection, prompt-injection defense, and API debugging.

Because the definition holds still, this year's result and last year's are measuring the same thing. That doesn't make scores automatically comparable, and we don't claim it does. Every result carries the version that produced it, so anyone comparing across versions knows it. Every change is reviewed before it ships for whether it moves what's being measured or shifts how hard each proficiency level is to reach. After a refresh, we watch the scores; if they move in ways the content change doesn't explain, that's drift, and it gets investigated. Keeping the core small is what makes that discipline sustainable.

If you maintain your own taxonomy

None of this requires our catalog. The split is portable. Decide which of your entries are durable capabilities and which are familiarity with a tool. Hold the first set to a high admission bar and let the second set expire on purpose. Put a person, a process, or an agent on a schedule to check whether the durable entries still describe the work.

The layer underneath

A taxonomy is a map of what skills exist. Who actually has them, and to what depth, is a separate layer, and it's the one most organizations are weakest on: the framework itself gets careful attention, while the data on who has each skill is a self-rating, a manager's impression, or a profile nobody has updated in years.

A newer approach is to infer skills from job titles, resumes, and activity data. It scales, and it's useful for seeing the rough shape of a workforce. But an inferred skill is a probability that someone has been exposed to something, not evidence they can do it, and any proficiency level attached to it is inferred from the same signals. It's still a proxy, just a more sophisticated one.

That matters because of what this data gets used for. Fairness conversations usually treat assessment as a hazard, something that might screen people out. That misses the alternative. Without measurement, opportunity gets allocated by proxy: where someone went to school, who noticed them, which projects they happened to land on, what an algorithm guessed from their title. Those proxies aren't neutral; they're just harder to audit. A good assessment lets someone demonstrate capability directly, which is often the only route available to the people the proxies skip. That's an argument for measuring, not for measuring carelessly. A badly built assessment relocates the same bias and attaches a number to it, which is worse than a hunch because it looks like evidence.

This is the real reason maintenance matters. A framework that has stopped describing the work is a badly built assessment by another route. Most organizations don't need another taxonomy. They need the part underneath it to be trustworthy, and a stale one is a set of decisions being made on bad information.

That's a narrow claim, not a modest one. Whether the framework is your own, a consultancy's, or an industry standard, something has to sit beneath it and verify that the capability is actually there. That layer is where rigor either lives or doesn't.

The organizations that get this right won't be the ones with the most detailed skills library. They'll be the ones whose measurement stayed honest while the work kept changing. We now have the tools to describe work at the granularity our field always wanted. The discipline is in keeping that description true.

The conversation worth having is where the capability layer fits your skills framework, and what becomes possible when measurement keeps pace with AI. Let's talk.

Category

Blog

United States Air Force DFAS

U.S. Defense Finance and Accounting Service Uses Workera to Upskill Employees and Develop Broad Technical Expertise

85%

average score improvement for continuous learning

1.7x

Best in-class learning velocity

Why you can't trust the signals you hire on
VIRTUAL

Why you can't trust the signals you hire on

Completion Was Never Proof: Why L&D Must Measure Real Capability Change

Blog

Completion Was Never Proof: Why L&D Must Measure Real Capability Change

by

Workera Team

Trending Updates

Pharma's AI Skills Are Strongest Where the Stakes Are Lowest

Customer Stories

Pharma's AI Skills Are Strongest Where the Stakes Are Lowest

by

Workera Team

Completion Was Never Proof: Why L&D Must Measure Real Capability Change

Blog

Completion Was Never Proof: Why L&D Must Measure Real Capability Change

by

Workera Team

August Product Release Round Up

Product updates

August Product Release Round Up

by

Workera Team

Transform your workforce and drive success