How Do We Measure Workforce-Wide AI Literacy?
Workforce-wide AI literacy is measured by combining role-based scenario assessments, manager observation and usage analytics - not by tracking training completion alone. Firms that triangulate these methods can identify genuine capability gaps and demonstrate credible progress to regulators and boards.
Key takeaways
- Training completion is an activity metric, not a literacy metric - the two must not be conflated.
- Effective measurement is segmented by role, since literacy requirements differ across trading, underwriting, operations and compliance.
- Scenario-based assessment (presenting realistic AI-related situations) reveals judgement gaps that self-reported surveys miss.
- AI tools can help analyse assessment data and usage patterns at scale, but human review remains essential for interpreting results and deciding action.
Every financial services firm rolling out AI training eventually faces the same question from a regulator, board member or risk committee: how literate is our workforce, really?
Attendance logs and completion certificates cannot answer that question. They tell you who sat through a module, not whether a trader, underwriter or compliance analyst can recognise AI risk, use a tool appropriately, or knows when to escalate rather than act alone.
As the government's AI Skills Compact commitments and regulatory expectations increase pressure to demonstrate genuine workforce capability, firms need a measurement approach that goes beyond activity tracking.
This article explains how that measurement works in practice - and where AI can help without replacing the judgement required to interpret the results.
Why measuring AI literacy has become an organisational priority
AI tools are already embedded across trading desks, underwriting teams, operations functions and customer-facing roles, whether formally sanctioned or adopted informally by individual staff.
Regulators and boards are increasingly asking firms to demonstrate that this adoption is safe, controlled and understood - not simply permitted.
The government's AI Compact adds a further layer of expectation: firms are being asked to show tangible commitment to workforce readiness, not just policy statements.
The risk of unmeasured adoption is straightforward. If a firm does not know how literate its workforce genuinely is, it cannot target training investment, cannot reassure regulators with evidence, and cannot track whether literacy is improving or deteriorating as AI tools change.
Measurement is what turns "we ran some training" into "we understand our capability gaps and are closing them."
Traditional approaches to measuring workforce capability
Financial services firms already have a toolkit for measuring workforce capability, developed over decades for compliance training, product knowledge and conduct risk. The same tools apply to AI literacy, with mixed results.
Self-assessment surveys ask staff to rate their own confidence and understanding. They are quick to deploy and easy to scale, but self-reported confidence often diverges significantly from actual capability - particularly with a new capability area like AI, where staff may overestimate or underestimate their own understanding.
Competency frameworks define expected knowledge and behaviours by role or grade. They provide a useful structure for what "good" looks like, but are only as good as the assessment method used to test against them.
Manager sign-off relies on line managers observing staff and confirming competence. This captures real-world behaviour that surveys miss, but is inconsistent across managers and dependent on the manager's own literacy.
Classroom or e-learning testing typically checks recall of training content through multiple-choice questions. This confirms attendance and basic knowledge retention, but rarely tests judgement in realistic, ambiguous situations - which is where AI-related risk most often arises.
Each method has genuine value. None is sufficient alone, which is why effective measurement combines several of them.
Where AI helps in measuring AI literacy
AI-assisted tools can meaningfully support the measurement process itself, without becoming the judge of what the results mean.
AI can help generate realistic, role-specific scenarios at scale - varying the details so that different cohorts are tested against situations relevant to their actual work, rather than one generic case study used organisation-wide.
Where assessments include open-ended responses, AI can support initial scoring by identifying patterns, flagging responses that suggest a misunderstanding of escalation routes, or clustering similar answers for human reviewers to examine more efficiently.
AI can also analyse usage data - which tools staff are actually using, how frequently, and in what context - to surface gaps between reported confidence and observed behaviour.
Across a large workforce, this reduces the manual effort of reviewing thousands of responses or usage logs individually. But the interpretation of what a result means, and what action it justifies, remains a human decision. A compliance officer or L&D lead still needs to decide whether a pattern reflects a training gap, a process gap or an isolated error.
Operational considerations for embedding measurement
Several practical issues determine whether a measurement approach produces genuinely useful results.
Role sensitivity matters more than a single score. A trader working with AI-assisted pre-trade analytics needs a different literacy profile to a customer service agent using an AI chatbot assistant. Measuring both against the same benchmark produces a misleading picture for at least one group.
Frequency should reflect risk and change, not a calendar default. A one-off assessment at rollout quickly becomes outdated as tools and use cases evolve. Reassessment after major tool changes, and periodic refreshers for higher-risk roles, keeps the picture current.
Triangulation reduces blind spots. Because self-reported confidence and actual capability often diverge, combining survey data, scenario results and manager observation gives a more reliable picture than any single method.
Results must feed governance reporting, not just L&D dashboards. Literacy data has genuine value to risk committees and boards, particularly where it demonstrates active management of AI-related conduct risk rather than passive training delivery.
Avoid a tick-box culture. Staff who sense that assessment exists only to generate a compliance statistic will disengage or game the process. Framing measurement as a tool for identifying support needs - not as a pass/fail test - maintains trust and produces more honest results.
Example
A London-based investment bank's fixed income trading desk introduces an AI-assisted pre-trade analytics tool. Rather than assuming training completion equals readiness, the Head of Trading Operations works with the Learning & Development lead to design a scenario-based assessment: traders are presented with a case where the AI tool's suggested trade conflicts with an internal risk limit, and asked how they would respond.
The assessment reveals that while 90% of traders completed the AI training module, only 60% correctly identified the need to escalate the conflicting recommendation rather than override the risk limit themselves.
This gap, invisible from completion data alone, prompts targeted follow-up coaching and a revised escalation protocol. The results are reported to the firm's risk committee as evidence of active AI literacy management, alongside the compliance officer's review of what the gap implies for wider desk-level controls.
FAQs
-
What is the difference between AI literacy and AI training completion?
Training completion measures whether someone attended or worked through a module. AI literacy measures whether they actually understand AI risk, use AI tools appropriately, and know when to escalate rather than act alone. A high completion rate can exist alongside significant literacy gaps, which is why the two metrics should never be treated as interchangeable.
-
How often should workforce AI literacy be measured?
This depends on the risk exposure of the role and the pace of change in the AI tools being used. Higher-risk, AI-intensive roles typically warrant more frequent reassessment, particularly after major tool rollouts or changes. A one-off exercise at launch is rarely sufficient; periodic reassessment, such as annually or following significant changes, keeps the picture current.
-
Can AI literacy be measured the same way across all roles?
No. A single organisation-wide score is misleading because literacy requirements differ significantly between, for example, trading, underwriting, operations, customer-facing and compliance roles. Segmenting measurement by role produces results that are more accurate and more actionable for targeting support.
Turn the Skills Compact into action
Get in touch for a free consultation on turning the Skills Compact into a practical AI skills plan for your teams.