A CEFR-modeled standard, not a survey.
Sophrosyne scores AI fluency the way language proficiency is scored — by evidence of what someone can actually do, staged from beginner to advanced so every employee has a clear next step.
Six stages. Each one auditable.
Scored against portfolio evidence — real work, not a self-reported survey.
Breakthrough
Personal productivity with general AI — research, drafting, summarizing.
An employee uses ChatGPT or Copilot for first drafts and summaries. No company data, no repeatable process — output quality depends entirely on the individual.
Waystage
Proprietary context — grounding AI in a curated corpus for reliable, citable work.
The same employee now grounds AI in the company's own documents, tickets, or product data, so answers are sourced and checkable rather than generic.
Threshold
Systems thinking — chaining tools into reproducible pipelines.
A workflow that used to take five manual steps now runs as one chained pipeline — documented, repeatable, and owned by the team rather than one person's habits.
Vantage
Internal data products — dashboards and scripted internal agents.
The team ships an internal tool that other people depend on daily — for engineering that might be a scoped internal agent; for finance or ops it might be a scripted reporting pipeline. Different artifact, same bar.
Advanced
Function-level AI integration — shipping features with evaluation and cost monitoring.
A product team ships a customer-facing AI feature with evaluation criteria and cost monitoring in place before launch; a marketing team ships a governed content-automation pipeline held to the same evaluation rigor. Different function, same bar.
Mastery
Multi-agent systems with explicit human oversight and governance.
An engineering or ops team runs a multi-agent system in production under a named governance model — who can approve what, and what happens when an agent is wrong.
A proprietary scoring model, built in-house.
The engine behind the Fluency Standard was built by our team — not licensed from a generic assessment vendor, and not a self-scored quiz. It's the same evidence-based model already proven inside higher education, now scoring employees at companies.
Then
Credit bureaus didn’t just describe risk — they standardized it.
Before a portable score existed, lending was a subjective, relationship-by-relationship judgment call. A common evidence-backed number let a lender act on a stranger’s risk as confidently as a known customer’s.
That standardization is what let consumer credit scale into a mass market, and let entire industries — credit cards, securitized lending, fintech underwriting — build on top of a number everyone trusted.
Now
The Fluency Score does for AI capability what the credit score did for risk.
AI fluency inside most companies today looks like lending before the credit score — a subjective call, inconsistent department to department, with no number a board can act on or compare against competitors. The Fluency Score turns a judgment call into a standardized, evidence-backed number.
That’s what lets a company move decisively on integration — knowing where to invest, benchmarking against peers, making its workforce’s AI capability a legible asset — while competitors still running on gut feel fall behind.
Evidence-based. Auditable by your board, not just by us.
This is the same portfolio-evidence approach the Standard was built on inside higher education — modeled on the same CEFR framework used worldwide for language proficiency.
- 01Function-by-function evidence collection
We collect real work product per function — documents, workflows, shipped tools — not survey answers. An engineer's evidence and a marketer's evidence look different; they're scored against the same rubric.
- 02Scored against the six-stage standard
Each piece of evidence is auto-graded against the named A1–C2 rubric the moment it's submitted, calibrated against human expert graders — not an algorithm guessing from a quiz.
- 03Aggregated into a stage distribution
Individual scores roll up into a distribution per function — where the company actually sits, not an average that hides the gap.
- 04Delivered as a named gap map
The output names which functions are behind, by how many stages, and what closing that gap requires — board-defensible, not vibes-based.
The four steps above are what a customer sees. Underneath, the scoring itself is held to the same bar as any real measurement instrument — not asserted, checked.
- 05Adaptive item selection
The exam is computerized-adaptive (CAT-style): each answer adjusts the difficulty of the next item, which is what compresses a reliable six-stage placement into a single sub-60-minute sitting instead of a long fixed-form test.
- 06LLM-assisted grading, calibrated to human experts
Open-ended scenario responses and portfolio evidence are scored against stage rubrics using LLM-assisted grading — calibrated against human expert graders from our university deployments, not used zero-shot. Human-graded evidence is the ground truth the model is checked against, not a fallback.
- 07Validation, not just scoring
We track inter-rater agreement between machine and calibrated human graders, test–retest reliability across quarterly re-scores, and item-level psychometric analysis (difficulty and discrimination) as the response corpus grows. A published validation study with an academic partner — inter-rater reliability coefficients, predictive validity against job performance — is a stated near-term objective, not a claim we're making today.
A board-ready roadmap — not a training-completion certificate.
The Assessment's deliverable is a single report your board and your function leads can both act on.
Current-stage distribution
Where every function actually sits on the A1–C2 scale today, evidence-scored, not self-reported.
Target stage per function
A realistic next stage for each function, set against what the role and the business actually need — not C2 for everyone.
Sequenced closing plan
Tooling, training, and governance steps in order, so the roadmap is something you can execute against, not a list of ideas.
Ongoing integration support
You're not implementing alone — support scales from email access up to a dedicated advisor, so teams actually act on the roadmap instead of filing it away.
The tracker behind the score.
Every subscription includes this dashboard: your teams' scores, the A1–C2 stage distribution, and quarter-over-quarter progress against the roadmap — so moving employees from beginner to advanced is something you watch happen, not something you take on faith.
This is what your dashboard looks like once your team is onboarded.
Where we fit — and why the alternatives leave a gap.
Generic AI-training vendors sell completion certificates. Large consultancies sell strategy decks. Neither leaves you with a score you can re-test.
| Sophrosyne | AI-training vendor | Large consultancy | Doing nothing | |
|---|---|---|---|---|
| What you get | An evidence-based score per function, plus a sequenced roadmap | Completion certificates for generic AI courses | A slide deck strategy with no scoring mechanism | No baseline — you find out you're behind when a competitor already isn't |
| How it's measured | Portfolio evidence, scored against a named 6-stage standard | Quiz completion, self-reported | Interviews and workshops, not scored | — |
| Re-testable | Yes — quarterly re-score against your own baseline | Rarely tracked after the course ends | Only if you re-engage | — |
| Provenance | Standard built and proven inside higher education before adapting to companies | Generic content licensed across industries | Varies by firm and team | — |
Know your company’s AI-fluency score before your competitors do.
Score your employees, track their progress, and get the roadmap to close the gap — subscription pricing, no consulting engagement required.
No sales call required · 12-month Tracker term · Response within one business day