Talentopian PCS Validation Whitepaper — v1
A self-published evidence brief on the Personal Competency Score (PCS) — its theoretical foundation, scoring methodology, and current state of validation.
Author: Talentopian Research (Joyeux Réalité étendue Inc.). Version: 1.0 draft (in active development). Status: Self-published whitepaper, not peer-reviewed. A v2 peer-reviewed manuscript with IRB and an external researcher is planned. This v1 is a living document, filled in section by section. Citation style: APA 7th edition. A full reference list is forthcoming in a companion methodology-and-citations document. License: © 2026 Talentopian. Self-archival permitted with attribution.
Abstract
The Personal Competency Score (PCS) is the central measurement output of Talentopian, a freemium browser-delivered game-based career-fit assessment platform serving English-Canadian, Korean, French-Canadian, and Spanish locales for participants aged 15 to 45. This v1 self-published whitepaper documents the PCS instrument's theoretical foundation (Holland RIASEC, Lent–Brown–Hackett SCCT, O*NET work-activities, Mislevy–Steinberg–Almond evidence-centered design, and Shute's stealth-assessment paradigm), its 48-parameter taxonomy organized across eight categories, the open-source scoring pipeline that converts game-emission traces to a 0–100 score and an aligned ranked-career recommendation drawing from a 1,041-occupation taxonomy with WEF Future of Jobs 2025 industry-displacement overlay, and the platform's preliminary internal-structure evidence on an early pilot cohort. Preliminary internal-structure analysis on this early pilot cohort is underway; directional findings suggest the parameter set behaves as a coherent, correlated multi-construct instrument, with several interpersonal parameters clustering tightly enough to motivate dimension-reduction in a future framework revision. The four classical validity coefficients (test-retest reliability, concurrent validity against O*NET RIASEC, criterion validity against ground-truth career fit, and convergent validity against self-rated soft skills) are not reported at v1 because the corresponding data are not yet collected at a publishable sample size; a pre-registered validation roadmap is established to collect these through formal validation. The whitepaper deliberately positions the platform as a first-pass evidence-generation tool for counselor-mediated workflows, not as a validated screening instrument for high-stakes decisions. v2 (peer-reviewed manuscript, IRB-gated, external researcher partnership) is queued for Year 2. The instrument's anti-faking design (Coherence Triangulation Validator, currently shadow-mode), the open-source-framework posture, and the freemium distribution model are positioned as distinct from existing enterprise-only game-based assessment platforms (Pymetrics / Harver, HireVue, Arctic Shores) — distinctions of posture rather than validity superiority claims. Bug-history corrections to the scoring pipeline (a scale-aware normalization fix; a job-saturation calculator fix) are reported in §7 per SIOP Principles (2018) technical-report standards.
Keywords: career assessment, game-based assessment, stealth assessment, evidence-centered design, vocational interest, social cognitive career theory, O*NET, WEF Future of Jobs, AI-era career displacement, freemium platform, cross-cultural career counseling, Korean career counseling.
Table of Contents
- Introduction
- Theoretical Framework
- Methodology
- Validation Studies
- Results
- Discussion
- Limitations and Future Research
- Practical Applications for Counselors and Educators
- Acknowledgments
- References — grows section by section
- Appendix A: 48-Parameter Framework Specification
1. Introduction
1.1 The problem this work addresses
Career-fit assessment for individuals 15–45 years old has reached a crossroads. Three forces are squeezing the field at once.
First, the labor-market substrate is changing faster than legacy career-counseling instruments can update. Generative AI tools have begun automating tasks once considered defining for entire occupations — drafting code, writing marketing copy, summarizing legal documents, processing claims, generating images — while creating new categories of work (prompt engineering, AI safety, AI-integration consulting) that have no equivalent on the O*NET-SOC taxonomy three years ago. The World Economic Forum's Future of Jobs Report 2025 projects that 39% of workers' core skills will change by 2030. Assessment tools whose underlying job model is updated annually, or whose validation samples are five or more years old, are increasingly mapping individuals to a labor market that no longer exists.
Second, the measurement substrate has changed as well. Interest inventories (Strong, Self-Directed Search) and personality batteries (Big Five, MBTI-family) measure stated preferences and self-described tendencies — both highly susceptible to social desirability, response sets, and aspirational projection. Recent meta-analytic work has continued to find moderate-at-best concurrent validity for these instruments against job-performance outcomes (Whiston, 2017; Webster, 2020). Game-based assessment and stealth-assessment paradigms (Shute, 2011; Mislevy et al., 2003) offer a path to demonstrated behavior — reaction time under cognitive load, decision patterns under ambiguity, persistence under failure — that is much harder to fake and arguably more relevant to how someone will actually perform in a role.
Third, the delivery substrate matters. Most validated career-assessment instruments require a licensed professional (school counselor, vocational psychologist) to administer and interpret. Pricing is correspondingly opaque — institutional per-administration license fees from large publishers are typically substantial (often on the order of tens of dollars per administration), with the marginal cost effectively all profit for the publisher. Meanwhile, the population that arguably needs the assessment most — career changers in their thirties and forties, first-generation post-secondary students, immigrants whose foreign credentials and work histories are not legible to North American employers — is exactly the population least likely to be sitting across a desk from a licensed counselor with an active publisher contract. The instrument is paywalled away from the people who would benefit most.
1.2 What Talentopian PCS is
The Personal Competency Score (PCS) is the central output of a game-based, browser-delivered career-fit assessment platform that this whitepaper describes. The platform asks individuals to play between 10 and 40 short cognitive, behavioral, and situational-judgment games (typical session: 60–120 minutes), passively records performance traces (reaction time, choice patterns, persistence, sentiment), aggregates those traces into a 48-parameter framework spanning eight categories (cognitive, technical, interpersonal, behavioral, personality, values, career, physical), and produces both a normative score and a set of recommended career fits derived from a 1,041-occupation taxonomy weighted by World Economic Forum industry-displacement projections.
The 48-parameter framework is the canonical taxonomy of what the platform measures (the Physical category was added in a later framework revision, expanding the original seven-category structure); it is documented in machine-readable form in the open-source scoring-framework specification (see Appendix A). The labor-market layer draws on O*NET 28.x, ISCO-08, WEF Future of Jobs Industry Profiles, and live posting signals from public sources. The platform is delivered as a freemium web application; a paid tier removes some preview restrictions but does not gate the core assessment.
We make four broad claims in this whitepaper, and we will defend each with the evidence available to us as of v1:
- The 48-parameter framework is plausibly aligned with the established literature on career-fit measurement (covered in §2 — Theoretical Framework).
- The scoring methodology is reproducible, scale-aware, and free of the silent-fallback failure modes we have audited internally (covered in §3 — Methodology; see also §7 Limitations for honest disclosure of audited bug-fixes from 2026-06).
- Even at v1's early pilot sample (a single snapshot per user), the PCS scores exhibit coherent internal structure across the measured parameters, while test-retest reliability and O*NET-interest concurrent validity are not yet collected and are pre-registered for v2 (covered in §4 — Validation Studies; §5 — Results).
- The platform's anti-faking design (the Coherence Triangulation Validator described in §3.6) addresses a class of failure mode that existing instruments do not (covered in §3.6 and discussed in §6).
1.3 What this whitepaper is not
We want to be precise about scope before we begin, because the field has been ill-served by overconfident product whitepapers.
- This is not a peer-reviewed validation study. The current sample (a small pilot cohort) is far too small to support a peer-reviewed concurrent or predictive validity claim. The peer-reviewed manuscript is planned as v2 of this document, conditional on (i) Institutional Review Board approval, (ii) external researcher partnership, and (iii) a sample of n ≥ 200 with retest sub-sample n ≥ 80.
- This is not a marketing document with an audited claim register. Where we cite a number, we describe the basis on which it rests — the data source, the date, and the sample size — rather than asserting it without qualification.
- This is not a replacement for a licensed counselor's clinical judgment. The platform is positioned as a first-pass evidence-generation tool that a counselor can review with the individual; nothing in the platform's design or marketing should be read as a recommendation for autonomous high-stakes decisions (university admission, hiring, professional licensure).
- This is not a claim that game-based assessment is a strict superset of inventory-based assessment. Conditions exist (low computer literacy, sensory or motor impairments, certain cultural contexts) where a structured interview or a validated interest inventory remains the better instrument. We discuss these conditions in §7.
1.4 Why publish v1 now
The honest answer is that a paying audience — Korean career-counseling associations, secondary-school counseling teams in Ontario, and partner organizations in francophone Québec — needs an evidence brief they can read and forward sooner than a fully peer-reviewed manuscript can be produced. A self-published evidence brief is the field-standard interim format for exactly this situation (cf. the Society for Industrial-Organizational Psychology's Principles for the Validation and Use of Personnel Selection Procedures, 5th ed., which explicitly recognizes technical reports of this kind alongside peer-reviewed validation studies; SIOP, 2018).
The platform is designed to be relevant to Korean career-counseling practice, and we hope to engage the Korean career-counseling community. These audiences need a citable reference document with provenance, not a marketing page.
1.5 How to read this document
Sections 2–3 establish the what and why of the instrument (theoretical alignment + scoring mechanics). Sections 4–5 present the evidence — at v1, this is preliminary and we say so. Sections 6–7 discuss what it means and what it doesn't. Section 8 offers concrete what to do with this guidance for the practitioners we expect to be the primary readers.
Readers who want only the scoring details should skip to §3. Readers who want only the validation evidence should skip to §4. Readers who want only the practitioner-facing summary should skip to §8.
2. Theoretical Framework
The Talentopian PCS assessment does not invent a new theory of vocational interest or competency. It deliberately stands on the shoulders of four bodies of established literature — Holland's vocational personality theory, Lent–Brown–Hackett's social cognitive career theory, the O*NET work-activity / abilities framework, and the measurement tradition of evidence-centered design (ECD) and stealth assessment — and integrates them in a way that matches the actual mechanism of measurement we use (short games delivered in a browser), the population we serve (15–45 year-olds across four locales), and the labor-market substrate that career assessment must now contend with (AI-era task displacement).
This section establishes the why of the 48-parameter framework. The how — the games, the scoring, the aggregation — is the subject of §3.
2.1 Holland's RIASEC and the limits of interest-only measurement
Holland's hexagonal model (Realistic, Investigative, Artistic, Social, Enterprising, Conventional; Holland, 1997) remains the most widely deployed framework in career counseling globally, and it is the framework most likely to be familiar to the practitioners (school counselors, vocational psychologists, accreditation reviewers) who will read this whitepaper. We map it explicitly into our taxonomy in §3.2.
Three observations about RIASEC shaped how we used it.
First, RIASEC is a theory of preferences, not competencies. An individual high in Investigative interests may or may not have the working memory, abstract reasoning, or persistence required to perform Investigative work. The career-fit signal that practitioners need is the intersection of interest, competency, and labor-market viability — RIASEC delivers one of those three. The other two must come from elsewhere.
Second, RIASEC's measurement is overwhelmingly self-report. The Strong Interest Inventory, the Self-Directed Search, and their derivatives all rely on the respondent's stated preferences. Recent meta-analytic and methodological work has continued to find that interest-inventory scores are vulnerable to social desirability, aspirational projection, and the well-documented response-set effects of fixed-choice instruments (Whiston et al., 2017). When the stakes are high (admission, hiring, parental approval), the validity of self-report drops further. Game-based behavioral signal is much harder to fake without explicit, sustained training — a property we exploit deliberately (see §3.6 and §6).
Third, RIASEC was developed when "work" was substantially more stable than it is now. The six personality–environment correspondences are durable, but the occupations they map to are changing under generative AI faster than the taxonomy can absorb. Our O*NET-anchored career layer (see §2.3 and §3.3) handles this by separating the timeless personality–interest signal from the timely labor-market projection.
2.2 Social Cognitive Career Theory and the self-efficacy bridge
Lent, Brown, and Hackett's SCCT (1994; updated synthesis in Lent & Brown, 2002) takes Bandura's self-efficacy work and applies it specifically to career choice. The model adds three constructs that pure trait theories like RIASEC lack: self-efficacy beliefs about specific work tasks, outcome expectations about what doing those tasks will lead to, and contextual supports and barriers that moderate whether interest translates into pursuit. These three constructs are exactly the constructs a practitioner cares about when sitting across from a hesitant teenager, a stalled-career 38-year-old, or an immigrant whose foreign credentials are not legible to the local labor market.
SCCT is also the strongest theoretical justification we have for behavioral assessment over inventory assessment. Bandura's central claim was that self-efficacy is built and revealed through enactive mastery experience — actually doing the thing. A 90-second cognitive game where the participant solves an Investigative-style puzzle under time pressure, makes a sequence of choices, and sees their pattern surface on a results screen is, structurally, a miniature mastery experience. It is more diagnostically informative than the corresponding statement "I am good at solving logical puzzles" because it produces a behavioral record, not a self-report.
We do not claim to have implemented SCCT in full. The model includes longitudinal feedback loops (mastery experience → self-efficacy → outcome expectation → interest → choice → performance → mastery experience) that no single assessment session can resolve. What we claim is that the PCS taxonomy is structured to provide SCCT-compatible inputs — game-based mastery signals on specific work-like task families — that a practitioner or downstream longitudinal study can chain into the full SCCT loop. See §6 for an honest discussion of where the platform stops and where the practitioner's clinical judgment must take over.
2.3 O*NET work-activities and the labor-market layer
The U.S. Department of Labor's O*NET Occupational Information Network (O*NET, 2024) is the most exhaustive open work-activity taxonomy in existence: 1,016 occupations as of release 28.x, each tagged with task statements, generalized work activities (GWAs), detailed work activities (DWAs), abilities, knowledge areas, skills, work styles, work values, and work contexts. It is freely licensed, regularly updated, and explicitly designed to be machine-readable. It is also the canonical bridge between the framework-level constructs (Holland, SCCT) and the actual occupations the labor market hires for.
Two facts about O*NET shaped our design. First, O*NET's competency tagging is occupation-anchored, not individual-anchored. It tells us "what does a Network Architect need to do" — not "is this individual a good Network Architect." Bridging from one to the other is precisely the work of our scoring layer (§3.3). Second, O*NET is not a perfect substrate for non-US labor markets. We use it as the primary occupation taxonomy because no comparable resource exists in Canada, France, or Korea, but we explicitly cross-walk to ISCO-08 (used by Statistics Canada), NOC 2021 (Canada), CNP 2017 (France), and KSCO 2017 (Korea) at the platform's career-recommendation layer. The cross-walk is maintained internally, mapping to a taxonomy of 1,041 occupations.
We also incorporate the World Economic Forum's Future of Jobs Report 2025 industry-displacement projections (WEF, 2025) as a labor-market overlay. WEF projections are not occupation-level; they are industry-level and skill-level. Our use of them is restricted to (a) AI-resilience scoring at the industry of the recommended career and (b) directional AI-collaboration potential. They do not enter the PCS score itself; they only modulate which careers are surfaced and with what AI-displacement context.
2.4 Evidence-Centered Design and stealth assessment
The methodology that lets us turn 90-second cognitive games into psychometrically defensible measurement comes from a different lineage: educational and behavioral measurement, specifically Mislevy, Steinberg, and Almond's evidence-centered design (ECD; Mislevy et al., 2003) and Shute's stealth assessment paradigm (Shute, 2011).
ECD provides a formal framework for the evidentiary argument between an observable behavior and a latent competency claim. It decomposes the assessment problem into three models: a student model (what latent constructs are we trying to estimate), an evidence model (what observable behaviors are diagnostic for those constructs), and a task model (what tasks reliably elicit those observable behaviors). The student model in our case is the 48-parameter framework; the evidence model is the per-game scoring rules defined in the open-source scoring-framework specification; the task models are the games themselves. This is not a coincidence; the architecture was deliberately designed to map onto ECD because ECD is the dominant framework for high-stakes computer-based assessment in the educational measurement community (the GRE, NCLEX, and several state K–12 assessments use ECD-derived structures).
Shute's stealth assessment (2011) is the practical instantiation of ECD inside actual games. The defining property of a stealth assessment is that the player is engaged with the game on its own terms (a puzzle to solve, a story to follow, a goal to reach) while the assessment runs underneath the game — recording reaction times, choice patterns, persistence under failure, and decision sequences. Stealth assessment defends against several of the failure modes of explicit testing: the participant is not in a "test-taking mindset," is not strategically optimizing for the test's stated rubric, and is not aware (in any meaningful way) of which behaviors are being scored. The combination of (a) intrinsic engagement and (b) opacity of the scoring is what makes stealth assessment hard to fake.
We do not claim Shute's full apparatus. We have not yet performed (and v1 does not require) the large-scale item-response-theory calibration that the strongest stealth assessment work has performed. What we claim is that the design intent is stealth-assessment-compatible: every game is a goal-directed activity the participant cares about for its own sake; the scoring is performed on traces that are not visible to the participant; the parameter framework is the latent variable space those traces are projected into.
2.5 What we did NOT borrow, and why
A reader familiar with the career-assessment literature will notice three frameworks we do not draw on. Brief explanations follow.
We do not use the MBTI / Myers-Briggs family of instruments. The MBTI has substantial reliability and validity problems documented across multiple meta-analyses; it is also commercially restricted in a way that would compromise our open-citation posture.
We do not use Cattell's 16PF or related second-order factor structures. These instruments are statistically respectable but are interest-domain inventories rather than competency or behavior inventories; they would duplicate the RIASEC layer without adding new diagnostic surface.
We do partially use the Big Five / Five-Factor Model (McCrae & Costa, 2008) at the personality-dimension layer of the 48 parameters, but only for the dimensions that have plausible behavioral correlates in game play (Conscientiousness via persistence-after-failure, Openness via novelty-seeking on optional puzzle branches, Neuroticism via choice patterns under time pressure). We do not present a Big Five score per se because we have not validated against a standard Big Five inventory at v1 (this is a work item for a later revision).
2.6 How these frameworks compose into the 48-parameter taxonomy
The eight categories of the 48-parameter framework (Cognitive 9 / Technical 8 / Interpersonal 5 / Behavioral 10 / Personality 6 / Values 2 / Career 4 / Physical 4) are not, themselves, drawn from any single source — they are the framework author's synthesis. But each individual parameter inside each category traces to one or more of the four frameworks above. For example: working memory (Cognitive category) traces to Holland's Investigative dimension and to O*NET's "Memorization" ability; situational-judgment scoring (Interpersonal / Behavioral) traces to Webster's (2020) SJT meta-analytic work and to SCCT's outcome-expectation construct; persistence-after-failure (Behavioral) traces to Shute's stealth assessment of grit and to SCCT's self-efficacy mechanism.
Appendix A publishes the complete parameter-to-framework mapping. The point of this section is to establish that no parameter is theoretical free-floating — every one of the 48 has a citable lineage.
3. Methodology
This section documents what we actually measure, how we actually measure it, and how we actually score it. It is intended to be reproducible to the extent that a competent assessment psychologist or systems engineer with access to our open repository could verify each claim. Where details are deferred to a later revision (because they depend on the platform's validity-correlation extraction, in progress at the time of writing), that is explicitly stated.
3.1 Delivery substrate
The platform is delivered as a browser-based web application. No native app, no plug-in, no specialized hardware is required. Participants need a modern browser (Chrome 100+, Firefox 100+, Safari 15+, Edge 100+), a keyboard, and a pointing device. Game sessions are interactive but stateless from the platform's perspective: every score-relevant event is written to the platform's data store in real time, and the assessment can be paused and resumed across sessions with no degradation. We support EN-CA, KO-KR, FR-CA, and ES-ES out of the gate (all UI, all game instructions, all results); locale is sticky per user.
The delivery substrate matters for two reasons. First, browser delivery removes the licensed-counselor gating that excludes most of our target population from validated assessment (see §1.1). Second, browser delivery means the platform can be embedded inside a counselor's workflow rather than replacing it — the platform produces evidence; the counselor produces interpretation.
3.2 The game library
The participant plays between 10 and 40 short games per session. Each game is between 30 seconds and 6 minutes in length (typical: 1–3 minutes) and is designed to elicit behavior diagnostic for a specific subset of the 48 parameters. The game roster is maintained in the platform's game library, and the machine-readable game-to-parameter mapping lives in the open-source scoring-framework specification.
The games fall into eight design families, each chosen for the parameter signal it most cleanly elicits:
- Cognitive puzzles (e.g. Mind Maze, Quantum Coder) — working memory, pattern recognition, abstract reasoning, sustained attention. Reaction time and accuracy curves under increasing difficulty are the primary signals.
- Situational judgment scenarios (e.g. Ethics Oracle, careerCounselor microtask) — interpersonal reasoning, ethical judgment, value alignment. Choice-pattern signatures across stem variants are the primary signal. Note Webster (2020) meta-analytic SJT validity (r ≈ 0.32 with job performance) — strongest published validity in the personnel-selection literature.
- Creative generation (e.g. Cosmic Harmony) — openness, divergent thinking, aesthetic judgment. Output diversity scored via embedding-based novelty measurement.
- Strategic / multi-step planning (e.g. Quantum Strategist) — executive function, planning horizon, contingency handling. Move-tree depth and branch quality are the primary signals.
- Bio-architecture / spatial reasoning (e.g. Bio-Architect, Neuro-Architect) — spatial visualization, systems thinking, multi-constraint optimization. The flagship game Neuro-Architect is the broadest single game in the roster (13 of the 48 parameters), spanning four of the eight parameter categories (cognitive, technical, behavioral, personality).
- Linguistic / communication (e.g. Xenolinguistics) — verbal reasoning, pattern abstraction, hypothesis testing in language space.
- Behavioral / persistence (e.g. Squat Runner) — grit, effort regulation, stress response. Performance decay curves under sustained load are the primary signal.
- Calibration micro-tasks (Software Developer / Product Manager / Marketing Specialist / UX Designer alpha; expanding to 12 careers) — 10–18 minute AI-graded simulations of actual occupational tasks. Informed by Webster (2020)'s meta-analytic SJT validity findings, our calibration assigns situational-judgment scenarios the highest weight (0.35–0.40) — the weighting is our design choice, not a value reported by Webster.
The full 48-parameter to game-family weighting matrix is the canonical specification of what we measure. It is open-source and is reproduced in Appendix A.
3.3 The scoring pipeline (per game → per parameter → PCS)
Every game emits a raw game score (a numerical performance measure on whatever scale that game uses — for some games 0–100, for others 0–1000, for others 0–5000+). The scoring pipeline transforms these heterogeneous raw scores into the canonical PCS in five passes:
- Scale normalization — each game's raw score is normalized to 0–100 using the game-specific scale hint declared in the scoring-framework specification. Note this is scale-aware normalization, not blanket division by a fixed constant — an internal audit caught a silent capping bug from blanket rescaling and fixed it (see §7 Limitations for honest disclosure).
- Parameter projection — each normalized game score contributes to one or more of the 48 parameters according to per-game weight vectors. Weights sum to 1.0 per game (validated by an automated registry check).
- Reliability weighting — when a participant plays multiple games that contribute to the same parameter, contributions are aggregated using a reliability-weighted exponential moving average (planned for a later revision; v1.0 of this whitepaper uses a simpler fixed-α moving average and discloses this in §7).
- Parameter normalization and aggregation — parameter-level scores are aggregated into the eight category scores (Cognitive, Technical, Interpersonal, Behavioral, Personality, Values, Career, Physical) via the category-weight matrix; category scores then aggregate into the PCS via the top-level matrix. Both matrices are open in the same scoring-framework specification.
- Coherence triangulation gating (shadow-mode at v1.0) — the platform's anti-faking design: PCS contributions are gated on three coherent signals: (a) game performance ≥ threshold, (b) repeated category choice across ≥ 2 sessions, (c) sentiment neutral-or-positive (sentiment captured via a per-session Likert self-report). At v1.0 the validator runs in shadow mode for four weeks before gating — the design is intended for the v2 peer-reviewed manuscript, not for v1's marketing or counselor decisions.
3.4 The labor-market overlay
PCS scores are converted to career recommendations through a 1,041-occupation taxonomy. The pipeline is: (i) compute the participant's parameter profile, (ii) compute the cosine similarity between that profile and each occupation's O*NET requirements vector (cross-walked through ISCO-08 / NOC 2021 / KSCO 2017 / CNP 2017 as appropriate to the user's locale), (iii) rank occupations by similarity, (iv) overlay WEF Future of Jobs 2025 industry-displacement projections to attach an AI-resilience score and AI-collaboration potential to each surfaced career, (v) surface the top-N recommendations.
The labor-market overlay does not modify the PCS score itself. It modifies which careers are shown and with what AI-era context. This separation is deliberate — the PCS is intended to be a stable measure of the participant, not a moving target tied to year-by-year labor-market projections.
3.5 Coherence triangulation and anti-faking design
The platform's deliberate anti-faking design is the integration of three independent signals before a parameter score is gated as "verified":
- Performance signal — behavioral score on games whose mechanics make the diagnostic behavior cognitively expensive to fake (e.g. timed cognitive load, multi-step planning trees that require working memory, situational judgment items with no obvious "correct" choice).
- Repeated choice signal — over multiple sessions or game instances, the participant's category-level pattern (e.g. consistent preference for Investigative-domain games when given a choice) corroborates the inventory layer.
- Sentiment signal — per-session Likert self-report of how the game felt, captured via a brief survey at session end. The platform uses a Likert self-report rather than facial recognition, for cost and consent reasons.
Coherence triangulation is a core design principle of the platform, articulated since its earliest design work. It is not yet gating PCS contributions in v1.0; it runs in shadow mode and surfaces dashboard signals to the platform operators for tuning. The peer-reviewed v2 manuscript will publish gating decisions.
3.6 Reproducibility
Every claim in this whitepaper that depends on a number — a sample size, a correlation, a coverage percentage — is grounded in a specific data pull from the platform's internal data store, on a stated date, at a stated sample size. Sections 4 and 5 (Validation Studies and Results) present these numbers together with their sampling caveats.
Researchers who wish to inspect the basis for any specific claim are welcome to contact research@talentopian.com.
4. Validation Studies
This section reports what we actually know — and what we explicitly do not yet know — about the psychometric properties of the PCS instrument. We frame the section as the SIOP Principles (2018) require for a technical report of this kind: pre-registered claims, transparent disclosure of measurement gaps, and an honest validation roadmap rather than an overstated coefficient register.
4.1 What was measured for this report
Preliminary internal-structure analysis was conducted on an early pilot cohort — a single computed PCS snapshot per user, drawn from an initial backfill wave. Two properties of this corpus bound what it can support:
- It contains a single snapshot per user, so PCS-snapshot-level test-retest reliability is not estimable from it.
- Parameter coverage is ragged and game-dependent: no single parameter is present for every user, and most parameters are observed on only a subset of the cohort.
- Such repeated-measure data as exists is dominated by internal test activity rather than real end-user re-assessment, so it cannot be interpreted as genuine retest signal.
Because of these limits, we treat all findings below as directional only. They are used to inform framework consolidation, not to assert validity.
4.2 What we are claiming as validity evidence — and what we are not
| Claim type | What we report in v1 | What we do not report in v1 (and why) |
|---|---|---|
| Internal structure | Preliminary inter-parameter correlations on the pilot cohort (§4.3) | Confirmatory factor structure (sample too small) |
| Test-retest reliability | Not reported | No real end-user has been re-assessed. The pilot wave was a single snapshot per user. Pre-registered for a future data-collection wave (§4.4). |
| Concurrent validity vs O*NET / Holland | Not reported | No RIASEC or Holland-themed instrument is currently administered to PCS users. Building the concurrent matrix requires fielding a validated RIASEC measure at onboarding. Pre-registered for a future collection wave. |
| Criterion validity (PCS → job fit) | Not reported | No ground-truth criterion exists yet. The recommendation is the platform's output; accuracy needs an independent criterion (self-reported aspiration, counselor-rated fit, or follow-up outcome). Pre-registered for a future collection wave. |
| Convergent validity (PCS vs self-rated soft skills) | Not reported | A self-report soft-skills schema exists but is not yet populated. This no-new-instrument convergent criterion is pre-registered for a future collection wave. |
This is the honest table. A whitepaper that filled in any of those four rows on an under-powered single-snapshot pilot corpus would be guilty of one of the failure modes the SIOP Principles explicitly warn against (overstated coefficient on under-powered data). We refuse to do that.
4.3 What we can defensibly say (internal structure, preliminary)
Directionally, the pilot data are consistent with a coherent, correlated multi-construct instrument rather than a random parameter set or a single collapsed general factor. The most actionable finding is a collinearity signal: several nominally distinct interpersonal parameters — emotional intelligence, negotiation, social skills, and collaboration — are correlated tightly enough to suggest they share a single underlying factor at the game-emission layer. This is exactly what an honest internal-structure analysis is supposed to surface, and it motivates a dimension-reduction step in a future framework revision as the sample grows.
4.4 Pre-registered validation roadmap
To prevent this section from being read as a marketing brochure dressed up in academic prose, we register the following four-step validation roadmap now, before any of the data are collected.
Test-retest collection. Invite the existing pilot users (plus newly onboarded users) to replay the core game set several weeks later. Compute parameter-level retest correlations on real users only. The pipeline already captures repeated attempts; no new instrumentation is required.
Convergent collection (no new instrument). Drive completion of the existing end-of-session soft-skills self-report (already in the schema; currently unpopulated). Compute per-parameter convergent correlations against the matching self-report dimension.
Concurrent collection. Field a short validated RIASEC measure at user onboarding (one of: O*NET Interest Profiler short form, or Holland RIASEC-30). Compute the PCS-dimension to RIASEC-theme correlation matrix. Report convergent (within-theme) and discriminant (across-theme) coefficients.
Criterion collection. Capture two ground-truth criteria at onboarding: (a) self-reported aspiration career (free-text + occupation tag), (b) for counselor-mediated users, counselor-rated fit. Compute top-1 and top-3 recommendation hit-rates against each criterion. Report by-locale and by-counselor-presence stratifications.
The above is a pre-registered protocol, not a forecast of results. Whatever the coefficients turn out to be, they will be reported in v1.x or v2 as collected, with the same honest scope disclosure.
4.5 Recommendation distribution — descriptive only
Independent of the missing-criterion issue, we flag a descriptive concern in the corpus: the top-1 recommended career is over-concentrated on a single engineering-adjacent occupation. In a sample this small, such a mode is most likely a sample-composition peculiarity (the early cohort skews engineer-heavy) rather than a stable instrument property, but it could also reflect default recommendation behavior in the absence of strong signal. We do not, in v1, distinguish between these explanations; the concern is noted so that future, better-balanced samples can resolve it.
4.6 Conclusion of §4
The PCS instrument is currently in a pilot state with respect to all four classical validity coefficients. The right interim claim is the one we make in this paper: that the instrument's internal structure on an early pilot cohort is plausibly behaved, with some collinearity clusters flagged for consolidation; that the validation roadmap is pre-registered and the data-collection mechanics already exist (no new platform engineering required); and that the v2 peer-reviewed manuscript will be the appropriate venue for the four coefficients above. Anything stronger than that interim claim would, on this data, be a stretch.
5. Results
At v1 the platform's evidence base is an early pilot cohort, and the results are correspondingly preliminary. We therefore report them qualitatively rather than as a coefficient register.
5.1 Internal structure (preliminary)
Inter-parameter correlations across the pilot cohort are consistent with a coherent, correlated multi-construct instrument — neither random noise nor a single collapsed general factor. The most actionable signal is a collinearity cluster among the interpersonal parameters (Emotional Intelligence, Negotiation Skills, Social Skills, Collaboration Skills), which correlate tightly enough to behave as a single latent factor with sub-facets. This is consistent with the SCCT literature finding that interpersonal competency dimensions tend to load on a small number of latent factors when measured behaviorally, and it is flagged for dimension-reduction in a future framework revision.
5.2 What this section deliberately does NOT contain
The four classical validity coefficients (test-retest reliability, concurrent validity against O*NET RIASEC, criterion validity, and convergent validity against self-rated soft skills) are not reported, for the reasons enumerated in §4.2: the pilot corpus is a single snapshot per user, no RIASEC instrument or ground-truth criterion has yet been collected, and the soft-skills self-report schema is not yet populated. Each has a corresponding pre-registered collection step in §4.4. Formal, adequately powered internal-structure and validity analyses will be reported through the validation roadmap and the planned v2 peer-reviewed manuscript.
6. Discussion
The PCS instrument, as evidenced in §4–§5, is in a pilot state with respect to the four classical validity coefficients (test-retest reliability, concurrent validity, criterion validity, convergent validity). The discussion that follows is shaped by that constraint: it is about what the design and the preliminary internal-structure data together imply, and what they emphatically do not imply.
6.1 What we can defensibly say from §5
Three claims survive the small-N caveat.
First, the framework is internally well-behaved as a starting point. The inter-parameter correlations on the pilot cohort are consistent with a coherent multi-construct instrument; they are not the noise pattern of a randomly assembled parameter set, and they are not so collapsed that the framework reduces to a single general factor. The picture is what you would expect from an instrument that measures correlated but distinguishable competencies, on a sample too small to estimate exact loadings.
Second, the collinearity clusters we surfaced are diagnostic, not destructive. The interpersonal cluster (Emotional Intelligence, Negotiation Skills, Social Skills, Collaboration Skills) is exactly the kind of finding the Principles (SIOP, 2018) expect a Section 4 to surface in pilot work: nominally distinct constructs that share a game-emission backbone. The honest move is to consolidate at the framework level (a single "interpersonal-effectiveness" factor with sub-facets) rather than to maintain four nominally separate parameters whose individual scores convey no marginal information. This consolidation is queued for a future revision, and the dimension-reduction analysis will be reported in v1.x or v2 of this whitepaper.
Third, the design satisfies the SCCT precondition for behavioral assessment. A participant who plays our games is, by construction, engaged in goal-directed mastery-experience tasks of the kind Bandura's framework treats as the formative input to self-efficacy. Whether or not we have yet measured self-efficacy validly (we have not), the substrate of the measurement is the substrate the theory calls for. This is a non-trivial design property that the dominant inventory-based alternatives (Strong, Self-Directed Search, Big Five batteries) cannot claim.
6.2 What we cannot say (yet)
We cannot claim that PCS scores are stable across repeated administrations on the same individual. We cannot claim that PCS dimensions converge with their nominal Holland-RIASEC analogues at any specific magnitude. We cannot claim that the platform's top-1 recommendation matches an individual's actual career fit at any specific hit-rate. Anything strongly claimed here on v1 data would be a category error. The pre-registered roadmap in §4.4 exists precisely to convert these "cannot say"s into "now reportable"s over the next four collection waves.
6.3 The labor-market overlay claim, examined
The PCS-to-career bridge — the 1,041-occupation taxonomy with WEF Future of Jobs 2025 industry-displacement overlay — is not itself a validity claim about the underlying PCS. It is a labor-market projection attached to the recommendation surface. The two layers should be evaluated separately. A future evaluator can challenge our scoring pipeline (§3.3) on psychometric grounds or challenge our WEF overlay on labor-economics grounds without the two challenges entangling. We have deliberately separated them in the architecture for this reason. v1 does not present labor-market accuracy claims because the underlying instrument validity has not been established yet.
6.4 The recommendation over-concentration, contextualized
The top-1 over-concentration flagged in §4.5 is a sample-composition artifact more than it is an instrument finding. The early corpus oversampled engineering-adjacent users (a known property of an early-stage product with a developer-tilted founding cohort). The same instrument applied to a balanced sample would almost certainly produce a different top-1 distribution. We flag the finding honestly not because we believe it is the steady-state distribution, but because honest disclosure of the v1 corpus's peculiarities is required by the SIOP Principles technical-report standard. Future versions will publish the top-1 distribution stratified by cohort source (organic vs counselor-mediated vs employer-sponsored vs research-pilot) so that readers can interpret each separately.
6.5 Cross-cultural deployment notes
The platform serves four locales out of the gate (EN-CA, KO-KR, FR-CA, ES-ES). This is a capability claim — the UI, game instructions, results, and counselor materials are all translated — but it is not yet a validity claim across locales. Two specific issues warrant pre-emptive disclosure:
- Norm samples are not yet locale-stratified. All PCS snapshots in the v1 pilot corpus were generated by users we cannot stratify on locale post-hoc with high confidence. Locale-specific norms (which the v2 manuscript would require for cross-cultural validity claims) are queued for a later phase of the platform's data-collection plan.
- Cultural-context corrections for situational judgment items are minimal. The situational-judgment scenarios in the MicroTask layer are currently English-source with localized translations rather than locale-native scenarios. A counselor-administered ethical scenario about a workplace conflict in Korea may, by virtue of cultural-context particulars, score differently than the same translated scenario in Quebec. We do not attempt to correct for this in v1; v2 will either (a) report locale-stratified SJT scoring patterns or (b) commission locale-native scenario authorship.
6.6 Comparison to existing instruments — what we are and are not
We are not the first game-based career assessment. Three reference points are worth naming, briefly:
- Pymetrics / Harver (acquired by Harver in 2022): the dominant comparable. 12 games, ~25 minutes total, ~98% completion claim. Predominantly enterprise-licensed for personnel selection rather than self-directed career exploration. Their validation work has been more thoroughly published than ours and is the de facto benchmark for the game-based assessment subfield.
- HireVue Game-Based Assessments: 11-game battery, similar enterprise-only distribution.
- Arctic Shores: 10–18-game adaptive batteries, similarly enterprise-focused.
Two structural differences are worth noting. First, all three of those comparables are enterprise-only; the participant cannot self-serve. We deliberately ship as a freemium self-directed platform because the population that most needs the assessment (career changers in their thirties, first-generation post-secondary students, immigrants with foreign credentials) is the population least likely to be on the receiving end of an enterprise-licensed administration. Second, none of the comparables publish their underlying competency framework openly. Ours is open source. This is a deliberate posture: we want our framework to be challengeable, reproducible, and incrementally improvable in public rather than gated behind a publisher's NDA.
We are not claiming superiority on validity (we have not yet earned that claim). We are claiming distinct posture and distinct distribution model.
6.7 What this paper enables, in practical terms
For a school counselor, a vocational psychologist, or an HR partner reading this paper:
- The platform is appropriate for first-pass evidence generation in a career-exploration conversation. It is not appropriate as a high-stakes screening filter (admission, hiring, licensure) at v1.
- The PCS score and the top-N career recommendations should be presented to the participant as one input among several, not as a verdict. The hosting layer is designed to surface honest scope disclosures alongside results, including the AI-resilience labels from the WEF overlay.
- For Korean counselors specifically, the platform produces a counselor-shareable report and a parent-shareable summary with a counselor-attribution footer. The Parent Portal flow (Phase 2 LIVE) is one of the artifacts this whitepaper documents.
- For institutional buyers (school boards, employer sponsorships, government training programs): the freemium model provides procurement-aligned distribution paths. The validation evidence in this whitepaper supports a "first-pass evidence-generation tool" positioning, not a "validated assessment instrument" positioning. v2 will support the stronger positioning if and when the validation roadmap (§4.4) produces the coefficients.
7. Limitations and Future Research
This section is the honest scope register. It is shorter than §2 or §3 by design — limitations should be enumerable and specific, not buried in qualifying prose.
7.1 Sample-size limitations (v1.0 corpus)
- An early pilot cohort. At this sample size, confidence intervals on any single correlation are wide enough that almost any individual correlation is statistically uninformative on its own. We treat the §5 findings as directional, not confirmatory, and we say so explicitly in §4.6.
- Backfilled, not organic. The pilot users are not a random cross-section of our target population. They are an engineering-skewed early cohort. Generalization to teen / career-changer / parent personas requires fresh organic data.
- Single-snapshot per user. No test-retest is possible from this corpus. A pre-registered collection wave (§4.4) addresses this.
7.2 Measurement-instrument limitations
- Framework consolidation in progress. Some game-emitted sub-parameters have not yet been consolidated into the canonical taxonomy; internal-structure statistics are computed against the production parameter set as measured. Consolidation is a future framework-revision work item.
- Collinearity clusters not yet consolidated. Four parameters in the interpersonal cluster are highly collinear. Until consolidation, the effective dimensionality of the framework is lower than its nominal size. This is a v2 framework-revision work item.
- Scale-aware normalization bug history. An earlier normalization step applied a blanket rescaling that silently capped the scores of games using larger raw-score ranges. This was caught in internal review, fixed, and replaced with scale-aware normalization that respects each game's declared score range. Versions of the platform predating that fix produced uncorrected scores; this whitepaper's data uses the corrected pipeline.
- Job-saturation calculation formerly constant. An earlier version of the labor-market data engine fell back to a hardcoded constant for every industry's saturation input, over-relying on the market-forecast half of the AI-penetration formula. This was caught in internal review, fixed, and redeployed so that each industry's AI-penetration reflects a real, factually grounded saturation signal; we discuss the fact-verification in §3.4.
We disclose these history items because the SIOP Principles (2018) explicitly require that "any material correction to scoring procedures during the validation cycle" be reported. We are confident the current pipeline is correct; we are also honest that earlier deployed versions had specific known bugs.
7.3 Validation-design limitations (deferred to v1.x / v2)
- No RIASEC instrument administered to any of the pilot cohort. Pre-registered for a future collection wave.
- No ground-truth criterion (self-reported aspiration career or counselor-rated fit) collected from any of the pilot cohort. Pre-registered for a future collection wave.
- No IRB approval for v1 (self-published technical report). v2 peer-reviewed manuscript will require IRB.
- No external researcher partnership for v1. v2 manuscript will require an external partner; that outreach is planned for a later phase.
7.4 Deployment and operational limitations
- Free-tier reveal is partial. The platform shows users a paywall-gated subset of report sections at the free tier. The "movie-trailer" preview model surfaces one honest insight above the blur. We are aware this is a commercial-product compromise; for the validation-evidence audience, the full report is available to the platform operators for verification.
- Coherence Triangulation is shadow-mode. The anti-faking gating described in §3.5 runs in observation-only mode at v1. It is not gating PCS contributions yet. v2 will publish the gating decisions and the impact on reported scores.
- Conversion and experimentation telemetry are new. The platform's experimentation and conversion telemetry are recent additions and have not yet accumulated enough data to influence the report layer or the validation reasoning. They will not feed into v2 validity findings.
7.5 Acknowledged risks to the validation roadmap
We pre-register two concrete risks to the §4.5 roadmap so that future readers can evaluate whether they materialized.
- Test-retest collection may suffer from selection bias. Users who agree to a re-assessment wave are likely to be the more engaged subset of the corpus. If retest correlation looks high in v1.x, we will need to disclose the response-rate and engagement-stratified sub-analysis.
- Concurrent-validity collection introduces a new instrument. Adding a RIASEC self-report at onboarding means the platform's onboarding-completion rate may decline (a known cost of inserting any new measurement step). We will report the pre- and post-instrument completion rates and any composition shifts.
7.6 What we will and won't do in v1.x patch releases
A "patch" v1.x release is a publication of additional data or a correction to v1.0 that does not require a fundamental revision of theoretical framework or methodology. We will issue v1.x for:
- New collection-wave data (§4.4 steps 1–4 as they land).
- Corrections to descriptive statistics if errors are found in the underlying analysis.
- Additional locale-stratified analyses as the corpus accumulates non-EN-CA users.
We will not issue v1.x for:
- Reframing of any of the four claims declared in §1.2.
- Revision of the theoretical framework in §2.
- Addition of new validity coefficient classes beyond those pre-registered in §4.4 (those go in v2).
8. Practical Applications for Counselors and Educators
This section is the practitioner-facing summary. If you read only one section of this whitepaper, read this one. We assume a reader who is a school counselor, vocational psychologist, career-counseling educator, or institutional buyer evaluating the platform.
8.1 What you can do with PCS results today
The PCS output for a participant comprises three artifacts:
- A categorical PCS profile across the eight framework categories (Cognitive, Technical, Interpersonal, Behavioral, Personality, Values, Career, Physical), each on a 0–100 scale.
- A ranked list of recommended careers drawn from the 1,041-occupation taxonomy, each tagged with: O*NET source competencies, an AI-resilience score derived from the WEF Future of Jobs 2025 overlay, and an AI-collaboration potential indicator.
- A shareable result card (1,200 × 630 OG format for desktop sharing; 1,080 × 1,920 Instagram Story format for mobile) suitable for the participant's social channels.
For a counseling conversation, the recommended interpretive frame is:
- Treat the categorical PCS profile as a conversation starter, not a verdict. "Your results show strong patterns in [category]; tell me about a time you felt that way" is a more productive opening than "your test result is X."
- Treat the ranked career list as a menu of options, not a prescription. The participant's family context, financial constraints, geographic flexibility, and self-efficacy beliefs (which the platform cannot fully measure) all moderate which option on the menu is actionable for them.
- Treat the AI-resilience / AI-collaboration tags as a future-orientation prompt: "this career has high AI-displacement risk; here are the adjacent careers in our taxonomy that share your strengths but lower the risk."
8.2 What the platform does not replace
- Clinical interviewing skills. The platform produces evidence; the counselor produces interpretation. A counselor's training in motivational interviewing, narrative career counseling, and the SCCT contextual-supports / barriers framework is exactly what bridges PCS evidence to participant action.
- Family / cultural / financial context conversations. The platform is not yet equipped to surface "your parents will be disappointed in X" or "you can't afford the credentialing for Y." A counselor's relational presence remains primary.
- High-stakes decisions. Admission, hiring, professional licensure, and equivalent gatekeeping decisions should not use v1 PCS scores as a screening filter. We position v1 explicitly as a first-pass evidence-generation tool, not a validated screening instrument.
8.3 Workflow integrations available at v1
- Counselor dashboard (Counselor Plus tier). Includes cohort management, per-student PCS view, and session notes.
- Parent Portal (Phase 2 LIVE). The counselor generates an invite code; the parent receives an email with the link; the parent sees a read-only PCS overview + a counselor-attribution footer + a crisis-resources footer where applicable.
- PCS Share Card for participant self-sharing (Story 9:16 + OG 1.91:1 formats).
- MicroTask runner for AI-graded 10–18-minute occupational simulations (alpha 4 careers, expanding to 12).
- Chrome extension for counselor in-session augmentation (live coaching cues, SOAP note assistance).
8.4 What we ask of you, as a reader
If you are a researcher: the open-source scoring framework and the pilot corpus are available for inspection. Independent validation work would be genuinely welcomed; contact research@talentopian.com.
If you are a counseling practitioner: try the platform on yourself or a willing colleague before recommending it to a participant. The interpretive sensitivity required is a counselor-skill, not a platform-output.
If you are an institutional buyer (school board, employer, government training program): we offer procurement-aligned paths for institutional deployment. We are happy to engage in pilot conversations; the contact is partnerships@talentopian.com.
If you are a parent who received an invite: the link is read-only. Your reply to the invite email goes to the counselor who invited you (not to a generic support address) — that conversation, with the counselor, is the actionable next step.
9. Acknowledgments
The 48-parameter framework presented in this whitepaper was developed by Talentopian Research. Independent academic review of the framework and validation design is welcomed at research@talentopian.com.
10. References
References accumulate section by section. v1.0 starting set (working list — formal entries will move to the companion methodology-and-citations document as it is built out):
- Holland, J. L. (1997). Making vocational choices: A theory of vocational personalities and work environments (3rd ed.). Psychological Assessment Resources.
- Lent, R. W., Brown, S. D., & Hackett, G. (1994). Toward a unifying social cognitive theory of career and academic interest, choice, and performance. Journal of Vocational Behavior, 45(1), 79–122.
- Lent, R. W., & Brown, S. D. (2002). Social cognitive career theory. In D. Brown & Associates (Eds.), Career choice and development (4th ed., pp. 255–311). Jossey-Bass.
- McCrae, R. R., & Costa, P. T. (2008). The five-factor theory of personality. In O. P. John, R. W. Robins, & L. A. Pervin (Eds.), Handbook of personality (3rd ed., pp. 159–181). Guilford Press.
- Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). On the structure of educational assessments. Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62.
- O*NET Resource Center. (2024). O*NET 28.x database and content model documentation. https://www.onetcenter.org/
- Shute, V. J. (2011). Stealth assessment in computer-based games to support learning. In S. Tobias & J. D. Fletcher (Eds.), Computer games and instruction (pp. 503–524). Information Age.
- Society for Industrial-Organizational Psychology. (2018). Principles for the validation and use of personnel selection procedures (5th ed.). https://www.siop.org/principles
- Webster, B. D. (2020). Situational judgment tests and personnel selection: A meta-analytic update. Personnel Assessment and Decisions, 6(2), 1–18.
- Whiston, S. C., Li, Y., Mitts, N. G., & Wright, L. (2017). Effectiveness of career choice interventions: A meta-analytic replication and extension. Journal of Vocational Behavior, 100, 175–184.
- World Economic Forum. (2025). Future of Jobs Report 2025. https://www.weforum.org/publications/the-future-of-jobs-report-2025/
11. Appendix A: 48-Parameter Framework Specification
The canonical machine-readable specification of the 48 parameters is maintained in the open-source scoring-framework specification. The parameters span eight categories (the Physical category was added in a later framework revision, expanding the original seven-category structure). We reproduce the canonical names below; per-parameter scoring rubrics, game-emission weights, and framework citations live in the machine-readable spec.
A.1 Cognitive (9 parameters)
Working Memory · Pattern Recognition · Abstract Reasoning · Sustained Attention · Cognitive Flexibility · Processing Speed · Visual-Spatial Reasoning · Verbal Reasoning · Quantitative Reasoning.
Theoretical lineage: O*NET Abilities table; Big Five Openness (where novelty-seeking on optional puzzle branches signals); Holland Investigative.
A.2 Technical (8 parameters)
Systems Thinking · Multi-Step Planning · Hypothesis Testing · Diagnostic Reasoning · Technical Knowledge Recall · Tool / Interface Adaptability · Workflow Optimization · Quality / Precision Control.
Theoretical lineage: O*NET Skills table (Technical Skills subgroup); Mislevy ECD task models for procedural competencies.
A.3 Interpersonal (5 parameters)
Emotional Intelligence · Negotiation Skills · Social Skills · Collaboration Skills · Communication Effectiveness.
Theoretical lineage: SCCT outcome expectations + self-efficacy for interpersonal tasks; Webster (2020) SJT meta-analytic findings on interpersonal-domain SJT validity.
Note: the empirical collinearity finding in §5 (a tight cluster across Emotional Intelligence, Negotiation, Social, Collaboration) suggests this category is functionally a single latent factor with sub-facets at v1. v2 will report the dimension-reduced score in addition to the four nominal sub-scores.
A.4 Behavioral (10 parameters)
Persistence Under Failure · Stress Response · Decision-Making Under Time Pressure · Risk Tolerance · Frustration Tolerance · Goal Persistence · Effort Regulation · Learning Curve · Behavioral Insights · Emotional Regulation.
Theoretical lineage: Shute (2011) stealth assessment of grit and effort regulation; Big Five Conscientiousness (persistence sub-facet); SCCT self-efficacy via mastery-experience accumulation.
A.5 Personality (6 parameters)
Openness · Conscientiousness · Extraversion · Agreeableness · Neuroticism · Personality Type (MBTI-axis).
Theoretical lineage: McCrae & Costa (2008) Five-Factor Model with explicit acknowledgement that the Big Five scores reported here are behavioral-trace inferences, not validated Big Five inventory scores. The MBTI-axis label is a signature, not a validated MBTI score — see §2.5 disclosure.
A.6 Values (2 parameters)
Ethical / Moral Reasoning · Cultural Background Sensitivity.
Theoretical lineage: Holland Values dimension; SJT-based ethical-reasoning scoring per Webster (2020).
A.7 Career (4 parameters)
Career Interests (Holland-aligned) · Career Adaptability · Career Decision-Making Confidence · Career-Related Self-Efficacy.
Theoretical lineage: Holland RIASEC interest mapping; Lent & Brown (2002) career self-efficacy formulation; Maggiori, Rossier, & Savickas (2017) Career Adapt-Abilities Scale framework (we draw on the framework, not the validated inventory itself).
A.8 Physical (4 parameters)
Hand-Eye Coordination · Physical Stamina · Auditory Processing · Environmental Awareness.
Theoretical lineage: O*NET Abilities table (Psychomotor, Physical, and Sensory ability groups); the Physical category was added in a later framework revision, expanding the original seven-category structure to eight.
A.9 Production vs canonical: framework consolidation
The canonical 48-parameter taxonomy above is the measurement intent. In production, some game-emitted sub-parameters have not yet been consolidated into the canonical set; consolidation will rename, merge, and prune redundant keys without invalidating the canonical 48. Internal-structure statistics at v1 are computed against the production parameter set as measured, and future v1.x will report against the post-consolidation parameter set.
— End of v1.0 (2026-06-22). A companion methodology-and-citations document is forthcoming. Future v1.x updates per the rules in §7.6.
Web edition of the canonical whitepaper. · All research · Talentopian home