1. Purpose & intended use
AssessFit provides pre-employment assessments that help employers compare candidates on job-relevant abilities, knowledge, work styles, and judgment. Scores are designed to be one structured input among several in a hiring decision — alongside interviews, work history, and references. They are not a medical, clinical, or diagnostic instrument, and AssessFit is designed so that the final hiring decision is always made by a person, not an algorithm.
This manual describes the methodology behind the platform as it is built today. It follows the structure recommended by the SIOP Principles and the AERA/APA/NCME Standards for documenting selection procedures (see References).
2. Assessment architecture
The library contains 100 tests built on a bank of 854 items, spanning five domains:
| Domain | Measures | Format |
| Cognitive ability | Numerical, verbal, logical, spatial & mechanical reasoning, working memory, processing speed, attention | Multiple-choice, adaptive |
| Job skills | 53 domain-knowledge tests (finance, IT, marketing, operations, healthcare support, trades, and more) | Multiple-choice, adaptive |
| Situational judgment | Ethics, prioritization, leadership, customer scenarios, crisis response | Scenario multiple-choice |
| Work styles | Conscientiousness, composure, integrity, teamwork, adaptability, and 10 further scales | Likert self-report with reverse-keyed items |
| Language | Business English (CEFR-aligned), written communication, proofreading | Multiple-choice |
Each multiple-choice test is a 9-item bank tagged by difficulty (3 easy / 3 medium / 3 hard) with stable item identifiers. Employers may combine up to five tests per role; the platform's 1,000+ occupation catalog recommends an evidence-oriented bundle per job title.
3. Test development
Items are authored against explicit rules, then machine-validated before release:
- One defensible key. Every item must have exactly one unambiguously correct answer and three plausible-but-wrong distractors. Items whose correctness could be argued are rejected at review.
- Concepts over trivia. Items test stable concepts, not version-specific or expiring facts.
- Cultural neutrality. Contexts are universal workplace situations; cognitive items are language-light where the construct allows.
- Balanced key positions. Correct answers are distributed nearly evenly across option positions (current bank: within ±1% of uniform), so position-guessing strategies confer no advantage.
- Structural validation. An automated validator enforces: unique item identifiers across the whole library, exact 3/3/3 difficulty composition, four options per item, in-range answer keys, and rejection of any item whose correct answer text duplicates another option.
4. Administration & adaptive design
Candidates receive a randomized two-stage adaptive form: stage one draws 3 medium items; candidates who answer at least 2 correctly advance to 3 hard items, others receive 3 easy items. No two candidates see identical forms, and item order is randomized. Each test carries a per-test time budget; the total session budget is shown to the candidate before starting.
Answer keys never leave the server: the client receives question text and options only, and all correctness decisions happen server-side at submission.
5. Scoring
- Multiple-choice tests are scored server-side from correct counts across both adaptive stages.
- Work-style scales are scored 0–100 with reverse-keyed items inverted before aggregation; unanswered statements score at the scale midpoint rather than being imputed favorably.
- Composite score. Employers set transparent weights across CV match, skills tests, and work-style measures; the composite is a weighted mean, and the weights are visible — no black-box factor is applied.
- Integrity verdict. Behavioral session signals (see §10) produce a separate integrity score and verdict; integrity never silently alters ability scores — it is reported alongside them for human review.
6. Norming
Percentiles are computed against empirical norm groups: real completed assessments for the same role template, using the midpoint percentile method. Two rules keep this honest:
- Every percentile is displayed with the true size of its norm group ("n = 40"), never an inflated aggregate.
- No synthetic or purchased norms are used. Where a norm group is still too small to be meaningful, raw scores are shown without percentile claims.
The platform-wide completion count is public and live: — assessments completed to date.
7. Reliability & item quality
Every item in the bank is monitored continuously against real candidate data:
| Statistic | Definition | Action threshold |
| Difficulty (p) | Proportion answering correctly | p > .95 flagged too easy; p < .20 flagged too hard |
| Discrimination (rpb) | Point-biserial correlation between item correctness and the candidate's test score | r < .05 at n ≥ 20 flagged weak; negative r at n ≥ 10 flags a suspect answer key for immediate review |
| Response time | Mean seconds per item | Outliers reviewed for ambiguity or leakage |
Flagged items are reviewed and retired or replaced; because tests are item banks, retirement does not interrupt administration.
Why we don't quote a coefficient alpha yet. Classical internal-consistency coefficients assume every candidate answers the same items; our adaptive forms deliberately violate that. Rather than publish an alpha computed on unsuitable data — or a number copied from a brochure — we publish the item-level statistics above, live. A platform-level reliability analysis will be published when the response matrix supports it, with its sample size.
8. Validity
Content validity
Each of the 1,000+ occupation titles in the catalog maps to one of 30 occupation profiles, and each profile's recommended test bundle was constructed by matching job-relevant constructs to tests (e.g., warehouse roles → logistics operations, safety, attention; care roles → patient scenarios, caregiving knowledge, integrity). Employers can inspect and edit every bundle — nothing is hidden.
Criterion validity — your own evidence
AssessFit ships a built-in predictive-validity dashboard: employers rate the on-the-job performance of the people they hired, and the platform computes the correlation between assessment scores and those ratings — overall and per test, always displayed with its sample size. This is a local validation study, the form of evidence the EEOC Uniform Guidelines treat as strongest, produced continuously and at no extra cost.
Commitment. As anonymized cross-employer outcome data accumulates, we will publish platform-level criterion-validity summaries with sample sizes and confidence intervals — in this manual, not in a press release. Benchmark from the meta-analytic literature: general cognitive ability predicts job performance at roughly r ≈ .3–.5 (Schmidt & Hunter, 1998).
9. Fairness
- Blind scoring by construction. AssessFit does not collect race, religion, gender, or age from candidates, and no such variable exists anywhere in the scoring pipeline.
- No face analysis. No emotion inference, no video scoring, no pseudo-scientific signals.
- Language-light cognitive items reduce construct-irrelevant language load where the construct allows.
- Employer responsibility. Under the EEOC Uniform Guidelines, adverse-impact analysis is conducted on the employer's applicant flow. AssessFit's structured, job-related scores make that analysis possible; employers remain responsible for monitoring their own selection rates, and our reports are designed to support that.
10. Integrity & security
Signals collected during a session (with candidate notice) include tab/window switches, paste and copy attempts, context-menu use, response-time anomalies (implausibly fast correct answers), and automation markers. Signals are weighted into an integrity score with three verdicts: clean (≥ 80), review (55–79), and high-risk (< 55), each listing its contributing flags so a human can judge.
- Answer keys exist only server-side; scoring is entirely server-authoritative.
- Invitation links are single-session tokens; candidate sign-in uses single-use, 15-minute magic links.
- Randomized adaptive banks mean leaked answers have limited value — and leakage shows up in item statistics (§7), triggering retirement.
11. Data protection
- All traffic is TLS-encrypted; the application enforces HTTPS and a canonical host.
- Candidate CV files are stored outside the web root and are never web-served to the public.
- Candidates can erase their profile and data via a self-service endpoint in their dashboard.
- Assessment responses are stored for scoring, norming, and item-quality monitoring as described above; aggregate statistics are anonymous.
12. Appropriate use & limitations
- Scores should never be the sole basis for a hiring decision.
- Percentiles from small norm groups are labeled and should be read cautiously.
- Work-style scales are self-report instruments; they describe typical tendencies, not guarantees of behavior.
- Integrity verdicts are decision-support flags, not accusations; a review verdict warrants a conversation, not an automatic rejection.
- Tests are delivered in English today; scores for candidates assessed in a second language partly reflect language proficiency. Use the dedicated language tests to measure that explicitly.
13. References & standards followed
- AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.
- SIOP (2018). Principles for the Validation and Use of Personnel Selection Procedures (5th ed.).
- EEOC et al. (1978). Uniform Guidelines on Employee Selection Procedures.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124(2), 262–274.
- ISO 10667 — Assessment service delivery.
Questions about methodology, or requests for the underlying statistics: research@assessfit.com. This is a living document — the version and date at the top change whenever the platform's methodology does.