Technical manual · v1.0 · July 2026

How AssessFit assessments actually work

The document procurement teams ask for: development methodology, administration, scoring, norming, reliability, validity, fairness, and security — written so every statement is verifiable against the running platform. Where a figure is measured live, we show the live figure. Where something is a commitment rather than a fact, it is labeled as one.

1. Purpose & intended use

AssessFit provides pre-employment assessments that help employers compare candidates on job-relevant abilities, knowledge, work styles, and judgment. Scores are designed to be one structured input among several in a hiring decision — alongside interviews, work history, and references. They are not a medical, clinical, or diagnostic instrument, and AssessFit is designed so that the final hiring decision is always made by a person, not an algorithm.

This manual describes the methodology behind the platform as it is built today. It follows the structure recommended by the SIOP Principles and the AERA/APA/NCME Standards for documenting selection procedures (see References).

2. Assessment architecture

The library contains 100 tests built on a bank of 854 items, spanning five domains:

DomainMeasuresFormat
Cognitive abilityNumerical, verbal, logical, spatial & mechanical reasoning, working memory, processing speed, attentionMultiple-choice, adaptive
Job skills53 domain-knowledge tests (finance, IT, marketing, operations, healthcare support, trades, and more)Multiple-choice, adaptive
Situational judgmentEthics, prioritization, leadership, customer scenarios, crisis responseScenario multiple-choice
Work stylesConscientiousness, composure, integrity, teamwork, adaptability, and 10 further scalesLikert self-report with reverse-keyed items
LanguageBusiness English (CEFR-aligned), written communication, proofreadingMultiple-choice

Each multiple-choice test is a 9-item bank tagged by difficulty (3 easy / 3 medium / 3 hard) with stable item identifiers. Employers may combine up to five tests per role; the platform's 1,000+ occupation catalog recommends an evidence-oriented bundle per job title.

3. Test development

Items are authored against explicit rules, then machine-validated before release:

  • One defensible key. Every item must have exactly one unambiguously correct answer and three plausible-but-wrong distractors. Items whose correctness could be argued are rejected at review.
  • Concepts over trivia. Items test stable concepts, not version-specific or expiring facts.
  • Cultural neutrality. Contexts are universal workplace situations; cognitive items are language-light where the construct allows.
  • Balanced key positions. Correct answers are distributed nearly evenly across option positions (current bank: within ±1% of uniform), so position-guessing strategies confer no advantage.
  • Structural validation. An automated validator enforces: unique item identifiers across the whole library, exact 3/3/3 difficulty composition, four options per item, in-range answer keys, and rejection of any item whose correct answer text duplicates another option.

4. Administration & adaptive design

Candidates receive a randomized two-stage adaptive form: stage one draws 3 medium items; candidates who answer at least 2 correctly advance to 3 hard items, others receive 3 easy items. No two candidates see identical forms, and item order is randomized. Each test carries a per-test time budget; the total session budget is shown to the candidate before starting.

Answer keys never leave the server: the client receives question text and options only, and all correctness decisions happen server-side at submission.

5. Scoring

  • Multiple-choice tests are scored server-side from correct counts across both adaptive stages.
  • Work-style scales are scored 0–100 with reverse-keyed items inverted before aggregation; unanswered statements score at the scale midpoint rather than being imputed favorably.
  • Composite score. Employers set transparent weights across CV match, skills tests, and work-style measures; the composite is a weighted mean, and the weights are visible — no black-box factor is applied.
  • Integrity verdict. Behavioral session signals (see §10) produce a separate integrity score and verdict; integrity never silently alters ability scores — it is reported alongside them for human review.

6. Norming

Percentiles are computed against empirical norm groups: real completed assessments for the same role template, using the midpoint percentile method. Two rules keep this honest:

  • Every percentile is displayed with the true size of its norm group ("n = 40"), never an inflated aggregate.
  • No synthetic or purchased norms are used. Where a norm group is still too small to be meaningful, raw scores are shown without percentile claims.

The platform-wide completion count is public and live: assessments completed to date.

7. Reliability & item quality

Every item in the bank is monitored continuously against real candidate data:

StatisticDefinitionAction threshold
Difficulty (p)Proportion answering correctlyp > .95 flagged too easy; p < .20 flagged too hard
Discrimination (rpb)Point-biserial correlation between item correctness and the candidate's test scorer < .05 at n ≥ 20 flagged weak; negative r at n ≥ 10 flags a suspect answer key for immediate review
Response timeMean seconds per itemOutliers reviewed for ambiguity or leakage

Flagged items are reviewed and retired or replaced; because tests are item banks, retirement does not interrupt administration.

Why we don't quote a coefficient alpha yet. Classical internal-consistency coefficients assume every candidate answers the same items; our adaptive forms deliberately violate that. Rather than publish an alpha computed on unsuitable data — or a number copied from a brochure — we publish the item-level statistics above, live. A platform-level reliability analysis will be published when the response matrix supports it, with its sample size.

8. Validity

Content validity

Each of the 1,000+ occupation titles in the catalog maps to one of 30 occupation profiles, and each profile's recommended test bundle was constructed by matching job-relevant constructs to tests (e.g., warehouse roles → logistics operations, safety, attention; care roles → patient scenarios, caregiving knowledge, integrity). Employers can inspect and edit every bundle — nothing is hidden.

Criterion validity — your own evidence

AssessFit ships a built-in predictive-validity dashboard: employers rate the on-the-job performance of the people they hired, and the platform computes the correlation between assessment scores and those ratings — overall and per test, always displayed with its sample size. This is a local validation study, the form of evidence the EEOC Uniform Guidelines treat as strongest, produced continuously and at no extra cost.

Commitment. As anonymized cross-employer outcome data accumulates, we will publish platform-level criterion-validity summaries with sample sizes and confidence intervals — in this manual, not in a press release. Benchmark from the meta-analytic literature: general cognitive ability predicts job performance at roughly r ≈ .3–.5 (Schmidt & Hunter, 1998).

9. Fairness

  • Blind scoring by construction. AssessFit does not collect race, religion, gender, or age from candidates, and no such variable exists anywhere in the scoring pipeline.
  • No face analysis. No emotion inference, no video scoring, no pseudo-scientific signals.
  • Language-light cognitive items reduce construct-irrelevant language load where the construct allows.
  • Employer responsibility. Under the EEOC Uniform Guidelines, adverse-impact analysis is conducted on the employer's applicant flow. AssessFit's structured, job-related scores make that analysis possible; employers remain responsible for monitoring their own selection rates, and our reports are designed to support that.

10. Integrity & security

Signals collected during a session (with candidate notice) include tab/window switches, paste and copy attempts, context-menu use, response-time anomalies (implausibly fast correct answers), and automation markers. Signals are weighted into an integrity score with three verdicts: clean (≥ 80), review (55–79), and high-risk (< 55), each listing its contributing flags so a human can judge.

  • Answer keys exist only server-side; scoring is entirely server-authoritative.
  • Invitation links are single-session tokens; candidate sign-in uses single-use, 15-minute magic links.
  • Randomized adaptive banks mean leaked answers have limited value — and leakage shows up in item statistics (§7), triggering retirement.

11. Data protection

  • All traffic is TLS-encrypted; the application enforces HTTPS and a canonical host.
  • Candidate CV files are stored outside the web root and are never web-served to the public.
  • Candidates can erase their profile and data via a self-service endpoint in their dashboard.
  • Assessment responses are stored for scoring, norming, and item-quality monitoring as described above; aggregate statistics are anonymous.

12. Appropriate use & limitations

  • Scores should never be the sole basis for a hiring decision.
  • Percentiles from small norm groups are labeled and should be read cautiously.
  • Work-style scales are self-report instruments; they describe typical tendencies, not guarantees of behavior.
  • Integrity verdicts are decision-support flags, not accusations; a review verdict warrants a conversation, not an automatic rejection.
  • Tests are delivered in English today; scores for candidates assessed in a second language partly reflect language proficiency. Use the dedicated language tests to measure that explicitly.

13. References & standards followed

  • AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.
  • SIOP (2018). Principles for the Validation and Use of Personnel Selection Procedures (5th ed.).
  • EEOC et al. (1978). Uniform Guidelines on Employee Selection Procedures.
  • Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology. Psychological Bulletin, 124(2), 262–274.
  • ISO 10667 — Assessment service delivery.

Questions about methodology, or requests for the underlying statistics: research@assessfit.com. This is a living document — the version and date at the top change whenever the platform's methodology does.