GS-1811 Criminal Investigation — existing vs. agent-built
Criminal Investigation
Federal criminal investigator (1811) series — investigates violations of federal statutes; typically FLETC-trained, firearms-qualified.
This is the one priority series with a real standardized federal pre-hire assessment. DEA, ATF, USPIS, USSS and others use variants of the Special Agent Entrance Exam (SAEE) covering Critical & Logical Thinking, Situational Judgment, Life Experience / Personality. FBI uses its own Phase I Special Agent Battery Test (SABT). Post-offer: polygraph, medical, physical fitness, background. FLETC CITP is post-hire training.
Job analysis anchored to O*NET / MOSAIC + 18 U.S.C., Federal Rules of Evidence. Every item carries signed task → KSAO → item provenance. Gated by licensed I-O psychologist sign-off before any applicant sees it.
Build this assessment →Eight-dimension comparison
Each axis is scored 0–100 using the rubric in comparisonMetrics.ts. Existing numbers come from the researched record; agent-built numbers reflect the swarm's uniform quality gates.
Chance to Compete Act (Pub. L. 118-19) three-prong test
- §3(2)(A) Allows applicants to demonstrate job-related skills
§3(2)(A) — the assessment must permit an applicant to demonstrate job-related knowledge, skills, abilities, or competencies, not merely attest to them.
Existing: partialAgent-built: yes - §3(2)(B) Is based on a job analysis
§3(2)(B) — content must be derived from a current job analysis that identifies the tasks and the KSAOs required to perform them.
Existing: yesAgent-built: yes - §3(2)(C) Is not principally reliant on a self-assessment
§3(2)(C) — the assessment "does not solely include or principally rely upon a self-assessment from an automated examination."
Existing: partialAgent-built: yes
Technical KSAO coverage gap
Of the 7 critical technical KSAOs identified in the job analysis for this series, how many does each assessment directly measure (not self-report)?
Per-dimension detail
No series-specific technical KSAOs are directly measured — assessment is non-technical or inferred.
14-agent swarm maps every task → KSAO → item; ≥90% of critical technical KSAOs directly measured, no self-report proxies.
DIF analysis: yes; adverse impact studied.
Per-item bias sensitivity review + pilot DIF (Mantel-Haenszel / logistic regression) + adverse-impact 4/5ths pre-deployment simulation.
content / criterion / construct validity evidence; no public technical report.
Auto-drafted 29 CFR §1607.15 packet: job analysis, content validity matrix (Lawshe CVR), technical report, signed audit log.
3 modalities: Multiple choice (reasoning), SJT, Biodata / personality inventory.
Blueprint enforces minimum 3 of 5 RFI modalities where content supports it (JKT + Work Sample + SJT / Simulation / SI).
Partial traceability — validation evidence exists but item-level provenance is not public.
Every item carries a cryptographically signed provenance record: task statement → KSA → criticality/frequency → item.
Three-prong test flags: (A) partial, (C) partial.
Three-prong compliance enforced at the blueprint gate: (A) skill demonstration, (B) job-analysis anchored, (C) ≤10% self-rating by weight.
not mobile-compatible; VPAT status unclear; FK grade 12.
Tailwind-responsive, Section 508 AA verified, Flesch-Kincaid target grade 9–11, accommodation paths pre-wired.
Aggregate score returned; sub-score breakdown and rationale generally not available to applicants.
Each score includes IRT theta + SEM + KSAO sub-scores + plain-English rationale (NIST AI RMF explainability).
Bias & fairness findings
Drawn from the researched record for this series. Severity follows EEOC / NCME Standards guidance.
- No public technical report
External defensibility review requires a public §1607.15 report. Agent-built runs auto-publish the technical report with the final assessment.
- Historical concern
Federal and state/local law-enforcement tests have faced repeated Title VII adverse-impact litigation (Ricci v. DeStefano, 557 U.S. 557 (2009); earlier NYPD, Boston FD cases)
- Historical concern
Biodata tests without fairness monitoring can encode socioeconomic correlates
