GS-0511 Auditing — existing vs. agent-built
Auditing
Performance and financial auditing per Generally Accepted Government Auditing Standards (GAGAS / Yellow Book).
Most agencies screen with Occupational Questionnaire + writing sample. GAO historically used a published analyst-track assessment for its own 0511 hires, but that is GAO-specific. There is no government-wide standardized audit-skills test.
Job analysis anchored to O*NET / MOSAIC + GAO Yellow Book (GAGAS 2024), FASAB. Every item carries signed task → KSAO → item provenance. Gated by licensed I-O psychologist sign-off before any applicant sees it.
Build this assessment →Eight-dimension comparison
Each axis is scored 0–100 using the rubric in comparisonMetrics.ts. Existing numbers come from the researched record; agent-built numbers reflect the swarm's uniform quality gates.
Chance to Compete Act (Pub. L. 118-19) three-prong test
- §3(2)(A) Allows applicants to demonstrate job-related skills
§3(2)(A) — the assessment must permit an applicant to demonstrate job-related knowledge, skills, abilities, or competencies, not merely attest to them.
Existing: partialAgent-built: yes - §3(2)(B) Is based on a job analysis
§3(2)(B) — content must be derived from a current job analysis that identifies the tasks and the KSAOs required to perform them.
Existing: partialAgent-built: yes - §3(2)(C) Is not principally reliant on a self-assessment
§3(2)(C) — the assessment "does not solely include or principally rely upon a self-assessment from an automated examination."
Existing: partialAgent-built: yes
Technical KSAO coverage gap
Of the 6 critical technical KSAOs identified in the job analysis for this series, how many does each assessment directly measure (not self-report)?
Per-dimension detail
1 technical competency directly measured: Writing (partially, where sample is used).
14-agent swarm maps every task → KSAO → item; ≥90% of critical technical KSAOs directly measured, no self-report proxies.
DIF analysis: no; no public adverse impact study.
Per-item bias sensitivity review + pilot DIF (Mantel-Haenszel / logistic regression) + adverse-impact 4/5ths pre-deployment simulation.
No published validity evidence; no public technical report.
Auto-drafted 29 CFR §1607.15 packet: job analysis, content validity matrix (Lawshe CVR), technical report, signed audit log.
2 modalities: Self-report questionnaire, Writing sample (agency-specific).
Blueprint enforces minimum 3 of 5 RFI modalities where content supports it (JKT + Work Sample + SJT / Simulation / SI).
Items anchored to PD language but not traceable to task-level job analysis or KSAO criticality data.
Every item carries a cryptographically signed provenance record: task statement → KSA → criticality/frequency → item.
Three-prong test flags: (A) partial, (B) partial, (C) partial.
Three-prong compliance enforced at the blueprint gate: (A) skill demonstration, (B) job-analysis anchored, (C) ≤10% self-rating by weight.
mobile support unknown; VPAT status unclear.
Tailwind-responsive, Section 508 AA verified, Flesch-Kincaid target grade 9–11, accommodation paths pre-wired.
Score is an aggregated self-rating; no per-competency sub-scores or rationale provided to applicant.
Each score includes IRT theta + SEM + KSAO sub-scores + plain-English rationale (NIST AI RMF explainability).
Bias & fairness findings
Drawn from the researched record for this series. Severity follows EEOC / NCME Standards guidance.
- No differential item functioning (DIF) analysis
NCME Standards §3.17 recommends per-item DIF analysis by protected class. Agent-built pipeline runs Mantel-Haenszel + logistic regression DIF at every pilot.
- No public adverse-impact (4/5ths) study
29 CFR §1607.4(D) requires impact records. Agent-built pipeline simulates the 4/5ths rule pre-deployment and logs results to the audit trail.
- Self-report vulnerability
Occupational Questionnaires are the exact pattern the Chance to Compete Act §3(2)(C) targets. They are known to be fakeable, produce ceiling effects, and fail to discriminate among qualified applicants.
- No public technical report
External defensibility review requires a public §1607.15 report. Agent-built runs auto-publish the technical report with the final assessment.
- Historical concern
The Chance to Compete Act of 2023 (Pub. L. 118-19) §3(2)(C) defines a Technical Assessment as one that "does not solely include or principally rely upon a self-assessment from an automated examination" — squarely aimed at this pattern
- Historical concern
OPM May 2025 Merit Hiring Plan directs agencies to reduce reliance on occupational questionnaires as the primary screening instrument
