White paper Transport infrastructure CRX-WP-0012 v1 · current Open

A Framework for Trustworthy Deployment: Governance, Human-Automation Teaming, and a Phased Implementation Roadmap for AI in Track Base Management

Paper 5, the capstone of a five-part series on artificial intelligence in railroad track base management.

Amy Chen 2 Aug 2026 17 MIN 2 EXHIBITS 1 SOURCES

Executive Summary

This paper converts four papers of analysis into one implementable structure. The framework comprises five pillars. The assurance pillar operationalizes the three-construct validation architecture and makes revalidation, drift monitoring, and label stewardship permanent funded functions, on the finding that the dominant hazards are quiet ones that a one-time certification cannot control. The human-automation teaming pillar allocates the four inspection functions between automation and qualified personnel, redefines the inspector role toward verification, interpretation, and remediation management, and solves the verification problem unique to substructure sensing, that the automation measures conditions no human can directly re-observe, through checkable correlates and sampled ground truth. The data governance pillar builds the shared, annotated, maintenance-history-linked corpus the model generation depends on, and places substructure data pipelines within the rail cybersecurity directive framework’s critical system discipline. The regulatory engagement pillar executes the five-step standing pathway, negotiates remediation criteria and response protocols before fielding, and preserves the additive, frequency-neutral positioning that keeps substructure deployment outside the unresolved automation dispute. The economics and accountability pillar runs the dual-purpose pilot that generates validation evidence and fleet economics in a single program, tracks benefits by the four categories established in Paper 3, and publishes condition targets on a defined cadence.

The framework is governed through a decision rights matrix distinguishing railroad decisions, regulator decisions, and jointly held decisions, and executed through a three-phase roadmap, pilot, corridor, and network, each phase carrying entry criteria, exit evidence, and an active hazard control set. The paper closes the three control gaps Paper 4 identified: an independent field validation program for deep learning ballast models, a tiered response protocol for AI-detected substructure conditions, and an investigation-grade decision record standard. A complete disposition table allocates all fourteen hazards to closed, mitigated and monitored, or explicitly accepted status. The framework aligns with federal doctrine throughout: the NIST AI Risk Management Framework supplies the trustworthiness characteristics and the govern, map, measure, manage structure, and the Department of Transportation’s own AI assurance and governance programs supply the departmental context any standing pathway will engage.

1. Design Principles

Five principles, each earned earlier in the series, govern the framework’s construction.

Technology neutrality. The framework specifies capabilities, interfaces, roles, and controls, never products or vendors. Paper 1 established that the sensing families are complementary and that no modality is sufficient alone; a framework bound to particular products would not survive the sensor generations its own drift controls anticipate.

Additive positioning. Substructure sensing measures what visual inspection cannot observe and substitutes for nothing. This stance, maintained since the foundational paper, is simultaneously a safety control (the geometry measurement layer remains an independent downstream check on substructure false negatives) and an institutional control (frequency-neutral program design avoids importing the unresolved automation dispute).

Permanent assurance. Paper 4’s central finding was that the dominant hazards, drift, label decay, and skill erosion, are quiet, and quiet hazards are controlled only by standing processes. Every assurance activity in this framework is specified as a funded recurring function, never a one-time gate.

Evidence before consequence. Thresholds attach to demonstrated performance consequences, remediation criteria are negotiated before fielding, and no phase of the roadmap advances without producing its defined evidence, applying the lessons of the GAO-documented disincentive and the geometry waiver record.

Federal alignment. The framework’s vocabulary and structure track the NIST AI Risk Management Framework, whose trustworthiness characteristics (valid and reliable as the base, with safe, secure and resilient, accountable and transparent, explainable, privacy-enhanced, and bias-managed above it) and whose govern, map, measure, and manage functions organize the assurance pillar; and it anticipates engagement with the Department of Transportation’s AI assurance effort, which is developing an initial AI and data assurance framework and common terminology across transportation modes, and with the department’s AI governance structure under Office of Management and Budget direction.

2. Pillar 1: Assurance

The assurance pillar operationalizes Paper 2’s validation architecture as a lifecycle.

Initial validation executes the three constructs in sequence: ground-truth-referenced validation of each sensor-derived index against the consolidated metric, across the ballast materials, climates, and traffic profiles of the intended territory, with stratified performance reporting so regional blind spots (hazard H4) are measured rather than assumed; performance-referenced validation tying index thresholds to demonstrated settlement and geometry consequences; and outcome-referenced validation accumulating through the pilot and corridor phases.

Closing the H1 gap. Paper 4 identified independent field validation depth for deep learning ballast condition models as the taxonomy’s weakest control. The framework closes it with an independent field validation program: validation excavations and ground truth sampling performed or witnessed by a party independent of the model developer, at sites the independent party selects, with results published to the shared corpus. The federal inspection program’s existing independent survey role is the institutional template, and the pilot phase budget carries this program as a defined line.

In-service assurance comprises four permanent functions: drift monitoring on input distributions with defined statistical triggers; scheduled revalidation and triggered revalidation on sensor fleet, material source, or practice changes, following the closed survey-model loop the pavement precedent demonstrated; label stewardship, with maintenance-history linkage mandatory on every corpus record so models are never trained across uncontrolled interventions; and performance monitoring comparing forecasts against subsequent geometry outcomes, the standing empirical check on the whole stack.

3. Pillar 2: Human-Automation Teaming

The teaming pillar applies the four-task decomposition of the FRA-sponsored teaming literature, data collection, data analysis, decision making, and action, and allocates deliberately.

Allocation. Data collection is automated by design; that is the technology’s purpose. Data analysis is automated with human verification at defined sampling rates and at every consequential flag. Decision making is human, informed by model outputs carrying confidence displays; the framework assigns no autonomous remediation authority to any model. Action remains with maintenance forces under existing authorities.

The inspector role is redefined toward verification, interpretation, and remediation management, and the redefinition is planned before deployment, with training, position descriptions, and proficiency requirements settled in the pilot phase, applying the workforce-transition-first principle carried since the foundational paper.

Verification without a human baseline. Substructure automation contributes measurements no inspector can re-observe directly, the hazard twist Paper 4 attached to H8. The framework’s answer is correlate-based verification: every consequential substructure flag is checkable against at least one of three observables, surface symptoms (mud spots, fouled cribs, drainage evidence), geometry behavior at the flagged location, or sampled ground truth under the independent validation program. Flags failing correlate checks at defined rates trigger model review, feeding the assurance pillar.

Trust calibration and skill retention. Confidence display with every output, explainable flag rationales consistent with the NIST explainability characteristic, verification sampling that audits accepted flags, and proficiency programs that maintain unassisted condition-assessment skill in the inspector workforce together control H7 and H8, with proficiency assessment reported on the same cadence as model performance so human and machine capability are monitored symmetrically.

4. Pillar 3: Data Governance

The shared corpus. Paper 1 specified the requirement and the gap: the federal inspection archive is the strongest foundation for a substructure training corpus but lacks an annotation standard, access terms, and stewardship. The governance pillar supplies the three: an annotation standard binding every record to the consolidated metric of Paper 2’s Step 1, to maintenance history, and to ground truth provenance; tiered access terms under which railroads, researchers, and suppliers train against common data with commercial and security interests protected; and a named steward accountable for label quality over time, closing H6’s stewardship gap.

Synthetic data discipline. Synthetic augmentation is admitted to training with disclosure and is excluded from validation; models are validated on field data only, and synthetic-to-field transfer performance is measured and published as pilot evidence, converting H4’s open question into a tracked quantity.

Security scoping. The framework treats substructure data pipelines, model infrastructure, and the corpus itself as candidate critical cyber systems under the rail cybersecurity directive framework, adopting the conservative scoping stance Paper 4 recommended: integrity verification on measurement archives, provenance controls on training data, and inclusion in the carrier’s cybersecurity implementation plan and assessment program, with the poisoning, manipulation, and exfiltration threat categories of federal AI risk doctrine explicitly in the assessment scope.

5. Pillar 4: Regulatory Engagement

The pathway. The pillar executes Paper 2’s five steps: metric consolidation through recommended practice; ground-truth-referenced validation; performance-referenced threshold setting; structured test programs that permit practice to vary; and conditional waivers maturing into rulemaking. The geometry record establishes the ladder’s viability and its approximate span, and establishes equally that the advisory step may fail to converge without halting the sequence.

Closing the H3 gap: the response protocol. Paper 4 found the response surface for AI-detected substructure conditions entirely unbuilt, and the East Palestine detection-response lesson, that a monitoring system’s protection is bounded by the thresholds and protocols around it, makes this the framework’s most safety-significant construction. The framework specifies a three-tier protocol, negotiated into test program terms before fielding: an advisory tier, in which flags below maintenance limits enter planning with no mandated action; a maintenance tier, in which flags exceeding maintenance limits require disposition, scheduled work or documented engineering justification, within a defined window; and a safety tier, applicable only where performance-referenced evidence has established a safety limit, in which flags require verification within a defined short window and protective action, restriction or remediation, upon confirmation. Every tier’s flag receives a verification disposition through the central desk model the geometry waiver record established, and every disposition enters the decision record. The tier boundaries are the maintenance-limit and safety-limit distinction of Paper 2, now carrying operational meaning.

Positioning discipline. Program designs under this pillar are frequency-neutral: they neither request nor depend on visual inspection reductions, keeping substructure deployment outside the contested automation dispute (H12), which this series carries as contested terrain to its end. Where an operator separately pursues frequency relief, that pursuit proceeds under its own docket and evidence, not under this framework.

6. Pillar 5: Economics and Accountability

The dual-purpose pilot. Paper 3 established that the value case is a substitution case and therefore testable. The pilot phase is designed to test it: substructure-informed territory against matched conventional territory, measuring maintenance spend, slow order hours, geometry quality trajectories, and defect recurrence, generating outcome-referenced validation evidence and fleet-scale economics in one program.

Benefit tracking follows Paper 3’s four categories, maintenance efficiency, asset life extension, defect and disruption avoidance, and long-horizon safety, with the assumption-driven cells of the value model converted to measured cells as pilot and corridor evidence accumulates, and with the safety category reported honestly against the derailment trend context the Congressional Research Service documents.

Condition-target accountability. Applying the California pavement lesson, the program publishes substructure condition targets, measured by its own consolidated metric, and reports achieved condition against target on a defined cadence, closing H11; a prediction capability without published targets reproduces the oversight outcome the state record documents.

7. Decision Rights

Exhibit 7. Decision Rights
7. Decision Rights
Decision Railroad Regulator Joint / Standards Body
Metric definition and classification bands Recommended practice body with FRA participation
Validation protocol design Proposes Approves within test program terms
Threshold values (maintenance tier) Sets from performance evidence Reviews in test program
Threshold values (safety tier) Sets or approves Evidence base jointly developed
Response protocol tiers and windows Proposes Approves in test program and waiver terms
Model deployment and revalidation acceptance Decides, per approved assurance plan Audits
Corpus annotation standard and access terms Steward under multi-party governance
Critical cyber system scoping Determines TSA reviews per directive framework
Phase advancement (Sections 8) Proposes with evidence Concurs for compliance-relevant phases
Workforce role definitions and proficiency standards Decides Consistent with teaming design guidance
Condition targets and reporting cadence Sets and publishes Receives
Accepted-risk register (Section 9) Maintains Receives
Author’s own analysis (A Framework for Trustworthy Deployment) Source record →

8. The Phased Roadmap

Phase 1: Pilot. Scope: one to three subdivisions selected by the recurring-defect targeting logic, sections where geometry degradation recurs after surfacing. Entry criteria: consolidated metric adopted for program use; assurance plan, teaming design, response protocol, and decision record standard documented; independent validation program funded; matched comparison territory designated; workforce transition plan executed. Exit evidence: ground-truth-referenced validation results with stratified performance; synthetic-to-field transfer measurements; correlate-verification statistics; first substitution economics from the matched comparison; zero unresolved safety-tier protocol failures; decision record demonstrated end to end. Active hazard controls: the full register, with H1, H3, and H14 controls exercised and evidenced.

Phase 2: Corridor. Scope: a full corridor or division, entered under a structured test program per Paper 2’s Step 4, with FRA engagement. Entry criteria: pilot exit evidence complete; test program terms approved, including response protocol and remediation criteria; corpus operating under its governance with the pilot’s data contributed. Exit evidence: performance-referenced threshold validation at corridor scale; outcome-referenced trends (defect recurrence, geometry quality, maintenance substitution) across two or more maintenance cycles; drift and revalidation functions demonstrated through at least one triggered revalidation; teaming proficiency results; published condition-target reporting initiated. This phase generates the waiver-grade record.

Phase 3: Network. Scope: network deployment under conditional waiver terms maturing toward rulemaking participation. Entry criteria: corridor exit evidence complete; waiver terms in force; accountability reporting established. Standing obligations rather than exit criteria govern this phase: permanent assurance functions funded and audited; annual disposition review of the hazard register; corpus stewardship sustained; and participation in the standards and rulemaking record as Paper 2’s Step 5 matures. Phase 3 has no completion; it is the operating state the framework exists to make trustworthy.

9. Hazard Disposition

Every Paper 4 register entry is allocated below. Closed means a specified control eliminates the gap; mitigated means controls reduce and monitoring bounds the residual; accepted means the residual is documented, owned, and reviewed.

Exhibit 9. Hazard Disposition
9. Hazard Disposition
# Hazard Disposition Framework mechanism
H1 False negative assessment Mitigated and monitored Independent field validation program (Pillar 1); geometry layer as downstream check; performance monitoring
H2 False positive over-flagging Mitigated Precision requirements; advisory tier routing; verification desk
H3 Miscalibrated response thresholds Closed as gap; residual mitigated Three-tier response protocol negotiated pre-fielding (Pillar 4)
H4 Training data unrepresentativeness Mitigated and monitored Stratified validation and reporting; shared corpus; synthetic data excluded from validation
H5 Distribution drift Mitigated and monitored Standing drift statistics; scheduled and triggered revalidation
H6 Label decay and spurious structure Closed as gap; residual mitigated Mandatory maintenance-history linkage; named steward; intervention-aware modeling
H7 Automation bias / misallocated trust Mitigated and monitored Confidence display; explainable rationales; verification sampling
H8 Inspector skill degradation Mitigated and monitored Task allocation retaining hands-on skill; proficiency programs; symmetric capability reporting; correlate-based verification design
H9 Remediation disincentive Mitigated Tiered limits; evidence-based thresholds; pre-fielding criteria
H10 Metric gaming Mitigated and monitored Fixed metric definitions; independent measurement audit
H11 Accountability gap Closed Published condition targets and cadence (Pillar 5)
H12 Automation dispute entanglement Accepted and bounded Frequency-neutral design; dispute remains contested terrain outside program scope
H13 Pipeline compromise Mitigated and monitored Critical-system scoping under directive framework; integrity and provenance controls
H14 Unassignable responsibility Closed as gap; residual mitigated Investigation-grade decision record: model version, inputs, output, confidence, verification disposition, responder action, logged for every consequential flag; decision rights matrix
Author’s own analysis (A Framework for Trustworthy Deployment) Source record →

Two residuals are formally accepted and carried on the accepted-risk register: the automation dispute itself (H12), which no program design can resolve, and the irreducible tail of H1, the possibility of a novel condition outside all validation coverage, bounded by the geometry layer and by the humility the register’s evidentiary-status column encodes.

10. Findings

Finding 5.1: The series’ analyses compose into a closed system. Every hazard identified in Paper 4 maps to a control built from Papers 1 through 3’s findings; every control gap identified is closed or explicitly accepted; and every evidence requirement of the standing pathway is produced by a defined roadmap phase. The framework’s completeness is a property of the series’ method, in which each installment’s open threads were carried forward rather than dropped.

Finding 5.2: The three constructed elements are the framework’s original contributions. The independent field validation program (closing H1), the three-tier response protocol for AI-detected substructure conditions (closing H3), and the investigation-grade decision record standard (closing H14) exist in no current rail standard; they are this framework’s additions, each built by analogy from documented precedent, the federal survey role, the geometry waiver verification model, and the accountability doctrine of federal AI risk management respectively.

Finding 5.3: Permanence is the framework’s cost center and its distinguishing demand. Drift monitoring, revalidation, stewardship, proficiency maintenance, and target reporting are recurring obligations; a deployment funded to certify once and operate indefinitely reproduces the quiet-hazard failure profile Paper 4 documented, and the roadmap’s Phase 3 is deliberately specified as standing obligations without completion.

Finding 5.4: Frequency neutrality is the load-bearing institutional choice. Keeping substructure deployment outside the contested automation dispute is what allows the standing pathway to proceed while that dispute remains unresolved; entangling the two would couple the program’s fate to a conflict its technology does not require.

11. Residual Research Agenda

The series closes by consolidating the items it deferred, in rough priority order: synthetic-to-field transfer evidence for deep learning ballast models, the pilot phase’s first measurable contribution to the open literature; fiber optic acoustic sensing and instrumented-particle methods as additional substructure modalities; comparative analysis of international fouling and maintenance limit regimes, including European infrastructure manager practice under separated ownership; quantitative cross-modality accuracy comparison on a common ground truth basis, which no current public dataset supports; insider threat and physical security of distributed sensing assets; rail asset management economics under investor ownership as distinct from the public agency precedent; and the eventual substructure-specific incident literature, which does not yet exist and whose first entries, when they come, should be met with the investigation-grade records this framework requires.

12. Limitations

The framework is a synthesis of documented evidence and this series’ analysis; it has not been exercised by an operating railroad, and its parameter values, sampling rates, response windows, revalidation schedules, target cadences, are deliberately left to program negotiation because the evidence to fix them generically does not exist. The decision rights matrix reflects the current United States institutional structure and would require adaptation elsewhere. The late-2025 regulatory developments noted in Paper 2 remain subject to primary-document verification, and any formal use of this framework in a proceeding should complete that verification first. The framework’s alignment with the Department of Transportation’s AI assurance effort is anticipatory; that effort’s initial framework and terminology were in development at this writing, and convergence should be revisited as departmental products publish.

13. Series Conclusion

Five papers ago, this series began from a single observation: the layer of the railroad that most governs its degradation is the layer its inspection regime measures least. The intervening installments established that the measurement problem is solved in the laboratory and the pilot fleet; that the path from measurement to meaning runs through metric consolidation, a redesigned validation architecture, and a demonstrated institutional ladder; that the value case is a testable substitution case whose magnitudes a single well-designed pilot can supply; and that the hazards, dominated by quiet failures and institutional traps, are controllable by standing processes that must be funded as permanently as the models they govern. This capstone assembled those results into five pillars, a decision rights structure, a three-phase roadmap, and a complete hazard disposition, and constructed the three elements no existing standard supplies: independent field validation, a tiered response protocol for conditions regulation does not yet classify, and the decision record a future investigator will need. The framework’s test is the pilot it specifies. The evidence that test produces, substitution economics from matched territory, stratified validation results, transfer measurements for synthetic training data, and the first operating record of the response protocol, will either bear out the series’ analysis or correct it, and the framework is built to accept the correction: its assurance pillar treats every model as provisional, and its accepted-risk register treats the framework’s own gaps the same way.

Bibliography

Tier 1: Government and Oversight Sources

California State Auditor. Report 2017-601. Sacramento, CA, 2017. https://information.auditor.ca.gov/reports/2017-601/chapters.html.

Congressional Research Service. East Palestine, OH, Train Derailment and Hazardous Materials Shipment by Rail: Frequently Asked Questions. CRS Report R47435. Washington, DC, 2023.

Congressional Research Service. Freight Rail Safety Issues in the 119th Congress. CRS Report R47911. Washington, DC.

Department of Homeland Security. “Ratification of Security Directives.” Federal Register, January 21, 2025.

Federal Railroad Administration. “Petition for Waiver of Compliance.” 83 Fed. Reg. 55450 (November 5, 2018). Docket No. FRA-2018-0091.

Federal Railroad Administration. “Track Geometry Measurement System (TGMS) Inspections.” Notice of Proposed Rulemaking. Federal Register, October 24, 2024. Docket No. FRA-2024-0032.

Federal Railroad Administration. ATIP: Automated Track Inspection Program (December 2025). TRB Quad Chart. Washington, DC, December 2025.

Federal Railroad Administration / John A. Volpe National Transportation Systems Center. Human-Automation Teaming in Track Inspection. Washington, DC: U.S. Department of Transportation.

National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Gaithersburg, MD, January 2023.

Transportation Security Administration. Security Directive 1580-21-01 series and Security Directive 1580/82-2022-01 series. Washington, DC, 2021-2025.

U.S. Department of Transportation, Highly Automated Systems Safety Center of Excellence. “AI Assurance in Transportation.” Workshop and whitepaper series. https://www.transportation.gov/hasscoe/highlights/AI-assurance.

U.S. Department of Transportation. Compliance Plans for OMB Memoranda M-24-10 (September 2024) and M-25-21 (September 2025) and associated DOT AI Strategy documents.

U.S. Government Accountability Office. Rail Safety: Federal Railroad Administration Should Report on Risks to the Successful Implementation of Mandated Safety Technology. GAO-11-133. Washington, DC, 2010.

Docket Filings

Association of American Railroads and American Short Line and Regional Railroad Association. Comments on TGMS Inspections NPRM. Docket No. FRA-2024-0032, January 2025.

Series Papers Integrated

Chen, Amy. Artificial Intelligence in Railroad Track Base Management (foundational paper) and Series Papers 1 through 4: Seeing Beneath the Rail; From Signal to Standard; The Economics of Knowing; When the Model Is Wrong. Each with its own complete bibliography of the Tier 1 through Tier 4 sources on which this capstone’s integrated claims rest, including the Selig and Waters substructure foundation, the FRA-sponsored ballast machine vision and scanning vehicle lineage, the GPR and InSAR validation literature, the vertical track deflection program record, the degradation prediction literature, and the California pavement management record.

A PDF of this paper is coming

The full text is on this page and nothing is held back. Leave your address and we will send the PDF as soon as it is rendered.

Request this paper

A Framework for Trustworthy Deployment: Governance, Human-Automation Teaming, and a Phased Implementation Roadmap for AI in Track Base Management

Give us an address and we will email a private link that works for seven days. Up to five papers a week. We do not sell or share addresses.

↑↓ to move ⏎ to open esc to close Type an identifier to jump straight to it