Engineering Resilient Reproducible Analytical Pipelines (RAP)
A Semantic-Based Self-Healing Framework for High-Velocity Heterogeneous Data Streams
DOI:
https://doi.org/10.31224/6466Keywords:
data engineering, reproducible analytical pipelines, autonomous agents, data provenance, schema drift, self-healing, BERT, telemetryAbstract
Mission-critical telemetry systems such as clinical monitoring systems all face critical limitations in data availability, veracity and velocity. High-frequency data pipelines break easily when upstream schemas shift, sensors fail or interfaces change. Traditional pipelines rely on brittle selectors or rigid schemas. When these fail, organizations experience data blackouts, delayed decision-making and loss of situational awareness at critical points. This research proposes a self-healing Reproducible Analytical Pipeline (RAP) framework designed to autonomously mitigate schema drift without manual intervention in real time. Leveraging a container-deployed Python ecosystem and BERT Large Language Model Processing, the framework replaces static schema changes with a dynamic semantic embedding- driven reconciliation. The framework uses a cross-domain generalizable model to work in various industries. Our framework introduces a domain-agnostic ingestion interface that implements validation and normalization logic. The approach enables a unified, cross-domain resilient data ingestion, while reducing pipeline fragility and ensuring the stability of critical, high-velocity analytical workflows in mission-critical environments. To validate this framework, it has been evaluated in a motorsport telemetry environment that shows a critical amount of schema drift. It was evaluated across 9 different hardware environments to validate the technology independence of the system.
Downloads
Downloads
Posted
Versions
- 2026-10-06 (4)
- 2026-02-16 (3)
- 2026-02-16 (2)
- 2026-02-12 (1)
License
Copyright (c) 2026 Tarek Clarke

This work is licensed under a Creative Commons Attribution 4.0 International License.