Process overview
How the research system works
An interactive walk through all 27 pipeline stages, grouped into six phases from raw sensor data to embedded deployment. Drag to rotate, or switch to the list view — either way, every color and dependency line below reads the same registry as the full pipeline page. Nothing here is illustrative or invented.
Phase 1 of 6
Data Foundation
Acquire, verify, and chronologically split the raw gas-sensor dataset.
Stage-by-stage notes
- Stage 00 — Environment and reproducibility: Captures Python/OS/package versions, git commit, and seed for every subsequent stage.
results/reproducibility/environment.json - Stage 01 — Dataset acquisition / integrity: Official UCI archive downloaded and SHA-256 verified against the published archive hash.
data/raw - Stage 02 — Dataset validation: Verified 13,910 samples, 128 features, 10 batches, 6 classes; 0 missing values.
results/reproducibility/dataset_validation.jsonresults/reproducibility/dataset_validation.md - Stage 03 — Chronological split definition: FIXED_ORIGIN (train batch 1, test 2-10), EXPANDING_WINDOW, and IID_DIAGNOSTIC protocols frozen in config.
configs/chronological_protocol.yaml - Stage 04 — Preprocessing: Standardization fitted on training data only per protocol; no leakage across chronological boundary.