Executed

UCI Gas Sensor Array Drift Dataset

Source: UCI Gas Sensor Array Drift Dataset (Vergara et al.). Collected over 36 months from a 16-sensor chemical array exposed to six gases at varying concentrations. This project does not redistribute the raw archive — download it directly from UCI and verify the hash below.

Verified structure

Observations
13,910
Features
128
Chronological batches
10
Gas classes
6
Missing values
0
Malformed rows
0

Batch composition (chronological)

BatchObservationsBatch archive SHA-256
Batch 1445f346beee8e0c5e31ac5961845b6d96a70dc1ccf799481592fb2f0d96a81952e6
Batch 21,24407f7e94a9bf4377240f9b230d20c750932fc785d4766b6a4b7fc370a035c825a
Batch 31,586a4c7a1a6744df32f0ca139c75c3b4dade7cc99f57a7785d28cb0f62577dd2061
Batch 4161bd76f86be34ffe89f46a27c6fc2b5b8c6ebfcf984dff4b3e5befc76d98034b7f
Batch 51973f95cb6e4a39a94bbacd2f1984f1754c9fb5eb3221812975ad70edbcea7abaec
Batch 62,30083348c504105a5aa5264d1209f2a5d7ba7d9c8bcae60290395f201ac4708ff95
Batch 73,6133168cb56d5c9bc29c36184e2e73c9f3474adb3c4a893b1ce0d88b6c47698cfdb
Batch 8294296346b932893ea18c513e23ac5b660f35e1c499fde8254d4804a2ce8a1b4ca7
Batch 9470e019f11f4fa8336ab1f41f7503eb456b3679da0a84eb8032fe6cef90f0547826
Batch 103,60030011067c7c05c2b01038f84c73af0c6f1820b622fa2b0ad89111cb4c649f791

Gas classes

Class IDGasTotal labeled observations
1Ethanol2,565
2Ethylene2,926
3Ammonia1,641
4Acetaldehyde1,936
5Acetone3,009
6Toluene1,833

Processed dataset hash

dc9dbcfc4c8eedceae4418d8f2096605ccb2b3bd554a3134f84c46d22b0615e6

Full validation artifact: results/reproducibility/dataset_validation.json

Why chronological, not random, splitting

Sensor drift means the feature distribution at batch 10 differs measurably from batch 1 (see drift evidence). A random train/test split mixes observations from all batches into both sets, letting a model implicitly learn future-batch characteristics it would never have access to in deployment — inflating apparent accuracy. This project treats FIXED_ORIGIN (train on Batch 1 only, evaluate on Batches 2–10 in order) as the primary protocol, and includes an explicit IID_DIAGNOSTIC split only to quantify how much random splitting overstates performance — never as a headline result. See Methodology for the full protocol definition.

Chronological split, visualized