UCI Gas Sensor Array Drift Dataset
Source: UCI Gas Sensor Array Drift Dataset (Vergara et al.). Collected over 36 months from a 16-sensor chemical array exposed to six gases at varying concentrations. This project does not redistribute the raw archive — download it directly from UCI and verify the hash below.
Verified structure
Batch composition (chronological)
| Batch | Observations | Batch archive SHA-256 |
|---|---|---|
| Batch 1 | 445 | f346beee8e0c5e31ac5961845b6d96a70dc1ccf799481592fb2f0d96a81952e6 |
| Batch 2 | 1,244 | 07f7e94a9bf4377240f9b230d20c750932fc785d4766b6a4b7fc370a035c825a |
| Batch 3 | 1,586 | a4c7a1a6744df32f0ca139c75c3b4dade7cc99f57a7785d28cb0f62577dd2061 |
| Batch 4 | 161 | bd76f86be34ffe89f46a27c6fc2b5b8c6ebfcf984dff4b3e5befc76d98034b7f |
| Batch 5 | 197 | 3f95cb6e4a39a94bbacd2f1984f1754c9fb5eb3221812975ad70edbcea7abaec |
| Batch 6 | 2,300 | 83348c504105a5aa5264d1209f2a5d7ba7d9c8bcae60290395f201ac4708ff95 |
| Batch 7 | 3,613 | 3168cb56d5c9bc29c36184e2e73c9f3474adb3c4a893b1ce0d88b6c47698cfdb |
| Batch 8 | 294 | 296346b932893ea18c513e23ac5b660f35e1c499fde8254d4804a2ce8a1b4ca7 |
| Batch 9 | 470 | e019f11f4fa8336ab1f41f7503eb456b3679da0a84eb8032fe6cef90f0547826 |
| Batch 10 | 3,600 | 30011067c7c05c2b01038f84c73af0c6f1820b622fa2b0ad89111cb4c649f791 |
Gas classes
| Class ID | Gas | Total labeled observations |
|---|---|---|
| 1 | Ethanol | 2,565 |
| 2 | Ethylene | 2,926 |
| 3 | Ammonia | 1,641 |
| 4 | Acetaldehyde | 1,936 |
| 5 | Acetone | 3,009 |
| 6 | Toluene | 1,833 |
Processed dataset hash
dc9dbcfc4c8eedceae4418d8f2096605ccb2b3bd554a3134f84c46d22b0615e6
Full validation artifact: results/reproducibility/dataset_validation.json
Why chronological, not random, splitting
Sensor drift means the feature distribution at batch 10 differs measurably from batch 1 (see drift evidence). A random train/test split mixes observations from all batches into both sets, letting a model implicitly learn future-batch characteristics it would never have access to in deployment — inflating apparent accuracy. This project treats FIXED_ORIGIN (train on Batch 1 only, evaluate on Batches 2–10 in order) as the primary protocol, and includes an explicit IID_DIAGNOSTIC split only to quantify how much random splitting overstates performance — never as a headline result. See Methodology for the full protocol definition.