Drift-Robust Explainable TinyML for Electronic-Nose Sensing
Can lightweight models retain useful predictive performance and interpretable behavior under chronological sensor drift while fitting resource-constrained TinyML hardware?
Electronic-nose sensors age and drift — a model trained on early sensor readings loses accuracy as time passes, even though nothing about the underlying chemistry changed. This project evaluates that decay honestly (in real chronological order, never shuffled), studies whether explanations of a model's predictions stay stable as sensors drift, and works toward a resource-aware path to a physical nRF52840 microcontroller — while reporting plainly what has and has not yet been executed.
Explore the Research →Explore the Digital TwinSee What's Verified
New to the project? Start here.
A recommended reading order, in roughly three minutes to three hours:
The 3MT Speech and Slide
The actual speech and single slide submitted for the FAU 2026 Three Minute Thesis competition, reproduced verbatim — not a paraphrase.
research-portal/public/3mt/Arun_Gharami_FAU_3MT_2026_Single_Slide.pptxImagine opening your refrigerator and smelling milk that has gone bad.
Before you read the date, your nose warns you.
Now imagine giving that ability to a machine.
An electronic nose uses chemical sensors and artificial intelligence to recognize patterns in gases. These systems could help monitor food, detect hazardous chemicals, and support environmental or health applications.
But there is a hidden problem.
The nose changes.
Electronic sensors age. Temperature and humidity affect them. Their signals slowly shift, even when the chemical being measured has not changed. Scientists call this sensor drift.
And when the sensor changes, an AI model trained yesterday may not understand the world tomorrow.
In this research, use an electronic-nose dataset collected over three years. Trained four baseline AI models on the earliest sensor readings and tested them as time moved forward.
The result was striking.
Between Month 3 and Month 36, every model lost roughly 35 to 49 percentage points of accuracy. In one lightweight model, accuracy fell from about 81 percent to 37 percent.
That means an AI system can look excellent when it is first built—and become unreliable simply because its sensors grow older.
So my research asks a different question:
Can we build small AI systems that remain trustworthy when their sensors and environment change?
Developing a drift-aware, explainable TinyML research pipeline. Instead of randomly mixing old and new sensor data, I evaluate models in chronological order, the way they would experience the real world. I examine not only whether a prediction is correct, but also whether the reasons behind that prediction stay stable as the sensors drift. Then I study which lightweight models are realistic candidates for resource-constrained embedded devices.
This matters because the future of AI is not only in giant data centers.
It is moving into factories, farms, wearable devices, environmental monitors, and small sensors that may operate for months or years.
For trustworthy AI, high accuracy on day one is not enough.
We need confidence that the system will still deserve our trust on day one thousand.
An electronic nose how to recognize when its world has changed.
Because when AI gains a sense of smell, it should not forget how to use it.
Source: docs/3mt/speech.txt
Complete Research Outline
Motivation and research gap
Electronic-nose classifiers are usually evaluated on randomly shuffled train/test splits, which can blend sensor behavior from different time periods and produce an optimistic estimate of real-world performance. This project instead evaluates strictly in chronological order, and treats explainability and TinyML deployability as first-class, separately-measured research questions rather than afterthoughts.
Research questions and hypotheses
10 research questions (RQ1–RQ10) and 10 hypotheses (H1–H10), each tied to explicit acceptance criteria and linked experiment IDs. See research/questions.yaml and research/hypotheses.yaml.
Dataset and chronology
UCI Gas Sensor Array Drift Dataset: 13910 samples, 128 features (16 sensors × 8 response characteristics), 10 chronological batches, 6 gas classes. Executed — full provenance on /dataset.
Data validation, frozen preprocessing, leakage controls
Preprocessing (standardization) is fit once on Batch 1 only and frozen; the chronological split is defined before any model is trained. Details on /methodology and configs/chronological_protocol.yaml.
Baselines, fixed-origin evaluation, expanding-window adaptation
Four classical models (Logistic Regression, Random Forest, RBF-SVM, MLP) evaluated under FIXED_ORIGIN (train once on Batch 1) and EXPANDING_WINDOW (periodic retraining) protocols. Executed — see /results.
Resource-aware explanations, fidelity, and stability
Permutation importance, intrinsic coefficients/impurity, and single-feature ablation — deliberately not SHAP or LIME. Fidelity and stability evidence is executed but predominantly unsupported or mixed, stated plainly rather than minimized. Executed — see /xai.
Quantization and export for nRF52840 / Cortex-M4F
Host-compiled (x86-64) numerical-equivalence export exists for two model candidates; the first two export attempts failed their preprocessing-equivalence criteria before a fused-architecture variant passed on host only. INT8 quantization has not been executed at any tier. See /tinyml.
Planned physical measurement
Flash, SRAM, on-device latency, and PPK2 energy all require a physical nRF52840 board and debug probe. Hardware blocked — no board has ever been detected. See /hardware.
Pareto analysis, reproducibility, limitations, next experiments
The accuracy/fidelity/resource Pareto analysis is blocked on physical hardware measurement and has not been produced. Reproducibility guarantees (seed, dataset hash, config hash, git commit per artifact) are documented on /reproducibility; the next planned experiment and every known limitation are stated on /professor-review.
Interactive Research View
The conceptual 3D pipeline below is the same evidence-driven digital twin used on /digital-twin — sample chamber → sensor array → drift detector → classical model → explanation → nRF52840 (currently hardware-blocked) → gateway → evidence registry. Every status shown is real; nothing here is a fabricated measurement.
Loading interactive 3D model…
Move between chronological batches
Select a test batch to see its real drift score (relative to Batch 1) and the lightest baseline model's real accuracy at that point in time — the same numbers behind the 3MT slide's 81%→37% figure, drawn directly from results/drift/global_drift_by_batch.csv and results/baselines/fixed_origin_by_batch.csv.
All four baseline models, accuracy by batch
Full data (text equivalent)
| Batch | Drift vs. Batch 1 (normalized Wasserstein) | MODEL-C1 accuracy |
|---|---|---|
| 2 | 0.579 | 81.4% |
| 3 | 0.821 | 62.9% |
| 4 | 0.996 | 54.0% |
| 5 | 1.164 | 35.0% |
| 6 | 0.587 | 46.7% |
| 7 | 0.602 | 43.1% |
| 8 | 0.621 | 32.7% |
| 9 | 0.680 | 34.3% |
| 10 | 0.412 | 37.4% |
New Doors This Work Could Open
These are potential application directions that this methodology could inform — they are not demonstrated outcomes, and each would require its own domain-specific validation.
Reliable environmental sensing
Chronological-drift-aware evaluation is directly applicable to any long-lived environmental sensor network (air quality, water quality) where recalibration is expensive or impossible.
Industrial gas and process monitoring
Industrial electronic-nose deployments run for years without replacement — a drift-honest evaluation protocol could inform maintenance/recalibration scheduling.
Food quality assessment
Spoilage and freshness detection is a natural electronic-nose application; this work's chronological protocol could validate whether a food-safety model still holds after sensor aging.
Portable diagnostics research
Breath- or odor-based diagnostic research faces the same sensor-drift risk; resource-aware explainability could support clinician trust in a constrained device.
Edge AI under distribution shift, broadly
The chronological-evaluation and resource-aware-explanation methodology is not electronic-nose-specific — it generalizes to any TinyML system where the sensor or input distribution changes after deployment.
Collaboration
This research is at a stage where specific kinds of help would directly unblock the next experiments:
- Experimental review — a second opinion on the chronological protocol, the resource-aware explanation methodology, or the fidelity/stability evaluation design.
- nRF52840 hardware access — a physical nRF52840 development kit and debug probe would unblock Stages 15–20 (the entire physical measurement chain), which have been architecturally ready and blocked on hardware access since Stage 13.
- Measurement guidance — experience with Nordic Power Profiler Kit II energy-measurement methodology, or with INT8 quantization-aware export for Cortex-M4F.
Reach out via the GitHub repository or the professor/reviewer page.
Evidence and Status
Every status below comes from the same evidence registry as the rest of this portal (configs/pipeline_stages.yaml) — nothing on this page is a separate or more favorable accounting.
By category
View Source on GitHubFull Pipeline RegistryRead the Paper DraftDownload 3MT Slide
3MT source materials: speech transcript · original slide (PPTX)