Drift-Robust Explainable TinyML for Electronic-Nose Sensing

Can lightweight models retain useful predictive performance and interpretable behavior under chronological sensor drift while fitting resource-constrained TinyML hardware?

Electronic-nose sensors age and drift — a model trained on early sensor readings loses accuracy as time passes, even though nothing about the underlying chemistry changed. This project evaluates that decay honestly (in real chronological order, never shuffled), studies whether explanations of a model's predictions stay stable as sensors drift, and works toward a resource-aware path to a physical nRF52840 microcontroller — while reporting plainly what has and has not yet been executed.

Explore the Research →Explore the Digital TwinSee What's Verified

New to the project? Start here.

A recommended reading order, in roughly three minutes to three hours:

  1. 1This page: the 3-minute version
  2. 2How it works, step by step
  3. 3The interactive digital twin
  4. 4The full system architecture
  5. 5Chronological evaluation methodology
  6. 6Every registered experiment
  7. 7Results dashboard
  8. 8How to reproduce this

The 3MT Speech and Slide

The actual speech and single slide submitted for the FAU 2026 Three Minute Thesis competition, reproduced verbatim — not a paraphrase.

WHEN AI LOSES ITS SENSE OF SMELL

Electronic noses can learn. Sensors can drift.

MONTH 381%baseline accuracy
SENSOR DRIFTthe world changes
while the model stays the same
MONTH 3637%same baseline, later sensors

Across 4 baseline models: 35–49 percentage-point accuracy loss (Month 3 → Month 36)

Research goal: build small AI that stays trustworthy as sensors change.

Arun Kumar Gharami | Ph.D. in Computer Engineering | Florida Atlantic University

Recreated for the web from the original single slide. Download the original slide (PPTX) · source: research-portal/public/3mt/Arun_Gharami_FAU_3MT_2026_Single_Slide.pptx

Imagine opening your refrigerator and smelling milk that has gone bad.

Before you read the date, your nose warns you.

Now imagine giving that ability to a machine.

An electronic nose uses chemical sensors and artificial intelligence to recognize patterns in gases. These systems could help monitor food, detect hazardous chemicals, and support environmental or health applications.

But there is a hidden problem.

The nose changes.

Electronic sensors age. Temperature and humidity affect them. Their signals slowly shift, even when the chemical being measured has not changed. Scientists call this sensor drift.

And when the sensor changes, an AI model trained yesterday may not understand the world tomorrow.

In this research, use an electronic-nose dataset collected over three years. Trained four baseline AI models on the earliest sensor readings and tested them as time moved forward.

The result was striking.

Between Month 3 and Month 36, every model lost roughly 35 to 49 percentage points of accuracy. In one lightweight model, accuracy fell from about 81 percent to 37 percent.

That means an AI system can look excellent when it is first built—and become unreliable simply because its sensors grow older.

So my research asks a different question:

Can we build small AI systems that remain trustworthy when their sensors and environment change?

Developing a drift-aware, explainable TinyML research pipeline. Instead of randomly mixing old and new sensor data, I evaluate models in chronological order, the way they would experience the real world. I examine not only whether a prediction is correct, but also whether the reasons behind that prediction stay stable as the sensors drift. Then I study which lightweight models are realistic candidates for resource-constrained embedded devices.

This matters because the future of AI is not only in giant data centers.

It is moving into factories, farms, wearable devices, environmental monitors, and small sensors that may operate for months or years.

For trustworthy AI, high accuracy on day one is not enough.

We need confidence that the system will still deserve our trust on day one thousand.

An electronic nose how to recognize when its world has changed.

Because when AI gains a sense of smell, it should not forget how to use it.

Source: docs/3mt/speech.txt

Complete Research Outline

Motivation and research gap

Electronic-nose classifiers are usually evaluated on randomly shuffled train/test splits, which can blend sensor behavior from different time periods and produce an optimistic estimate of real-world performance. This project instead evaluates strictly in chronological order, and treats explainability and TinyML deployability as first-class, separately-measured research questions rather than afterthoughts.

Research questions and hypotheses

10 research questions (RQ1–RQ10) and 10 hypotheses (H1–H10), each tied to explicit acceptance criteria and linked experiment IDs. See research/questions.yaml and research/hypotheses.yaml.

Dataset and chronology

UCI Gas Sensor Array Drift Dataset: 13910 samples, 128 features (16 sensors × 8 response characteristics), 10 chronological batches, 6 gas classes. Executed — full provenance on /dataset.

Data validation, frozen preprocessing, leakage controls

Preprocessing (standardization) is fit once on Batch 1 only and frozen; the chronological split is defined before any model is trained. Details on /methodology and configs/chronological_protocol.yaml.

Baselines, fixed-origin evaluation, expanding-window adaptation

Four classical models (Logistic Regression, Random Forest, RBF-SVM, MLP) evaluated under FIXED_ORIGIN (train once on Batch 1) and EXPANDING_WINDOW (periodic retraining) protocols. Executed — see /results.

Resource-aware explanations, fidelity, and stability

Permutation importance, intrinsic coefficients/impurity, and single-feature ablation — deliberately not SHAP or LIME. Fidelity and stability evidence is executed but predominantly unsupported or mixed, stated plainly rather than minimized. Executed — see /xai.

Quantization and export for nRF52840 / Cortex-M4F

Host-compiled (x86-64) numerical-equivalence export exists for two model candidates; the first two export attempts failed their preprocessing-equivalence criteria before a fused-architecture variant passed on host only. INT8 quantization has not been executed at any tier. See /tinyml.

Planned physical measurement

Flash, SRAM, on-device latency, and PPK2 energy all require a physical nRF52840 board and debug probe. Hardware blocked — no board has ever been detected. See /hardware.

Pareto analysis, reproducibility, limitations, next experiments

The accuracy/fidelity/resource Pareto analysis is blocked on physical hardware measurement and has not been produced. Reproducibility guarantees (seed, dataset hash, config hash, git commit per artifact) are documented on /reproducibility; the next planned experiment and every known limitation are stated on /professor-review.

Interactive Research View

The conceptual 3D pipeline below is the same evidence-driven digital twin used on /digital-twin — sample chamber → sensor array → drift detector → classical model → explanation → nRF52840 (currently hardware-blocked) → gateway → evidence registry. Every status shown is real; nothing here is a fabricated measurement.

Loading interactive 3D model…

Move between chronological batches

Select a test batch to see its real drift score (relative to Batch 1) and the lightest baseline model's real accuracy at that point in time — the same numbers behind the 3MT slide's 81%→37% figure, drawn directly from results/drift/global_drift_by_batch.csv and results/baselines/fixed_origin_by_batch.csv.

Drift vs. Batch 1 (normalized Wasserstein, median across features)0.579
MODEL-C1 (Logistic Regression) accuracy at this batch81.4%

All four baseline models, accuracy by batch

Full data (text equivalent)

BatchDrift vs. Batch 1 (normalized Wasserstein)MODEL-C1 accuracy
20.57981.4%
30.82162.9%
40.99654.0%
51.16435.0%
60.58746.7%
70.60243.1%
80.62132.7%
90.68034.3%
100.41237.4%

New Doors This Work Could Open

These are potential application directions that this methodology could inform — they are not demonstrated outcomes, and each would require its own domain-specific validation.

Reliable environmental sensing

Chronological-drift-aware evaluation is directly applicable to any long-lived environmental sensor network (air quality, water quality) where recalibration is expensive or impossible.

Industrial gas and process monitoring

Industrial electronic-nose deployments run for years without replacement — a drift-honest evaluation protocol could inform maintenance/recalibration scheduling.

Food quality assessment

Spoilage and freshness detection is a natural electronic-nose application; this work's chronological protocol could validate whether a food-safety model still holds after sensor aging.

Portable diagnostics research

Breath- or odor-based diagnostic research faces the same sensor-drift risk; resource-aware explainability could support clinician trust in a constrained device.

Edge AI under distribution shift, broadly

The chronological-evaluation and resource-aware-explanation methodology is not electronic-nose-specific — it generalizes to any TinyML system where the sensor or input distribution changes after deployment.

Collaboration

This research is at a stage where specific kinds of help would directly unblock the next experiments:

  • Experimental review — a second opinion on the chronological protocol, the resource-aware explanation methodology, or the fidelity/stability evaluation design.
  • nRF52840 hardware access — a physical nRF52840 development kit and debug probe would unblock Stages 15–20 (the entire physical measurement chain), which have been architecturally ready and blocked on hardware access since Stage 13.
  • Measurement guidance — experience with Nordic Power Profiler Kit II energy-measurement methodology, or with INT8 quantization-aware export for Cortex-M4F.

Reach out via the GitHub repository or the professor/reviewer page.

Evidence and Status

Every status below comes from the same evidence registry as the rest of this portal (configs/pipeline_stages.yaml) — nothing on this page is a separate or more favorable accounting.

16/ 27 stages
2/ 27 stages
2/ 27 stages
6/ 27 stages
1/ 27 stages

By category

Software research pipeline (baselines, drift analysis)
Executed
Explainability (fidelity/stability mixed)
Executed
TinyML export / quantization
Not executed
Physical MCU deployment
Hardware blocked
Latency / energy (physical)
Not measured
Manuscript
DRAFT_EVIDENCE_ONLY
Hugging Face dataset/model repos
In preparation — see /huggingface

View Source on GitHubFull Pipeline RegistryRead the Paper DraftDownload 3MT Slide

3MT source materials: speech transcript · original slide (PPTX)