Paper

Evidence-controlled manuscript pipeline. Results, Discussion, and the final Abstract are intentionally left undrafted — see paper/RESULTS_DRAFT.md, which states plainly: “No prose result claims are drafted until verified artifacts and the claim-evidence matrix support them.”

Section status

SectionStatus
AbstractPlanned
IntroductionPlanned
Related WorkPlanned
Research GapPlanned
MethodologyRunning
DatasetRunning
Chronological Evaluation ProtocolRunning
ModelsRunning
Resource-Aware ExplainabilityPlanned
FidelityPlanned
StabilityPlanned
TinyML DeploymentPlanned
Energy MeasurementPlanned
ResultsPlanned
DiscussionPlanned
Threats to ValidityPlanned
ReproducibilityRunning
ConclusionPlanned

Claim–evidence matrix

Every candidate claim is linked to its experiment ID, dataset/config hash, git commit, and result artifact — status SUPPORTED or UNSUPPORTED, decided by the artifact, not by narrative convenience.

IDClaimMetricStatus
C001The independently acquired dataset passed structural validationvalidationSUPPORTED
C002Raw features have measured distribution differences relative to batch 1normalized_wassersteinSUPPORTED
C003All evaluated classical models have lower B10 than B2 accuracyaccuracySUPPORTED
C004Expanding-window retraining improves mean accuracy for all evaluated modelsaccuracy_changeSUPPORTED
C005IID diagnostic accuracy exceeds mean fixed-origin future accuracy for all evaluated modelsaccuracy_gapSUPPORTED
C006Global drift has a consistent statistically supported association with performanceSpearmanUNSUPPORTED
C007Some batch-2 predictive features combine above-median importance with below-median driftpermutation_importance_macro_f1SUPPORTED
C008Global permutation-importance (macro-F1) rankings and single-feature-ablation local explanations were generated for all four evidence-selected FIXED_ORIGIN models across all nine evaluation batches, with no unjustified explanation method forced onto an incapable modelpermutation_importance_macro_f1SUPPORTED
C-XAI-FID-01Top-ranked features produced greater predictive degradation than matched random features for every tested model/method/batch/K combination.selected_minus_random with CIUNSUPPORTED
C-XAI-FID-02Explanation fidelity was constant across chronological drift batches.batch-conditioned bootstrap CIUNSUPPORTED
C-XAI-FID-03Explanation fidelity was equivalent for correct and misclassified predictions.category-conditioned local fidelityUNSUPPORTED
C-XAI-STAB-01Top-K explanation rankings remain more similar across adjacent batches than widely separated anchor comparisons.adjacent_minus_far_anchor_jaccard_ci_low_gt_0UNSUPPORTED
C-XAI-STAB-02Explanation change increases as measured sensor-distribution shift increases across all eligible models.explanation_distance_vs_input_change_spearman_ci_low_gt_0UNSUPPORTED
C-XAI-STAB-03Explanation stability differs between correct and misclassified predictions.correct_minus_misclassified_distance_ci_excludes_0UNSUPPORTED
C-XAI-STAB-04Physical sensor-level attribution change increases systematically with chronological input drift across eligible models.sensor_distance_vs_input_change_spearman_ci_low_gt_0UNSUPPORTED
C-XAI-STAB-05Stage-10 fidelity and Stage-11 stability are positively associated across matched model/method analyses.matched_fidelity_stability_spearman_ci_low_gt_0UNSUPPORTED
C-XAI-COST-01Local explanation methods differ materially in host computational overhead relative to corresponding baseline inference.At least one matched local-method overhead ratio differs by >=2x: True.SUPPORTED
C-XAI-COST-02A substantial fraction of local ablation cost is explained by repeated model evaluations.Additional-call criterion=True; model-level Spearman rho=1.000, but positive bootstrap CI is not established with N=4 models.UNRESOLVED
C-XAI-COST-03Intrinsic extraction has lower recurring host cost than perturbation-based generation within comparable global scope.Intrinsic extraction is below permutation p05 for both applicable models: True.SUPPORTED
C-XAI-COST-04No single method dominates simultaneously on fidelity, stability, and host computational cost.Stage-10/11 method coverage is incomplete for a fully matched dominance test; no universal composite score was constructed.UNRESOLVED
C-EMBED-FP32-01MODEL-C1 standalone FP32 implementation reproduces the frozen reference within all Stage-13 criteria.all_preregistered_mandatory_criteriaUNSUPPORTED
C-EMBED-FP32-02MODEL-C4 standalone FP32 implementation reproduces the frozen reference within all Stage-13 criteria.all_preregistered_mandatory_criteriaUNSUPPORTED
C-EMBED-FP32-XAI-01MODEL-C1 local coefficient explanation is reproduced within separately frozen explanation criteria.all_preregistered_mandatory_criteriaUNSUPPORTED
C-EMBED-C1-REPAIR-01At least one prospectively defined all-FP32 explicit StandardScaler representation satisfies all frozen C1 criteria.all_frozen_mandatory_criteriaUNSUPPORTED
C-EMBED-C1-REPAIR-XAI-01The selected repaired C1 representation preserves local coefficient explanations within frozen criteria.all_frozen_mandatory_criteriaUNSUPPORTED

Full matrix: paper/claim_evidence_matrix.csv

Figures (38)

confusion matrix evolution

confusion matrix evolution · source

dataset timeline

dataset timeline · source

feature batch drift heatmap

feature batch drift heatmap · source

fid 01 deletion vs random

fid 01 deletion vs random · source

fid 02 chronological fidelity

fid 02 chronological fidelity · source

fid 03 error conditioned

fid 03 error conditioned · source

fid 04 sensor groups

fid 04 sensor groups · source

figure 10 class error heatmap

figure 10 class error heatmap · source

figure 11 fixed vs expanding

figure 11 fixed vs expanding · source

figure 12 iid vs chronological

figure 12 iid vs chronological · source

figure 13 drift vs performance

figure 13 drift vs performance · source

figure 14 feature drift vs importance

figure 14 feature drift vs importance · source

figure 15 complexity comparison

figure 15 complexity comparison

figure 6 accuracy

figure 6 accuracy · source

figure 7 macro f1

figure 7 macro f1 · source

figure 8 balanced accuracy

figure 8 balanced accuracy · source

figure 9 degradation

figure 9 degradation · source

global drift trajectory

global drift trajectory · source

lat 01 baseline host inference

lat 01 baseline host inference · source

lat 02 local explanation

lat 02 local explanation · source

lat 03 local overhead

lat 03 local overhead · source

lat 04 global total

lat 04 global total · source

lat 05 computational counts

lat 05 computational counts · source

lat 06 host latency distribution

lat 06 host latency distribution · source

lat 07 fidelity vs host cost

lat 07 fidelity vs host cost · source

lat 08 stability vs host cost

lat 08 stability vs host cost · source

lat 09 pre hardware evidence map

lat 09 pre hardware evidence map · source

sensor drift vs importance

sensor drift vs importance · source

stab 01 topk overlap

stab 01 topk overlap · source

stab 02 rank stability

stab 02 rank stability · source

stab 03 sensor stability

stab 03 sensor stability · source

stab 04 family evolution

stab 04 family evolution · source

stab 05 explanation vs input

stab 05 explanation vs input · source

stab 06 explanation vs performance

stab 06 explanation vs performance · source

stab 07 error conditioned

stab 07 error conditioned · source

stab 08 fidelity vs stability

stab 08 fidelity vs stability · source

top20 drifting features

top20 drifting features · source

top20 stable features

top20 stable features · source

Tables (6)