Archived

This content is available here for research, reference, and/or recordkeeping.

Author ORCID Identifier

https://orcid.org/0009-0002-7986-003X

Date Available

7-27-2026

Year of Publication

2026

Document Type

Doctoral Dissertation

Degree Name

Doctor of Philosophy (PhD)

College

Engineering

Department/School/Program

Electrical and Computer Engineering

Faculty

Henry Dietz

Faculty

Simone Silvestri

Abstract

Artificial intelligence (AI) has demonstrated tremendous success in handling complex tasks across multiple domains, a leap largely attributed to complex deep neural network (DNN) architectures. However, the increasing demand for real-time processing, low latency, limited connectivity, and enhanced privacy necessitates edge intelligence, where computation is performed directly on resource-constrained embedded devices rather than relying on cloud infrastructures. While optimized TinyML techniques such as pruning, quantization, and neural architecture search (NAS) have made edge deployment feasible by significantly reducing computation and memory demands, achievable performance remains severely limited by hardware constraints. Furthermore, models deployed in highly dynamic real-world environments suffer from severe performance degradation due to domain shifts and non-Identically and Independently Distributed (non-IID) data distributions, leading to catastrophic forgetting when naively fine-tuned. This dissertation addresses these critical challenges by establishing novel, resource-efficient strategies tailored for both model inference and autonomous on-device adaptation.   To overcome foundational computational and memory bottlenecks, this research introduces highly compact, integer-only Keyword Spotting (KWS) architectures tailored for extreme-edge deployment. As a baseline, we proposed integer-based feature extraction utilizing Short-Time Fourier Transforms (STFT). To significantly reduce memory footprint and latency, we instead suggest using Wavelet Packet Transforms (WPT). This approach balances time-frequency precision while bypassing the hardware limitations of standard microcontrollers.   To tackle the critical issue of post-deployment domain shifts, this work presents a robust Domain-Incremental Continual Learning (DICL) framework designed specifically for embedded systems. The framework incorporates multi-stage denoising, utilizing discrete wavelet transform and spectral subtraction techniques to ensure the extraction of reliable audio features in unpredictable and noisy environments. It enables efficient on-device adaptation by continually updating the complete quantized model using a lightweight rehearsal buffer of initial training data. To capture real-time environmental shifts, the system dynamically identifies "effective samples" during runtime. These samples are filtered based on classification confidence and latent-space class prototype similarity, assigned pseudo-labels, and integrated into the retraining mini-batch. This drift-aware methodology successfully mitigates catastrophic forgetting, allowing the model to acclimate to new acoustic domains and maintain robust performance even in highly degraded signal-to-noise ratio (SNR) conditions.   Building upon the optimized feature extraction pipelines, this research introduces the efficient Dual-Feature-Map (DFM-KWS) architecture to maximize classification accuracy without inflating the parameter count. By deploying a dual-branch feature extractor utilizing broadcasted residual blocks (BCResBlocks), the model independently processes the broad energy context of Log-Mel spectrograms alongside the fine-grained acoustic texture of Mel-frequency cepstral coefficients (MFCCs). This unified feature set is fused through grouped depthwise and pointwise convolutions, allowing the DFM-KWS model to drastically reduce computational overhead and parameter count while surpassing state-of-the-art architectures in 12-class keyword classification tasks.   Central to our adaptive framework is a dual-feature-based anomaly and concept drift detection mechanism designed to maintain model reliability in non-stationary environments. By monitoring the latent projections of the DFM-KWS model, the system calculates the distance between incoming audio samples and the global class centroids within a learned latent space. This dual-feature input, incorporating both spectral and cepstral characteristics, provides a high-fidelity representation of the input signal. When the distance from the global centroid exceeds a pre-defined threshold, the system flags the occurrence as a potential anomaly. To distinguish between transient noise spikes and genuine concept drift, we employ a temporal hit-buffer that tracks the frequency of these anomalies. If an anomaly persists beyond a specific temporal threshold, the system confirms the presence of concept drift and triggers an automatic, on-device recalibration of the model’s baseline, effectively isolating environmental changes from stable keyword characteristics.   Ultimately, this dissertation bridges the gap between computational efficiency and real-world applicability in the TinyML paradigm. By holistically optimizing the feature extraction pipeline, the classifier architecture, and drift-aware continual learning strategies, this research provides a comprehensive foundation for deploying autonomous, highly robust, and lifelong-learning AI solutions directly on low-power embedded devices.

Digital Object Identifier (DOI)

https://doi.org/10.13023/etd.2026.363

Archival?

Archival

Share

COinS