Archived

This content is available here for research, reference, and/or recordkeeping.

Author ORCID Identifier

https://orcid.org/0009-0008-7985-6309

Date Available

7-30-2026

Year of Publication

2026

Document Type

Doctoral Dissertation

Degree Name

Doctor of Philosophy (PhD)

College

Engineering

Department/School/Program

Computer Science

Faculty

Qiang Cheng

Abstract

High-throughput transcriptomic and proteomic technologies have enabled opportunities for studying biological processes and disease mechanisms. However, extracting meaningful biological information from these high-dimensional datasets remains challenging due to limited sample sizes, biological heterogeneity, measurement noise, and the absence of biological annotations. In particular, many molecular datasets lack temporal information required for circadian analysis, making the study of circadian rhythms difficult. Moreover, disease diagnosis and biomarker discovery from transcriptomic data often rely on complex machine learning models whose predictions are difficult to interpret and may not generalize well across independent datasets and experimental platforms. These limitations motivate the development of artificial intelligence methods capable of uncovering hidden biological structure and producing biologically meaningful insights from large-scale omics data.

This dissertation develops artificial intelligence frameworks for two fundamental biological discovery problems: circadian rhythm recovery, and interpretable early disease diagnosis and biomarker discovery. For circadian rhythm recovery, unsupervised deep learning approaches are developed to infer biological time directly from transcriptomic and proteomic measurements without requiring temporal annotations or predefined rhythmic markers. The proposed methods accurately recover circadian organization across diverse datasets and enable the identification of rhythmic genes and proteins associated with biological processes and disease states. Application of these approaches to Alzheimer's disease datasets reveals widespread alterations in molecular rhythmicity and provides new insights into disease-associated circadian dysregulation.

For early disease diagnosis, interpretable artificial intelligence frameworks are developed to identify compact molecular signatures and derive explicit diagnostic formulas from transcriptomic data. Applied to hepatocellular carcinoma and lung adenocarcinoma, the proposed methods identify small sets of informative biomarkers and achieve strong predictive performance across multiple independent datasets while maintaining transparency and interpretability. Unlike many existing machine learning approaches, the resulting models provide simple symbolic formulas that facilitate both accurate disease classification and biological interpretation.

Overall, this dissertation demonstrates how artificial intelligence can be used to recover latent biological structure from unlabeled molecular data and provide interpretable formulas based on minimal set of biomarkers for disease diagnosis. The proposed frameworks contribute new computational approaches for circadian biology and early disease identification based on biomarker discovery, while providing general strategies for extracting biologically meaningful information from high-dimensional omics datasets.

Digital Object Identifier (DOI)

https://doi.org/10.13023/etd.2026.374

Archival?

Archival

Funding Information

This research was supported through my advisor by the National Artificial Intelligence Research Resource (NAIRR) Pilot NSF OAC 240219, and Jetstream2, Bridges2, and Neocortex Resources since 2025 till present. Additional support was provided by the National Institutes of Health (NIH) under Grant R21AG070909,  P30AG072946, and R01HD101508-01, and the National Science Foundation (NSF) under Grants IIS 2327113, and ITE 2433190 during years 2023-present.

Share

COinS