Master DNA Large Language Models (DNA-LLMs): Leveraging AI and NLP for Genomic Sequence Analysis in 4 weeks through hands-on, project-based online training with DSTC.
Bioinformatics & Computational Biology
Module-by-module breakdown of DNA Large Language Models (DNA-LLMs): Leveraging AI and NLP for Genomic Sequence Analysis, from foundations to a certified capstone project.
Adaptation
โข Why DNA is not language: no words, no sentences, weak compositional grammar
โข Tokenisation choices โ single nucleotide, k-mer, byte-pair โ and their consequences
โข Context window against genomic distance, and the enhancer problem it creates
Models
โข DNABERT, Nucleotide Transformer, HyenaDNA and Evo compared on context length
โข Attention against state-space architectures for very long sequences
โข Pretraining corpora and the species bias they carry
Prediction
โข Promoter, enhancer and splice site identification
โข Chromatin accessibility and expression prediction, and Enformer as the reference point
โข Non-coding variant effect prediction where laboratory data is scarce
Practice
โข Task heads, fine-tuning strategy and class imbalance in genomic labels
โข Attention and attribution maps, and the weakness of reading biology from them
โข In silico mutagenesis as a more defensible interpretation method
Evaluation
โข Chromosome-level splits, since random splits leak through sequence homology
โข Comparison against position weight matrices and CNN baselines
โข Experimental validation such as MPRA before a regulatory claim is made
e-Certificate and e-Marksheet issued on successful completion.