Identify the sequence determinants of cell-type-specific translational control
I am training a convolution- and transformer-based model to predict per-nucleotide ribosome footprint density and RNA abundance directly from sequence, across a curated atlas of 78 human and 68 mouse cell types, using multi-species data as a form of evolutionary regularization. Two attribution methods — in silico mutagenesis and integrated gradients — will be used to extract the sequence grammar the model learns, and predicted variants will be tested experimentally in luciferase reporter constructs across multiple cell lines.
- Nucleotide-resolution Ribo-seq and RNA-seq prediction
- Resolves initiation from elongation
- Model and trained weights released publicly