Skip to content

About

Protein secondary structure predicition using ESMFold embeddings and deep learning

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

minor-project

Protein secondary structure prediction using ESMFold embeddings and a small downstream classifier. Final-year minor project for my undergrad. Currently a scaffolding exercise — code is being developed locally and will land here in pieces.

The setup

The classical pipeline for predicting per-residue secondary structure (helix / sheet / coil, or the finer 8-class DSSP scheme) trains a convolutional or recurrent network directly on the amino-acid sequence. That works, but it leaves a lot of free signal on the table: a protein language model like ESMFold has already learned a rich embedding for each residue based on the full sequence context, evolutionary co-variation, and structural priors.

The project: take residue-level embeddings out of ESMFold, freeze them, and train a small head (a 1D-CNN or a 2-layer transformer) to do the structure classification. The hypothesis is that the embeddings carry enough structural information that the head can be tiny, train in minutes, and still beat baseline LSTM-on-one-hot models on standard benchmarks (CB513, TS115, CASP12).

Status

What's here right now: nothing but LICENSE and .gitignore. I'm developing on a local machine with the dataset checked out, and the embedding-extraction step takes long enough that I haven't bothered pushing intermediate state. Real commits will start landing once the training loop converges on something honest enough to share.

If you've stumbled here looking for a working ESMFold + secondary-structure pipeline — this isn't it yet. Come back later, or just use SecondaryStructure-ESM as a starting point.

Planned layout

data/                  Pre-computed ESMFold embeddings (cached, gitignored — too big)
notebooks/             Exploratory analysis, baselines
src/
  ├── embeddings.py    ESMFold extraction wrapper
  ├── dataset.py       Per-residue label loader (DSSP-derived)
  ├── model.py         The downstream head
  └── train.py         Training loop with W&B logging
results/               Per-checkpoint metrics, comparison tables

Why this and not a transformer-from-scratch

I considered training a transformer end-to-end on raw sequences as the project. Two reasons against:

  1. The compute budget (a single 4090 borrowed for evenings) doesn't really cover end-to-end training of a sequence model big enough to be interesting.
  2. The intellectually honest answer to "how would you predict secondary structure today" is not "ignore the existing protein LM"; it's "stand on its shoulders and ask what you can do with embeddings as a frozen feature."

Reframing the question that way also makes the project tractable in the time I have.

About

Protein secondary structure predicition using ESMFold embeddings and deep learning

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors