A lightweight Chemical Translation Service using a curated subset of 13 million compounds from the PubChem database
- Please refer to the documentation for information regarding the API
- PubChem - database used for compound data
- RDKit - library used for SMILES to InChIKey conversion
- Wishart Research Group - Chemical Taxonomy using ClassyFire
- Go
- SQLite
- JavaScript/HTML/CSS
- Docker
- Playwright
- Locust (load testing)
- Run unit tests with
go test ./... - Run E2E playwright tests with
cd playwright && npm test
- CTS-Lite is containerized with Docker
- The GitHub Actions workflow will automatically build and deploy the complete application image upon any push or merge to the main branch
- Note: to build the docker image locally, you must have the database built and stored as
dataset/compounds.dbcd dataset && go run cmd/build-db/build-db.go cts-lite.csv compounds.db
- To create the csv dataset, simply run the
create_csv_dataset.shscript found undercmd/ - To update the dataset used by production, make sure you elect to push to S3 at the end of the script
- Then, the next time the app is deployed via GitHub Actions (push/merge to main), the latest dataset will be downloaded from S3 and the database will be rebuilt
- To create a local instance of compounds.db (SQLite database used by the app), run the build-db module like so:
cd dataset && go run cmd/build-db/build-db.go cts-lite.csv compounds.db