GitHub: github.com/LeonardFH/sleipnir-hcon
PyPI: pypi.org/project/sleipnir-hcon
A fast, interpretable, dictionary-based framework for heat of formation (HOF) prediction from SMILES strings.
SleipnirHCON (pronounced "SLAYP-neer-HCON") predicts heat of formation using a [PLACEHOLDER: method summary - reference-state composition model with bond-overlap corrections / group-additivity scheme / etc.]. It requires no 3D conformers, no DFT calculations, and no GPU acceleration.
Named after the eight-legged horse of Odin, reflecting the software's intended speed and its ability to traverse wide molecular property spaces.
The package implements heat of formation prediction for organic molecules containing C, H, O, N, and common heteroatoms (S, F, Cl, Br, P, I).
The framework is designed with extensibility in mind, allowing additional molecular property predictors to be added in future versions.
| Metric | Value |
|---|---|
| Mean Absolute Error | [PLACEHOLDER] kJ/mol (pooled, across [PLACEHOLDER] molecules, [PLACEHOLDER] independent datasets) |
| Inference Speed | [PLACEHOLDER] molecules/second (single core) / [PLACEHOLDER] molecules/second ([PLACEHOLDER] cores) |
| Parameters | [PLACEHOLDER] (fully interpretable [PLACEHOLDER] coefficients) |
| Hardware | Standard laptop CPU (no GPU required) |
A full account of the method, validation, and benchmark results is available as a preprint:
Haasbroek, L. F. (2026). SleipnirHCON: [PLACEHOLDER: full paper title]. ChemRxiv. DOI: 10.XXXX/chemrxiv-2026-XXXXX
For detailed performance across specific datasets, convergence behaviour, stability analyses, and [PLACEHOLDER], please refer to the paper.
pip install sleipnir-hconfrom sleipnir import train_hof, predict_hof, predict_hof_batch
import pandas as pd
from sklearn.metrics import mean_absolute_error
# Train a dictionary on your own dataset
weights = train_hof(
data_path="trainingdata.csv", # columns: SMILES, HOF
output_path="my_weights.pkl",
filter_cocrystals=True, # Recommended for pure crystals
filter_hcon=True, # Recommended for H,C,O,N only
verbose=True
)
# Predict a single molecule
hof = predict_hof("CCO", weights_path="my_weights.pkl")
# Predict a batch of molecules
smiles_list = ["CCO", "CC", "c1ccccc1", "O"]
results = predict_hof_batch(smiles_list, weights_path="my_weights.pkl")SleipnirHCON explicitly challenges the assumption that high-accuracy heat of formation prediction requires deep learning, 3D conformers, or expensive quantum calculations.
The paper demonstrates that a physically motivated linear model with fewer than [PLACEHOLDER] parameters can achieve competitive accuracy while being:
- Transparent - the [PLACEHOLDER: coefficients] represent [PLACEHOLDER: physical interpretation]. Their relative magnitudes provide chemical insight into which [PLACEHOLDER] contribute most to [PLACEHOLDER].
- Fast - microsecond-scale inference on commodity hardware.
- Stable - dictionaries transfer across independent datasets and converge rapidly.
- Diagnostic - the model can identify systematic biases in [PLACEHOLDER] datasets.
Important note: [PLACEHOLDER: any caveats about the interpretation of the fitted coefficients, analogous to the Hofvarpnir note about the reference volume + corrections being a paired system.]
The training and evaluation data used in the paper may be obtained from the following publicly available sources:
-
[PLACEHOLDER Author (Year)]: [PLACEHOLDER full citation]. DOI: 10.XXXX/XXXXX
- Dataset: [PLACEHOLDER repository link]
-
[PLACEHOLDER Author (Year)]: [PLACEHOLDER full citation]. DOI: 10.XXXX/XXXXX
- Dataset: [PLACEHOLDER repository link]
-
[PLACEHOLDER Author (Year)]: [PLACEHOLDER full citation]. DOI: 10.XXXX/XXXXX
These datasets are available as Supporting Information with their respective papers or via the linked public repositories.
For optimal accuracy, we recommend training separate dictionaries for each chemical family:
- HCON only (C, H, N, O) - best overall performance
- HCON + F - fluorine-containing molecules
- HCON + Cl - chlorine-containing molecules
- HCON + S - sulfur-containing molecules
- HCON + P - phosphorus-containing molecules
Avoid mixing different heteroatom types (e.g., S and Cl together) in a single training run, as this can degrade prediction accuracy.
For molecules containing rare halogens (Br, I), the HCON-only dictionaries are recommended, as there is insufficient data to train reliable halogen-specific [PLACEHOLDER].
SleipnirHCON handles co-crystals (SMILES strings containing a dot, e.g., "CCO.O=C(O)C") using [PLACEHOLDER: mass-weighted averaging / stoichiometric mixing / etc.] of the predicted heat of formation of each component.
For datasets containing a large number of co-crystals, improved accuracy can be achieved by training separate dictionaries on co-crystal data only. For datasets with only a few co-crystals, the pure-trained dictionaries provide reliable estimates.
For detailed co-crystal performance, see the paper.
[PLACEHOLDER: one or two paragraphs analogous to the Polymorphs section in the Hofvarpnir README, covering any structural or phase subtleties specific to HOF prediction.]
If you use SleipnirHCON on your own dataset, I invite you to share your results.
Email: leonardfhaasbroek@gmail.com
Please include:
- MAE, RMSE, R2
- Number of molecules
- Dataset description and source (if public)
Results will be posted here (with your permission).
Hi there,
I built SleipnirHCON because heat of formation prediction should be fast, transparent, and accessible. I'm glad you found it.
If you need to get in touch: leonardfhaasbroek@gmail.com
This project is distributed under the BSD 3-Clause License.
If you use this software or method in your research, please use the following citation format:
Haasbroek, L. F. (2026). SleipnirHCON: Fast dictionary-based heat of formation prediction (Version 0.1.0) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.XXXXXXX
Leonard F. Haasbroek
leonardfhaasbroek@gmail.com
