A python library to work with molecules. Built on top of RDKit.
Project description
Molecular Manipulation Made Easy
Datamol is a python library to work with molecules. It's a layer built on top of RDKit and aims to be as light as possible.
- 🐍 Simple pythonic API
- ⚗️ RDKit first: all you manipulate are
rdkit.Chem.Mol
objects. - ✅ Manipulating molecules often rely on many options; Datamol provides good defaults by design.
- 🧠 Performance matters: built-in efficient parallelization when possible with optional progress bar.
- 🕹️ Modern IO: out-of-the-box support for remote paths using
fsspec
to read and write multiple formats (sdf, xlsx, csv, etc).
Try Online
Documentation
Visit https://doc.datamol.io.
Installation
Use conda:
mamba install -c conda-forge datamol
Quick API Tour
import datamol as dm
# Common functions
mol = dm.to_mol("O=C(C)Oc1ccccc1C(=O)O", sanitize=True)
fp = dm.to_fp(mol)
selfies = dm.to_selfies(mol)
inchi = dm.to_inchi(mol)
# Standardize and sanitize
mol = dm.to_mol("O=C(C)Oc1ccccc1C(=O)O")
mol = dm.fix_mol(mol)
mol = dm.sanitize_mol(mol)
mol = dm.standardize_mol(mol)
# Dataframe manipulation
df = dm.data.freesolv()
mols = dm.from_df(df)
# 2D viz
legends = [dm.to_smiles(mol) for mol in mols[:10]]
dm.viz.to_image(mols[:10], legends=legends)
# Generate conformers
smiles = "O=C(C)Oc1ccccc1C(=O)O"
mol = dm.to_mol(smiles)
mol_with_conformers = dm.conformers.generate(mol)
# 3D viz (using nglview)
dm.viz.conformers(mol, n_confs=10)
# Compute SASA from conformers
sasa = dm.conformers.sasa(mol_with_conformers)
# Easy IO
mols = dm.read_sdf("s3://my-awesome-data-lake/smiles.sdf", as_df=False)
dm.to_sdf(mols, "gs://data-bucket/smiles.sdf")
How to cite
Please cite Datamol if you use it in your research: .
Compatibilities
Version compatibilities are an essential topic for production-software stacks. We are cautious about documenting compatibility between datamol
, python
and rdkit
.
See below the associated versions of Python and RDKit, for which a minor version of Datamol has been tested during its whole lifecycle.
datamol |
python |
rdkit |
---|---|---|
0.7 |
[3.8, 3.9] |
[2021.09, 2022.03] |
0.6 |
[3.8, 3.9] |
[2021.09] |
0.5 |
[3.8, 3.9] |
[2021.03, 2021.09] |
0.4 |
[3.8, 3.9] |
[2020.09, 2021.03] |
0.3 |
[3.8, 3.9] |
[2020.09, 2021.03] |
CI Status
The CI run tests and perform code quality checks for the following combinations:
- The three major platforms: Windows, OSX and Linux.
- The two latest Python versions.
- The two latest RDKit versions.
main |
|
---|---|
Lib build & Testing | |
Code Sanity (linting and type analysis) | |
Documentation Build |
Changelogs
See the latest changelogs at CHANGELOG.rst.
License
Under the Apache-2.0 license. See LICENSE.
Authors
See AUTHORS.rst.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file datamol-0.7.12.tar.gz
.
File metadata
- Download URL: datamol-0.7.12.tar.gz
- Upload date:
- Size: 278.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.1 CPython/3.10.6
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | 9f9ec5380d71dddc0e15b095e2e013c705eb4b5979b07370ca176095688d3f11 |
|
MD5 | 72b7cdbe17b074f76bcff275583fd47b |
|
BLAKE2b-256 | 1c0e84e95f269b36b4302e3edeaa0c24c524f7edd26b8a30ded2a46179bc383e |