v1.0.0 — Released May 2026

Molecular modeling
at scale

ChemLink is a modular orchestration platform for computational chemistry on HPC clusters. Automate docking campaigns and molecular dynamics simulations with a single unified CLI — from GPU-accelerated search to trajectory analysis.

ChemLink
2.2×
GPU speedup (3 nodes)
100%
Docking success rate
6
MD simulation types
MIT
License
Project status

Health overview

Current version
v1.0.0
Build status
Passing
GPU acceleration
CUDA 13 · sm_89/90/120
SLURM Job Arrays
Active
MPI support
Active (GROMACS)
Shared storage
NFS v4 (Linux, manager node)
Why ChemLink

Purpose-built for HPC

Running molecular docking campaigns and dynamics simulations on HPC infrastructure is complex, fragile, and deeply manual. ChemLink abstracts the orchestration layer so your team can focus on chemistry, not plumbing — reducing experiment setup from 45 minutes to under 5.

01

Automated orchestration

Reduces experiment setup time from 45 minutes to under 5. One unified CLI handles job submission, monitoring, and post-processing for both docking and dynamics workflows.

02

GPU-accelerated docking

Full pipeline: active site detection with fpocket, grid preparation with AutoGrid4, and conformational search with AutoDock-GPU targeting NVIDIA architectures sm_89, sm_90, and sm_120.

03

Molecular dynamics with GROMACS

Structured MD simulation from topology preparation to trajectory analysis, supporting six biological system types with CUDA and MPI acceleration via GROMACS 2025.

04

Distributed via SLURM

Both pipelines operate in single-node or distributed mode using SLURM Job Arrays. Scales docking campaigns across multiple GPU nodes with automatic work partitioning.

05

Fault-tolerant execution

Individual ligand or trajectory failures are isolated, logged with a full traceback, and skipped without stopping the campaign. A complete success/failure breakdown is reported at the end of every run.

06

Reproducible & traceable

Every run is logged and traceable. Prometheus and Grafana monitoring track cluster health in real time, and Conda-pinned environments guarantee reproducibility across nodes and users.

Architecture overview

ChemLink is built in five decoupled layers — CLI, pipelines, steps, adapters, and HPC infrastructure — so each scientific tool can be updated or swapped without touching the orchestration logic.

ChemLink layered architecture diagram
ChemLink layered architecture diagram
Changelog

Release notes

v1.0.0
Latest
2026-05-22
Initial stable release — full docking & dynamics pipelines
  • Molecular docking pipeline: active site detection (fpocket), grid preparation (AutoGrid4), GPU-accelerated search (AutoDock-GPU) with 100% success rate on 1,000-ligand campaigns
  • Molecular dynamics pipeline: GROMACS 2025.4 with CUDA + MPI support, covering 6 biological system types from topology preparation to trajectory analysis
  • Unified CLI (chemlink docking / chemlink dynamic / chemlink doctor) operating in single-node or SLURM Job Array distributed mode
  • SLURM Job Array support: automatic ligand-set partitioning across GPU nodes with fault-tolerant per-ligand error isolation
  • 2.2× speedup validated with 3-node GPU configuration; experiment preparation reduced from 45 min to <5 min
Documentation

Reference guide

Installation
Configuration
Commands
Flags & options
HPC cluster guide
Installation
Requirements

ChemLink runs on Linux HPC nodes with NVIDIA GPUs. The installer handles all scientific tool compilation and Conda environment creation.

  • Linux (Ubuntu 22.04+ recommended)
  • NVIDIA GPU with CUDA 12.x drivers
  • CUDA Toolkit 12.x, OpenMPI, cmake (for --full install)
  • SLURM (optional, for distributed mode)
Minimum tested hardware

The benchmarks reported below were obtained on the following configuration. Any node that meets or exceeds these specs should replicate the published performance.

ComponentManagerworker1worker2
CPUIntel Core Ultra 9 285 (24 cores / 24 threads)Intel Core i9-10900 est. (10 cores / 20 threads)Intel Core i7-12700 (12 cores / 20 threads)
RAM32 GB64 GB32 GB
GPURTX 5060 Ti 16 GB (Blackwell, sm_120)RTX 3080 10 GB (Ampere, sm_86)RTX 3060 LHR 12 GB (Ampere, sm_86)
StorageSK Hynix NVMe SSD 1 TBSamsung NVMe SSD 1 TBSamsung NVMe SSD 1 TB
NetworkIntel GbE + Wi-Fi 7Intel I219-LM GbE (1 Gbps)Intel I219-LM GbE (1 Gbps)
OSUbuntu 24.04 LTSUbuntu 24.04 LTSUbuntu 24.04 LTS
CUDA12.x12.x12.x
Quick install
bash# Installs ChemLink in /opt/chemlink with Conda envs bio + mgl_legacy
curl -fsSL https://raw.githubusercontent.com/PipeJF9/chemlink/main/install.sh | bash
Full install (with scientific tools)
bash# Compiles fpocket, AutoGrid4, AutoDock4, AutoDock-GPU, GROMACS 2025.4
# Estimated: 45–90 min — requires CUDA Toolkit, OpenMPI, cmake
curl -fsSL https://raw.githubusercontent.com/PipeJF9/chemlink/main/install.sh | bash -s -- --full
Installer options
OptionDescription
--fullCompile all scientific tools + GROMACS
--with-gromacsCompile GROMACS only
--dir PATHInstall directory (default: /opt/chemlink)
--version TAGGit branch/tag (default: main)
--skip-condaSkip Conda environment creation
Verify installation
bashchemlink doctor          # check environment, GPU, and dependencies
chemlink docking --help  # docking pipeline options
chemlink dynamic --help # dynamics pipeline options
Configuration

ChemLink pipelines are configured via CLI flags at runtime. The chemlink doctor command verifies that the environment is correctly set up before running a pipeline.

Conda environments
EnvironmentContentsUsed by
bioPython 3.10, ACPYPE, AmberTools, OpenBabel, RDKit, pdbfixer, biopythonDocking & dynamics preparation
mgl_legacyPython 2, MGLTools, pythonshAutoDock4 file preparation
Environment variables
VariableDescription
GMXRCPath to GROMACS environment script (auto-set by installer)
CUDA_VISIBLE_DEVICESOverride visible GPUs for docking runs
SLURM_ARRAY_TASK_IDSet automatically by SLURM for distributed array jobs
Shared storage

The lab uses two independent shared storage systems with distinct purposes:

SystemWhere it runsMount pointPurpose
NFS v4Linux kernel on manager node/nfs/chemlinkChemLink computation — code, Conda envs, inputs, intermediates, and results. Mounted on every compute node; required for all pipeline runs.
SMB/CIFS (Samba)OpenMediaVault NAS nodeper-user sharePersonal and team storage — accessible from Windows, macOS, and Linux. Used for backups, raw data, and researcher files. Not involved in ChemLink pipelines.
Commands
Core pipelines
CommandDescription
chemlink dockingRun the molecular docking pipeline (fpocket → AutoGrid4 → AutoDock-GPU)
chemlink dynamicRun the molecular dynamics pipeline (GROMACS 2025.4)
Utility
CommandDescription
chemlink doctorCheck environment, GPU availability, and all dependencies
Example: docking workflow
bash# Single-node docking run
chemlink docking \
  --receptor inputs/protein.pdb \
  --ligands  inputs/ligands/ \
  --out      results/docking/

# Distributed: SLURM Job Array across 4 GPU nodes
chemlink docking \
  --receptor /nfs/chemlink/protein.pdb \
  --ligands  /nfs/chemlink/ligands/ \
  --nodes    4 \
  --slurm
Example: dynamic workflow
bash# Protein-ligand MD simulation with GROMACS
chemlink dynamic \
  --system  protein-ligand \
  --input   inputs/complex.pdb \
  --out     results/md/ \
  --gpu     --mpi-ranks 8
Flags & options
chemlink docking
FlagDefaultDescription
--receptor PATHrequiredReceptor PDB file
--ligands PATHrequiredLigand file or directory (.pdbqt / .sdf)
--out PATH./docking_outOutput directory
--nodes N1SLURM node count for distributed mode
--slurmfalseSubmit via SLURM Job Array
--dry-runfalseValidate inputs without running
chemlink dynamic
FlagDefaultDescription
--system TYPErequiredprotein | protein-ligand | membrane | protein-membrane | ligand | custom
--input PATHrequiredInput structure file (.pdb / .gro)
--out PATH./md_outOutput directory
--gpufalseEnable CUDA GPU acceleration
--mpi-ranks N1MPI rank count for parallel run
--slurmfalseSubmit via SLURM
HPC cluster guide
Cluster requirements

ChemLink is designed for a SLURM-managed cluster with NFS shared storage. All nodes must be reachable via passwordless SSH from the head node. The cluster uses two independent storage systems: NFS v4 served by the Linux kernel on the manager node (exports /nfs/chemlink, mounted on all compute nodes — this is what ChemLink uses), and a dedicated OpenMediaVault NAS node that serves SMB/CIFS/Samba for personal and team file storage accessible from any OS. The two are fully independent and serve different purposes.

HPC cluster deployment diagram
bash# Verify cluster connectivity
chemlink doctor --check cluster

# Test NFS mount on all nodes
chemlink doctor --check nfs
SLURM Job Arrays

Pass --slurm and --nodes N to any pipeline to submit a SLURM Job Array. ChemLink automatically partitions the ligand set or MD replicas across array tasks.

bash# Docking campaign: 1000 ligands across 10 GPU nodes
chemlink docking \
  --receptor /nfs/chemlink/targets/cdk2.pdbqt \
  --ligands  /nfs/chemlink/ligands/set_1000/ \
  --nodes    10 \
  --slurm \
  --partition gpu
GPU architecture targets

AutoDock-GPU and GROMACS are compiled for sm_89 (Ada Lovelace), sm_90 (Hopper), and sm_120 (Blackwell) architectures. Run chemlink doctor to confirm your GPU is detected.

Monitoring

The cluster ships with Prometheus + Grafana for real-time node and GPU monitoring. Access the Grafana dashboard on the head node at port 3000.

Performance tips
  • Use NFS for all input/output paths — avoid local-only paths in multi-node runs
  • Set CUDA_VISIBLE_DEVICES when running multiple jobs on the same node
  • For RTX 5000 / Blackwell GPUs (sm_120), CUDA 13.0+ is required — use the provided Docker image
  • Run chemlink doctor before every campaign to catch config issues early
Open source

Contribute

ChemLink is built in the open at the Laboratorio de Química y Biología Computacional of Universidad del Norte. We welcome bug reports, documentation improvements, and new pipeline contributions.

Fork on GitHub

Fork the repository, create a feature/* or fix/* branch, and open a pull request against develop. All contributions go through code review.

Fork repo

Report issues

Found a bug or have a feature request? Open an issue with your OS, CUDA version, GPU model, and the full error output from chemlink doctor.

Open issue

Improve docs

Documentation source lives in /docs in the main repo. Corrections, translations, and new guides are all welcome.

Edit docs

Add a pipeline

ChemLink's modular architecture makes it straightforward to add new simulation backends or analysis stages. See the developer guide in docs/Desarrollo.md.

Developer guide
Students

Development team

Samuel Matiz García
Samuel Matiz García
Student · Lab. Q. y B. Computacional UniNorte
Juan Felipe Santos Rodríguez
Juan Felipe Santos Rodríguez
Student · Lab. Q. y B. Computacional UniNorte
Camilo Andrés Navarro Navarro
Camilo Andrés Navarro Navarro
Student · Lab. Q. y B. Computacional UniNorte
Tutors

Academic advisors

Daniel José Romero Martinez
Daniel José Romero Martinez
Tutor · Universidad del Norte
Augusto Salazar Silva
Augusto Salazar Silva
Tutor · Universidad del Norte
Credits

Collaborators

Edgar Alexander Márquez Brazon
Edgar Alexander Márquez Brazon
Professor · Computational Chemistry Lab
Oscar Andrés Saurith Coronell
Oscar Andrés Saurith Coronell
Script developer
José Elías Samur Benitez
José Elías Samur Benitez
Script collaborator
Francesco Genaro Rosa Chedraui
Francesco Genaro Rosa Chedraui
Script collaborator
MIT
©2026

License — MIT

ChemLink is released under the MIT License. You are free to use, modify, and distribute this software in any project — commercial or otherwise — provided the copyright notice and permission notice appear in all copies.

View full license text →

Ready to model?

Install ChemLink and run your first docking campaign or MD simulation in minutes.

Read the docs ★ Star on GitHub