Skip to content

Package LoQI for pip installation with a Python API and CLI - #6

Merged
FilippNikitin merged 8 commits into
mainfrom
packaging/pip-installable-loqi
Sep 15, 2026
Merged

FilippNikitin merged 8 commits into
mainfrom
packaging/pip-installable-loqi

Conversation

@bkalita23

@bkalita23 bkalita23 commented Sep 4, 2026 •

Copy link
Copy Markdown

Makes LoQI installable with pip, with load_model() and generate_conformers() as the Python entry points and loqi sample / loqi download as CLI commands.

  • Builds the loqi and megalodon packages with Hatchling, including inference configs, existing model resources, and license notices.
  • Defines minimum inference dependencies and optional train, aimnet, and dev extras.
  • Resolves registered or local checkpoints, verifies downloaded files by SHA-256, and caches them for reuse.
  • Uses PyTorch Geometric scatter fallbacks so the released LoQI models run without compiled PyG extensions.
  • Shares molecule preparation and sampling between the Python API, CLI, and existing sampling script. Keeps public entry points stable and uses concise private helper names and documentation.
  • Adds Python 3.11/3.12 unit CI and scheduled or manual sampling tests.

Validation: all 57 tests pass against the installed wheel on CPU with Python 3.12.12 and PyTorch 2.14.0+cpu, including released-checkpoint sampling and reproducibility. Six fixed-seed conformers have exactly identical coordinates before and after the refactor. Computational expressions and parsed configuration values are unchanged. Ruff and formatting checks pass. Both wheel and source distributions build; rebuilding the wheel from the source archive produces identical contents, with configs, model resources, and all three license files present.

Replace the setuptools configuration (dynamic version from a missing VERSION
file, exact pins including CUDA-specific torch_scatter/torch_sparse wheels) with
a hatchling build of distribution "loqi" 0.1.0 that ships the megalodon and
loqi import packages.

Runtime dependencies are inference-only floors: torch>=2.8, torch_geometric>=2.6,
rdkit>=2025.3, lightning>=2.4, omegaconf>=2.3, einops>=0.8, numpy>=1.26,
scipy>=1.11, tqdm>=4.66. Training/preprocessing packages move to the [train]
extra, AIMNet2 to [aimnet], pytest/ruff to [dev]. License metadata covers
LICENSE.md (MIT) and megalodon_licence/ (Apache-2.0 NVIDIA code and third-party
notices). requirements.txt now mirrors the floors without the PyG wheel index,
and the README installation section describes uv/pip installs from GitHub
(PyPI publication is a later step). Adds pytest (slow marker deselected by
default) and ruff (line length 120, new code only) configuration.
….scatter

Add megalodon.scatter, a drop-in module for the torch_scatter functions the
models use (scatter, scatter_sum/scatter_add, scatter_mean, scatter_softmax)
implemented with torch_geometric.utils.scatter and torch_geometric.utils.softmax.
Those fall back to pure torch (index_add_/scatter_reduce) when the compiled
extension is absent, so neither torch_scatter nor torch_sparse has to be built
for the installed torch/CUDA combination. Signatures follow torch_scatter (dim
defaults to -1, reduce="add" is accepted); the unused out= argument raises.

The 16 modules that imported from torch_scatter now import from
megalodon.scatter. torch_sparse is only referenced lazily inside the EQGAT
edge message-passing path (eqgat_denoising_model.py), which the LoQI configs
do not use, so it stays optional.
src/loqi is a thin public layer over the megalodon code:

- registry: MODELS table for the released loqi (diffusion) and loqi_flow
  (flow matching) checkpoints from the KiltHub record doi:10.1184/R1/31441570,
  and checkpoint_path(), which downloads with the standard library into
  $LOQI_CACHE_DIR or ~/.cache/loqi, verifies SHA-256, renames atomically and
  re-uses verified files; local checkpoint paths are accepted as-is.
- configs: loqi.yaml and loqi_flow.yaml bundled as package data, copies of the
  training configs with sample.node_distribution and the dataset/AIMNet2 paths
  set to null (the node-count prior is never used for conformer generation).
- featurize: SMILES validation with canonical round-trip, hydrogen addition,
  stereo edge construction, replication and atom-aware batching, ported from
  scripts/sample_conformers.py; RDKit legacy stereo perception is applied via a
  context manager instead of a global setting.
- api: load_model() (lightning load_from_checkpoint with map_location and,
  where supported, weights_only=False; CUDA if available) and
  generate_conformers(), which seeds torch/numpy, samples, drops non-finite
  samples and returns RDKit molecules with explicit hydrogens and up to
  n_conformers conformers plus the integer property loqi_failed.
- cli: 'loqi sample' and 'loqi download' (argparse), installed as the loqi
  console script.

scripts/sample_conformers.py now calls the package for loading, featurisation
and sampling; --config is optional (bundled config by default) and --ckpt also
accepts a registry name. AIMNet2 optimisation and iRMSD pruning stay in the
script and are not part of the API.
Unit tests (no checkpoint needed): megalodon.scatter numerics against
index_add_ references including mean over empty segments and per-segment
softmax sums, registry table sanity and download/verification/caching via
file:// URLs, SMILES validation and hydrogen handling, stereo edge
construction, bundled config loading, and the CLI parser and early exits.

tests/test_e2e.py is marked slow and deselected by default (pyproject addopts).
It loads the cached loqi checkpoint on CPU, generates 3 conformers each for
ethanol and aspirin, and checks atom counts, conformer counts, canonical SMILES
round-trip, finite coordinates and bond lengths, seed reproducibility and the
CLI SDF output.
On push to main and pull requests: Python 3.11 and 3.12, uv venv, CPU torch
from the PyTorch index, editable install with the dev extra, ruff check and
pytest -m "not slow". A second job on a weekly schedule and workflow_dispatch
caches ~/.cache/loqi, downloads the verified checkpoint and runs the slow
end-to-end test.
Add Quick start (generate_conformers/load_model usage and the loqi command),
Checkpoints (KiltHub source, MIT license, SHA-256 verified cache location) and
Environment (verified package versions, no compiled PyG extensions, timings)
sections. Point the usage notes at the package instead of PYTHONPATH and
describe the optional --config / registry-name --ckpt of
scripts/sample_conformers.py. Training documentation is unchanged.
Rename private CLI, SMILES parsing, and scatter helpers for clarity.
Shorten docstrings and package usage instructions, remove redundant
comments, and retain public APIs and license notices.

Validation: all 57 tests pass against the installed wheel, including
CPU sampling and reproducibility. Six fixed-seed conformers match the
previous implementation exactly. Ruff and formatting checks pass.
The source distribution rebuilds a wheel with identical contents.
@bkalita23 bkalita23 changed the title Make LoQI pip-installable: floors instead of pins, pure-torch scatter, loqi API and CLI Package LoQI for pip installation with a Python API and CLI Sep 11, 2026
@FilippNikitin
FilippNikitin marked this pull request as ready for review September 15, 2026 15:22
@FilippNikitin
FilippNikitin merged commit 57cd41b into main Sep 15, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants