Package LoQI for pip installation with a Python API and CLI - #6
Merged
Merged
Conversation
Replace the setuptools configuration (dynamic version from a missing VERSION file, exact pins including CUDA-specific torch_scatter/torch_sparse wheels) with a hatchling build of distribution "loqi" 0.1.0 that ships the megalodon and loqi import packages. Runtime dependencies are inference-only floors: torch>=2.8, torch_geometric>=2.6, rdkit>=2025.3, lightning>=2.4, omegaconf>=2.3, einops>=0.8, numpy>=1.26, scipy>=1.11, tqdm>=4.66. Training/preprocessing packages move to the [train] extra, AIMNet2 to [aimnet], pytest/ruff to [dev]. License metadata covers LICENSE.md (MIT) and megalodon_licence/ (Apache-2.0 NVIDIA code and third-party notices). requirements.txt now mirrors the floors without the PyG wheel index, and the README installation section describes uv/pip installs from GitHub (PyPI publication is a later step). Adds pytest (slow marker deselected by default) and ruff (line length 120, new code only) configuration.
….scatter Add megalodon.scatter, a drop-in module for the torch_scatter functions the models use (scatter, scatter_sum/scatter_add, scatter_mean, scatter_softmax) implemented with torch_geometric.utils.scatter and torch_geometric.utils.softmax. Those fall back to pure torch (index_add_/scatter_reduce) when the compiled extension is absent, so neither torch_scatter nor torch_sparse has to be built for the installed torch/CUDA combination. Signatures follow torch_scatter (dim defaults to -1, reduce="add" is accepted); the unused out= argument raises. The 16 modules that imported from torch_scatter now import from megalodon.scatter. torch_sparse is only referenced lazily inside the EQGAT edge message-passing path (eqgat_denoising_model.py), which the LoQI configs do not use, so it stays optional.
src/loqi is a thin public layer over the megalodon code: - registry: MODELS table for the released loqi (diffusion) and loqi_flow (flow matching) checkpoints from the KiltHub record doi:10.1184/R1/31441570, and checkpoint_path(), which downloads with the standard library into $LOQI_CACHE_DIR or ~/.cache/loqi, verifies SHA-256, renames atomically and re-uses verified files; local checkpoint paths are accepted as-is. - configs: loqi.yaml and loqi_flow.yaml bundled as package data, copies of the training configs with sample.node_distribution and the dataset/AIMNet2 paths set to null (the node-count prior is never used for conformer generation). - featurize: SMILES validation with canonical round-trip, hydrogen addition, stereo edge construction, replication and atom-aware batching, ported from scripts/sample_conformers.py; RDKit legacy stereo perception is applied via a context manager instead of a global setting. - api: load_model() (lightning load_from_checkpoint with map_location and, where supported, weights_only=False; CUDA if available) and generate_conformers(), which seeds torch/numpy, samples, drops non-finite samples and returns RDKit molecules with explicit hydrogens and up to n_conformers conformers plus the integer property loqi_failed. - cli: 'loqi sample' and 'loqi download' (argparse), installed as the loqi console script. scripts/sample_conformers.py now calls the package for loading, featurisation and sampling; --config is optional (bundled config by default) and --ckpt also accepts a registry name. AIMNet2 optimisation and iRMSD pruning stay in the script and are not part of the API.
Unit tests (no checkpoint needed): megalodon.scatter numerics against index_add_ references including mean over empty segments and per-segment softmax sums, registry table sanity and download/verification/caching via file:// URLs, SMILES validation and hydrogen handling, stereo edge construction, bundled config loading, and the CLI parser and early exits. tests/test_e2e.py is marked slow and deselected by default (pyproject addopts). It loads the cached loqi checkpoint on CPU, generates 3 conformers each for ethanol and aspirin, and checks atom counts, conformer counts, canonical SMILES round-trip, finite coordinates and bond lengths, seed reproducibility and the CLI SDF output.
On push to main and pull requests: Python 3.11 and 3.12, uv venv, CPU torch from the PyTorch index, editable install with the dev extra, ruff check and pytest -m "not slow". A second job on a weekly schedule and workflow_dispatch caches ~/.cache/loqi, downloads the verified checkpoint and runs the slow end-to-end test.
Add Quick start (generate_conformers/load_model usage and the loqi command), Checkpoints (KiltHub source, MIT license, SHA-256 verified cache location) and Environment (verified package versions, no compiled PyG extensions, timings) sections. Point the usage notes at the package instead of PYTHONPATH and describe the optional --config / registry-name --ckpt of scripts/sample_conformers.py. Training documentation is unchanged.
Rename private CLI, SMILES parsing, and scatter helpers for clarity. Shorten docstrings and package usage instructions, remove redundant comments, and retain public APIs and license notices. Validation: all 57 tests pass against the installed wheel, including CPU sampling and reproducibility. Six fixed-seed conformers match the previous implementation exactly. Ruff and formatting checks pass. The source distribution rebuilds a wheel with identical contents.
loqi API and CLI
FilippNikitin
marked this pull request as ready for review
September 15, 2026 15:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes LoQI installable with pip, with
load_model()andgenerate_conformers()as the Python entry points andloqi sample/loqi downloadas CLI commands.loqiandmegalodonpackages with Hatchling, including inference configs, existing model resources, and license notices.train,aimnet, anddevextras.Validation: all 57 tests pass against the installed wheel on CPU with Python 3.12.12 and PyTorch 2.14.0+cpu, including released-checkpoint sampling and reproducibility. Six fixed-seed conformers have exactly identical coordinates before and after the refactor. Computational expressions and parsed configuration values are unchanged. Ruff and formatting checks pass. Both wheel and source distributions build; rebuilding the wheel from the source archive produces identical contents, with configs, model resources, and all three license files present.