Metadata integration and quality assessment of research software.
This repository contains the data pipeline that powers the Research Software Observatory—a platform for monitoring and assessing the quality and FAIRness of research software in the life sciences.
It consolidates software records, resolves duplicates, and precomputes the quality and FAIRness statistics displayed in the Observatory’s interface.
-> 📄 Documentation
The ETL runs in modular stages, which can be executed independently or orchestrated end-to-end through the unified CLI command rsetl.
- Transformation – Fetches raw records from source collections and standardizes them.
- License normalization – Maps license strings to SPDX identifiers.
- Blocking and recovery – Groups related software records from normalized data.
- Metrics removal (optional) – Filters low-information OpenEBench metrics.
- Conflict detection – Identifies inconsistent or duplicate records.
- Simplification – Reduces block complexity for later processing.
- Conversion to JSONL – Formats data for large-scale or LLM-based steps.
- Disambiguation – Uses heuristics and AI-assisted agreement scoring to resolve conflicts.
- Human integration – Incorporates curator decisions from Git-based annotations.
- Merge – Produces final, merged software entries and updates the database.
- FAIRsoft scores and statistics – Computes FAIR compliance metrics and aggregated statistics stored in the database to support visualization and longitudinal monitoring.
- Similarity – Embeds tool descriptions and precomputes the top-10 nearest neighbours per tool to power "similar software" recommendations.
Each execution creates a versioned run directory under data/integration/runs/<run_id>/ with a manifest file tracking inputs, outputs, and environment metadata.
# install in editable mode
pip install -e .
# set up environment variables (MongoDB + API tokens)
cp .env.example .env # then fill in the values
# .env is auto-loaded; every variable the pipeline reads is documented there
# run full integration
rsetl run
All intermediate and final files are automatically stored in timestamped directories, and a latest symlink always points to the most recent run.
The pipeline ships as a Docker image published to GitHub Container Registry. CI (.github/workflows/build_image.yml) builds and pushes ghcr.io/inab/research-software-etl on v* tags and published releases (and :latest on manual runs of the default branch), mirroring how the importers (ghcr.io/inab/*-importer) are deployed.
On the VM, docker-compose.vm.yml defines two one-shot services off that single image, each triggered by host cron:
# full integration pipeline (twice weekly)
docker compose -f docker-compose.vm.yml run --rm rsetl-full
# web-availability refresh (daily)
docker compose -f docker-compose.vm.yml run --rm rsetl-webavailability
Both services read credentials and collection overrides from .env (env_file); the image never bakes it in. data/ is mounted so run outputs and the cross-run curator history survive --rm containers. See .env.example for the full list of variables and Dockerfile for the image build.
Clean architecture with four layers (adapters → application → domain → infrastructure):
src/adapters/ # CLI entry points (rsetl) and the scheduler
src/application/ # Use cases (workflows) and services (domain logic + I/O)
src/domain/ # Pydantic models and repository protocols
src/infrastructure/ # MongoDB adapter/repositories, API clients, config
scripts/ # One-off utilities (outside the architecture rules)
data/integration/runs/ # Versioned outputs per run (git-ignored)
See the Development Guide for details.