Conversation
|
oh this is great to have, thanks! |
bernt-matthias
left a comment
There was a problem hiding this comment.
Excellent. Just a few comments.
| key_points: | ||
| - Data managers are tools to be run by admins of a Galaxy instance. | ||
| - They automate reference data collection and preparation and they write data table (.loc file) records. | ||
| - In addition to a regular tool wrapper xml file, data managers require several xml config files that define the interaction of the data manager with the Galaxy framework and with the data tables they are supposed to populate. |
There was a problem hiding this comment.
Strictly only one xml file more compared to tools that use the data table.
| # What are data managers and why are they needed? | ||
|
|
||
| Many tools run with two kinds of input data: some experimental data specific to | ||
| the tool run (like, e.g., sequencing data) and another dataset, which stays the |
There was a problem hiding this comment.
| the tool run (like, e.g., sequencing data) and another dataset, which stays the | |
| the tool run (like, e.g., sequencing data) and other data, which stays the |
|
|
||
| One possible solution (which was actually used in the early days of Galaxy) is to have | ||
| Galaxy server admins collect and prepare commonly used data, store it on the server and | ||
| record the location in a so-called .loc file. |
There was a problem hiding this comment.
| record the location in a so-called .loc file. | |
| record the location in a so-called .loc file. These files are simple tab separated files storing data like paths and metadata. |
| > An in-depth, technical explanation of this matter is provided at | ||
| > <https://docs.galaxyproject.org/en/latest/dev/data_managers.html> | ||
| > and, when in doubt, that material should be considered the reference document | ||
| > for data managers. |
There was a problem hiding this comment.
Maybe we can also xref the admin docs: https://docs.galaxyproject.org/en/latest/admin/data_tables.html
There was a problem hiding this comment.
That page is already referenced multiple times from the provided page.
| With more and more tools requiring data of very different formats, this approach becomes increasingly | ||
| unmanageable because, for each tool, an admin has to research where to obtain the data, | ||
| or how to calculate it, and whether it needs some reformatting or other pre-processing before being usable | ||
| by tools. |
There was a problem hiding this comment.
Another advantage is more consistency between different Galaxy servers.
There was a problem hiding this comment.
The reproducibilty aspect is now mentioned at the end of the section.
|
|
||
| # How does a Data Manager communicate with Galaxy? | ||
|
|
||
| 1. It declares itself a Data Manager: |
There was a problem hiding this comment.
| 1. It declares itself a Data Manager: | |
| 1. It declares itself a Data Manager: | |
| which is essentially a normal tool, but uses `tool_type="manage_data"` |
| </data_managers> | ||
| ``` | ||
|
|
||
| This file declares column names for a single Data table (`cat_database`) that |
There was a problem hiding this comment.
Would it be better to stick to one example? ... But the example is also a good one...
There was a problem hiding this comment.
Left as is for now - though I agree that there is maybe room for improvement here.
Sth to address in an update perhaps.
| Galaxy to record the `extra_files_path` folder name both in the value column and in the | ||
| name column of the `cat_database` Data table. | ||
|
|
||
| It also wants to store the paths to the extracted `database_folder` and `taxonomy_folder` so that tools that later want to use that data can discover |
There was a problem hiding this comment.
so that -> such that? (at multiple places?)
|
|
||
| The definitions of the `database_folder` column hold two types of instructions for Galaxy: | ||
|
|
||
| 1. The `<move>` element says that Galaxy should take (see the `<src>` element) the data that lives where the `${database_folder}` item of the `data_manager_json` output says it lives and move it to a destination `CAT/${database_folder}` under the base path indicated by `${GALAXY_DATA_MANAGER_DATA_PATH}` (which itself is the configured cached data storage path of the Galaxy instance). |
There was a problem hiding this comment.
| 1. The `<move>` element says that Galaxy should take (see the `<src>` element) the data that lives where the `${database_folder}` item of the `data_manager_json` output says it lives and move it to a destination `CAT/${database_folder}` under the base path indicated by `${GALAXY_DATA_MANAGER_DATA_PATH}` (which itself is the configured cached data storage path of the Galaxy instance). | |
| 1. The `<move>` element says that Galaxy should take (see the `<src>` element) the data that lives where the `${database_folder}` item of the `data_manager_json` output says it lives and move it to a destination `CAT/${database_folder}` under the base path indicated by `${GALAXY_DATA_MANAGER_DATA_PATH}` (which itself is the admin configured data storage path of the Galaxy instance). |
| > If you're following the standard layout of data managers with the tool xml | ||
| > file in a subfolder, you need to run planemo from the parent folder to have | ||
| > it discover all required files beyond the tool xml, but point it to the tool | ||
| > xml to test or serve like, e.g.: `planemo serve data_manager/my_dm.xml` |
There was a problem hiding this comment.
I always run from the IUC repo root folder. Which is then maybe a bit easier for the serve of DM + tool.
There was a problem hiding this comment.
provided now as an example, too
paulzierep
left a comment
There was a problem hiding this comment.
need to continue later
| --- | ||
| layout: tutorial_hands_on | ||
|
|
||
| title: Understanding Galaxy Data managers |
There was a problem hiding this comment.
| title: Understanding Galaxy Data managers | |
| title: Understanding Galaxy Data Managers |
| because: | ||
|
|
||
| - some of that data is complicated to gather from public sources or needs some pre-processing | ||
| - it leads to unnecessary copies of data that would better be reused across user accounts |
There was a problem hiding this comment.
| - it leads to unnecessary copies of data that would better be reused across user accounts | |
| - it leads to unnecessary copies of data (that can easily be more then 100 GB of size) that would better be reused across user accounts |
|
|
||
|  | ||
|
|
||
| In principle, admins could automate some of this work through scripts, but it would be nice to not have each admin reinvent the wheel. |
There was a problem hiding this comment.
| In principle, admins could automate some of this work through scripts, but it would be nice to not have each admin reinvent the wheel. | |
| In principle, admins could automate some of this work through scripts, but it would be nice to not have each admin reinvent the wheel. Furthermore, since it is in the intererst of researchers to have the reference data for their tool available, they should be able so support admins in the data collection task trough an open process. |
There was a problem hiding this comment.
Adds too much advanced perspective, imo. Lets keep things simple here
|
|
||
| 1. The `<move>` element says that Galaxy should take (see the `<source>` element) the data that lives where the `${database_folder}` item of the `data_manager_json` output says it lives and move it to a destination `CAT/${database_folder}` under the base path indicated by `${GALAXY_DATA_MANAGER_DATA_PATH}` (which itself is the configured cached data storage path of the Galaxy instance). | ||
|
|
||
| 2. The first `<value_translation>` element says that Galaxy should not write the Data manager tool-provided value for `database_folder` directly, but instead first translate it to `${GALAXY_DATA_MANAGER_DATA_PATH}/CAT/${database_folder`. If you compare the resulting string with the `<move>` instructions, you will see that it will now be the same as the ultimate path to the folder after Galaxy has moved it. |
There was a problem hiding this comment.
Still struggling to understand value_translation. Why is here https://github.com/natefoo/tools-iuc/blob/bd4a77513afed69f727038cdd11647daddd41c2f/data_managers/data_manager_bwa_mem_index_builder/data_manager_conf.xml#L14 the copy target different from the value translation?
Also there is a missing closing parenthesis in ${GALAXY_DATA_MANAGER_DATA_PATH}/CAT/${database_folder
There was a problem hiding this comment.
So, in the case of bwa mem, the DM creates the index dir at ${out_file.extra_files_path}/${fasta_file_name}, see:
and records ${fasta_file_name as the path in:
the move operation you're linking to then moves everything under ${out_file.extra_files_path} including the index dir to the destination <target base="${GALAXY_DATA_MANAGER_DATA_PATH}">genomes/${dbkey}/bwa_mem_index/v1/${value}</target>, but the data table needs to record the entire path to the index dir (named after ${fasta_file_name) so it appends $path to the value_translation.
There are almost always several ways to achieve the same kind of result so probably one could also rewrite this one in a way that target and value_translation are identical.
There was a problem hiding this comment.
Thanks for pointing out the missing parenthesis.
wm75
left a comment
There was a problem hiding this comment.
Perhaps ready for merging before the tool-dev workshop?
| > An in-depth, technical explanation of this matter is provided at | ||
| > <https://docs.galaxyproject.org/en/latest/dev/data_managers.html> | ||
| > and, when in doubt, that material should be considered the reference document | ||
| > for data managers. |
There was a problem hiding this comment.
That page is already referenced multiple times from the provided page.
|
|
||
|  | ||
|
|
||
| In principle, admins could automate some of this work through scripts, but it would be nice to not have each admin reinvent the wheel. |
There was a problem hiding this comment.
Adds too much advanced perspective, imo. Lets keep things simple here
|
|
||
| 5. A `tool-data` folder | ||
|
|
||
| Here, tools can provide samples of the .loc files they are going to work with. |
| With more and more tools requiring data of very different formats, this approach becomes increasingly | ||
| unmanageable because, for each tool, an admin has to research where to obtain the data, | ||
| or how to calculate it, and whether it needs some reformatting or other pre-processing before being usable | ||
| by tools. |
There was a problem hiding this comment.
The reproducibilty aspect is now mentioned at the end of the section.
| </data_managers> | ||
| ``` | ||
|
|
||
| This file declares column names for a single Data table (`cat_database`) that |
There was a problem hiding this comment.
Left as is for now - though I agree that there is maybe room for improvement here.
Sth to address in an update perhaps.
| > If you're following the standard layout of data managers with the tool xml | ||
| > file in a subfolder, you need to run planemo from the parent folder to have | ||
| > it discover all required files beyond the tool xml, but point it to the tool | ||
| > xml to test or serve like, e.g.: `planemo serve data_manager/my_dm.xml` |
There was a problem hiding this comment.
provided now as an example, too
| > e.g. by uploading the tool to a Galaxy Toolshed, you must *never* change its layout again in a later version of that tool, or in a new tool reusing the Data Table! | ||
| > No renaming of columns, no reshuffling, and no addition of new columns! You need to keep that layout frozen! | ||
| > | ||
| > It is, therefor, important to consider a new Data Table's columns very carefully. |
There was a problem hiding this comment.
maybe add thw new sectio in the iuc guidelines maybe xref the new iuc standard somewhere galaxy-iuc/standards#93
There was a problem hiding this comment.
both are very good suggestions, thanks!
| 'name': '<extra_files_path>', | ||
| 'taxonomy_folder': '<extra_files_path>/a_taxonomy', | ||
| 'value': '<extra_files_path>' |
There was a problem hiding this comment.
Should we use extra_files_path name and value in the example?
There was a problem hiding this comment.
We should explain the name and value columns
There was a problem hiding this comment.
Kind of addressed now by referencing the IUC column name standards above.
Maybe you're right about this not being the perfect example for using name and value.
I'd leave this for future improvements.
@SaimMomin12 @bernt-matthias finally!