Open, machine-readable datasets related to Turkmenistan.
tm-data collects, organizes, and shares Turkmen data for use in software, research, data analysis, education, and machine learning.
| Dataset | Contents | Format | Documentation |
|---|---|---|---|
| Geography | Administrative divisions, settlements, alternative names, and search helpers | GeoJSON, PostgreSQL, MySQL, SQLite, SQL Server | Guide |
| Poetry | 855 works by 13 Turkmen poets | MySQL 8+ SQL, JSON — see Guide | Guide |
| Stories | Page-level text extracted from Turkmen stories and prose | JSON, SQLite, MySQL SQL | Guide |
| Colors | Turkmen color names, English translations, HEX values, and categories | MySQL 8+ SQL, JSON, CSV | Guide |
| Dictionary | 18,674 searchable Turkmen headwords with pronunciations, definitions, and examples | JSON | Guide |
| Given names | 447 Turkmen names with meaning, origin, gender, variants, and provenance status | JSON | Guide |
Each dataset documents its structure, import process, limitations, and available sources in its own directory.
Flat utility datasets also provide deterministic UTF-8 CSV exports for users who do not need JSON or SQL. See the CSV export guide for coverage, representation rules, and regeneration commands.
See DATA_FORMAT.md for shared encoding, field naming, date, and export conventions.
catalog.json lists all dataset directories and data files,
available formats, content-based versions, per-file record counts (or null
when unknown), and documentation/provenance references. See the
catalog guide for the schema and counting rules.
After changing datasets, regenerate and verify the inventory:
go run ./tools/catalog
go run ./tools/catalog --checkRepository-wide duplicate and format validation is available in
tools/validate/main.go. It checks JSON, JSONL, and
CSV datasets, including GeoJSON, using only the Go standard library:
go run ./tools/validate/main.goSpecific files or directories can be passed as arguments. The validator exits
with status 1 when it finds duplicate records, duplicate JSON keys, invalid
UTF-8, or malformed JSON, JSONL, or CSV. It is original repository tooling,
uses only the Go standard library, and does not contain or download external
data; dataset provenance remains documented with each dataset.
Strict UTF-8 and modern Turkmen Latin alphabet validation is available in
tools/validation:
go run ./tools/validation -text "Türkmenistanyň paýtagty Aşgabat."Machine-readable, conservative Turkmen text normalization rules and fixtures
are documented in utils/normalization. A
reference implementation provides UTF-8 validation, BOM handling, LF line
endings, and Unicode NFC without changing spelling or script:
go run ./tools/normalization -text $'A\u0308new\r\n'
go run ./tools/normalization --checkSchema, required-field, uniqueness, range, and cross-file reference checks are
defined in tools/integrity/rules.json and run with:
go run ./tools/integrityClone the repository:
git clone https://github.com/turkmenos/tm-data.git
cd tm-dataFollow the import instructions for the dataset you want to use. SQL files use UTF-8; choose a connection encoding and collation that preserve Turkmen characters such as ä, ç, ň, ö, ş, ü, ý, and ž.
Contributions may include new datasets, corrections, sources, translations, or additional export formats.
- Place data in an appropriately named directory and use a clear, machine-readable structure.
- Document the source URL, retrieval date, and redistribution rights.
- Describe the schema, format, import steps, and known limitations in the dataset README.
- Preserve the Turkmen alphabet in UTF-8 and check for duplicates where possible.
- Do not include private, sensitive, or non-redistributable data.
See ROADMAP.md for the planned datasets, data quality improvements, export formats, and usage documentation.
- Proverbs, sayings, and riddles with topic labels
- Turkmen given names with gender, meaning, and origin
- District codes, postal codes, and telephone codes
- Holidays, historical dates, and cultural heritage sites
- Thematic vocabulary for animals, plants, food, occupations, and family relationships
- Turkmen suffixes and additional language-processing data
- GeoJSON and CSV exports for geography, plus JSON and CSV exports for other datasets
- Automated schema, encoding, duplicate, and integrity checks
For any new dataset, reliable provenance, redistribution rights, and verification status are more important than record count alone.
General provenance guidelines are available in SOURCES.md. Detailed sources are documented within each dataset when available.
The repository's original code and independently created material are provided under the MIT License. External data, source material, and literary works may have separate licenses or copyright restrictions. Review each dataset's documentation before use or redistribution.
