ids4c is a C++ library and toolset for parsing and matching IDS (Ideographic Description Sequence) expressions for Chinese characters, with database import and IWDS-based fuzzy unification.
The project provides:
- ids4c: a GTKmm desktop application;
- ids4c-cli: a command-line query and database-import tool;
- ids4c.pyd: a Python extension module built with pybind11;
- C++ static libraries and public headers for integration into other applications.
The project source code is licensed under the Apache License, Version 2.0; see LICENSE. Upstream IDS data, IWDS data, and syntax documentation have their own licenses and usage terms; see Data Sources and References.
- Parse ordinary IDS expressions, Bai-style IDS extensions, and query expressions;
- Preserve structure operators, components, strokes, suffixes, variables, and query parameters;
- Clone IDS trees, access child nodes, and compare tree structure;
- Calculate stroke counts from IDS expressions and use the database stroke cache.
- Exact IDS matching and structural matching;
- Component-search expressions:
<search=...>,<any=...>, and<except=...>; - The component wildcard ⬚, variables such as
<var=...>, and Unicode escapes\uXXXX/\UXXXXXXXX; - Stroke conditions
<stroke=...>and remaining-stroke conditions<residue=...>; - Approximate stroke ranges with
~, for example<stroke=21~1>and<residue=12~2>; - Multiple HVExtract alternatives and same-IDS expression grouping;
- First-level component indexing for
<search=...>candidate filtering, without changing final match semantics; - Same-IDS uniqueness markers (
{glyph}) are respected by default; - U+1F504 (🔄) replacement queries;
- Optional filtering of overlay structures represented by ⿻;
- Unicode-block, private-use, abstract-glyph, and custom-range filtering;
- Detailed match paths, equivalent-query indexes, and preprocessing-rule information.
When IDS data is imported, ids4c keeps both the original IDS expressions and derived query caches. The HVExtract cache is used by structural and component queries; the stroke-neutral composition cache allows components with lowercase stroke suffixes to participate in composition lookup without changing the data returned by raw_ids(). A first-level component index for raw, HV, and stroke-neutral expressions is also stored in SQLite. It narrows search candidates conservatively; final matches still use the normal matcher. Missing or older indexes are rebuilt or migrated when the database is loaded. Reimporting private data or rebuilding the cache regenerates these derived caches from the raw IDS.
Database queries can use the following IWDS unification levels:
none: no IWDS fuzzy unification;srcseparation: Source Code Separation;lv1: Source Code Separation plus component-level unification;lv2: additional IWDS component unification.
Unification is performed during query preprocessing. The resulting equivalent expressions are then passed to the IDS matching layer. Detailed results can retain the equivalent expression and match paths for use by higher-level applications.
The complete description of the specialized query syntax is in docs/query-syntax.md. The document covers <search>, <any>, <except>, stroke and residue conditions, variables, Unicode escapes, U+1F504 replacement queries, and the ▥/▤ structure extensions.
include/ids4c/ Public C++ headers
src/core/ IDS tree, parser, and core matching logic
src/db/ SQLite database, IDS import, and IWDS import
src/cli/ Command-line application
src/app/ GUI application entry point
src/ui/ GTKmm user interface
src/python/ pybind11 Python module
docs/ API and user documentation
tests/ CTest tests and test data
third_party/ Third-party source used directly by the project
Basic requirements:
- CMake 3.20 or newer;
- A compiler with C++17 support (the build uses
-std=c++17or the toolchain equivalent); - SQLite 3;
- pkg-config;
- Boost. The CLI requires program_options; the GUI requires filesystem and system.
Additional target dependencies:
- GUI: GTKmm 3;
- Python binding: CPython development files and the pybind11 CMake package;
- CLI: Boost.Program_options.
On Windows, MSYS2 MinGW is recommended. Python, pybind11, SQLite, GTKmm, Boost, and CMake should target the same architecture and toolchain. The Python extension is a native CPython module and normally must be rebuilt for a different Python version, architecture, or toolchain.
Python binding is enabled by default. If a compatible Python and pybind11 installation is available:
cmake -S . -B build `
-DIDS4C_BUILD_GUI=ON `
-DIDS4C_BUILD_CLI=ON `
-DIDS4C_BUILD_PYTHON=ON `
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config ReleaseIf CMake cannot locate Python or pybind11 automatically, specify them explicitly:
cmake -S . -B build -G "MinGW Makefiles" `
-DIDS4C_BUILD_GUI=ON `
-DIDS4C_BUILD_CLI=ON `
-DIDS4C_BUILD_PYTHON=ON `
-DPython3_EXECUTABLE=E:/msys64/mingw64/bin/python.exe `
-DPython3_ROOT_DIR=E:/msys64/mingw64 `
-Dpybind11_DIR=E:/msys64/mingw64/lib/cmake/pybind11 `
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config ReleaseIf the Python development environment is unavailable or the module is not needed:
cmake -S . -B build `
-DIDS4C_BUILD_GUI=ON `
-DIDS4C_BUILD_CLI=ON `
-DIDS4C_BUILD_PYTHON=OFF `
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config ReleaseMain CMake options:
| Option | Default | Description |
|---|---|---|
| IDS4C_BUILD_GUI | ON | Build the GTKmm GUI |
| IDS4C_BUILD_CLI | ON | Build the CLI |
| IDS4C_BUILD_PYTHON | ON | Build the pybind11 module |
| IDS4C_ENABLE_LTO | OFF | Enable LTO/IPO for Release and RelWithDebInfo |
Runtime database files are intentionally not bundled with the public source repository. Prepare them by importing the upstream IDS and IWDS data as described below.
On Windows, the main build outputs are usually:
ids4c.exe GUI
ids4c-cli.exe CLI
ids4c.pyd Python extension module
The Python module is not a standalone application. Python must be able to find ids4c.pyd, and the runtime DLL directory must be on PATH:
$env:PATH = "E:/msys64/mingw64/bin;$env:PATH"
$env:PYTHONPATH = "build"
python -c "import ids4c; print(ids4c.parse('⿰亻言').text)"Install the built targets, headers, and documentation with:
cmake --install build --prefix installThe installation includes the enabled applications, static libraries, headers, Python module, and documentation.
ids4c uses SQLite databases. Database name NAME maps to:
db/NAME.sqlite
For example, --database yibai0 opens db/yibai0.sqlite. If db/unifiable.sqlite exists, IWDS unification data is loaded from it as well.
Database files are generated by importing IDS source data. The recommended exact stroke-matching base is YiBai's ids_lv0.txt:
New-Item -ItemType Directory -Force db
.\build\ids4c-cli.exe `
--database yibai0 `
--import path/to/ids_lv0.txt `
--format yibaiImport IWDS XML:
.\build\ids4c-cli.exe --import-iwds path/to/iwds.xmlPrivate IDS extensions can be added or reimported:
.\build\ids4c-cli.exe `
--database yibai0 `
--import-private path/to/private.dat `
--format default
.\build\ids4c-cli.exe `
--database yibai0 `
--reimport-private path/to/private.dat `
--format defaultImport errors include the source line, IDS index, character position, and other diagnostics when available. Import formats are default and yibai.
CLI arguments, database import, filters, output formats, and examples are documented in docs/cli.md.
Minimal query:
.\build\ids4c-cli.exe --database yibai0 --query "⿰贝⬚"Python API documentation is available in English and 中文. Minimal example:
import ids4c
query = ids4c.parse("⿰贝⬚")
database = ids4c.Database("yibai0")
for glyph in database.match(query):
print(glyph)The Python extension depends on the CPython ABI, architecture, and toolchain selected at build time. Binary distributions generally need separate builds for their target Python versions and platforms.
After configuring and building:
ctest --test-dir build -C Release --output-on-failureTests cover IDS database import, private-data import and reimport, import diagnostics, IWDS import, and selected query behavior. The complete upstream datasets are not test fixtures; validate with the target database before distribution.
See CONTRIBUTING.md for the build, test, documentation, and issue-reporting workflow. The project history is summarized in CHANGELOG.md.
The primary IDS data source is YiBai's IDS repository. The exact stroke-matching baseline is ids_lv0.txt:
According to the upstream description, ids_lv0.txt distinguishes all stroke differences, ids_lv1.txt merges selected stroke-level differences, and ids_lv2.txt merges selected variants considered mandatory to unify. Distributions containing databases generated from this data must comply with the upstream MIT license and other notices.
IWDS XML data comes from:
iwds.xml is used to build the unification data in unifiable.sqlite. The IWDS repository also contains schema files, images, and document-generation sources, which may have separate copyright and usage terms.
The IDS syntax reference is:
The bai-ids documentation repository is released under the MIT license. Always follow the current upstream terms for syntax and data usage.
third_party contains third-party source used directly by this project. Follow the license and copyright notices in each dependency and its upstream project. The repository's NOTICE file summarizes the bundled notices and upstream data sources. Binary distributions should retain applicable third-party licenses and NOTICE files.
Some code was generated, refactored, or debugged with assistance from OpenAI Codex. The project maintainer reviewed, modified, and tested the final code. Codex is not claimed as a copyright holder or project contributor; attribution follows the repository history, license, and maintainer statements.
Issues, test cases, and improvements are welcome. For IDS, IWDS, or upstream-data issues, include:
- the original query expression;
- database name and import source;
- expected and actual results;
- CLI --explain output or detailed JSON;
- a minimal reproducible IDS data sample.