ucharsetversion Documentation on ocaml.org

Character classes for Unicode-aware lexers and regex engines

A faster, smaller Set.Make(Uchar). Cost follows the run count, not the cardinal: the Alphabetic property is 147,421 codepoints in 761 runs, held in 12KB against 5.6MB, and the whole codespace is 40 bytes against 42.4MB. Alongside the usual set algebra there is a builder for accumulating a class from many fragments, a compiled two-level bitmap trie for membership in an inner loop, partition refinement for derivative classes during DFA construction, and a stable packed encoding for embedding generated tables in source. Surrogates are excluded by construction.

Tags unicode charset codepoint interval set regex lexer dfa
AuthorMichael Thomas <mthomas180@gmail.com>
LicenseMIT
Published
Homepagehttps://github.com/enetsee/ucharset
Issue Trackerhttps://github.com/enetsee/ucharset/issues
Documentationhttps://enetsee.github.io/ucharset/
MaintainerMichael Thomas <mthomas180@gmail.com>
Availablearch != "x86_32" & arch != "arm32"
Dependencies
Source [http] https://github.com/enetsee/ucharset/releases/download/v0.2.1/ucharset-0.2.1.tbz
sha256=d63ee03f167eb6f6e3b3198f4691ec7668836de0316d6e46ba078a42b75ca342
sha512=1d0bd0ceda50c94f9df140f29f6aa7c433e7baa41ce43573d5e31c48479c4af29c9683c5c4a9a9a2df3f840887cd11141304b1b5a6343c6e24d62878588ee506
Edithttps://github.com/ocaml/opam-repository/tree/master/packages/ucharset/ucharset.0.2.1/opam
No package is dependent