termwidth

termwidth

Grapheme cluster segmentation and terminal cell-width measurement for Crystal, implemented entirely in Crystal with generated three-stage Unicode tries for width, pictographic, and UAX #29 grapheme-break properties.

Beyond plain segmentation, the shard models how real terminals render text: configurable width rules, emoji/VS15/VS16/keycap/flag handling, and a calibration profile format that captures a terminal's actual behavior (including the widths no rule can explain).

  • Unicode 17.0.0 data tables, pinned and verified against GraphemeBreakTest.txt
  • Zero-allocation iteration over clusters (Bytes slices, no strings)
  • Automatic printable ASCII run coalescing
  • No runtime dependencies or native build step

Installation

Add the dependency to your shard.yml:

dependencies:
  termwidth:
    github: shpeckman/termwidth

or point at a git remote once the shard is published. Then:

shards install

The generated Unicode tables ship with the shard. Installation only compiles Crystal code and does not download data, run a postinstall hook, or require a C toolchain.

Requires Crystal >= 1.21.0.

Usage

require "termwidth"

Segmenting and measuring

seg = TermWidth::Segmenter::NARROW   # standard Unicode widths
seg = TermWidth::Segmenter::WIDE     # ambiguous-width codepoints take 2 cells
seg = TermWidth::Segmenter.new(TermWidth::WidthProfile::CAUTIOUS)

seg.each("ab日é ") do |cluster, width|
  # cluster : Bytes (a slice of the input), width : UInt8 (cells)
end

seg.each_span("ab日é ") do |span, width, ascii_run|
  # like each, but printable ASCII is coalesced into runs (ascii_run == true)
end

seg.measure("ab日é ") # => 6  (total cells)
seg.count("é🇺🇸")   # => 2  (clusters, not codepoints)
seg.codepoint_width(0x65E5_u32)   # => 2

each/each_span accept Bytes or String and never allocate: clusters are slices of the input.

Width profiles

A WidthProfile bundles a set of Rule flags with two Category masks describing measurements a terminal makes that no rule reproduces:

profile = TermWidth::WidthProfile::UNICODE
profile.rules       # TermWidth::Rule flags
profile.volatile    # categories whose width varies by terminal state
profile.lead_in     # categories whose width depends on the preceding character
profile.rule_names  # => "clusters,pair-flags,lone-ri-wide,vs16-widens,modifier-merges,keycap-wide"

Rules: Clusters (UAX #29 clusters vs. codepoints), AmbiguousWide, PairFlags, LoneRiWide, Vs16Widens, Vs15Narrows, ModifierMerges, KeycapWide, SpacingWidens.

Profiles can be inferred from terminal measurements — feed WidthProfile.infer a set of Samples (text plus observed width, optionally at screen edge and after a leading character) and it picks the best rule set and blames unexplainable widths on volatile/lead_in categories:

samples = TermWidth::WidthProfile::CALIBRATION.map do |text|
  TermWidth::WidthProfile::Sample.new(text, measured_width_of(text))
end
profile = TermWidth::WidthProfile.infer(samples)

TermWidth::WidthProfile.categories(text) reports which Category flags (Ambiguous, Flag, Zwj, Vs16, Combining, Conjunct, Jamo, …) a piece of text falls into.

Low-level pieces

TermWidth::Props.lookup(0x1F600_u32)  # packed property byte from the trie
TermWidth::Props.width(prop)          # 0, 1, or 2
TermWidth::Props.ambiguous?(prop)
TermWidth::Props.pictographic?(prop)

cp, used = TermWidth::Utf8.decode(bytes, 0)  # strict, replaces malformed input
TermWidth::Utf8.encoded_length(cp)

TermWidth.classify(bytes)      # Category bitmask for a cluster
TermWidth.classify_cp(cp)      # same, for a single codepoint
TermWidth.segment(bytes, from, rules, dst)  # raw batched segmentation

TermWidth::Tables::UNICODE_VERSION reports the pinned Unicode version.

Development

crystal spec                    # run the spec suite
crystal tool/verify_unicode.cr  # check segmentation against GraphemeBreakTest.txt
crystal tool/fetch_ucd.cr       # fetch the pinned UCD files into tool/ucd
crystal tool/gen_unicode.cr     # regenerate src/termwidth/tables.cr

The batched kernel is checked against an independent span-building reference (spec/support/reference.cr) on the Unicode conformance data, randomized text, and malformed byte sequences.

License

MIT

Repository

termwidth

Owner
Statistic
  • 0
  • 0
  • 0
  • 0
  • 0
  • 28 minutes ago
  • September 28, 2026
License

MIT License

Links
Synced at

Tue, 29 Sep 2026 19:33:58 GMT

Languages