An OpenType text shaper for Zig, ported from HarfBuzz and checked against its test suite, for fonts decoded by zig-font.
  • Zig 93.7%
  • Python 4.3%
  • Nix 2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Jeffrey C. Ollie 7de03de0dd
All checks were successful
test / test (push) Successful in 24m24s
test / docs (push) Successful in 11m40s
Update zig-font, zig-font-config and zig-font-renderer
All three packages have dropped the zig_ prefix from their names, and
the dependency keys follow: font, font_config and font_renderer.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01UkN5DesmFtC4GPiqTrmkn8
2026-10-10 17:48:16 -05:00
.forgejo/workflows Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
LICENSES Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
src Add HarfBuzz's Hebrew, Thai and Hangul shapers 2026-10-07 23:26:10 -05:00
tests Add HarfBuzz's Hebrew, Thai and Hangul shapers 2026-10-07 23:26:10 -05:00
tools Add HarfBuzz's Universal Shaping Engine 2026-10-07 23:11:13 -05:00
.gitignore Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
build.zig Update zig-font, zig-font-config and zig-font-renderer 2026-10-10 17:48:16 -05:00
build.zig.zon Update zig-font, zig-font-config and zig-font-renderer 2026-10-10 17:48:16 -05:00
build.zig.zon.nix Update zig-font, zig-font-config and zig-font-renderer 2026-10-10 17:48:16 -05:00
flake.lock Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
flake.nix Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
package.nix Add an OpenType shaper ported from HarfBuzz 2026-10-07 21:31:43 -05:00
README.md Rename the package to font_shaper 2026-10-10 17:42:10 -05:00
REUSE.toml Annotate the fuzz regression inputs for REUSE 2026-10-07 22:18:35 -05:00

zig-font-shaper

A Zig 0.17 library that shapes text: given a run of text and a font decoded by zig-font, it picks the glyphs the font draws the text with and works out where each one goes, applying the font's GSUB and GPOS features — ligatures, contextual forms, kerning, mark attachment, cursive joining — the way HarfBuzz does.

It is a port of HarfBuzz's OpenType shaper, kept close enough to the original that it gives the same glyphs, clusters and positions, to the unit. That is checked against HarfBuzz's own shaping tests, which run as part of the test suite: of the 5,639 cases imported from HarfBuzz 13.2.1, 5,633 pass, 3 fail (see HarfBuzz's tests), and 3 are skipped because they ask for things outside the OpenType shaper: HarfBuzz's fallback shaper, and synthetic slant and emboldening.

The API reference is generated from the doc comments, and is published from the default branch to https://jeff.jcollie.page/zig-font-shaper/. To read it locally, run zig build docs-serve, which serves it at http://127.0.0.1:8000/; use -Ddocs-port=N for a different port. The pages have to be served rather than opened from disk, because the viewer fetches its data at runtime and a browser refuses to do that from a file:// page.

Quick start

const std = @import("std");
const font = @import("font");
const shaper = @import("font_shaper");

pub fn example(gpa: std.mem.Allocator, bytes: []const u8) !void {
    var f = try font.Font.parse(gpa, bytes, .{});
    defer f.deinit();

    // The face positions in font units unless told otherwise.
    var face: shaper.Face = .init(gpa, &f, &.{});

    var buffer: shaper.Buffer = .init(gpa);
    defer buffer.deinit();
    try buffer.addUtf8("office", 0, null);
    // Script and direction from the text; set them to override.
    buffer.guessSegmentProperties();

    try shaper.shapeBuffer(gpa, &face, &buffer, &.{.init("liga", 1)});

    for (buffer.info.items, buffer.pos.items) |info, pos| {
        // info.codepoint is now the glyph ID, info.cluster the byte offset
        // of the text it came from.
        std.debug.print("glyph {d} from {d}: advance {d}, offset {d},{d}\n", .{
            info.codepoint, info.cluster, pos.x_advance, pos.x_offset, pos.y_offset,
        });
    }
}

To shape many runs alike, make a shaper.Plan once with Plan.init(gpa, &face, .{ .direction = .ltr, .script = .latin, ... }) and call shaper.shape.shape(&plan, &buffer) for each; shapeBuffer makes one plan per call.

A variable font is shaped at a point in its design space by giving the face normalized coordinates, which shaper.variations.normalize works out from axis values the way HarfBuzz does:

var coords: [8]font.F2Dot14 = @splat(.{ .raw = 0 });
const axes = if (f.fvar) |fvar| fvar.axes.len else 0;
shaper.variations.normalize(&f, &.{.{ .tag = std.mem.readInt(u32, "wght", .big), .value = 700 }}, null, coords[0..axes]);
var face: shaper.Face = .init(gpa, &f, coords[0..axes]);

To use it from another project, run zig fetch --save git+https://git.jcollie.dev/jeff/zig-font-shaper.git and import the module in build.zig:

const shaper = b.dependency("font_shaper", .{ .target = target, .optimize = optimize });
exe.root_module.addImport("font_shaper", shaper.module("font_shaper"));

With zig-font-renderer

The font_shaper_text module joins the shaper to zig-font-renderer and zig-font-config, replacing the renderer's unshaped fallback_layout:

const text = @import("font_shaper_text");

var list = try loader.resolveName("sans-serif");
defer list.deinit();
var layout = try text.layout(gpa, &list, 24, "Hello, world", .{});
defer layout.deinit(gpa);
for (layout.runs) |run| try renderer.drawRun(gpa, &surface, run, x, y, .{});

The text is split into items wherever the face that has its characters, or its script, changes, and each item is shaped on its own with the text around it as context. layout.clusters says, for each glyph, which byte of the text it came from. Items are laid out in the order they come in the text: text that mixes right-to-left and left-to-right scripts on one line needs the Unicode bidirectional algorithm to order its items first, and that is not done here.

On the command line

zig build run -- FONT-FILE TEXT shapes a line of text and prints the glyphs in the format hb-shape --no-glyph-names prints them, with a subset of hb-shape's options (--features, --variations, --direction, --script, --language, --font-size, --unicodes and the output ones), so the two can be compared directly. The tool is zig-shape in the package.

What it shapes

The pipeline is HarfBuzz's, stage by stage: Unicode properties and grapheme clusters, dotted circles before stray marks, mirroring for right-to-left text, normalization against what the font can draw, the feature map with its stages and pauses, GSUB, advances from hmtx and HVAR (or gvar's phantom points), GPOS or else the legacy kern table, mark zeroing and fallback mark positioning, default ignorables, and the glyph flags that say where text may be broken without reshaping.

Every lookup type of GSUB and GPOS is applied, contextual and chained ones included, with lookup flags, mark filtering sets, feature variations, and Device and VariationIndex adjustments in variable fonts. Features can be set over ranges of clusters, with HarfBuzz's feature string syntax (kern=0, +liga[3:5], aalt=2).

Script and language select the font's script and language system with HarfBuzz's tables of OpenType script and language tags, BCP 47 tags included.

Shapers: all of HarfBuzz's OpenType shapers. The default shaper covers Latin, Greek, Cyrillic, CJK and the other scripts without a shaper of their own. The Arabic shaper, for Arabic and Syriac, does their joining forms, mark reordering, stch stretching, and shaping from the Unicode presentation forms when a font has no GSUB. The Indic shaper, for Devanagari, Bengali, Gurmukhi, Gujarati, Oriya, Tamil, Telugu, Kannada and Malayalam, finds syllables and base consonants and reorders reph and pre-base forms, for old- and new-specification fonts. The Khmer and Myanmar shapers reorder their syllables likewise. The Universal Shaping Engine takes Tibetan, Sinhala, Javanese, Balinese, Mongolian and the many other complex scripts, and Indic fonts with dev3-style tags. The Thai shaper splits SARA AM, for Lao too, and positions marks with the private-use glyphs of fonts without Thai GSUB; the Hebrew shaper composes the presentation forms older fonts need; the Hangul shaper composes or decomposes syllables to suit the font. The Zawgyi shaper leaves Zawgyi-encoded Myanmar to the font.

Not done: Apple's morx, kerx and trak tables; the kern table's state-machine subtables; hinting-dependent positioning (anchor points and Device tables at a pixel size), since the shaper works in font units scaled, not at a ppem; and glyph extents for bitmap and COLR glyphs, which only fallback mark positioning uses.

Unicode data comes from uucode, at Unicode 18; HarfBuzz 13.2.1 is at Unicode 17, which can only matter for characters new in 18.

How it works

The source follows HarfBuzz's layout so that a reader of one can find their way in the other: Buffer.zig is hb_buffer_t, Map.zig is hb_ot_map_t, apply.zig is the matching machinery of hb-ot-layout-gsubgpos.hh, gsub.zig and gpos.zig the lookup subtables, shape.zig is hb-ot-shape.cc, and so on. Functions keep HarfBuzz's names in their doc comments.

Agreeing to the unit takes reproducing more than the algorithms. Values are scaled the way HarfBuzz scales them, in 16.16 fixed point and rounding each value rather than the sum; roundf is HarfBuzz's, which rounds halves up rather than away from zero; axis values are normalized through floating point as HarfBuzz does; the cmap subtable is chosen as HarfBuzz chooses it, symbol and Arabic PUA remappings included; and where a font's data is ambiguous — duplicate glyphs in a coverage table the Arabic fallback builds — HarfBuzz's binary search is reproduced to land on the same entry.

Memory. A Buffer owns its glyphs, and a Plan owns what it worked out; both take an allocator and free with deinit. A Face borrows its font and coordinates, and uses its allocator only for scratch space. Shaping does no I/O.

Hostile fonts are what the fuzzer is for: it edits fonts and shapes with them, looking for panics, leaks and hangs. The buffer limits its length and the number of operations it allows as HarfBuzz does, and nested lookups are limited in depth; a font that hits a limit gets whatever shaping was done up to it.

Building and testing

The toolchain comes from Nix, as does hb-shape from the same HarfBuzz release the tests were imported from.

$ nix develop
$ zig build test --summary all        # unit tests, HarfBuzz's tests, the renderer glue
$ zig build hbtest -- -v arabic       # HarfBuzz's tests in detail, filtered by name
$ zig build test --fuzz=1M            # a bounded fuzzing run
$ zig build coverage                  # kcov report in zig-out/coverage
$ zig build docs                      # API reference in zig-out/docs
$ zig build run -- font.ttf 'text'    # shape a line, as hb-shape prints it
$ nix build .#zig-font-shaper         # the package, tests included, in the sandbox

After changing a dependency in build.zig.zon, regenerate the Nix expression for it:

$ nix develop -c zon2nix --17 --nix=build.zig.zon.nix build.zig.zon

HarfBuzz's tests

tests/hb holds HarfBuzz's shaping tests — its own, the Adobe Annotated OpenType Specification tests, and the Unicode text rendering tests — with the fonts they use. tools/import_hb_tests.py copies them out of the HarfBuzz source that nixpkgs builds hb-shape from, and runs every case through that hb-shape twice: once to confirm it agrees with the expected output recorded in HarfBuzz, and once with --no-glyph-names, whose output is what is kept, so the tests compare glyph IDs and need no glyph names. The handful of cases that release does not reproduce with its own font functions are left out, with the reason, in tests/hb/dropped.txt. The AAT and platform-shaper suites are not imported.

$ nix develop -c python3 tools/import_hb_tests.py

tests/hb/known-failures.txt lists the cases expected to fail: a .dfont collection, which zig-font does not read, and two that print the extents of bitmap and COLR glyphs. The test fails if any other case fails, or if a listed one passes; zig build hbtest -- --update rewrites the list.

Generated tables

Tables are translated from HarfBuzz's generated sources by scripts in tools: the BCP 47 to OpenType language tags (gen_ot_tag_table.py), the Arabic presentation forms (gen_arabic_table.py), the legacy Arabic PUA encodings (gen_arabic_pua.py), the character categories of the Indic, Khmer and Myanmar shapers and of the Universal Shaping Engine (gen_indic_table.py, which compiles HarfBuzz's own tables into a small C++ program to read them out, and so wants a compiler), and the vowel sequences that get a dotted circle (gen_vowel_constraints.py). Each takes the HarfBuzz source as its argument; run zig fmt on what it writes.

The syllable scanners HarfBuzz compiles with Ragel are translated by gen_ragel_machine.py, which takes one hb-ot-shaper-*-machine.hh and a name. Ragel's tables carry over unchanged; its driver loop is written once, in src/shapers/ragel.zig, and each machine's actions become data for it. The table of canonical compositions is generated at build time from uucode by tools/gen_compose.zig.

Where this lives

git clone https://git.jcollie.dev/jeff/zig-font-shaper.git

It is on Radicle as rad:z44vbkcWYHEAZUUDiGA7Gc8vniqfb. A Radicle repository can only be found by its identifier, so that is all a peer needs to fetch it:

rad clone rad:z44vbkcWYHEAZUUDiGA7Gc8vniqfb

License

The code is MIT-licensed, apart from what is ported or translated from HarfBuzz, which also carries HarfBuzz's own "Old MIT" license (MIT-Modern-Variant) and copyright lines, file by file. HarfBuzz's test data keeps its licenses: HarfBuzz's own under MIT-Modern-Variant, the Adobe and Unicode suites under Apache-2.0. The project follows the REUSE specification; reuse lint checks it.

References cited

These are kept in the zig-font-shaper Zotero collection.