Markdown versions of all docs pages are available by appending .md to any docs URL.
PDF export
Stitch a docs section into one paginated PDF with the book output format and render-pdf.mjs.
A docs section can ship a downloadable PDF alongside the site. The module
provides the Hugo half β a book output format template that stitches every
page under a section into one long print document, plus the print stylesheet.
The consumer provides the Node half that paginates it and prints it.
Important
This pipeline is proven on a flat site (ambientmesh.io, one book) and on a versioned one (kgateway.dev, 14 chunks merged into a 1,828-page PDF). A book stitches whichever page opts in plus that page’s own subtree, so version scoping falls out of where the opt-in lives rather than needing any version logic of its own, and the version printed on the cover and in the running footer is read from that same tree. Several version trees of one product can opt in, each producing its own book. The constraint that remains is on the download link, not on the books: see Known limitations.
How the pieces split
The split is not arbitrary, and knowing it saves an afternoon of wondering why importing the module did not give you a PDF.
| Piece | Lives in | Why |
|---|---|---|
docs/list.book.html, docs/single.book.html, _partials/docs/book-document.html | This module | Ordinary layouts, so module.mounts carries them |
assets/css/print-book.css | This module | Ordinary asset, linked by the book document itself |
The outputFormats.book block | This module | Hugo merges a module’s top-level outputFormats into the importing project |
book in a page’s outputs front matter | Your repo | Per-page opt-in. List the whole set β outputs replaces the defaults |
HUGO_PARAMS_BUILDBOOK=true | Your repo’s PDF workflow | Per-build opt-in. Front matter cannot say “only when a PDF is being made” |
scripts/render-pdf.mjs | Fetched from this repo | See Fetching the renderer |
playwright, pdf-lib | Your repo’s package.json | Node resolves node_modules relative to the invoking project, not to wherever the script was downloaded |
1. The output format
Nothing to do. The module declares it, in its consumer-facing hugo.toml:
[outputFormats.book]
mediaType = "text/html"
baseName = "book"
isHTML = trueHugo merges a module’s top-level outputFormats into the importing project, the
same way it merges the module’s [module] block, so importing this module is
enough. baseName is what produces book.html next to the section’s
index.html.
It is inert until a page asks for it: defining an output format renders nothing
on its own, and no [outputs] default includes book. A consumer that publishes
no PDF is unaffected.
Note
This page used to say the opposite β that Hugo does not merge a module’s
outputFormats, so the block had to be repeated in every consumer config. That
was wrong, and expensive: solo-io/docs carried 25 identical copies (three
configs for each of eight products, plus the all-products preview), and
forgetting one did not fail the PDF, it failed the entire build, because the
outputs front matter that selects the format is a property of the page and so
applies to every config that builds that content:
ERROR error building site: assemble: failed to create page from pageMetaSource
/latest: failed to resolve output formats [html book]:
OutputFormat with key "book" not foundVerified on Hugo 0.160.1 by deleting the block from one consumer config and
building: the site builds and book.html renders. Control: deleting it from the
module config as well reproduces the error above. If you are on an older Hugo
and see that error with no block in your own config, declare it yourself β a
project-level block wins over the module’s.
Declaring it in your own config is still required in one case: a project build
that passes --config, which replaces Hugo’s default config lookup so the
module’s hugo.toml is not read. That is why this repo’s own
hugo-{oss,enterprise}*.toml fixture configs keep a copy.
2. Opt a page in
---
title: Documentation
outputs: ["html", "rss", "markdown", "llms", "book"]
---Warning
List every output the page should have, not just html and book. Hugo’s
outputs front matter replaces a page’s default outputs rather than adding
to them, so outputs: ["html", "book"] silently drops that page’s .md, RSS
and llms.txt. Copy [outputs] section out of your config and append book.
Nothing fails, and the damage is narrow enough to survive review: only the
version root loses them, while every page below it keeps its .md, so what
breaks is Copy-as-Markdown and llms discovery on exactly one page per product.
Check it directly on the built output:
ls public/<product>/<version>/index.md public/<product>/<version>/llms.txtCompare against a version tree that did not opt in β that is the control that makes the missing files obvious.
A section page (one with its own _index.md) renders through
list.book.html; a leaf page renders through single.book.html. Both exist
because Hugo resolves an output format’s template per page kind.
Warning
A leaf page that opts in when no single.book.html is reachable silently
falls back to the site’s normal HTML template instead of erroring. You get a
book.html that looks plausible and is not a book document at all. The tell is
a missing paged.polyfill.js script tag β see Verifying.
3. Ask the build for a book
Front matter selects the output format, but it cannot say when. It is static, so an opt-in applies to every build of that page, and a book is the whole version tree stitched into one file. Books are therefore off by default, and a build that wants one says so:
HUGO_PARAMS_BUILDBOOK=true hugo --config=hugo-<product>.tomlSet it in the PDF workflow, not in a site config. Setting it in the config is the same as not having the switch at all.
The cost of not having it, measured on solo-io/docs: 12 book documents and 92 MB on
top of a 4.3 GB site, about 7% of the build time on the largest product, and 12
publicly reachable URLs each carrying a complete unstyled duplicate of a manual
with no noindex and no canonical, because a book document deliberately skips
baseof.html and all the normal head chrome. It also forces the link checker to
exclude book.html, whose relative links do not resolve from its own location
until prepare_book.py rewrites them.
Two ways to read the switch, since Hugo lowercases every config key: the
environment variable HUGO_PARAMS_BUILDBOOK and a TOML buildBook = true both
land as params.buildbook, which is what the theme reads.
Note
Forgetting it does not produce a broken PDF, which is the failure mode worth
knowing. The book document still renders β cover, table of contents, and no
chapters β and WeasyPrint paginates that perfectly happily and exits 0.
prepare_book.py fails on a zero-chapter book for exactly this reason, and its
error names this variable.
4. Add the dependencies
npm install --save-dev playwright pdf-lib
npx playwright install chromiumThe script imports from playwright, not @playwright/test. A repo that
already has the test package still needs this one.
The second command is separate on purpose. Installing the playwright package
does not download a browser binary, so a machine that has never run Playwright
gets a launch failure rather than a PDF. A repo whose Playwright browsers are
already installed for a test harness needs nothing extra, which is why this step
is easy to miss locally and then fail in CI.
Paged.js is loaded from a CDN by the book document itself, so it is not an
npm dependency. A pagedjs entry in package.json is unused weight.
5. Fetching the renderer
scripts/render-pdf.mjs is not a module mount. module.mounts covers only
layouts, assets and data, so the file rides along in this repo purely as
fetchable content.
Fetch it pinned to the version your go.mod already requires, so the go.mod
bump stays the single version pin and there is no second Makefile variable to
drift out of sync with it:
RENDER_PDF_VERSION := $(shell awk '/solo-io\/docs-theme-extras/ {print $$3}' go.mod)
RENDER_PDF_SCRIPT := .pdf-tools/render-pdf-$(RENDER_PDF_VERSION).mjs
pdf:
@mkdir -p .pdf-tools
@test -f $(RENDER_PDF_SCRIPT) || curl -fsSL \
https://raw.githubusercontent.com/solo-io/docs-theme-extras/$(RENDER_PDF_VERSION)/scripts/render-pdf.mjs \
-o $(RENDER_PDF_SCRIPT)
PDF_PROD_HOST=https://example.com node $(RENDER_PDF_SCRIPT)
The version is part of the cached filename, so a pin bump fetches a new copy instead of reusing a stale one.
6. Run it
Build the site first β the script serves the built public/ directory itself.
Most repos wrap what follows in a make pdf target; see
Generating a PDF locally for that and for the
WeasyPrint equivalent.
| Variable | Required | Default | What it does |
|---|---|---|---|
PDF_PROD_HOST | yes | β | The site’s real origin, for rewriting internal links |
PDF_BOOK_PATH | no | /docs/book.html | Single-document book |
PDF_BOOK_PATHS | no | β | Comma-separated, ordered. Chunked book; see below |
PDF_OUTPUT | no | public/downloads/docs.pdf | Output path |
PDF_PROD_HOST is required rather than defaulted, deliberately. Internal links
in the book resolve relative to wherever the document is loaded from β the
script’s own throwaway local server β so without rewriting them against the real
origin they would point at a dead http://127.0.0.1:<port>/ URL once the PDF is
downloaded and the server is gone. Failing loudly beats shipping a PDF full of
links to nowhere.
Both shipped consumers read PDF_PROD_HOST out of the site config rather than
hardcoding it twice, which keeps one origin to change when a domain moves:
PDF_PROD_HOST := $(shell yq '.params.themeExtras.prodHost' hugo.yaml)
Note
If your live site links a specific filename, pass PDF_OUTPUT explicitly on
every invocation that feeds a real build. Do not rely on the default.
The output lands directly in public/, which a plain make build does not
touch, so the PDF is in the same tree that gets deployed. Order the target so
Hugo runs first and the renderer second, since the renderer reads the built
public/ rather than producing it.
The dev server never sees it
hugo server renders in memory, so there is no public/ tree for the renderer
to read and no public/downloads/ for the dev server to hand back. A download
link is a 404 during local preview unless the PDF is written to static/
instead, which the dev server serves verbatim:
serve:
hugo --gc --minify
PDF_OUTPUT=static/downloads/docs.pdf $(MAKE) render-pdf
hugo server
For that override to reach the script, the render-pdf target has to leave
PDF_OUTPUT overridable. A recipe that assigns it inline on the node command
wins over the environment and silently discards the caller’s value, so declare
it as a ?= variable instead:
PDF_OUTPUT ?= public/downloads/docs.pdf
render-pdf:
... PDF_OUTPUT=$(PDF_OUTPUT) node $(RENDER_PDF_SCRIPT)
The PDF written this way is a snapshot as of server startup. Content edited during the session does not reach it until the target is rerun.
Nothing links to the PDF for you
Generating the file is the whole of what this pipeline does. No layout, card, or
sidebar entry points at the result, so a site that renders a PDF and never adds
a link ships a file reachable only by guessing its URL. Add the download link
yourself, and point it at the same path PDF_OUTPUT writes to.
Deploying it
The PDF exists only if the deploy build runs the renderer. A hosting provider
configured to run bare hugo produces a site with no PDF in it, no matter how
the Makefile is wired. Point the build command at one target that covers the
whole job instead:
ci-build: hugo-install
npm ci
npx playwright install chromium
$(MAKE) build HUGO=$(abspath bin/hugo)
Pinning Hugo inside that target, rather than in the provider’s own settings,
earns the extra lines twice over: the version stops living in a dashboard nobody
reviews, and the same command reproduces the deploy locally. One caveat on such
an installer target is that Hugo ships the extended build as a tarball for Linux
and as a .pkg for macOS, so it works in a build image and not on a
contributor’s Mac.
What the script does
- Serves the built
public/on a throwaway local server. - Opens each book path.
- Waits for every
.mermaidelement to gaindata-processed, so pagination never runs against a half-rendered diagram. - Drives Paged.js manually. The book document sets
auto: falseprecisely so pagination happens after the DOM is final rather than racing it. - Prints to PDF, rewriting internal links and building an outline (bookmark) tree from the chapter structure.
Chunking a large docset
Paged.js has a real ceiling, but it is set by the size of the stitched HTML
rather than by the page count of the result. An earlier version of this page put
it at “150β200 pages”, which measurement disproves: kgateway.dev’s reference
chunk is 584 KB of HTML and paginates to 363 pages without complaint, and its
2.8 MB traffic-management chunk succeeds too, while the full 7.1 MB tree never
finishes. Watch input size, not output pages.
The ceiling does look inherent to monolithic CSS Paged Media rendering rather than being a Paged.js defect. WeasyPrint, a completely separate implementation, slows down on the same document in the same way. See Choosing a rendering engine.
Above that size, generate one book per top-level section and merge. Each chunk
root sets bookChunkRoot: true in addition to opting in:
outputs: ["html", "rss", "markdown", "llms", "book"]
bookChunkRoot: truebookChunkRoot makes the opted-in page’s own title render as the first
chapter and TOC entry before recursing into its children, instead of starting
silently at the children the way a true book root does. Without it, a merged
multi-chunk PDF loses its section groupings and reads as a flat run of
subsections with nothing marking which section each came from.
Then pass the chunks in order:
PDF_BOOK_PATHS=/docs/a/book.html,/docs/b/book.html node render-pdf.mjsEach is rendered independently, then merged with pdf-lib, with each chunk’s
outline page indices offset by the running page total so the bookmark tree stays
continuous.
PDF_BOOK_PATH (singular) keeps working for a single-document book and takes a
fast path that skips the merge entirely.
Tip
Set outputs and bookChunkRoot with cascade, not by hand on every
section. kgateway-oss first set both directly on each of its 14 chunk roots.
A cascade block on the version root instead pushes both onto every direct
child automatically, and a section added later picks them up with no content
edit:
cascade:
- target:
path: "/docs/envoy/latest/*"
outputs: ["html", "rss", "markdown", "llms", "book"]
params:
bookChunkRoot: trueoutputs is a reserved front-matter field and stays at the top level.
bookChunkRoot is a custom one, read as .Params.bookChunkRoot. Hugo still
routes an unnested custom key into Params, so the params: block is not
strictly required, but writing it out says which of the two fields is which.
That single path-segment glob matches direct children only β it does not reach two levels down. This is plain Hugo, not a module feature.
PDF_BOOK_PATHS still has to be an explicit ordered list maintained by hand.
Hugo has no query for “every page that opted into an output format”, so a new
section needs one manual addition there even with the cascade in place.
Choosing a rendering engine
Paged.js is not the only way to turn the book document into a PDF, and the
alternatives get suggested often enough to be worth recording. These numbers come
from one afternoon’s spike against real kgateway.dev content on 2026-08-27,
against docs-theme-extras v0.3.3. Two chunks were used as the benchmark:
reference (584 KB of stitched HTML, 246 tables) and traffic-management
(2.8 MB, 77 chapters).
Two of these three are in production and the third was never enabled. Read the status row first β the rest of the table is a comparison of capabilities, not a menu of supported options.
| Paged.js + Chromium | WeasyPrint 69.0 | Pandoc 3.10.2 + TeX Live | |
|---|---|---|---|
| Status | Shipping. Drives render-pdf.mjs; kgateway.dev publishes with it (make pdf) | Shipping. Drives the split/merge pipeline; solo-io/docs publishes every product PDF with it | Never enabled. Evaluated once, in the spike above, and abandoned. No supported path uses it and none is planned |
print-book.css | Used as authored | Used as authored, zero unsupported-property warnings | Discarded; CSS has no role in a LaTeX pipeline |
string-set running headers, @bottom-* boxes, counter(page) | Yes | Yes | Reimplement in a LaTeX template |
Repeats <thead> when a table splits | No | Yes | Yes, via longtable |
reference chunk | 363 pages | 404 pages, 12s | No PDF produced |
traffic-management chunk | Renders | 595 pages, 25s | No PDF produced |
| Whole tree as ONE document (7.1 MB, 227 chapters) | Never completes (inherited claim, not re-measured here) | 1,879 pages, ~90s | Not reached |
| Extra runtime dependency | Chromium | Pango, GLib | TeX Live, rsvg-convert |
| Client-side JS, for example mermaid | Renders it | Needs a pre-render step | Needs a pre-render step |
Warning
Pandoc is not a supported engine, and the row above is the whole story. Every attempt in the spike ended without a PDF, so there are no page counts or timings to compare against β the two “No PDF produced” cells are not gaps in the measurements, they are the result. See Pandoc did not produce a PDF below for the five configurations that were tried and where each one stopped. Choosing it is a project, not a configuration change.
Which of the two shipping engines applies to you depends on which pipeline
your site is wired into, not on a preference: an OSS site that curls
render-pdf.mjs from a make pdf target is on Paged.js, and a product built by
solo-io/docs’s pdf-export.yml is on WeasyPrint. Nothing selects between them at
runtime, and there is no engine setting to change.
WeasyPrint is the closest substitute, and it is the only one that removes the
chunking requirement. It consumed print-book.css without a single
unsupported-property warning, including every paged-media feature the stylesheet
leans on, and it rendered the whole tree as one document where Paged.js cannot.
Rendering the whole tree in one pass is what makes the difference, because four of the chunked pipeline’s compromises exist only because of chunking:
| Chunked, Paged.js | One document, WeasyPrint | |
|---|---|---|
| Table of contents | None; each chunk’s own TOC is dropped as incomplete | Complete, whole book |
| Page numbers | Restart per chunk, hence the “Section N” footer label | Continuous, 1 to 1,879 |
| Bookmark outline | 14 top-level, hand-built via pdf-lib | 1,688 entries, generated by the renderer |
| In-PDF jumps | 1,386; cross-section links fall back to web URLs | 4,074 |
The renderer generating its own outline is worth noting on its own: the
hand-rolled PDF outline-dictionary code in render-pdf.mjs can be deleted
rather than ported.
The one prerequisite is globally unique ids. Hugo only guarantees heading
ids unique within their own source page, so the stitched tree carries 78
duplicated ids, before-you-begin alone appearing 112 times. Chunked, that was
survivable because every fragment lookup was scoped to the target chapter; in a
single document there is nothing to scope to, so ids have to be rewritten with
their owning chapter as a prefix, and links rewritten to match. That work ports
out of the browser cleanly. A ~150-line Python pass over the stitched HTML with
lxml does it in 0.3 seconds, leaving zero duplicate ids and zero dangling
jumps.
Its one quality win independent of chunking is repeating table headers across page breaks.
Memory, and why a big book is still split
One document does not mean one render. Peak memory tracks output pages, not
input bytes, at a steady ~1.6 MB per page measured across cuts of the
gloo-mesh-enterprise book from 347 pages up to 3,481. That book is ~6,500 pages,
so a single render needs ~11 GB and a 16 GB GitHub runner is torn down
mid-render. The failure is unhelpful: the runner dies, so the job reports only
The runner has received a shutdown signal and exit 143, with no traceback and
no partial output.
So prepare_book.py --max-part-bytes (default 2 MB) cuts the prepared document
into parts, the caller renders them one at a time, and merge_book.py
reassembles the result. Input size is only a proxy for pages, and a leaky one:
| Content | Pages per MB | 2 MB part |
|---|---|---|
| Ordinary prose | ~250 | ~500 pages, ~0.8 GB |
| Table-dense reference | ~620 | ~1,240 pages, ~2.0 GB |
The default is set for the table-dense case, since that is what actually constrains it.
Splitting is unconditional, and that is deliberate. A small book yields one part and a merge that is effectively a copy, so its output is unchanged. A conditional split would mean two code paths, with the rarely-exercised one belonging to the largest, slowest, least-frequently-run build.
Nothing is lost in the split, because WeasyPrint writes internal links as jumps
to named destinations and emits one for every element id, whether or not
anything links to it. Since the ids are already unique document-wide, the merged
file has one global namespace. The single gap is that WeasyPrint drops
<a href="#x"> when x is not in the part being rendered, logging
No anchor #x for internal URI reference, so every jump is rewritten to a
pdfjump: URI that survives as an ordinary link annotation and becomes a real
jump again at merge time.
Two consequences worth knowing:
- Parts must render sequentially. Page numbers are baked in during layout,
so each part needs the previous part’s page count, supplied as
@page :first { counter-reset: page N }. A bare@pageresets the counter on every page, andcounter-resetonhtmlorbodyis ignored. - A chapter larger than the target is cut between its direct children, never
inside a table or
details. Continuation slices are markedpdf-chapter-contso they do not start a new page.
On the gloo-mesh-enterprise book this took peak memory from >15 GB to 2,134 MB and total render time from 32+ minutes, never finishing, to 5m20s across 10 parts.
Page numbers in the table of contents
Splitting costs one more thing than links, and it is not obvious: the printed contents page.
CSS Paged Media has an answer β target-counter(attr(href url), page) on a TOC
link prints the page its target landed on, and WeasyPrint implements it. It
works only while the book is one document, and a book long enough to want a
printed contents page is exactly the book that had to be cut up. After the cut,
every chapter the TOC points at lives in a different document from the TOC, and
target-counter has nothing to count.
So the numbers come from the finished article instead:
- Render every part and merge, with
merge_book.py --page-map pages.json. The merged file is the first moment any destination’s page is knowable. number_toc.py pages.json --manifest book.parts.txtfinds the part holding the TOC, writes the numbers into its empty.pdf-toc-pagespans, and prints that part’s path.- Re-render only that part, at the same page offset it had before.
- Merge again.
The second render cannot invalidate the numbers it is printing, because
.pdf-toc-page is flex: 0 0 3em β a fixed-width column, so an empty box and a
1234 box take identical space. No title rewraps, the TOC keeps its length, and
every chapter after it stays where it was. That invariant is the whole design, so
number_toc.py --expect-pages/--assert-pages checks it rather than assuming it,
and fails the build if the count moved.
Reading destinations back is two linear passes. pypdf’s get_page_number()
scans the page list per call, which on this book would be ~2,900 destinations
against ~6,500 pages, or roughly 19 million comparisons.
Numbering starts at the contents page
The cover is unnumbered, which is the usual convention for a manual and also means the number a reader reads off the footer is the number the contents page printed against that chapter. Two settings have to agree for that:
print-book.cssblanks all four margin boxes on@page pdf-cover, a named page that only the cover element uses.@page :firstcannot do this job β it means the first page of the document being rendered, and the book is rendered as one document per part, so it would blank a footer somewhere in the middle of the book for every part after the first.- The caller renders the first part with
counter-reset: page 0rather than1, and passesmerge_book.py --page-mapthe matching--first-page 0. Without the second half, the page map holds physical positions while the footers hold printed ones, and every line of the contents page is one out.
The bookmark tree has to be rebuilt after a split
WeasyPrint derives the PDF bookmark tree from heading levels, per document. Each part is its own document, so each part’s tree is nested against the shallowest heading that part happens to contain rather than against the book. Concatenate those trees and the bookmark panel is correct until the first part boundary and flat afterwards. In the gloo-mesh-enterprise manual that meant “Get started”, “About” and “Setup” nested properly and then 56 more entries at the top level, most of them third- and fourth-level headings whose parents were in an earlier part.
The part HTML still knows the real answer, because the levels there are
absolute: the book layout emits a chapter at h2 plus its depth, and
utils/shift-headings.html pushes each page’s own headings down to match. So
merge_book.py --outline-from book.parts.txt drops the imported per-part trees,
reads the headings back out of the HTML, and builds one tree over the merged
file.
Two things this depends on, both easy to break by accident:
A heading’s id is almost never on the
<h*>element. Hextra’s heading render hook emits it on an empty offset anchor span inside the heading:<h5>Before you begin<span class="hx:absolute hx:-mt-20" id="β¦"></span> <a href="#β¦" class="subheading-anchor"></a></h5>So a descendant id is used when the element has none. Reading only the element’s own id finds 459 headings in the gloo-mesh-enterprise book β one per chapter and not one inside a page β which produces a plausible-looking panel missing 84% of its entries. A chapter’s title heading comes from the layout rather than from Goldmark and has neither, so it borrows its
<section>’s id; the contents heading is not in a chapter at all, which is why the layout gives itid="pdf-contents"explicitly.Pass
--outline-fromto both merges. The second merge (after the contents page is re-rendered with its numbers) rewrites the same file, so leaving it off puts the flat trees straight back into the artifact people download.
One limit carries over rather than being introduced here: heading levels stop at
h6, so a chapter nested five or more deep and the body headings inside it all
land at h6 and become siblings in the panel. utils/shift-headings.html caps
there because HTML has nothing deeper, and WeasyPrint’s own tree had the same
ceiling.
Fonts the renderer needs
Beyond fonts-dejavu-core and fonts-liberation (diagram SVGs ask for
Helvetica, and Liberation Sans is the metric-compatible stand-in), the emoji
font matters more than it sounds like it should. Comparison tables in this
content use β
and β, and under WeasyPrint they came out unreadable.
WeasyPrint cannot draw a color font. Not badly β at all. Rendered side by side in the container below, all three kinds embed into the PDF and all three leave the glyph box blank:
| Font | Technology | What WeasyPrint 69 draws |
|---|---|---|
| Noto Color Emoji | CBDT (bitmap) | nothing; the advance is reserved and no ink is placed |
| Noto COLRv1 | COLRv1 (vector) | nothing |
| Twemoji Mozilla | COLRv0 (vector) | nothing |
| Noto Emoji | glyf (outline) | the glyph, in the inherited text color |
So install the monochrome outline font, Noto Emoji, and get the color back a different way β see Color emoji below.
Installing it is not sufficient on its own. A hosted GitHub runner image already ships the color font, and Pango keeps choosing it, so the color font has to be rejected outright:
<!-- /etc/fonts/conf.d/99-no-color-emoji.conf -->
<?xml version="1.0"?>
<!DOCTYPE fontconfig SYSTEM "fonts.dtd">
<fontconfig>
<selectfont>
<rejectfont>
<pattern>
<patelt name="family"><string>Noto Color Emoji</string></patelt>
</pattern>
</rejectfont>
</selectfont>
</fontconfig>Warning
Do not verify this with fc-match. fc-match "sans-serif:charset=2705"
answers Noto Emoji even in the configuration that ships color glyphs,
because Pango resolves emoji fallback by script tag rather than by that query.
Render the characters and read back the embedded font instead:
printf '%s' '<meta charset="utf-8"><p>✅</p>' > probe.html
weasyprint probe.html probe.pdf
python3 -c "from pypdf import PdfReader; print([str(f.get_object()['/BaseFont']) for f in PdfReader('probe.pdf').pages[0]['/Resources']['/Font'].values()])"Color emoji come from the HTML, not the font
Because no color font renders, the color is reapplied to the characters
instead, before the renderer sees them. prepare_book.py --color-emoji wraps
each emoji it knows a meaning-bearing color for in
<span class="pdf-emoji" style="color:β¦">. An outline glyph honors color; a
bitmap one does not, which is the reason the monochrome font is the one
installed rather than a workaround around it.
The map lives in EMOJI_COLOURS in that script and covers the colored circles
and squares (π΄ π π‘ π’ π΅ π£ π€ β« βͺ and the square set) plus the status
marks (β
β β β β β β π« β β βΉ). Values are GitHub Primer colors, so a
table of status dots in the PDF reads the way the same table reads on the
website. Anything outside the map is left to the monochrome font.
Three details worth knowing before you change it:
- Tinting, not substituting. The character stays in the PDF’s text layer, so it still copies, searches and reads out. A CSS shape in its place would look cleaner and lose all three.
- Noto Emoji’s hatching survives, and that is a feature: π‘ is dotted, π’ is diagonally hatched and π΄ is vertically striped, so the distinction does not rest on hue alone for a color-blind reader.
- βͺ and β¬ are gray, not white. White on a white page is an invisible glyph, which is worse than the monochrome one it replaced.
The flag is opt-in because it is a WeasyPrint workaround. A Paged.js consumer renders in Chromium, which draws the real color font, and there the tint would repaint emoji that are already correct.
Diagram SVGs need two fixes, and neither one reports an error
prepare_book.py --fix-svgs public rewrites the built SVGs under public/
(never the sources under assets/) to work around two WeasyPrint bugs. Both are
specific to Excalidraw exports, and neither is anything wrong with the drawings:
every one of these files is correct in a browser.
| Symptom in the PDF | Cause |
|---|---|
Gloo Mesh resources prints as Gloo Mesh resources, legends overlap | Excalidraw writes font-family="Helvetica, Segoe UI Emoji". That second family does not exist on Linux, and WeasyPrint resolves the SPACE character through it, giving the space a wildly wrong advance |
| Connectors and most labels are simply gone, leaving loose badges and a few clipped words | <mask>. Excalidraw punches a hole in a connector where its label sits, with a two-rect luminance mask applied as <g mask="url(#mask-β¦)">. WeasyPrint 69 renders that group as very nearly nothing |
The mask fix removes the mask reference and leaves the <mask> element in
place, inert. The only thing lost is the hole, so a connector draws through its
own label instead of stopping short of it β legible, and what an unmasked
Excalidraw export looks like anyway. It is guarded on the svg-source:excalidraw
marker comment, because a hand-authored mask is usually load-bearing and
revealing content an author meant to hide is a worse failure than the one being
fixed.
Warning
WeasyPrint logs nothing for either of these. The render step’s ^ERROR:
check catches an image it cannot load at all; it cannot catch one it renders
incorrectly. A published manual can ship a gutted diagram with the job green,
which is why this runs unconditionally rather than waiting for an upstream
release. When a diagram looks wrong, render the single SVG on its own and
compare β a one-page probe is far faster than re-rendering the book:
printf '%s' '<meta charset="utf-8"><img src="arch.svg">' > probe.html
weasyprint probe.html probe.pdf--fix-svg-fonts is the flag’s former name and still works, which matters for a
historical build that fetches an older prepare_book.py. It now does both fixes.
Pandoc did not produce a PDF from this content at all
This is why the status row above says never enabled, and why the Pandoc column has no page counts to compare: there was never an output to count.
Five configurations
were tried, and each fix surfaced the next failure: a missing rsvg-convert, then
emoji that pdflatex cannot typeset, then LaTeX’s nested-list depth limit
(“Too deeply nested”), then the svg package wanting -shell-escape, then a
This can't happen (vertbreak) internal error inside a table. The HTML-to-LaTeX
conversion itself always succeeded, in a few seconds. The wall is the LaTeX
compile meeting real documentation content. None of this proves Pandoc cannot be
made to work, and a custom template with Lua filters and content sanitizing
probably would, but that is a project rather than a swap, and it starts by
throwing print-book.css away.
Note
Treat the timings as orders of magnitude, not benchmarks. They were taken in Docker Desktop on a laptop, where repeat runs of identical input ranged from 87s to 428s purely on VM contention. ~90s is the steady-state figure on an unloaded machine; a CI runner deserves its own measurement before anyone promises a number. The structural results (page counts, link counts, duplicate ids, which CSS is honored) are stable and repeatable; the clock is not.
Generating a PDF locally
Which command you run depends on which engine your site is wired into β see the status row in Choosing a rendering engine.
Paged.js sites (kgateway.dev)
There is a make target, because the renderer is a Node script the repo already
curls:
make pdf # Hugo first, then render-pdf.mjs over the built public/make serve runs the same thing with PDF_OUTPUT redirected into static/, so
the download link resolves during local preview instead of 404ing β see
The dev server never sees it.
WeasyPrint sites (solo-io/docs)
There is no make target for this one. The pipeline lives in
.github/workflows/pdf-export.yml, and that workflow is the source of truth β
the steps below reproduce it rather than replace it, so check them against the
workflow if the two ever disagree.
Run it in a container. Not for isolation, but because the fonts are part of the output: the renderer picks up whatever emoji font the host provides, and the wrong one silently prints β /β as invisible specks (see Fonts the renderer needs). A container that matches CI is the only way to see locally what will actually publish.
FROM ubuntu:24.04
ENV DEBIAN_FRONTEND=noninteractive
RUN apt-get update -qq && apt-get install -y -qq --no-install-recommends \
libpango-1.0-0 libpangoft2-1.0-0 libharfbuzz0b \
fonts-dejavu-core fonts-liberation \
fontconfig curl ca-certificates python3-pip poppler-utils
RUN pip install --quiet --break-system-packages weasyprint lxml cssselect pypdf
RUN curl -fsSL -o /usr/local/share/fonts/NotoEmoji.ttf \
"https://github.com/google/fonts/raw/main/ofl/notoemoji/NotoEmoji%5Bwght%5D.ttf"
COPY 99-no-colour-emoji.conf /etc/fonts/conf.d/
RUN fc-cache -fThe quick check: does this page look right?
This is the question you actually have most of the time, and it needs none of the splitting or numbering machinery:
HUGO_PARAMS_BUILDBOOK=true hugo --config=hugo-<product>.toml
python3 -m http.server 8000 --bind 127.0.0.1 --directory public &
weasyprint "http://127.0.0.1:8000/<product>/<version>/book.html" out.pdfWithout HUGO_PARAMS_BUILDBOOK=true there is no book.html to serve at all β see
Ask the build for a book.
Serve the tree rather than opening the file directly β WeasyPrint has no notion of a site root, so root-relative image and stylesheet URLs only resolve over HTTP.
Two things will look wrong, and both are expected here: the table of contents has no page numbers, and on a large book this either takes many minutes or exhausts memory, because nothing has split it. Neither is a bug in your page.
The full pipeline
Reproduce the published artifact β split, continuous page numbers, working cross-references, numbered contents β with the four stages the workflow runs, in order:
# 1. Prepare: unique ids, deferred jumps, emoji color, SVG fixes, split.
python3 prepare_book.py public/<product>/<version>/book.html \
public/<product>/<version>/book.html https://docs.solo.io \
--strict --color-emoji --fix-svgs public
# 2. Render each part IN ORDER, telling each one where its page numbers start.
# Sequential is not an optimization choice β a part cannot know its first
# page number until every earlier part has been rendered and counted.
# The FIRST part starts at 0, because the cover is unnumbered.
echo "@page :first { counter-reset: page $NEXT; }" > offset.css
weasyprint -s offset.css "http://127.0.0.1:8000/<product>/<version>/$NAME.html" "$PDF"
# 3. Merge: record where every destination landed, and rebuild the bookmark
# tree from the part HTML. --first-page must match the first part's offset.
python3 merge_book.py out.pdf pdf-parts/*.pdf --page-map pages.json \
--first-page 0 --outline-from public/<product>/<version>/book.parts.txt
# 4. Number the contents, re-render only that part, merge again.
python3 number_toc.py pages.json --manifest public/<product>/<version>/book.parts.txt
# ...re-render the part it names, then merge once more.Stage 2’s loop over book.parts.txt and stage 4’s re-render are the fiddly
parts; copy them out of the workflow’s Render PDF step rather than retyping
them.
Checking what you produced
The failures worth catching do not announce themselves β a diagram that vanished, emoji that rendered blank, a contents page of blanks, a bookmark panel that goes flat halfway down. These read the finished file:
# Emoji resolved to the outline font, not a color one.
python3 -c "from pypdf import PdfReader; print([str(f.get_object()['/BaseFont']) \
for f in PdfReader('out.pdf').pages[0]['/Resources']['/Font'].values()])"
# Bookmark nesting. Every top-level entry should be a top-level SECTION; a run
# of deep headings here is the per-part flattening described above.
python3 -c "
from pypdf import PdfReader
r = PdfReader('out.pdf')
print([i.title for i in r.outline if not isinstance(i, list)])"
# Look at a page instead of guessing.
pdftoppm -png -r 100 -f 12 -l 12 out.pdf pagemerge_book.py already fails on any cross-reference that resolves to nothing,
and number_toc.py fails on any contents entry with no page β so a clean run of
the full pipeline is itself a check. A WeasyPrint ERROR: line, though, does
not stop it: the renderer logs an unrenderable image and exits 0, which is
how a diagram goes missing from a green build. The workflow greps its own log for
^ERROR: and fails; do the same locally.
Naming the version on the cover
The cover and the running footer print the version of the tree the book walked,
resolved by utils/book-version.html from that tree’s own params.versions
entry. Most products need no configuration at all β the entry’s version already
holds a real number:
[[params.versions]]
version = "2.13.x" # printed: "Version 2.13.x"
linkVersion = "latest" # served at /latest/A product that instead puts the URL segment in version has nothing printable,
and its book comes out labelled “Version latest”:
[[params.versions]]
version = "latest" # printed: "Version latest" β useless on paper
dropdown = "2026.8.0 (latest)" # the real number, but this is a UI label
linkVersion = "latest"Add releaseVersion, which wins over version and is read by nothing else:
[[params.versions]]
version = "latest"
releaseVersion = "2026.8.0" # printed: "Version 2026.8.0"
dropdown = "2026.8.0 (latest)"
linkVersion = "latest"Warning
Do not fix this by correcting version instead. That field is not
display-only. assemble-assets.py in solo-io/docs names asset directories
assets/<product>/<version>, and reuse.html locates them by matching each
URL segment against .version β so on a tree served at /latest/, renaming
the field breaks the match, the resolved version comes back empty, and every
{{< reuse >}} snippet silently falls back to the unversioned asset path.
reuse.html and rebase.html also substitute .version into content for the
OSSβenterprise version remap. None of it fails loudly.
releaseVersion is deliberately not parsed out of dropdown. That string is a
UI label: it carries a (latest) suffix, and a hidden entry sets it to a single
space.
It changes the print label only, not the download URL. The link in the
Copy-as-Markdown menu resolves {version} from the URL segment, because it has
to match the release asset the PDF workflow publishes β and that asset is named
from the version directory. So a tree at /latest/ keeps a stable
β¦-latest.pdf URL while its cover names the actual release. Those two
coordinates answer different questions and are meant to differ.
Linking to the published PDF
Set params.pdfDownload and the Copy-as-Markdown menu on every docs page gains a
Download all docs (PDF) item, next to Print. For a site publishing to
solo-io/docs-pdfs, one line is the whole configuration:
[params.pdfDownload]
distribution = "enterprise"The release-asset URL shape is supplied by the partial, so it is not repeated per
consumer. Setting distribution is what turns the item on. Override the shape
only if you publish somewhere else:
[params.pdfDownload]
urlTemplate = "/downloads/docs.pdf"{product} comes from params.pdfDownload.product, falling back to
params.currentProduct and then params.folder.
{version} depends on whether the site has versions at all:
- A versioned tree substitutes the version segment of the current URL.
- An unversioned (flat) site has no segment to read, so
lateststands in. An unversioned docs set is by definition the current one, and the publishing workflow labels such an assetlatestfor the same reason, so the link and the asset agree by construction rather than by convention. - Either kind, with
params.sectionsregistered, prefixes the section:<section>-latest. Two parallel doc sets therefore cannot collapse onto one URL. - A consumer that wants none of this writes a
urlTemplatewith no{version}placeholder in it, which is left untouched.
With neither distribution nor urlTemplate the item does not render at all, so
a site that publishes no PDFs needs no change β and that is deliberate rather
than incidental. A site can build a book and still want no link: the book is what
makes a PDF publishable and says nothing about whether one was published, so
defaulting the URL for every consumer would hand such a site a link to a release
asset that does not exist. Read that as the reason the default stays off rather
than as a description of anyone’s setup β all six consumers configure
pdfDownload today.
Warning
[params.pdfDownload] is a fully-qualified TOML table header, so every bare
key = value line after it belongs to it until the next header. Dropped in
above the loose keys of a [params] block, it silently swallows
currentProduct, folder and everything else that follows β and the symptom is
not a build error, it is a download URL with a product name like
Docs%20framework%20test%20fixture in it, because {product} fell through to
the next candidate. Put it after the loose keys, immediately before the next
table header.
The item follows the book output format, not the version. It asks the book
root for its output formats rather than reading a separate flag, which is the
same opt-in that makes a PDF publishable in the first place β so the menu follows
the build and there is nothing to keep in sync:
{{ $vr := partial "utils/version-root.html" . }}
{{ $bookRoot := $vr.docsSection }}
{{ if and (not $bookRoot) (not $vr.isVersioned) }}
{{ $bookRoot = .FirstSection }}
{{ end }}
{{ with $bookRoot }}{{ if .OutputFormats.Get "book" }}β¦{{ end }}{{ end }}On a versioned site the book root is the version root, so the item appears
for a version that builds a book. On an unversioned site
version-root.html matches no version segment and leaves docsSection empty,
so the root falls back to .FirstSection β the page’s top-level section, which
on a flat site is the book root. Before 0.3.8 there was no fallback, and the
item was unreachable on such a site no matter how pdfDownload was configured.
The fallback is reached only when no version was found, because inside a
versioned subtree .FirstSection returns the product rather than the version
root. Non-docs sections resolve to themselves and carry no book output format,
so the if still keeps the item off /blog/ and the marketing pages.
Warning
Set params.pdfDownload in every config that builds the product, not just
the production one. This is now the only part of the pipeline with that
requirement β the book output format moved into the module β and it is the
worse of the two to get wrong: a missing block does not fail the build, the item
just silently disappears from preview and local builds.
Note
The build knows the book is produced; it cannot know the PDF has been uploaded. A version enabled between two nightly runs shows a link that 404s until the next one, so dispatch the PDF workflow when you enable a version rather than waiting for the schedule.
A build-time existence check was prototyped and rejected. resources.GetRemote
with method: head does work, returning nil on a 404 without downloading the
file, but it requires [security.http] methods to be widened to permit HEAD in
every consumer, and caches.getresource defaults to maxage = -1, so a cached
“missing” answer would never expire.
Cutting a new version
Copying the current tree to a numbered directory carries the book opt-in with
it, so the frozen version keeps its download link and the rolling tree keeps
publishing under the same tag. One edit in the publishing workflow’s own registry
finishes the job. For solo-io/docs that registry is .github/products.yaml, and
the field is pdf.versions on the product:
- name: gloo-mesh-enterprise
pdf:
versions: [latest, 2.13.x] # 2.13.x added by the cutThat list, not the front matter, is what the workflow renders. It builds its matrix from these entries and never scans content for the opt-in, so a frozen tree that has every other piece of the plumbing still produces no PDF until it appears here.
For a product served at /latest/:
- Copy
latestto the numbered directory, which the release does anyway. - Add that number to
pdf.versions. - Dispatch the workflow, so the frozen tree’s link works before the next scheduled run.
The old latest asset is overwritten with the new version’s content, and that is
the intent. The tag comes from the version path, so /latest/ always holds the
current manual and the frozen copy is where the old one gets archived.
A product whose directories are all numbered needs two changes instead of one:
step 2 replaces its entry rather than adding one, and the promoted tree needs
book added to its outputs by hand. The opt-in lives on whichever tree is
current, so a development directory created by copying an older one does not
inherit it.
When a version retires, remove its entry and leave the release alone. Removing an entry stops future renders and deletes nothing, so an archived PDF stays downloadable after its content directory is gone. That is also how to stop paying for an archive you want to keep: each listed version costs a Hugo build per run, since the render gate can only compare the stitched book once the site is built.
Warning
Do not leave the opt-in on a frozen tree you never list. The download link
renders from the output format, so the page advertises an asset that was never
published, and the workflow cannot catch it β it checks links against assets
only for the versions in its matrix. Either list the version or delete book
from the copied tree’s outputs.
Note
A frozen numbered tree needs no releaseVersion and no params.versions entry
to print a correct cover. utils/version-root.html accepts a segment shaped
like X.Y.x as a version on its own, so a 2.13.x tree prints “Version
2.13.x” from the path alone.
Verifying the output
The book document deliberately skips baseof.html and all normal docs chrome β
no navbar, no sidebar, no in-page TOC. It is a print artifact, self-contained
from <!DOCTYPE html> down. So a quick check on the built book.html:
| Check | Expected |
|---|---|
paged.polyfill script present | yes β its absence means the fallback described above |
print-book stylesheet linked | yes |
sidebar-container present | no β site chrome means you got the wrong template |
Warning
Do not check a book.html produced while hugo server is running against the
same publishDir. A dev server can write a LiveReload script into the output
and race a static build. Stop the server, delete the output directory, and
rebuild before inspecting or shipping.
Known limitations
Two trees sharing a version segment need params.sections to stay apart.
copy-markdown.html fills {version} from utils/version-root.html’s
currentVersion, which is the version segment of the URL and nothing above
it. A product whose version trees nest under a section β agentgateway’s
/agentgateway/kubernetes/latest/ and /agentgateway/standalone/latest/ β has
two trees whose segment is latest. A registered params.sections prefixes the
segment (kubernetes-latest, standalone-latest) and the collision goes away.
Without one, both trees still resolve to the same download URL, and opting both
in publishes two books while linking only one of them. So register the sections,
or opt in one tree per version segment.
A tree served at /latest/ needs releaseVersion to print a real number.
The cover and footer read releaseVersion off the tree’s own params.versions
entry, falling back to version. Most products need nothing: their version
already holds the real number, so gloo-mesh-enterprise’s latest tree prints
“Version 2.13.x” with no extra configuration. But agentregistry, kagent and
agentgateway set version = "latest" literally and keep the number only in
dropdown, so without releaseVersion their covers read “Version latest” β see
Naming the version on the cover.
A tree with no params.versions entry falls back to its URL segment.
utils/version-root.html treats a segment matching X.Y.x, X.Y.Z, latest
or main as a version even with nothing configured for it, so an unregistered
tree prints the segment itself rather than nothing. Correct, but it is the raw
segment: a main tree’s cover reads “Version main”.
The Paged.js page ceiling is a property of the renderer, not of this module. Chunking is the workaround, not a fix.
PDF_BOOK_PATHS is hand-maintained. No Hugo query returns every page that
opted into an output format, so the list stays manual. The mismatch runs both
ways: a section added after the cascade is in place gets
a book.html and is still missing from the PDF until it is listed, and a page
that should stay out of the PDF still gets a book.html built for it. The
second case costs build time and nothing else.