Content tree
The content tree is a semantic document, not markup. Everything downstream consumes these types.
Most callers never build one. Markdown is the usual input, and the frontend produces this. The types are here for a host whose source is already structured, such as a CMS or a docx converter, which constructs a Book in Rust directly.
The tree serializes, internally tagged so the shape maps one-to-one onto mdast. That is an output only: fleuron manuscript.md --dump-tree reads back what the frontend did, and nothing parses one back into a Book.
The dump, abridged:
{ "metadata": { "title": "Gulliver's Travels", "author": "Jonathan Swift", "extra": { "language": "en", "year": "1726" } }, "sections": [ { "source": "chapter-01.md", "title": "PART I. A VOYAGE TO LILLIPUT.", "blocks": [ { "type": "heading", "level": 1, "inlines": [{ "type": "text", "value": "CHAPTER I." }], "attributes": { "id": "opening", "classes": ["grand"] } }, { "type": "paragraph", "inlines": [ { "type": "text", "value": "My father had a small estate in " }, { "type": "emphasis", "children": [{ "type": "text", "value": "Nottinghamshire" }] }, { "type": "text", "value": "." } ], "position": { "line": 7, "column": 1 }, "span": { "start": 143, "end": 216 } } ] } ]}metadata
Section titled “metadata”title and author are the two the engine reads: the title for running heads, both for the PDF’s document information. extra is a string map the frontend owns: language, ISBN, subtitle, anything. language is the one key layout reads, for the hyphenation patterns. The engine does not read the rest. A stylesheet can read it.
All three are optional. A book with no metadata lays out.
sections
Section titled “sections”A section is a chapter or a file. It is the unit of markdown input and the unit of source attribution for diagnostics. Sections are in reading order, and a section starts a new page.
| field | |
|---|---|
source | The file the frontend read, such as chapter-01.md. Diagnostics name it. |
title | A title supplied outside the body, from frontmatter title:. Implies heading level 1. |
attributes | The classes and the id that a sheet reaches the section by, as section.front and section#preface. |
blocks | The section’s blocks, in reading order. |
position | Where in source the section began. |
span | The bytes of source the section was read from: the heading that opened it to the end of its last block. |
Blocks
Section titled “Blocks”type is the tag. Every block takes an optional position, span and attributes.
| type | |
|---|---|
heading | level is 1 to 6, the range markdown defines. The parser rejects a level outside that range. inlines is the heading’s text. |
paragraph | inlines. The unit line layout breaks. |
blockquote | blocks, not inlines. Blockquotes nest. |
code_block | text and an optional info. text is the block’s own text, with the newlines and the indentation the author wrote. info is the word after the opening fence, carried for a painter that reads it. |
thematic_break | ---. A scene break, set as space or an ornament depending on the stylesheet. |
page_break | \pagebreak. The content after it starts on a new page. It holds no text. |
column_break | \columnbreak. The content after it starts in the next column. It holds no text. |
image | url and alt. The string in url is the name the image is matched under, and it does not have to be a real URL, since the engine neither resolves it nor decodes the image. alt is not laid out, reaches painters and accessibility tools unchanged, and is not optional. |
list | ordered, start, tight and items. See Lists. |
table | head and body, each a list of rows. See Tables. |
A list holds items. An item holds blocks, as a blockquote does, so a list can be a block of an item of another list.
{ "type": "list", "ordered": true, "start": 7, "tight": true, "items": [ { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "That he shall help our workmen." }] }] }, { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "That he shall survey our dominions." }] }] } ]}| field | |
|---|---|
ordered | true for a list whose items are numbered. |
start | The number of the first item. A list that starts at 1 leaves the field out. |
tight | true where the source has no blank line between the items. A tight item has no p element in the stylesheet. |
items | The items, in reading order. |
blocks | The content of one item. An empty item has none. |
An item takes an optional position, span and attributes, the same as a block.
Tables
Section titled “Tables”A table holds rows, and a row holds cells. A cell holds blocks, as a blockquote does. A cell that the frontend reads from markdown holds one paragraph.
{ "type": "table", "head": [ { "cells": [ { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "Pocket" }] }], "align": "left" }, { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "Found" }] }] } ] } ], "body": [ { "cells": [ { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "The right fob" }] }], "align": "left" }, { "blocks": [{ "type": "paragraph", "inlines": [{ "type": "text", "value": "A watch" }] }] } ] } ]}| field | |
|---|---|
head | The header rows, which name the columns, in reading order. |
body | The body rows, in reading order. |
cells | The cells of one row, from the left. A row with fewer cells than the table has columns leaves the rest of them empty. |
blocks | The content of one cell. An empty cell has none. |
align | left, center or right: the alignment that the delimiter row wrote on the column of the cell. A column with no alignment leaves the field out. |
A row and a cell each take an optional position, span and attributes, the same as a block.
Inlines
Section titled “Inlines”| type | |
|---|---|
text | value. Plain Unicode, with the entities decoded by the frontend. |
emphasis | children. Italic, in the built-in sheet. |
strong | children. Bold, in the built-in sheet. |
strikethrough | children. A rule through the text, in the built-in sheet. |
code | value. Literal, with no markup inside, monospace and never hyphenated. |
link | url and children. The engine lays out the text. The url reaches the painters that can express one. |
span | children. A run a sheet names, with no meaning of its own. It takes the style of the element around it, in the built-in sheet. |
note | blocks, not inlines. A footnote. See Notes. |
break | No fields of its own. The line ends at the break, and the text after it stays in the same block. |
Text runs are not elements as far as CSS is concerned. They take the style of the inline or block around them, and never count towards :first-child. A break is not an element either.
Where the engine reads an inline tree as one string, a break is a newline. string-set with content() and target-text() read a heading that way.
A span is what a sheet reaches part of a heading by. A chapter opening written as a number over a title is one heading of two spans. A rule on each one sets the two runs at their own sizes, and string-set on one of them puts that run in the running head.
A note is an inline that holds blocks. It stands where its reference was written. The engine puts its blocks at the foot of the page that reference lands on.
{ "type": "note", "blocks": [ { "type": "paragraph", "inlines": [{ "type": "text", "value": "Forty pounds is the whole of the author's patrimony." }] } ]}A note holds no mark of its own. The cascade decides the number the reference prints and the number beside the note. counter-reset: note restarts the numbering. See Footnotes.
The text of the element around a note leaves the note out. A running head taken from a heading with a note in it holds the heading alone.
Naming a node
Section titled “Naming a node”attributes is what a sheet names one node by: classes, any number of them, and id, at most one. The names are as CSS spells them, without the . or the #. A node that has neither leaves the field out.
{ "attributes": { "id": "frontispiece", "classes": ["map"] } }Specificity counts them in buckets of their own: .map outranks img, and #frontispiece outranks .map.
The frontend reads them from an attribute line. A host with a structured source of its own sets them on the tree it builds. An id on two nodes warns naming both, and both still match.
A text run has the field like every other node, and a sheet reaches nothing through it.
Node identity
Section titled “Node identity”Every node has an id, and the engine assigns it, so input cannot collide ids or forge a diagnostic origin. Ids are never serialized, so the ids in a tree built by hand are unassigned until Book::assign_node_ids numbers them from 1 in document order. Numbering is pre-order: a node before its children, sections in reading order.
Call it once, after building the tree. Calling it again renumbers. A session numbers what it is handed, so content set through a session arrives numbered.
Source positions
Section titled “Source positions”position is a 1-based line and column into the markdown the frontend read the node out of, exactly as its parser reported them. Paired with the section’s source, it is what a diagnostic points at: chapter-01.md:12:3.
Positions are diagnostic data and never layout input, so a missing one never fails a run. A node with no position degrades to the bare file name, and a node with neither still warns, without a location.
Where a node was read from
Section titled “Where a node was read from”span is the other half: the bytes of source the node was read from, markup and all. A source and the text of the nodes read from it are different bytes, because markup is not text, so the span is the node’s extent rather than a letter-by-letter map. A byte of the file lands on the node written there, not on a letter of it. The node that contains another was read from a stretch that contains its own, and the innermost nodes tile the file, so one byte is one node.
Two questions are answered from it, and they are the two halves of a cursor’s way onto a page and back:
Book::node_at(source, byte) | The node one byte of one source was read into, innermost first: a byte of prose answers with the run it was typed into, a byte of markup with the construct it opens, a byte between two blocks with the section around them. |
Book::source_of(node) | The source a node was read from, and the bytes of it. |
A cursor becomes a node, and the runs of the display structure that name that node are on the page it is set on. A run under the pointer goes the other way. Only the sections read from the source asked about are looked at, so one file’s cursor is answered by one file’s nodes.
Both answers are about the book as it stands. Ids renumber whenever the book is set or one of its sources replaced, so a host that keeps one across an edit asks again rather than reusing it.
A node the engine synthesized, or one from a tree built rather than parsed, was read from nothing, and both questions answer with nothing rather than guessing.
Which pages a node’s content is on is a question about the book once it is laid out, so a session answers that one rather than the tree.