Skip to content
API

Markdown mapping

The frontend reads markdown into the content tree. Every construct below either maps to a node or warns.

let (sections, warnings) = fleuron_markdown::to_sections(&text, "chapter-01.md", &Options::default());
let book = fleuron_markdown::assemble(fleuron_markdown::frontmatter(&text), sections);

to_sections works on one source, not one book. A novel may arrive as one file to split into chapters, or as one file per chapter. Composing sources into a book is assemble, a step of its own, so the caller orders them and decides the metadata.

Sections

A section is what the fragmenter opens a page on. Options::sections sets where one begins.

policy
Sections::AtHeading(level)A heading at that level or shallower opens a section. AtHeading(H2) cuts at # and ## alike.
Sections::WholeThe source is one section. A file per chapter.

Content before the first heading gets a section of its own.

The two arrangements produce the same book. A single file split at its chapter headings and one file per chapter read whole give the same tree, section for section and block for block. What differs is each section’s source, and the positions inside it, which count from the top of whichever file the prose was read from.

Metadata

A source’s frontmatter belongs to that source. What it means depends on whether the source is the book or a chapter of one.

frontmatter reads the --- block at the top of a source. title and author are the named fields. Every other scalar joins extra, which the engine passes through and style may read. Values are scalars: a line that is not key: value is not metadata.

---
title: Pride and Prejudice
author: Jane Austen
year: 1813
---

A source read whole is one chapter, so its title: becomes that section’s title rather than the book’s. Nothing lays a section title out yet. It stays in the tree unused.

The class: and id: of a source read whole become the section’s attributes. class: takes several classes, separated by spaces. The following example names a preface that a sheet reaches as section.front or section#preface:

---
title: Preface
class: front
id: preface
---

That leaves book metadata for the caller. assemble(metadata, sections) takes it as an argument, so a library or WASM host passes whatever it already knows and the CLI passes what its flags named. The frontend never takes the book’s title from whichever file came first.

A book with no metadata lays out. The engine reads three fields: title and author reach the PDF’s document information, and so does extra["language"], which also chooses the hyphenation patterns. Everything else in extra is passed through for whoever wants it.

Several sources

Each file is read on its own, under its own name, and the sections concatenate in whatever order the caller composes them. Metadata is decided once, by the caller, and handed to assembly:

use fleuron::content::Metadata;
use fleuron_markdown::{Options, Sections, assemble, to_sections};
// A file per chapter, so nothing is cut at a heading.
let reading = Options {
sections: Sections::Whole,
..Options::default()
};
let mut sections = Vec::new();
let mut warnings = Vec::new();
for (name, text) in chapters {
let (read, complaints) = to_sections(&text, &name, &reading);
sections.extend(read);
warnings.extend(complaints);
}
let book = assemble(
Metadata {
title: Some("The Levant Papers".into()),
author: Some("E. Marsh".into()),
..Metadata::default()
},
sections,
);

assemble numbers the whole tree, so it runs once over every section rather than once per file. Node ids are document-order over the book, and a file read on its own has no position in it.

Warnings accumulate the same way, and each one already names the file it came from.

Blocks

markdowncontent tree
# … ######heading, at that level
a paragraphparagraph
> …blockquote, nesting
a fenced or an indented blockcode_block
---thematic_break
\pagebreak on a line of its ownpage_break
\columnbreak on a line of its owncolumn_break
![alt](url)image, as a block
- item or 1. itemlist, with an item for each marker
a tabletable, with a row for each line and a cell for each column
{.class #id} on a line of its ownthe classes and id of the block under it

An image is a block in the vocabulary and inline in markdown. One written on a line of its own is that block. One written among prose becomes a block directly after the paragraph it appeared in, alt text and all. Moving it changes the page, so it warns.

The string in ![](…) is the name the image is matched under, and it does not have to be a real URL. The host supplies the bytes under that name, the engine reads the header for the intrinsic size, and a name nothing is supplied for is a warning and a gap.

Lists

A list in markdown becomes a list in the tree. Each item holds blocks, so a list that is indented under an item becomes a block of that item.

A list whose markers are numbers is ordered. start is the number on its first item, so the following list counts from 7:

7. That the said man-mountain shall help our workmen.
8. That he shall deliver an exact survey of our dominions.

A list with no blank line between its items is tight. A blank line between any two items makes the list loose, and each item of a loose list holds paragraphs. In a tight item, the text is not in a p element, as HTML writes <li>text</li>. So li > p in a stylesheet selects only the paragraphs of a loose list.

An attribute line above a list names the list, not its first item.

Tables

A table in markdown becomes a table in the tree. The row above the delimiter row is the header row, and each row under it is a body row. Each cell holds its prose as one paragraph. An image written in a cell stays in that cell.

| Where it was found | What it was |
|:---|---:|
| The right fob | A watch |
| The left fob | A purse |

The delimiter row gives each column an alignment: :--- is left, :---: is center, and ---: is right. Every cell of the column carries it as align. The engine applies it to the cell as text-align, and a text-align rule in a stylesheet takes priority over it.

A row with fewer cells than the header row gets empty cells at its end. A table has no syntax for a cell that spans two columns or two rows.

Code blocks

A fenced block and an indented block both become a code_block in the tree. The block holds text rather than inlines, because a code block has no markup inside it. The word after the opening fence becomes info. The engine carries info for a painter that reads it and makes nothing of it, so there is no syntax highlighting.

The text keeps the newlines and the spaces the author wrote. A code block breaks its lines at its own newlines and nowhere else. Nothing in it hyphenates and no line of it is justified. A line wider than the measure runs past the measure, and the engine warns with the line and column of the block.

The following example sets three lines, with the second one indented by four spaces:

```sh
if [ -f build.toml ]; then
build build.toml
fi
```

The engine has no tab stops, so the frontend replaces a tab with the spaces that reach the next four-column stop.

A stylesheet reaches a code block as pre. The block holds text rather than an inline code, so pre code selects nothing. code is the inline code span. The built-in stylesheet gives pre a monospace family and space above and below it. See the CSS subset.

Page and column breaks

A line with nothing but \pagebreak on it starts a new page. A line with nothing but \columnbreak on it starts the next column. The following example starts a new page between two scenes:

The door closed behind him, and the house was quiet.
\pagebreak
Three weeks passed before the next letter came.

A stylesheet reaches the two breaks as pagebreak and columnbreak. The built-in stylesheet gives them break-after: page and break-after: column, and a rule in a stylesheet can change either one. An attribute line names one break, so a rule can reach that break and no other. The following example keeps a break with the class draft from starting a new page:

{.draft}
\pagebreak
pagebreak.draft { break-after: auto }

The line must hold nothing else. \pagebreak in a paragraph of prose is prose. Two breaks with nothing between them start one new page. On a page of one column, a column break starts a new page.

A break at the end of a chapter does not change the page that the next chapter starts on. Under the built-in stylesheet, the next chapter still starts on a right-hand page.

Naming a block

A stylesheet reaches a kind of element by its name: every img, every blockquote. A class or an id names one of them instead, this image rather than every image.

A line with nothing but a brace run names the block written under it. It takes any number of classes and at most one id, in any order:

{.epigraph}
> Man is the only animal that blushes.

Every block in the vocabulary takes one, the scene break included. Leave a blank line between the line and a paragraph, since two lines of prose are one paragraph.

Two blocks also take the run written after them. A heading takes it at the end of its line, and an image alone on its line takes it directly after the url:

# Chapter One {.opening #ch1}
![a map of Lilliput](plate.jpg){.plate}

Nothing else takes a trailing run. On a blockquote it stays inside the quote’s last paragraph, as prose. After --- the line is prose too, and no longer a break.

A line that names nothing becomes plain prose, and warns. This happens when the source ends under it, or when another attribute line follows it. A brace run that is not classes and an id does the same. {key=value} is prose and a warning, like every other construct the vocabulary has no room for.

An id names one element. A book that writes one twice lays out with both matching, and warns naming both places.

Dialect::attributes is the switch for the attribute line and for the span. It is on under Dialect::fleuron(), which is the default, and off under Dialect::common_mark(), where a brace run is prose.

Naming a run

A brace run written directly after a bracketed run names that run. The run becomes a span, an inline element with no meaning of its own. A class or an id on it is what a rule reaches, so a sheet can set part of a heading on its own.

The following example writes a chapter opening as a number over a title, in one heading:

# [Chapter One]{.number} [The Road]{.title}

The following rules set the two runs at their own sizes:

.number { font-size: 24pt }
.title { font-size: 12pt }

The whole heading is one heading to string-set, to a running head, and to its slug. To take one run instead, set the string on the span. The following rule puts the title alone in the running head:

h1 span:last-child { string-set: chapter-title content(text) }

The brace run must follow the bracket directly. [words] {.class} is prose. A heading takes the brace run at the end of its line, except where that run closes a bracketed run. That run is the span’s.

A brace run that is not classes and an id stays plain text, the brackets included, and warns.

A link can name a heading in its own source or in another source. target-counter() and target-text() then print the page and the text of that heading. The same link works in Obsidian, with no change to the source.

A link has two parts, divided by #. The part before # names a source, and the part after # names a heading in that source. The following links in Chapter 1.md all name the heading # The Hunter in Chapter 3.md:

[the hunter](Chapter%203.md#the-hunter)
[the hunter](Chapter%203.md#The%20Hunter)
[[Chapter 3#The Hunter]]
[[Chapter 3#The Hunter|the hunter]]

The last two links are wikilinks, which need the wikilinks switch.

The source part

The engine compares the source part with the name that the host gives to to_sections. It compares them the way Obsidian finds a note:

  • The engine ignores case, and .md is optional.
  • The engine decodes percent escapes, so %20 is a space.
  • A path that starts with ./ or ../ is relative to the folder of the source that holds the link.
  • Any other path matches the end of a source name. If two sources match, a source in the folder of the link comes first.

A vault is the folder that holds the notes of Obsidian. A host that reads a vault names each source by its path from the root of the vault. A link with no source part, such as #notes, names the source that holds the link. A link with no # part, such as [[Chapter 3]], names the start of the source.

The heading part

The heading part takes one of three forms:

formexample
an id written in the source#hunt
the slug of the heading#the-hunter
the text of the heading#The Hunter

A slug is the text of a heading in lowercase, with hyphens between the words. The engine makes it in four steps:

  1. It lowercases the text of the heading, without the markup.
  2. It changes each run of characters that are not letters or digits to one hyphen.
  3. It removes a hyphen at either end.
  4. If an earlier heading in the same source has the slug, or an id in the source is the slug, the engine adds -2, then -3, and so on.

A heading with a written id has no slug. A heading with no letter or digit in its text has no slug either.

The engine matches the text form the way Obsidian matches a heading. The engine ignores case. It also ignores the characters below in the link and in the heading:

! " # $ % & ( ) * + , . : ; < = > ? @ ^ ` { | } ~ / [ ] \ _

If two headings in a source have the same text, the text form names the first. [[Notes#Part Two#Sources]] names the first heading Sources that comes after the heading Part Two and is at a deeper level.

If a link has no source part and no heading in its own source matches, the link names an id written in another source.

The following example is a book of two sources. The first source is one.md:

# Chapter One
See [the notes](#notes).
# Notes
# Notes

The second source is two.md:

# Chapter Two
See [the notes](#notes) and [[one#Notes]].
# Notes

The headings in one.md have the slugs chapter-one, notes, and notes-2. The headings in two.md have the slugs chapter-two and notes. The link in one.md names the first Notes in one.md. The first link in two.md names the Notes in two.md. [[one#Notes]] names the first Notes in one.md.

A slug is not an id

A slug and the text of a heading are not CSS ids. A stylesheet cannot select a heading by its slug. In the example above, #notes in a stylesheet matches no heading. An id written in a source, such as {#hunt}, is a CSS id. A link can also name it.

A link to a web address, such as https://example.com, names nothing in the book. target-counter() and target-text() print nothing for it, and the engine does not warn. In the PDF, the link opens the address.

A link to a source or a heading that the book does not have also prints nothing. In the PDF, its text is not a link. For that link, the engine warns with the line and column where the link is.

Inlines

markdowncontent tree
texttext, entities decoded
*em*emphasis
**strong**strong
~~struck~~strikethrough
`code`code
[text](url)link
[text]{.class}span
[^label], with [^label]: elsewherenote, holding the blocks of the note
two spaces or \ at the end of a linebreak

~~struck~~ needs the gfm dialect, which is off by default. The dialects below say how to turn it on. The built-in stylesheet draws a rule through it.

A line wrapped in the source is a space in the tree. The shaper never sees the markdown’s ragged column.

Two spaces or a backslash at the end of a line is a hard break. The line ends there, and the next line stays in the same paragraph. Authors use it for verse, an address, and the closing of a letter. The following example keeps two lines of a sonnet as two lines:

Shall I compare thee to a summer's day?\
Thou art more lovely and more temperate:

A heading written with # holds one line of the source, so it has no hard break. A heading underlined with = or - can hold one. The following example is a chapter heading of two lines:

Chapter One\
The Voyage to Lilliput
======================

A rule on h1::first-line styles the text before the break. The text after the break takes the style of h1. The whole heading is one heading to string-set, to a running head, and to its slug.

Footnotes

A footnote is written the way GitHub writes one: [^label] where the reference belongs, and [^label]: the note on a line of its own. The label names the note and is not printed. The following example writes a note on the sum the author was given:

I got forty pounds,[^purse] and a promise of thirty pounds a year.
[^purse]: Forty pounds is the whole of the author's patrimony.

The note becomes a note inline where the reference was written. A note holds blocks, so a note of two paragraphs is two paragraphs, and a note can hold a list or a quotation. Indent the lines under the first one to keep them in the note.

The reference and the note can be written in either order. A note can be written far from the paragraph that refers to it. The engine puts the note at the foot of the page its reference lands on. The CSS subset says how to style the note and the area at the foot of the page.

Two things warn. A note that no reference names is kept where it was written. Where two notes are written under one label, the reference takes the second, and the first is kept where it was written.

What warns

The manuscript below uses every construct in the tables above, the ones with no counterpart included. Constructs the content tree has a node for are laid out. The rest is laid out as prose and reported under the page.

The Levant Papershe letter reached Marsh on a Tuesday, in the second post,1 andTit was not what he had been waiting for. It began Dear sir andended without a name, which is the whole of the difficulty.It closed, as the last one had:Yours in haste,A friend at SmyrnaNothing in the file said where the ship had gone.And nothing said who had asked.❦ShipSailedPortThe Hesper3 MarchSmyrnaThe Ariel9 MarchBeirutMarsh wrote down the two things he knew:1. the Hesper sailed from Smyrna2. nobody saw her come inif [ -f manifest.toml ]; then read-manifest manifest.tomlfiWhat the vocabulary has no room for is set as prose and reported.Marsh was certain almost certain, and wrote at once to the shipping officethat afternoon.1. The second post reached the office at four, and the clerk read it first.1
page 1
One manuscript through the frontend, read as GFM so that strikethrough is recognized. Take a construct out, or put another in, and the report changes with it.

These constructs have no counterpart in the content tree. The frontend lays them out as prose and reports the file, line and column. Prose is never dropped.

constructbecomes
a definition listone paragraph per entry
superscript, subscriptplain text
mathplain text
an image among prosea block after the paragraph
html, a task markernothing. There is no prose in them
an attribute line naming nothinga paragraph of its own
a brace run that is not classes and an ida paragraph of its own
a brace run after ] that is not classes and an idplain text, the brackets included

Dialects

Options::dialect is a set of switches, so a host’s departures from CommonMark are configuration.

switch
frontmatterA leading --- block is metadata rather than a scene break. On by default.
attributesAn attribute line names the block under it, a heading or an image takes the run written after it, and [text]{.class} is a span. On by default.
tablesTables, written as GitHub writes them. On by default.
gfmStrikethrough and task lists.
wikilinks[[wikilinks]], which become links.
smart_punctuationDashes and curly quotes at parse rather than in the manuscript.
breaksA line of nothing but \pagebreak or \columnbreak starts a new page or the next column. On by default.
footnotes[^label] and [^label]:, as GitHub writes them. On by default.

Dialect::fleuron(), Dialect::common_mark(), Dialect::gfm() and Dialect::obsidian() are the four named combinations. fleuron() is the default, and the other three are named after whose markdown they read. common_mark() is the strict one: frontmatter and nothing else, so a brace run, a table, and \pagebreak are prose.

A switch that is off turns its syntax into prose rather than an error. [[Another Note]] without wikilinks is four brackets and a title.

Reading a source twice

Cache keys a source’s sections by its name and a hash of its bytes, which is what a host re-rendering on every keystroke wants. The key is the name and the hash, not a node id: ids are assigned in document order over the whole book and renumber whenever a section is added or removed.

Parsing is deterministic, so two readings of the same bytes give the same tree.