joey castillo's personal website

In Search of Lost Type

Originally published at Crowd Supply.

With the Open Book Touch campaign drawing to a close, I thought it would be worth stepping back and talking about what got me started on all of this: books, and the formats backing them. Today’s post is about a rabbit hole I fell down and a format that I love for Open Book Touch, but keep in mind, EPUB is still supported; it’s just my second favorite format for the device.

When I wrote in our campaign text that Open Book Touch had — finally — gained EPUB support, that was a hard-earned finally. And yet, Open Book’s affinity for plain text files from the beginning wasn’t about me being lazy and not wanting to implement EPUB support. It was about the fact that I truly like plain text! Like, imagine the number of technologies involved in reading an EPUB. It presents as a binary file, which has to be unzipped; then, inside, there’s the blizzard of angle brackets that is the package file, whose spine tells you how to collate a whole host of XHTML files, which you need to render using CSS rules… I mean, it works, and we use it, but if you presented it to an alien civilization discovering the smoking husk of our powerful culture, they’d need a lot of context to even start to read, “It was the best of times, it was the worst of times…”

By contrast, plain text is such a gift. Somehow in the last 60 years or so, we came up with a durable, universally agreed-upon format for encoding all the characters used to render all the languages of the world. It started with ASCII in 1963, encoding 95 printable characters, which somehow was enough for a whole lot of use cases. At the time we mostly applied them to teletypes, clacking out characters over hotlines in the hopes of staving off nuclear armageddon, but technically at that point, we had a format that could convert A Tale of Two Cities into a string of bytes (even if storage for the several hundred kilobytes involved would have cost as much as a large car or a small house). And now, with UTF-8, we can say the same for works by Leo Tolstoy or Laozi.

My thinking, in prioritizing plain text as a format for storing literary works: it’s just about the right format for a whole lot of books, and it’s kind of a delight that in the 20th century we invented a series of formats that makes it possible. Although there are some subtleties.

Character Study

For the Open Hardware Summit in Berlin this year, I printed an anthology of short stories to give to everyone in the conference swag bag. To typeset it, I used a cursed amalgamation of formats: Markdown, transformed to LaTeX, then typeset to PDF for eventual printing and binding. At every step along the way, I ran into challenges. For example, take this line, from In Another Country by Ernest Hemingway. At this point in the story, a character reveals that his wife has just died. To which the narrator replies (in plain text):

“Oh—” I said, feeling sick for him. “I am so sorry.”

Having said that, that’s not exactly what’s in the text. In the printed book, the word “so” is emphasized:

“Oh—” I said, feeling sick for him. “I am so sorry.”

ASCII (and by extension UTF-8) doesn’t really offer a way to italicize this. For the purposes of this web page, the plain text is encoded like this:

“I am <em>so</em> sorry.”

That’s well and good for reading in a browser, but it doesn’t exactly spark joy as plain text. But hey, that’s why we invented Markdown! In Markdown, I rendered that line as:

“I am *so* sorry.”

I won’t deny: this looks a bit better. Still, it’s not a solution for archival storage of literary texts, if only because it’s ambiguous: asterisks, after all, are real typographic marks that have appeared in printed books for hundreds of years. Who’s to say whether an asterisk is meant to be printed, or meant to emphasize a word?

Anyway. Once my Markdown rendered out to LaTeX, the ambiguity was gone — but with it went the readability:

“I am \emph{so} sorry.”

\emph{Woof.} In all these cases, plain text has the ability to show some version of the emphasis. But there’s a tradeoff: the more human-readable versions are more ambiguous, and the less ambiguous versions make it harder for a human reader to parse.

Setting the Scene

A similar issue manifests in any book with page breaks or semantically meaningful whitespace. In Anzia Yezierska’s The Lost “Beautifulness”, all the residents of the tenement marvel at the protagonist’s newly painted kitchen, and the hopes it gives each for their own lot:

“Amen!” breathed Hanneh Hayyeh. “May we all forget from our worries for rent!”

But the next line takes place some time after, as Hanneh Hayyeh washes the laundry for the woman who employs her. It’s a gap in time, but also a gap in space; these next lines appear after a short vertical whitespace gap:

Mrs. Preston followed with keen delight Hanneh Hayyeh’s every movement as she lifted the wash from the basket and spread it on the bed. Hanneh Hayyeh’s rough, toil-worn hands lingered lovingly, caressingly over each garment. It was as though the fabrics held something subtly animate in their texture that penetrated to her very finger-tips.

Markdown doesn’t have the capacity to draw this gap, as such. It can draw a horizontal line! But it can’t leave blank space. To get the gap working on the web, I had to break the kayfabe and insert a

<br>

In LaTeX, I dropped a

\vspace{1.5em}

But neither of these things is a human-sized, semantically meaningful thing. They’re repurposing the same plain text that appeared on the page, to add something that’s not actually plain text.

A similar issue manifested in the reprinting of E. M. Forster’s The Machine Stops. Forster breaks the novella’s text into three parts. Markdown can sort of fake that with a…

## Part I: THE AIR-SHIP

…and while LaTeX of course has syntax for this…

\chapter{Part I: THE AIR-SHIP}

…again, the syntax doesn’t quite spark joy. So maybe I was wrong, way back up top; maybe UTF-8 doesn’t give us everything we need to represent a book in plain text.

Unless…

Reading Between the Lines

The first 128 characters of the UTF-8 standard are identical to the old ASCII standard, full stop. 95 of those are the printable characters that make up the bulk of English language text: letters, numbers, punctuation marks. The remaining 33 are control codes which, in the Teletype and modem era, contained instructions for the machines handling these streams of text. Of those 33 control codes, exactly three survive today in general use: the tab (move ahead some spaces), the line feed (move down a line), and the carriage return (return to the beginning of a line).

A few others still have uses in some places: printers still obey the “Form Feed” control character, and UNIX man pages use the backspace in much the way classic teletypes did, to “overprint” an underline by emitting an underscore, a backspace, and then the actual character. But what of those others? What did the makers of ASCII anticipate when they wrote this 1960s standard?

What of their legacy did we inherit?

Red Shift

My first sense of enchantment with the nooks and crannies of ASCII came from the U+000E <SO> and U+000F <SI> characters. These were, in the Teletype era, designed to shift the two-color printer ribbon: “out” to red, for emphasis, and then back “in”, to black, for normal text.

My thought: what a perfect way to mark italic emphasis. We’re using this ancient code, to some extent, for its exact purpose.

The joy of this is that modern text readers generally don’t render these control codes at all — but they remain completely valid UTF-8! So a text file containing those characters reads fine, and a renderer like Open Book Touch, knowing that they mean emphasis, can repurpose these invisible control codes to italicize text.

In perhaps a slightly ahistorical move, I’m letting these stack up for increasing levels of emphasis, ␎italic␏ and ␎␎bold␏␏ to match Markdown’s convention — but the upshot is that marking text in this way requires no printable characters in the stream, just control characters.

Chapter and Verse: The Separators

ASCII also gifted us this cluster of control characters in its legacy:

  • U+001C <FS> - File Separator
  • U+001D <GS> - Group Separator
  • U+001E <RS> - Record Separator
  • U+001F <US> - Unit Separator

Originally, these were used to separate files on a long run of tape, or records in a database. But when I saw them, all I could think was that this is the perfect way to separate books into chapters. In-band, in a single stream of text, a <RS> record separator breaks a book into chapters, and a <US> signals a scene break. Above that, <GS> group separators separate parts, and <FS> file separators separate volumes. My taxonomy:

  • <FS> - A standalone work that could be shelved as its own physical book
  • <GS> - A major named division: part, biblical book, something that groups chapters
  • <RS> - The standard reading unit: the chapter
  • <US> - An in-flow break: a section or scene

This maps astonishingly well to a vast swath of literature, and allows for a scan of the book to yield a semantically meaningful table of contents: Fitzgerald’s The Great Gatsby gets RS separators for chapters and US for scene breaks. Nothing above. Meanwhile Proust’s In Search of Lost Time gets FS separators for volumes, GS for parts, RS for chapters and US for scene breaks. The King James Bible: FS for testaments, GS for books, RS for chapters.

And — again — if you open a plain text book in a plain-text-oriented program, odds are it’ll simply skip them. You just see the title of the chapter, or the title of the volume.

The Escape Hatch

Books also contain formatting like block quotes, features like footnotes, and occasionally inline images. For this, there’s no way around polluting the stream with some printable characters; the ASCII committee couldn’t have anticipated everything we needed. But they did anticipate that we might need to say something to the effect of, “This next character shouldn’t be treated the same way.”

That’s where U+0010 <DLE> comes in. The “Data Link Escape” character was meant to switch out of data mode — the printable characters that represent part of a book — and into a “Control Mode” where the next run of characters represents a control sequence. That makes this next part simple:

  • Block quotes: ␐>This is a blockquote.
  • Footnotes: ␐^[1]: This is a footnote.
  • Images: ␐![480x320 An inline image](path/to/image.png)

Again: ASCII has given us the opportunity to completely disambiguate these situations. A line beginning with a > symbol — entirely possible in a book — won’t inadvertently render as a block quote, because the ␐ character is required for it to render that way.

Type Regained

I’m calling this format .text, which is already an alternative extension for plain text — but then again this format is plain text. There are some other details to the .text spec, and I’ll be putting forward a final draft of it in the coming weeks. But the desire to store books as plain text has been part of Open Book’s DNA from day one. In much the same way a microcontroller of the 2020s has the power of a computer from the 1990s, I think it’s worth considering how a lean, capable format like UTF-8 — ASCII, all grown up — can punch above its weight as an archival format for the great cultural works of our time.

Put another way: I’d rather the aliens find e-books in the rubble bearing our legacy as UTF-8 text, and not a Babel-like cacophony of MOBIs and encrypted AZW3s.

Open Book will read your EPUBs. But I hope it also serves to popularize plain text books: books that use the full capabilities of what UTF-8 text has to offer.