A typeface from one letter
A tiny model, trained on 2,806 open-source typefaces, that draws a whole alphabet from one letter, right here on your CPU. The first version lost to a lookup table. The second reads your letter, checks its own work, and says how sure it is.
Type designers don’t start at A.
They start with a handful of letters, usually H and O, and fight with them for weeks. How thick is a stroke? Where does it get thin? Are there serifs, and what shape are they? How wide is a round letter next to a square one? Does anything lean? Once those few letters are right, a surprising amount of the alphabet is already decided. The B is mostly an H that learned to curve.
So I wondered how much of a typeface hides in a single letter. If I showed a computer one A, could it guess the B?
I built a tiny model that tries. Draw a letter below. Any letter, any style. It reads your drawing and writes the rest of the alphabet in the same hand.
This figure is interactive.
All of that ran on your own computer. There’s no server behind this page and nothing you drew went anywhere. The whole model is about 200 KB, smaller than one photo off your phone, and it runs on the same CPU that’s rendering this sentence.
The rest of this post takes it apart, including the version that didn’t work, with the real models running inside every picture.
A typeface is a few decisions
Look at the letters it gives back and you’ll see it copy things it was never shown. Draw a heavy A and the O comes back heavy. Give the A little feet and the T grows feet too. Lean it, and everything leans.
That’s because a typeface isn’t twenty-six separate drawings. It’s a few decisions, made once and repeated twenty-six times:
- Weight: how much ink a stroke carries.
- Contrast: whether strokes swell and thin, like a pen held at an angle, or stay even, like a marker.
- Serifs: the little feet, and their shape: sharp, slab, round, none.
- Width: squashed and tall, or wide and round.
- Slant: upright or leaning.
A typeface is a few decisions, repeated twenty-six times.
An A shows you nearly all of them. The model’s whole job is to read those decisions off one letter and replay them on the other twenty-five. To do that it first had to see a lot of typefaces.
Twenty-eight hundred typefaces
Fontsource packages open-source fonts, including nearly all of Google Fonts, as npm packages. I wrote a script that walked the npm registry for every one of them, a little over two thousand families, and kept the ones under an open licence: the SIL Open Font License, Apache or MIT. From each family I took the regular, the bold and the italic, where they exist. Then every typeface gets drawn the same way:
- Render A to Z at four times the final size, so curves come out smooth.
- Line them up. Every typeface is scaled so its capital H is exactly 21 pixels tall and sits on the same baseline, then each letter is centred.
- Shrink to 32 × 32 by averaging each 4 × 4 block of pixels into one, which gives soft grey edges, the same trick antialiasing uses.
- Throw out the impostors. Barcode fonts, icon fonts, a font where every letter is a black bar, and fonts with missing letters.
One surprise: dozens of the “different” fonts were identical. Google’s Noto family ships a separate package for each writing system it covers, and every one of them carries the same Latin letters. I hashed every rendered alphabet and kept one copy of each.
That left 2,806 alphabets: 1,130 sans-serif, 590 serif, 514 display, 393 handwriting, 155 monospace and 24 that fit no label. About seventy-three thousand little pictures of letters. I set 140 of them aside before training and never let the model see them, so there’d be an honest test at the end.
A letter is a list of numbers
Before the model can do anything clever, the letter has to become something a computer can do arithmetic on. That part is easy. A 32 × 32 letter is 1,024 pixels, and each pixel is one number: 0 for paper, 1 for ink, something in between at the soft edges. Read them row by row and the letter is just a list.
This figure is interactive.
So to the model, every letter is a point in a space with 1,024 directions, one per pixel:
That space is enormous, and almost all of it is static: pick 1,024 random numbers and you get television snow, never a letter. Real letters live on a very thin sheet inside it, and the letters of one typeface sit close together on that sheet. The model’s job is to learn the shape of the sheet.
The obvious model
My first model did the obvious thing. Here it is on one page. The A on the left is a real drawing, and everything to its right is what that model did with it.
An encoder boils the 1,024 pixels down to 32 numbers, the style. A
decoder keeps a tiny sketch of every letter and paints it in whatever
style it’s handed, turning 136 knobs, a scale and a shift for each of its
inner pictures, that the 32 numbers set
Training is a game played seventy-three thousand letters at a time. Pick a typeface. Show the model one of its letters. Ask it for all twenty-six. Score the answer against the real ones, and nudge every number inside the model a little toward a better score. Then pick another typeface.
Keeping score
To learn, a model needs a number that says how wrong it was. It gets one for every pixel of every letter, and they’re added up.
The score is cross-entropy. If a pixel should be ink and the model said “ink, with probability ”, the cost is . Saying 0.9 costs about 0.1. Saying 0.5, a shrug, costs 0.7. Saying 0.01, confidently wrong, costs 4.6. Drag the guess along the curve to feel it.
This figure is interactive.
Over a whole alphabet that’s
and the first model added one more term, a leash that pulled every style toward the middle of its 32-dimensional space so that no two typefaces sat too far apart. Keep the leash in mind. It comes back.
A map of every typeface
Every one of the 2,806 typefaces now boils down to 32 numbers. You can’t
look at 32 dimensions, but you can squash them onto a page and try to keep
neighbours as neighbours. That’s what t-SNE does
The colours come from Fontsource’s own labels. The model never saw them.
This figure is interactive.
Nobody told it what a serif is, but the serif faces float off on their own island, away to the right. The sans-serifs make up the big continent, which makes sense: there are twice as many of them and they differ from each other in subtler ways. Handwriting gathers along the bottom edge, next to the italics, because to a model that only sees pixels a script and an italic are both letters that lean. The display faces, the loud ones made for posters, scatter everywhere, because each of them is weird in its own way.
Those 32 numbers aren’t labelled, but some directions through them matter much more than others. Line up all 2,806 styles and ask which direction they spread out along the most, then the next, at right angles to the first. That’s principal component analysis, and it gives the sliders below.
This figure is interactive.
Then I looked at what each slider did and named them afterwards. The first four directions turn out to be weight, width, slant and serifs: almost exactly the list of decisions at the top of this post, rediscovered from pixels. Press average and you get the typeface at the centre of everything, a plain, medium-weight sans-serif. The most typical typeface is the one nobody would notice.
It’s a lovely map. It’s also the problem.
It lost to a lookup table
Draw something odd on the first model and you get back something ordinary. A wobbly hand-drawn A came back as a tidy typeface that already exists. So I ran the most embarrassing test I could think of. Take the 140 typefaces neither model ever saw. For each, find the training typeface whose A looks most like it, pixel for pixel, and hand back its other twenty-five letters. No learning at all. A lookup.
The lookup won.
This figure is interactive.
Scored by ink overlap, the share of inked pixels that the guess and the real letter agree on, over the twenty-five letters it wasn’t shown:
| From one A | Real typefaces | Typefaces nobody made |
|---|---|---|
| The average typeface | 0.47 | 0.38 |
| Nearest-A lookup | 0.68 | 0.50 |
| First model | 0.62 | 0.47 |
| Second model | 0.63 | 0.56 |
The second column is typefaces bent into new ones, slanted or widened or heavied, so that nothing in the training set matches them. Ignore the last row for now.
The first model wasn’t reading the A. It was recognising it. Squeezing a style into 32 numbers throws away everything that doesn’t fit the shapes it already knows, the leash drags whatever’s left toward the middle, and it had only ever seen typefaces that exist. So it learned to file your A next to its nearest neighbours and hand back a blend of the neighbourhood. That’s a lookup table with extra steps, and a worse one.
The fix took three ideas: one old, and two from the last year.
Read, don’t summarise
The second model never squeezes your letter into a summary. It reads your A the same way the first one started to, with filters: small 4 × 4 grids of weights that slide across the letter, two pixels at a time. At each stop a filter lays its 16 weights over the 16 pixels underneath, multiplies each pair, and adds the products up. That one sum becomes one pixel of a new, smaller picture. Press slide and watch it build.
This figure is interactive.
Written down, the value at row , column of the new picture is
where is the letter, the filter’s sixteen weights and one more number on top. Each sum then goes through a gentle bend, , which lets big values through and quietens small and negative ones. Without a bend, stacking layers would be pointless: a stack of sums is just one big sum.
Nobody designed these filters. They started as random numbers, and training shaped them into detectors: green weights reward ink under them and orange ones penalise it, so a filter that’s green on the left and orange on the right lights up wherever ink stops at a right-hand edge.
Two layers of this turn your A into an 8 × 8 grid of patches, 64 little descriptions of 48 numbers each, every one saying what’s going on in its corner of the letter: a stroke edge here, a foot there, empty paper over there. And then, instead of squeezing those 64 patches into a style, the model keeps all of them.
Every letter it draws is also an 8 × 8 grid of patches, starting as its own
sketch of that letter. To paint its B, each patch of the B asks every
patch of your A a question, and takes a mix of the answers weighted by how
well each one matched. That’s attention
Here is what this patch of the B is looking for, each is what a patch of your A offers, and is what it hands over if chosen. The fraction is a softmax: it turns the matches into weights that add up to one, so the patch spends its attention like a budget. Point at the B below and see where the budget went.
This figure is interactive.
Watch where the glow goes. It isn’t spread evenly over your A, and it isn’t
parked on the patch in the same spot as the one asking. It lands on ink: the
strokes, the joins, the feet. Across typefaces the model never saw, 60% of
the attention paid to your A goes to the patches with ink in them, which are
only 30% of the letter. It’s measuring your strokes, not recognising your
typeface, and it’s the same trick the few-shot font models of the last few
years use
Check your own work
The second idea comes from a 7-million-parameter model called TRM
So the second model doesn’t draw your alphabet once. It runs one block of looking-then-thinking four times. And before every pass it paints its own copy of your A, the letter it can check, and reads that copy side by side with the original. The patches of every other letter can then look at two things: what your A looks like, and where the model’s own A is still wrong.
Written as one line, with every letter’s patches after pass :
The matters: each pass adds a correction to what was there rather than starting over. Watch it work.
This figure is interactive.
Almost all of the work happens on the first pass. Before it, the model’s copy of your A is just its own sketch of an A, overlapping yours by 0.30. After one pass it’s at 0.69; after four, 0.73. The later passes polish: watch the right-hand panel and you’ll see them fix single pixels at the edges of strokes, and the attention shifts with them. On the first pass the letters spend 46% of their attention on your A; from the second on, less than a third, and the rest on where the model’s own copy is still wrong.
That’s less than I hoped for. Running more passes than it was trained with makes it slightly worse, not better, so this model hasn’t learned to “think longer” the way TRM does on puzzles. A letter isn’t a puzzle; once the model has seen your A, there isn’t much left to work out.
Training scores every pass, not just the last, so each pass has to be a decent answer on its own. That’s called deep supervision, and it’s what makes a looping model like this stable enough to train.
Typefaces nobody made
The third idea is about the training data, and it’s the one that matters most. The first model only ever saw typefaces that exist, so remembering them was a winning strategy. The second sees typefaces that don’t.
Most of the time, before showing it a typeface, training bends the whole alphabet the same way: leans it by a random amount, squeezes or stretches it, makes every stroke heavier or lighter. Every letter gets the same bend, so it’s still a consistent typeface. It’s just not one that anybody designed, and no amount of remembering will find it.
This figure is interactive.
Against an alphabet like that, the only way to score well is to actually measure the slant of the A, the weight of its strokes and the width of its bowl, and carry them over. Recognising stops paying. Reading starts paying.
Training also scribbles on the letter it shows the model: hard black and white pixels instead of smooth grey edges, a pixel or two off centre, a few strays. That’s what a letter drawn on the pad looks like, and the first model had never seen one.
Knowing what it doesn’t know
An A tells you a lot about an H and very little about a Q’s tail. The first model hid that: it painted a faint half-tail and said nothing. The second says so.
Besides the letter, it gives every letter a number between 0 and 1: how
much of this letter it expects to get right. That idea comes from Jev
Honest has a precise meaning. Group every letter it said it was 70% sure of, and about 70% of each should be right. Here’s how close it comes, on typefaces it never saw:
It comes close. Where it says 84% it gets 82%; where it says 56% it gets 57%; it’s a shade overconfident at the top and about right everywhere else. Averaged over every typeface, it’s least sure of J (46%), M and W (56%) and Q (62%): the tails some designers drop below the baseline and some don’t, and the letters that come in every width, none of which an A can tell you about. It’s surest of T and Z (72%) and B and D (71%). The hero at the top of the page names the three letters it’s surest of and the three it’s least sure of, for whatever you draw.
Nudging every number
Training either model is then a loop:
- Pick 16 typefaces, bend most of them into ones nobody made, and for each, one or two letters to show the model and eight to ask for.
- Run the model and add up the score.
- Work out, for each of the 203,922 numbers inside it, which way to move it to lower the score. Calculus does this for every number at once; the method is called backpropagation.
- Move them all a tiny step that way, and go back to 1.
That’s gradient descent. Ten thousand rounds took two and a half hours on four CPU cores, and four thousand more with an extra term in the score that punishes grey, uncommitted ink took another hour. Here’s what two typefaces the model never trained on looked like along the way, read from their A alone.
This figure is interactive.
For the first fifty steps it’s television snow. Then a grey smudge fills the box where letters live, the same smudge for every letter, because “some ink, roughly here” is the cheapest bet. By step two hundred it can write HANDGLOVES, cleanly, but identically in both typefaces: it has learned the alphabet before it has learned that typefaces differ. By four hundred the italic leans, earlier than the first model managed, because bending typefaces in training made slant pay early. Serifs and the thin hairlines of contrast come last, sharpening over the remaining thousands of steps.
How small is small
| Parameters | 203,922 |
| File size | 203 KB, one byte each |
| Work per letter | about 21 million multiply-adds |
| A whole alphabet | about a second on one CPU core |
| Training | 14,000 steps, 3½ hours, 4 CPU cores |
| Data | 2,806 typefaces × 26 letters × 1,024 pixels |
The language models people argue about have billions of parameters. This one has about two hundred thousand. It knows exactly one thing.
Each weight is stored as a single byte. After training, every block of weights is scaled so its largest value fits in −127 to 127 and rounded, which makes the file four times smaller. On typefaces it never saw, the rounded model scores 0.6337 against the original’s 0.6336. Rounding cost nothing.
In the browser it’s a few hundred lines of plain JavaScript loops in a
background thread
Where it breaks
Here’s the table from the start of the second half again, with the bottom row filled in, and a row for showing it a second letter:
| From the letters shown | Real typefaces | Typefaces nobody made |
|---|---|---|
| Nearest-A lookup | 0.68 | 0.50 |
| First model, from A | 0.62 | 0.47 |
| Second model, from A | 0.63 | 0.56 |
| Second model, from A and R | 0.65 | 0.59 |
On typefaces nobody made, which is what anything you draw is, the second model wins clearly: it reads the slant, weight and width off your A and carries them over, where the lookup can only hand back the nearest thing that exists. On real typefaces it still loses to the lookup. Two and a half thousand typefaces is a crowded room; for most real typefaces there’s another one very like it, and the lookup answers with letters a human designed, crisp to the last pixel, while the model blurs wherever it’s unsure. A model this small, at this size of letter, doesn’t beat a good memory at remembering. It beats it at everything else.
By kind of typeface, from one A:
| Kind | Ink overlap |
|---|---|
| Sans-serif | 0.77 |
| Serif | 0.69 |
| Monospace | 0.67 |
| Display | 0.58 |
| Handwriting | 0.25 |
Handwriting is still where it falls over. A handwritten A tells you about the pen but almost nothing about the B, because in handwriting every letter is its own little gesture. Show it a cursive A and you get a faint, leaning ghost of a sans-serif. Neat printed handwriting, the kind taught in schools, it copies fine; anything messier turns to grey mush.
The looping helped less than I hoped, as you saw: almost everything happens on the first pass. I can’t tell you which of the other ideas did the most, because I didn’t have the hours to train it again without each one. The confidence turned out to be the most useful thing to show: when it says it’s unsure of your J, believe it.
It has no lowercase, no numbers and no punctuation, and at thirty-two pixels tall nothing it draws would survive being printed. But it fits in 200 KB, it runs on the processor you’re reading this on, and from one scribbled letter it gets the weight, the width, the slant and the feet of a typeface that never existed, and tells you which letters it’s guessing at. That’s a lot of a typeface to find in one A.