Work · Case study · 03
Printer recommendations & assistant
Explainable “what else will work” recommendations across 6,657 printers in OpenPrinting’s Foomatic database, plus a deterministic natural-language assistant. Everything runs at build time; the browser loads a 6 KB shard, not a 24 MB dataset.
Final report ↗PR #224 · recommendations ↗PR #230 · assistant ↗PR #236 · documentation ↗
Architecture
The system, at a glance.
01 — Claims and evidence
| Claim | Measurement | Context | Verify |
|---|---|---|---|
| Perfect-score saturation eliminated | 86.8% → 0% | share of recommendations scoring exactly 1.0, before and after the scoring redesign; measured by the project’s evaluation pipeline | final report ↗ |
| Scores now track evidence | 0.078 → 0.851 | correlation between supporting evidence and score | final report ↗ |
| Misleading explanation claims | 1,470 → 0 | final validation reaches zero false claims | final report ↗ |
| Scale | 6,657 printers · 463 dimensions | the full Foomatic printer set and the engineered feature space; 415 dimensions stay active once obsolete drivers are excluded. One shard each, median 3.2 KB | PR #236 · quality doc ↗ |
| Recommendations pull request | +4,013 lines · 34 files | 36 commits; open for upstream review | PR #224 ↗ |
| Assistant pull request | +11,200 lines · 70 files | 50 commits; screen recording in the PR; open for upstream review | PR #230 ↗ |
| Documentation | +1,091 lines | data formats, pipeline architecture, evaluation, regeneration, UI contract | PR #236 ↗ |
02 — The problem
OpenPrinting’s Foomatic database records 6,657 printers with their drivers and capabilities. You could look a printer up. You could not ask the database what to do next: “my printer is discontinued, what is similar?” or “find me a colour laser printer with good Linux support.”
The constraint shaped everything: the OpenPrinting website is a statically exported Next.js application on GitHub Pages. There is no request-time backend and no database behind these features. The expensive work has to happen at build time, and the browser’s job has to stay small.
03 — How it works
From Foomatic XML to a 3 KB shard
- foomatic-db XML is parsed and normalised into 6,657 printer records during the site’s generate step.
- Each printer is encoded as a feature vector in a 463-dimension space: driver families, command sets, PostScript and PCL levels, colour, mechanism type, resolution tiers and support grade. Drivers that foomatic-db marks obsolete are then excluded, so a superseded driver never counts as shared evidence; that leaves 415 dimensions active in the published build.
- Candidates are ranked by IDF-weighted cosine similarity, so sharing a rare driver (necp6, 8 printers) counts for far more than sharing postscript (1,746 printers).
- The score is damped by how much evidence the pair actually shares, then multiplied by penalties for capability conflicts in type, colour and extreme resolution gaps.
- Results below a minimum score are dropped, the top 10 are kept in deterministic score-then-id order, and human-readable explanations are generated from the shared attributes.
- Output is one JSON shard per printer, median 3.2 KB. A printer page fetches its own record and recommendations instead of a 24 MB aggregate.
- The assistant is a deterministic local pipeline over the same artifacts: normalisation, entity resolution, intent classification, a typed query against local data, a typed response. It is not a language model connected to an API, so it cannot invent printers outside the catalogue.
Do the expensive work before deployment. Keep the browser’s job small.
04 — Decisions
Build time, not request time
Everything is generated during the site build. No backend, no external API, and no data leaves the static site. This is what makes the feature deployable on GitHub Pages at all, and it is why every printer page loads a few kilobytes instead of the whole dataset.
An engineered similarity pipeline, not a trained model
With no ground-truth “replacement printer” dataset there was nothing to train against, and a black box could not have explained its own recommendations. An explicit pipeline can: every recommendation shows the evidence behind it.
Rare evidence should matter more
Common features are not necessarily informative. IDF weighting makes sharing a driver used by eight printers count far more than sharing one used by 1,746.
Similarity is not enough
A recommendation supported by one weak signal should not look as confident as one supported by many independent signals. Evidence damping and conflict penalties encode that directly.
A deterministic assistant that reuses the recommendation artifacts
The assistant does not recompute or re-rank; it uses the same shards, scores and shared features as the printer pages, so “similar” means one thing everywhere. Being deterministic, it is reproducible and grounded in the data.
Unknown is not false
Foomatic records duplex support for no printer at all. Answering “no” to “does this printer support duplex?” would be wrong. The assistant reports a data gap instead of guessing; missing colour information does not mean monochrome.
05 — What went wrong
Stated on purpose.
The first scoring model looked right and was wrong
The first version produced too many perfect-looking recommendations: 86.8% of scores saturated at 1.0. The cause was sparse data. Cosine similarity made two printers that shared a generic driver look identical, because that one shared feature was most of what either record contained. IDF weighting, evidence damping and capability-conflict penalties took saturation to 0% and moved the evidence-to-score correlation from 0.078 to 0.851. The evaluation harness now fails if any documented metric drifts.
06 — Limitations
- The metrics measure internal consistency and scoring behaviour, not human-labelled recommendation accuracy. Foomatic has no ground-truth replacement dataset to measure against.
- Capabilities that Foomatic does not record (duplex, for example) cannot be recommended on; the assistant says so rather than guessing.
- All three pull requests are open for upstream review as of this page’s last verification; the final report and the PRs are the primary sources until they merge.
07 — What I’d do next
- Land the upstream review on #224, #230 and #236.
- A small hand-labelled set of known replacement pairs, to put a human-judged number next to the internal-consistency ones.
Last verified 2026-09-29 · numbers reported as measured, with their context and source
