Spotting address changes at a glance with jsdiffr
A tiny HTML diff that makes data cleaning reviewable
This packages uses the jsdiffr package. Visit repo →
If you’ve ever standardised a column of addresses — expanding St to Street, fixing a postcode, normalising Ave — you know the hard part isn’t the transformation. It’s reviewing it. A before/after table of 5,000 rows is unreadable; your eye can’t find the one character that changed.
jsdiffr ships two HTML helpers, diff_to_html() and diff_html(), that turn a diff into colour-coded markup: red strike-through for what was removed, green for what was added. It’s a one-liner, and it works inline in Quarto, R Markdown, or a Shiny app.
One pair, one line
old <- "12 Baker St, London, NW1 6XE"
new <- "12 Baker Street, London, NW1 6EX"
cat(diff_to_html(diff_words(old, new)))12 Baker StStreet, London, NW1 6XE6EX
At a glance: St became Street, and the postcode swapped its last two characters — a typo you’d never catch scanning two text columns.
A whole column, side by side
diff_html() takes two vectors and renders an Original / Updated / Changes table. Pass view = FALSE to get the HTML back instead of opening a browser.
original <- c(
"350 Fifth Avenue, New York, NY 10118",
"1600 Pennsylvania Ave NW, Washington DC",
"Flat 4, 22 Acacia Avenue, Manchester M14 5TP"
)
cleaned <- c(
"350 5th Avenue, New York, NY 10118",
"1600 Pennsylvania Avenue NW, Washington, DC 20500",
"Flat 4B, 22 Acacia Ave, Manchester M14 5TP"
)
cat(diff_table(original, cleaned, method = diff_words))| Left | Right | Changes |
|---|---|---|
| 350 Fifth Avenue, New York, NY 10118 | 350 5th Avenue, New York, NY 10118 | 350 Fifth5th Avenue, New York, NY 10118 |
| 1600 Pennsylvania Ave NW, WashingtonDC | 1600 Pennsylvania Avenue NW, Washington, DC 20500 | 1600 Pennsylvania AveAvenue NW, Washington, DC 20500 |
| Flat 4, 22 Acacia Avenue, Manchester M14 5TP | Flat 4B, 22 Acacia Ave, Manchester M14 5TP | Flat 44B, 22 Acacia AvenueAve, Manchester M14 5TP |
Every row tells you exactly what your cleaning step did — and, just as importantly, what it didn’t touch.
Hide the noise you don’t care about
Reviewing a St → Street expansion across thousands of rows? That’s expected; you don’t want it lighting up red and green. ignore_words renders changes made up entirely of those words as plain context, so only the surprising edits stand out:
cat(diff_table(
original, cleaned,
method = diff_words,
ignore_words = c("ave", "avenue", "5th", "fifth")
))| Left | Right | Changes |
|---|---|---|
| 350 Fifth Avenue, New York, NY 10118 | 350 5th Avenue, New York, NY 10118 | 350 5th Avenue, New York, NY 10118 |
| 1600 Pennsylvania Ave NW, WashingtonDC | 1600 Pennsylvania Avenue NW, Washington, DC 20500 | 1600 Pennsylvania Avenue NW, Washington, DC 20500 |
| Flat 4, 22 Acacia Avenue, Manchester M14 5TP | Flat 4B, 22 Acacia Ave, Manchester M14 5TP | Flat 44B, 22 Acacia Ave, Manchester M14 5TP |
Now the abbreviation expansions fade into the background and the genuine changes — the added ZIP code, 4 → 4B — are what catch your eye.
Why it’s worth reaching for
- Reviewable cleaning. Diffs make an address-standardisation pass auditable instead of a black box.
- Zero setup. No JS toolchain —
diff_to_html()returns a string you can drop straight into a report orshiny::renderUI(). - Tunable signal.
ignore_wordsand thediff_*granularity (diff_chars,diff_words,diff_sentences) let you dial in exactly which changes matter.
Two functions, a couple of vectors, and a column of addresses becomes something you can actually read.