← Blog Posts
September 10, 2026 8 min read AI Claude Code Salesforce

Identity Resolution in a Shoebox

A relative handed me a folder of family history. 120 files, 72 megabytes, no index.

Roughly 95 of them are scanned images. Census pages from 1870 through 1940, handwritten in fading county clerk cursive. Ship manifests. Naturalization index cards. Death certificates. WWI and WWII draft registrations. Photographs of headstones. A Belgian birth record. A 1924 Canada-to-Detroit border crossing index. Twenty-one of the files are family group sheets saved as RTF, a couple are PDFs, and one is a GEDCOM export.

It’s the kind of gift that sits on your desktop for a year. Mine was well on its way.

Not because it isn’t interesting. Because opening it means opening 120 files one at a time, transcribing what each one says, and then figuring out which of several identically-named men it actually belongs to. The transcription is tedious. The second part is the one that stops you.

The pile is not the problem

Any single document in that folder is legible in about two minutes. A naturalization Declaration of Intention is a form. It has a name, a birthplace, a birth date, a ship, a port, an arrival date, a street address, and a physical description. My great-uncle’s says he was five foot nine, ruddy, blonde, blue eyes, 170 pounds, a laborer living on Bellevue Avenue in Detroit, and that he renounced allegiance to Albert I, King of the Belgians.

That’s not hard. That’s just slow.

The hard part is that the folder contains at least three pairs of people who share a name, and the filenames don’t tell you which is which.

There are two Victor Hernalsteens. One born in 1895 in Welle, Belgium, who arrived in New York in 1920. One born in 1925 in Detroit. A dozen files are named some variation of “Victor Hernalsteen,” and they split across two different men, an uncle and a nephew, thirty years apart.

There are two Peter Hernalsteens, born four years apart, to two different fathers who were brothers.

There’s a Cyriel and a Cyril, uncle and nephew again.

And that’s before the spelling. My great-grandfather is Petrus on his Belgian paperwork and PETROS in block capitals on the 1913 manifest that admitted him at Port Huron. Dudley Clark is enumerated as “Clarke Dudly” in 1870. And on a single 1924 border crossing index, the family surname is typed “Hernalsteens” on one line and “Hernalstin” four lines below it. Same page, same clerk, same afternoon.

None of these records share a key. There is no identifier that appears on both a 1924 border crossing and a 1942 draft card. There is a name, a date, and a place, and the name is unreliable, the dates disagree, and the places are spelled however they sounded.

I have built this system before

About two hours into reading the folder, I realized I was doing my day job. Same work, fewer status meetings.

This is identity resolution. It’s the same problem I’ve spent years solving in Data 360 (formerly Data Cloud): many source records, no shared key, names that vary by system, and a requirement to decide which ones describe the same human being. The mechanics I’d draw on a whiteboard for a client are exactly the mechanics this folder demands.

You match on a composite. Name plus birth date plus birth place plus a known relationship. Any one of those alone produces garbage, because “Victor Hernalsteen, born in Belgium” describes two men. Together they resolve.

You accept that matching happens at write time, not read time. Once I decided that the 1924 Declaration of Intention and the 1929 Petition for Naturalization and the 1942 draft card all belong to the Victor born in 1895, that decision gets recorded. It doesn’t get re-derived every time somebody opens the folder, which is precisely the failure mode the folder was already in.

And you need a survivorship rule, because sources contradict each other.

Which source wins

The family group sheets say my great-uncle’s wife was born on 27 December, either 1904 or 1905. Her naturalization index card says 27 October 1904.

That’s a real conflict and it needs a rule, not a coin flip. In Data 360 you’d settle this before configuring a single matching rule; here I settled it after the fact, which is the wrong order and cost me a second pass. The rule here is source authority: a naturalization index card is a contemporaneous government record generated from documents she presented. A family group sheet is somebody’s later transcription of what they’d been told. The card wins.

Sometimes the rule doesn’t resolve it. One of my great-aunts has a birth date of 19 February 1897 on the family sheet and 26 September 1897 carved on her headstone. Both are the kind of source you’d normally trust. A headstone is engraved by people who knew her, but decades after the fact and often from memory. That one stays open, flagged, waiting on a Belgian civil record to break the tie.

That’s the part most people get wrong about survivorship. It isn’t a ranking you apply once. It’s a per-field decision, and the honest output includes the ones you couldn’t settle.

Two files that lie

My favorite thing in the whole folder is a pair of files named Victor_J_Hernalsteen.jpg and Helen_Hernalsteen.jpg.

They look like portraits. Everything about the naming says portrait. They’re both passenger manifests, and neither one has a single face on it.

One is a passenger manifest for a flight from London to New York in July 1955, destined for Idlewild. The other is a list of US citizens aboard the Queen Elizabeth out of Southampton in the late 1940s, which incidentally records the exact date and the exact court where each of them was naturalized. They’re two of the most information-dense documents in the folder, and their filenames point away from that.

A human working through this alphabetically opens those two, sees no faces, and moves on.

Worth admitting: the first pass got one of them wrong. My working summary described that 1955 flight manifest as a 1927 ship crossing out of Antwerp, confidently and in detail, and I only caught it by going back and opening the scan. Which is the survivorship rule turned back on my own notes. A tidy summary of a messy pile is still a source, and it doesn’t outrank the document.

Then the enrichment pass

Resolution tells you who these people are. It doesn’t tell you anything the folder doesn’t already contain.

So the second pass went outward: Find A Grave records, funeral home obituaries, ship histories for the vessels named on the manifests. The Kroonland, the Lapland. Names that were just words on a form until you look up what they were and when they ran.

That’s the same shape as the Data 360 pattern too. You resolve first, then enrich, and never the other way around. Enriching before you’ve resolved means attaching a stranger’s obituary to your great-uncle with real confidence and no way to notice.

The enrichment is also what turned a list of dates into something I’d want to read. A family leaves Flemish Brabant and East Flanders between 1913 and 1927. The father crosses the Canada-US border in 1913, ahead of everyone. Six children follow over the next fourteen years, most of them routed through Canada rather than straight into New York. They land in Detroit, and the naturalization paperwork tracks them from laborer to a job at an aluminum plant on Mack Avenue.

None of that is in any single file. All of it is in the folder.

What came out

A document that says who everyone is, which record proves it, where the sources disagree, and what’s still unresolved.

That last section is the one I care about most. It lists the questions the folder can’t answer: the birth date that needs a third source, the two 1861 daughters on one family sheet who are probably one girl entered twice, the Belgian parish registers for Hekelgem and Welle that are largely online and entirely unworked.

An answer with no open questions attached would have been the wrong output. The folder is full of contradictions, and a summary that resolved all of them cleanly would be lying about how much it actually knew.

The part I keep thinking about

I’ve explained identity resolution to a lot of clients. Match rules, composite keys, survivorship, resolve at write time. It’s abstract enough that people nod and don’t retain it. That’s part of why I built a simulator you can actually play with. Its companion, the data harmonization visualizer, shows the step before matching: getting every source into the same shape.

Then a relative hands you a shoebox and it’s all right there. Two men named Victor. A name spelled four ways across two languages. A headstone and a family sheet that disagree by seven months. Two files named like portraits that turn out to be manifests.

Same problem. Older paper. It turns out the reason it’s hard in a warehouse isn’t the technology.

Someone in my family spent years pulling those records together. The least I could do was read them properly.