All work

A family history from handwritten archives: eight generations found with AI

A family asked me to find ancestors they only knew from stories. AI read handwritten parish registers and censuses, and only what a document confirmed went into the tree. The work in the registers took one day. The result: eight generations documented back to the late 18th century, in a tree that belongs to the family rather than a platform.

Context

The client was my own family. What we knew about our ancestors came from stories, a few photos and letters, and our origins rested on a single family legend. The task: find ancestors on every line and living relatives, and put it all into a tree that can be opened and extended.

Everything older than the 1920s is handwritten church registers: old spelling, the Julian calendar, and a different hand for every clerk. The registers are digitised but not indexed, so an entry has to be found by paging through a file scan by scan. The work in the registers took one day: in that time the agent went through the files of several parishes and took the line back to the late 18th century.

What was built

Searching handwritten registers with AI. The agent opens a file and reads the title page and the index by year first. It works out the offset “scan number = leaf + k” and goes straight to the right year. Then it reads the whole year: births, marriages, deaths. Each entry found is transcribed into modern spelling, and the agent notes the leads: godparents, marriage sureties, social estate, parish. To prove a family had no other children, it went through every birth in the parish for 15 years in a row and every death for 16.

Strips instead of spreads. A script cuts a single column out of each spread — “names of the born” or “parents” — so dozens of spreads can be scanned in minutes. The full scan is opened only when a strip has a match.

A search log. For every file the log records which leaves were read, what was found and what was not. Empty results go into a “don’t repeat” list, so no one pays for the same file twice.

The tree on my own server. For the client I deployed Gramps Web, an open-source family tree app. It runs on my own cluster and is installed from a Helm chart: web server, background workers and queue in one pod. The trees came over from a genealogy platform as GEDCOM. Its export was broken: 371 Cyrillic letters split in half at line boundaries, about 300 line breaks without a level, 19 references to photos that did not exist. A script repaired the file before import.

AI works with the live tree. The agent connects to Gramps Web through an MCP server that runs locally in a container with no open ports. It enters what has been confirmed itself: people, events, sources, notes. Its rights are limited by the role of its account in the app, not by a config flag: lower the role and writing is closed, even if someone forgets the setting.

Two backups. Every night an exact copy of the database is taken through the SQLite online backup API, without stopping the app, and a script restores it. A second CronJob exports the tree to Gramps XML and commits it to a private repository. The export timestamp is blanked, so a commit appears only when something changed, and git log -p becomes a day-by-day log of edits to the tree. Scans and photos are read by the app from the same folder the export picks up, so there is no second copy. Files over 50 MB stay out of git and are logged.

Composites for the cards. Each find gets one image: the file cover, the index, the leaf with the entry, an arrow showing where the entry is, and a source caption. The builder checks by itself that all the text fits. The same register page used to be uploaded to the card of every person it mentioned. Now each card gets one composite with people tagged on it, and 32 duplicates were removed.

Scale: 100 people, 37 families, 187 events. Eight generations are backed by documents, the ninth comes from a directory. About 12 major hypotheses: 3–4 confirmed by documents, 6–7 refuted.

What it looks like

A parish register entry and its transcript: AI read the handwriting and checked the transcript against the scan. Scan blurred, names and places hidden
A parish register entry and its transcript: AI read the handwriting and checked the transcript against the scan. Scan blurred, names and places hidden, mobile version
A parish register entry and its transcript: AI read the handwriting and checked the transcript against the scan. Scan blurred, names and places hidden
The tree in numbers and one line upwards: every link has a document or an “unconfirmed” mark. Living generations hidden
The tree in numbers and one line upwards: every link has a document or an “unconfirmed” mark. Living generations hidden, mobile version
The tree in numbers and one line upwards: every link has a document or an “unconfirmed” mark. Living generations hidden

Why this way and not another

AI searches first, the document decides. AI finds candidates fast, but it can be confidently wrong. In the first session it called a candidate great-grandfather “almost certain” and built a chain of generations on top of him. After that came a separate document, “Hypotheses vs facts”, splitting everything into “confirmed by a document”, “from the family” and “hypothesis”, and a rule: a link enters the tree only with a reference to a scan. Links without a direct document carry a confidence percentage and the name of the document that would settle them. The cost: some branches stay open until that document arrives.

One day in the registers instead of trawling databases. At first the search followed where the family legend about our origins pointed, and went through dozens of databases with no result. Once the agent switched to reading the registers cover to cover, one day was enough to find a baptism entry in the register of a district town and take the line back to the late 18th century. The same register showed that a “distant” relative was in fact the great-grandfather’s own sister. The lesson for the method: a person would page through such registers for weeks, AI reads them in full in a day — so reading the primary source cover to cover beats searching indexes.

The tree belongs to the family, not the platform. A platform is handy for matching with other people’s trees, but it has a daily view limit, a one-day block after bulk browsing, limited photo storage, and a broken export. The main copy is now our own, and the platform stays a place for matches. The cost: the server and the backups are now on me.

XML in git, not just dumps. A dump is tied to the app version and unreadable in a diff. XML opens in any version of Gramps, and every edit shows up as a line. The cost: git history can’t be scrubbed and the tree holds living people, so the repository stays private for good.

How the search itself works and where AI goes wrong is in the note “Finding ancestors in handwritten registers”. The second tree, back to 1715, is a separate case: “A family tree to 1715 from confession lists”.

Other case studies

Case studies and personal projects