RedactCheck · Yunseon Hong

What redaction tools miss in a subject access request

A practitioner pointed me at a distinction the tools do not make. I went and read the definition, then built a small tool, and met the same problem from the other side.

The tools look for identifiers. The job is about people.

In August I wrote to fifteen organisations that handle subject access requests and asked one question: how do you redact at the moment? Most did not reply, which is a fair answer to a cold email. One reply came from someone who does this work for clients, and it changed what I was building.

They redact by hand. The software they had looked at was built around the American idea of PII, a list of identifier types, and did not follow the wider European definition of personal data. They suggested I go and look at that distinction myself, so I did.

It sounds like a lawyer’s footnote until you try to write code against it. PII, as most tools understand it, is a list of things with a fixed shape: a national insurance number, an email address, a phone number, a date of birth. You can find those with pattern matching and a good dictionary of names. Personal data, in Article 4(1) UK GDPR, is any information relating to an identified or identifiable natural person. Not a list of identifiers; any information. A sentence in an email saying that a colleague was late again and it was not the first time is personal data about that colleague. It contains nothing a pattern matcher would stop on.

There are two questions, and a detector only answers the first

The same reply drew a line I have kept coming back to. Structured data can be automated. Unstructured material, meaning email and chat, needs a person, because there are two judgements to make and not one. The first is whether something is personal data at all. The second is which person each piece of it relates to.

A subject access request lives almost entirely in the second question. Article 15(4) UK GDPR and DPA 2018 Schedule 2, Part 3 are the reason the redaction exists: the requester gets their own data, and the other people in the same file do not have theirs handed over with it. So it is not enough to know that a sentence is about a person. You have to know which person, and whether that person is the one who asked.

A tool that returns a list of names and numbers has answered the first question, partly, and has not started on the second. That is not a criticism of any particular product. It is what a detector is.

What gets missed is the part that is not built out of identifiers

I asked what the gap looked like in a real file. The answer was the unstructured material: the informal comment in an email, the opinion in a chat message, the personal data that has no identifier attached to it. They had looked at a number of tools and had not found one that did that part of the job.

Take a line like “my manager said she would not have hired him.” There are two people in it, both identifiable to anyone who knows the team, and it is personal data about both. There is no name, no number, nothing with a shape. A tool that works from a list of identifier types returns nothing, and it returns nothing confidently, which is worse than returning nothing at all.

This is not something a longer list fixes. You can add job titles, add pronouns, add every first name in the country, and the sentence above still has to be read by someone who knows who the manager is. That is the gap, and I do not think it closes from the tool side.

Then I built a small one, and the file pushed back even on the easy part

Having heard that, I built the narrow thing anyway, to find out what the easy part actually costs. It is a page where you open a PDF, type the requester’s name, and get a list of the other people it can find. It runs in the browser, it only looks for identifiers and capitalised names, and the result screen says what it did not check. It does not attempt the second question above. I would rather say that on the page than have someone discover it with a deadline running.

Even at that narrow job, the file had opinions. The first version joined lines together before looking for names, which seemed harmless until I ran it over a letter with a signature block. Marcus Odell on one line and Peter Vance on the next came out as a single four-word person, so two people in the file became one entry that matched nobody in it, and both dropped out of the list at once. I now treat a line break as a hard boundary. The price is that a name split across two lines, a Dr Helen at the end of one line and Whitmore at the start of the next, is missed. I chose the miss over the merge, because a miss is one person and a merge is two.

Tables were the next argument. The PDF library I use will tell me where a line ends but not where a column does, so in a test file the name in one cell picked up the first word of the cell beside it and became Peter Vance Neighbour. Surname-first entries, WHITMORE, Helen, have exactly the shape of an address, Bramhall Lane, Stockport, so I only accept the surname-first form when the surname is in capitals, and accept that this drops some real people.

Every one of those is a decision about where a person starts and ends, made by a regular expression, and in each case I had to pick which way to be wrong. One more decision came out of the same lesson. When the tool drops a match because it looks like the requester, it prints the names it dropped. Quietly removing a match is exactly how a third party disappears from a list, and if the match was wrong, the only person who can tell is the one reading the screen.

There is no accuracy figure on that page. I do not have one I could defend, and a number without the file it was measured on would be decoration.

Where that leaves the work

The practitioner who wrote to me still redacts by hand, and after building the small thing I understand that choice better than I did before I started. The identifiers are the part software can take. The comment in the email about the manager is the part it cannot, and it is the part that gets an organisation into trouble when it goes out unredacted.

If you do this work and your experience is different, or the same, I am still asking the one question I asked in August. It is on this page, and a reply of two lines is a real answer.