Copy and paste was the first automation

Nothing in a duplicated file records where it came from, so an error does not just spread. It spreads into a structure that has forgotten how it was assembled.

3 min read

The way you find it is that one entity’s figure looks wrong.

You open the file, you trace it one cell at a time, and eventually you land on a rate, or a reference, or a cut-off, that is not what it should be. Ten minutes of work. Then the cold part arrives, which is the realisation that this file did not originate anywhere. It was a copy. Somebody built one for the first entity and then made eleven more from it, in an afternoon, years ago, and every one of them carries whatever was in the original at the moment it was duplicated.

I have done that afternoon of tracing at a company running multi-entity reporting through acquisitions, and again in an accounting department with a portfolio of clients on the same template. The error itself was trivial both times. Finding the descendants was the job.

Copies have no record of their parents

This is the structural fact and it does not get fixed by being careful.

A real automation, whatever else is wrong with it, keeps one definition. There is a place the logic lives, and when the logic is wrong, there is a place to go and fix it. Duplication produces the opposite: twelve independent definitions that were identical on the day they were made and have been drifting apart ever since, because each one has been edited separately by different people for different reasons.

Nothing in the artefact records the lineage. The eleventh file does not know it came from the first. There is no list. You reconstruct the family by opening files and recognising the shape, and if somebody renamed a tab in one of them, you will miss it.

So the error does not just spread, it spreads into a structure that has deliberately forgotten how it was assembled.

The ratio nobody set

Here is the part I find genuinely interesting, and it is why I do not think this is a discipline problem.

In every process I have seen, the copy is the cheapest operation available. It is one keystroke and it always succeeds. Verification of a copy is one of the most expensive operations available, because it requires opening the descendant, understanding what it was for, and comparing it to something.

That ratio between how easy it is to propagate and how hard it is to check is not set by anybody’s carefulness or by any policy. It is set by the interface. Which means the rate at which mistakes spread through an organisation is partly a property of a piece of software from the 1980s, and any process design that does not account for it is designing against a force it has not noticed.

Careful people copy. Careful people copy more, actually, because copying an already-working file is the conscientious move. You are not inventing anything new. You are reusing something that has been checked.

Why forbidding it fails

I spent a while as the person who wanted it banned. Write it once, reference it everywhere, no duplicates.

I was right about the structure and completely wrong about people, and here is why. The reason copying wins is that it is the only operation that always works. A linked or referenced version introduces a failure mode the user cannot repair: the link breaks, and the cell shows an error, and the file is now useless to somebody who has a deadline and no way of fixing it.

A stale copy fails differently. It produces a number. The number is wrong, but the wrongness is silent and deferred and, crucially, somebody else’s problem later, whereas a broken reference is your problem now, in front of a person waiting for the file.

Given those two failure modes, a non-technical user under time pressure chooses the survivable one every single time, and they are not being lazy. They are correctly optimising for the failure they can live through.

When I inherit an estate of files now, I do not start by looking for errors. I start by looking for siblings, for the files that have the same shape as each other. The error is rarely the interesting part. The interesting part is the count of places it managed to reach.

All essays