When Labour Becomes Data

When Labour Becomes Data

We were asked on a podcast recently about the ethics of using blue-collar workers to train robotic models.

It is a good question because the obvious answers are not especially useful.

One answer is that this is exploitation: people are being paid to generate the data that may eventually automate their own jobs. Another is that this creates a new category of work and gives economic value to skills that have historically sat outside the digital economy.

Both can be true.

Robotics has a data problem that is quite different from the one we had with language models. You can scrape a very large part of the public internet to train a language model. You cannot scrape the internet to teach a robot how to unload a truck, handle a drill, fold a shirt, navigate a crowded kitchen or recover when a box slips out of its hands.

For that, you need people doing things. And many of those people will be drivers, warehouse workers, technicians, factory workers, construction workers, farm workers and others who have spent years developing physical skills that are almost invisible precisely because humans are so good at them.

This creates an odd economic inversion.

For the last thirty years, the technology industry placed enormous value on knowledge that could be expressed through a keyboard. Now it is discovering that knowing how to move a pallet, repair a machine or handle an awkward object may also be valuable data.

The difficult question is not whether people should help train machines. Humans have always built tools that subsequently changed the demand for human labour.

The difficult question is how this particular market gets constructed.

If someone is filmed without properly understanding why, paid very little, given no meaningful choice and has no idea where the data eventually ends up, then we have not invented a new ethical problem. We have simply recreated an old labour problem with GPUs attached.

But there is an equally strange argument on the other side: that blue-collar workers should somehow be protected from participating in this economy because the technology might ultimately affect their jobs.

That sounds benevolent, but it can become paternalistic very quickly.

A skilled worker may possess twenty years of tacit knowledge that suddenly has commercial value. Why should the correct response be to exclude that worker from the market?

At Clairva, we have to think about this quite practically because this is not an abstract debate for us. We work with people generating physical-world data.

The minimum standard seems fairly straightforward. People should know what they are doing and why. They should consent to it. They should be paid properly. Privacy cannot be an afterthought. The rights associated with the resulting data should be clear.

None of this makes the automation question disappear.

A worker could be compensated perfectly fairly for generating training data and still contribute to a technology that reduces demand for that category of work ten years later. There is no clever consent form that resolves that. But the alternative is not a world in which robotics stops. More likely, the data gets collected somewhere else. Or simulated. Or produced inside a small number of companies with enough capital to build their own capture infrastructure.

Which brings us to the part we find more interesting.

The physical AI economy may turn human experience itself into infrastructure. Not just photographs, text or video, but the accumulated practical knowledge of how people interact with the physical world.

Once that happens, the important question is not simply whether using that knowledge is ethical.

It is who gets paid for it.

The worst outcome would be quite familiar: we discover that millions of people were sitting on an enormously valuable resource, build an industry around extracting it, and somehow conclude that the people who created the resource were the least important part of the value chain.

Back to Journal