If a machine learns from watching a person perform a task, what exactly has that person contributed? The video is easy to identify. The knowledge inside it is much harder to define.
Take a mechanic diagnosing an engine fault. Much of what makes the demonstration useful may sit in a few seconds of judgement: what the mechanic notices first, where he places his hand, what sound he listens for, what he ignores and what he checks again. A camera records the sequence, somebody annotates it, and a model learns from it. At some point the original footage may no longer matter very much. But the useful thing did not begin with the camera, the annotation system or the model. It began with the mechanic and the years of experience that made those few seconds meaningful.
This creates a problem that the first generation of AI data did not have to confront in quite the same way. The arguments so far have largely been about things people created: books, music, photographs, films, journalism, software. The next argument may be about things people know how to do.
That is a less tidy category. Copyright can tell us who owns a photograph. Privacy law gives us rules around personal information. Consent can establish that somebody agreed to be recorded. None of these, on their own, quite answers what happens when the real value being captured is a person’s behaviour, skill or accumulated experience, and when that value is being captured specifically because a machine can learn from it.
This is where REN began.
REN is the name we have given to a research effort inside Clairva around rights, provenance and attribution in human-generated AI data. The name came after the questions. The questions are still more important.
What does meaningful consent look like when someone is not simply appearing in a dataset, but contributing behaviour that may help create machine capability? Can provenance remain attached to data as it moves from raw capture to annotation to training? Should some forms of attribution survive that journey? If a skilled worker contributes knowledge that materially improves a system, is a one-time payment always the right economic arrangement? And where, for that matter, does ordinary behaviour end and valuable know-how begin?
We do not have neat answers to these questions, and I am wary of anyone who claims they do. The technology is moving too quickly and the categories are still being formed. But the questions are becoming harder to avoid because the industry is moving from content that already exists to behaviour that has to be deliberately captured.
That shift matters enormously in markets like India, Indonesia, the Philippines, Vietnam and the Middle East. These are places full of environments, work patterns, languages and practical knowledge that are poorly represented in existing training sets. As AI companies look for more diverse and more realistic real-world data, these markets will become increasingly important sources of it.
There is a straightforward commercial model for that future. Find large pools of contributors, pay them to perform tasks, capture the behaviour, package the data and sell it upstream. In many cases that may be perfectly fair. In others, it may turn out to be an extremely efficient way of moving knowledge from one part of the world to another while leaving very little trace of where it came from.
That is the part we think deserves more thought.
At Clairva, we are interested in whether rights can travel with data more intelligently than they do today; whether permission can be expressed with more precision; whether provenance can remain visible through the life of a dataset; and whether contributors should, in some circumstances, remain connected to the value they helped create. Not as a vague promise of fairness, but as something that can eventually be represented in systems, contracts and commercial practice.
We are deliberately not publishing a grand framework around REN yet. The industry has enough frameworks, manifestos and diagrams already, many of them produced before the underlying problem was properly understood. We would rather spend time with the uncomfortable questions first.
What does seem increasingly clear is that rights will not remain a legal footnote to AI data. Buyers will want greater confidence in what they are acquiring. Enterprises will want to know whether important systems have been trained on data whose origins can withstand scrutiny. Contributors will begin asking more informed questions about how their behaviour is being used. And as human-generated data becomes more valuable, it will become harder to pretend that the humans behind it are incidental.
For the past decade, AI learned largely from what humanity chose to publish. The next decade may be about what humanity knows how to do.
Those are not the same thing, and the rights around them may not be either.
That is the problem REN is trying to understand.
P.S. REN is also a deliberate choice of name. In Chinese philosophy, rén (仁) is usually translated as humaneness: the quality of being human in our dealings with other people. Confucius placed it somewhere near the centre of a good society.
For a research effort asking what happens to people when their knowledge becomes machine intelligence, the name seemed difficult to improve upon.
