Many AI models wear a 'free to use' label that their training data never granted

Researchers traced over 232,000 chains from datasets to AI models to apps. In 62% of chains, at least one step had no real license, yet a clear 'free to use' label appeared downstream anyway.

Why it matters

When you build on an AI model, its license is supposed to tell you what you are legally allowed to do. This study shows those labels often hide the true rules set by whoever made the original data. That gap is not just theory: one company, Anthropic, settled a lawsuit for $1.5 billion in July 2026 over books used to train its AI. If you ship an app built on a model with a false license, the original owner could later demand you take it down or pay. The authors say the safest step today is to check where a model's data really came from before you trust its label.

Who's behind it: James Jewitt and colleagues, Queen's University, Canada (with Huawei's Centre for Software Excellence in Canada).

Summary by the Lemma AI · how we grade

Read the original paper (arxiv.org)