Data ownership: who owns what your AI learns from
What happens to your data when it flows through AI tools and vendors, which rights you keep, which you may be giving away in the terms, and how to keep control.
The short answer
When your business uses AI, data flows in three directions, and the terms decide who may use each. The inputs you provide, documents, emails, records, prompts. The outputs the system produces, drafts, classifications, summaries, code. And whatever the vendor learns from either, if their terms allow them to use your data to improve their models. Ownership of your inputs stays with you in almost every arrangement; what varies is the licence you grant. Consumer terms frequently permit training on inputs, which means your customer correspondence or contracts may become part of a model others use. Business terms usually prohibit it and limit retention. The difference is a contract clause, and it matters more than any feature. Beyond the vendor relationship, the data you prepare for AI, cleaned records, labelled examples, prompts and rules, evaluation sets, logs, is an asset in its own right and should live in your own storage, in open formats, independent of any vendor. Ask every vendor four questions before any data flows: do you train on our data, how long do you keep it, where is it processed, and can we export everything we put in and everything we built.
The three kinds of data and what to check
| Data | What it is | What to check in the terms | What to keep yourself |
|---|---|---|---|
| Inputs | Everything you send: documents, records, prompts, conversations | Training excluded; retention limited; processing location; sub-processors; deletion on exit | Originals in your own systems |
| Outputs | What the system produces | Rights assigned to you; no reuse by the vendor; indemnities where offered | Stored in your systems, not only in the tool |
| Derived assets | Cleaned data, labels, prompts, rules, evaluation sets, fine-tuned models, logs | Exportable in open formats; not locked to the vendor’s platform; who owns a fine-tuned model | In your repository and storage, versioned |
Keeping control
- Classify your data before choosing tools: public, internal, confidential, personal, special category.
- Use business terms for anything beyond public data, with training excluded, retention limited and a processing agreement signed.
- Read the four answers from each vendor and record them in your vendor register.
- Keep derived assets in your own storage: prompts and rules in your repository, evaluation sets versioned, cleaned data in your database, logs exported.
- Prefer open formats for everything you build, so it can move.
- Store outputs in your systems of record, not only inside the tool.
- Review yearly as vendors change terms, and on every new tool before adoption.
Why derived assets matter most
Any business can rent the same model. What makes your AI useful is what you built around it: the cleaned data that describes your business, the labelled examples that define your categories, the prompts and rules that encode your way of working, the evaluation set that says what good looks like. Those took months of your team’s attention. If they live only inside a vendor’s platform, changing vendor means starting again. If they live in your repository and storage, the model behind them is a choice you can revisit.
What this means for you
Your data stays yours, but the terms decide what a vendor may do with it: insist on training exclusion, limited retention, known processing location and full export, under business terms, before anything confidential flows. Keep the assets you build around AI, cleaned data, examples, prompts, rules, evaluation sets and logs, in your own storage and formats, so the model remains a choice and the investment remains yours.
Frequently asked questions
If we use an AI tool, does the vendor own our data?
Ownership stays with you in almost every case; what varies is what the vendor may do with it. Consumer terms often grant the vendor the right to use your inputs to improve their models, which means your customer emails or contracts may become training material. Business terms usually exclude that and limit retention. Read the clause, or ask for it in writing; it is the single most important line in any AI agreement.
Who owns what the AI produces for us?
Under most business terms, the outputs are yours to use, and vendors assign whatever rights they can. Two cautions: outputs that closely reproduce third-party material may carry that material's rights, and in some jurisdictions purely machine-generated content has limited copyright protection. For business documents and code this rarely matters in practice; for creative work sold as your own, take advice.
What data should we be most careful to keep?
The data you invest in preparing for AI: cleaned and structured records, labelled examples, the prompts and rules that encode how your business works, the evaluation sets that define quality, and the logs of what the system did. Those are the assets that make your AI useful and yours, and they should live in your own storage, in open formats, independent of any single vendor.