← Back to AI Legal Lab
Insight
Generative AI for Legal WorkAI Service Legal

Generative AI, copyright and training data checklist

Hello, this is Legal Agent.

"Japanese copyright law has a carve-out for AI training, so anything goes once it's for training." This comes up often in copyright questions about generative AI, and it is not quite wrong. Japan's Copyright Act does contain provisions relevant to information analysis and machine learning that can permit use of a work within limits. But the leap to "anything goes" is dangerous, because training, model provision, output generation and output use each raise different legal questions.

Training data and output are different problems

Using text, images or code to train, evaluate or refine a model is a different question from using the AI's output in a business or publishing it externally. On the training side, the Copyright Act's provision for uses that are not for the purpose of enjoying the thoughts or feelings expressed in a work can permit use within a certain scope, but the assessment turns on the purpose and manner of use and whether it unreasonably harms the copyright holder's interests, so it has to be judged case by case rather than assumed. On the output side, the question is whether the generated text or image resembles an existing copyrighted work and what material the user fed into the prompt. Between the two sits a service's own terms of use: whether the provider uses input data to improve its model, whether training use can be opted out of, who owns rights in the output, and what indemnity applies if a third party alleges infringement. "Legally permitted," "contractually permitted," "explainable to a client" and "an acceptable business risk" can each have a different answer.

The same data that creates value creates the conflict

Feeding a company's own contracts, know-how and documentation into an AI raises efficiency significantly, which is exactly why the underlying data, and any third-party copyrighted material or client confidential information mixed into it, cannot be assumed free to use just because it is technically available. Being publicly viewable online does not override a website's terms of use or technical restrictions, and a client's own documents sitting in the company's files are not automatically cleared for AI training; using them that way can breach the client contract's purpose limitation or confidentiality obligation. On the output side, prompts that name a specific artist or existing character raise the dependence and similarity questions more sharply.

Classify the data before deciding anything else

The practical starting point is classifying data by source: company-created, client-provided, public web data, purchased datasets. From there, a company should trace how each was obtained and fix the purpose of use, since training a foundation model, running internal retrieval-augmented search, summarising for one client and referencing something only within a single prompt carry different risk profiles. Vendor terms on storage and training reuse should be checked before uploading client materials, and a company should be able to explain, after the fact, how a given output was produced. Without that record, there is no way to answer an infringement question, let alone win the argument.

Keywords
Copyright & training data
Browse all keywords

Related articles

Articles connected to this topic.

Insight / 2026.07.10 Verifying Statutory Citations and Sources in Generative AI Outputs Insight / 2026.07.07 Explaining Legal Terms Plainly with Generative AI Insight / 2026.07.06 Deepfake Impersonation Advertising
View AI Legal Lab articles