Generative AI, copyright and training data checklist
Hello, this is Legal Agent.
The idea that Japanese copyright law grants an unlimited carve-out for AI training is common, but that conclusion goes too far. While Japan's Copyright Act includes provisions on information analysis and machine learning that permit using works within defined boundaries, assuming that "anything goes" overlooks key distinctions. Model training, model provision, output generation, and output use each present distinct legal questions.
Distinct legal issues in training and output
Using text, images, or code to train, evaluate, or refine a model involves different legal considerations from using AI output in business operations or publishing it externally. On the training side, the Copyright Act allows using works where the purpose is not to enjoy, or enable others to enjoy, the thoughts or feelings expressed, within certain parameters. However, that exception depends on the purpose and manner of use, as well as whether the use unreasonably harms the copyright holder's interests, which requires case-by-case evaluation. On the output side, the core questions are whether the output reproduces protected creative expression with similarity and dependence on an existing work. Prompt materials and the production process matter to that assessment; resemblance in style alone does not automatically establish infringement. Between input and output sit the platform's terms of service: whether the vendor uses inputs to improve its models, whether training opt-outs exist, who holds rights in generated outputs, and what indemnities apply if an infringement claim arises. What is legally permissible, contractually allowed, explainable to a client, and commercially acceptable often lead to different answers.
High-value internal data and contractual limits
Feeding proprietary contracts, internal know-how, and company documentation into an AI system can drive productivity, but that utility does not make the underlying data free to use. Public availability on the web does not override website terms of service or technical access restrictions. Similarly, client files held internally are not automatically cleared for AI training; processing them through external models can violate contractual purpose limitations or confidentiality obligations. On the output side, prompts that cite a specific artist or existing copyrighted character sharpen questions of similarity in protected expression and dependence on an existing work.
Data classification and process documentation
A practical compliance approach starts with categorising data by origin: company-created assets, client-provided materials, public web data, and purchased datasets. Organizations should document how each dataset was obtained and define its permitted use, since training a foundation model, running internal retrieval-augmented generation, preparing client deliverables, or referencing text within a single prompt carry distinct risk profiles. Vendor storage and model-training terms should be checked before uploading client materials. Maintaining a clear record of how outputs were generated is essential, because without documentation, assessing an infringement claim or explaining the workflow becomes significantly harder.