← Back to AI Legal Lab
Insight
Generative AI for Legal WorkAI Service LegalLegal Outsourcing

Training Dataset License Agreements: Dividing Rights Between the Data Provider and the Developer

Hello, I'm Noriaki Asato, Representative Attorney at LegalAgent.

"If we provide our data to an AI developer, who owns the rights in the trained model built from it?" This is one of the questions that companies providing data want answered in contract negotiations. Developers, for their part, ask: "May we sell the finished model externally without the individual consent of the data provider?" and "If the data provided contains personal information or another company's copyrighted works, which party is responsible?" In license agreements for training datasets, two issues are often confused: the fact that the Copyright Act broadly permits data use in some situations, and the question of how far the parties can restrict conditions of use by contract. If negotiations proceed without sorting this out, the parties end up with a contract under which, after development, the permissible scope of use is unclear.

Limiting the Purpose of Use and the Relationship with Article 30-4 of the Copyright Act

The first thing the data provider should make clear is the scope of purposes for which the other party may use the data provided. The scope of the license differs greatly depending on whether it is described broadly as "use for AI training under this Agreement" or narrowed to "use for training a specific model developed by the counterparty under this Agreement." A broad formulation like the former leaves room for the developer to divert the same data to training models for other projects or other customers. If the agreement is limited like the latter and does not permit diversion to other projects, unauthorized diversion can be pursued as a breach of contract.

Whether redistribution of the data or provision to third parties is permitted should also be specified independently of the limitation on purpose. If the developer envisions a business of providing or selling the data it receives to outside parties, either as is or after processing, the contract should spell out concretely whether redistribution is permitted and, if so, the consideration and how restrictions extend to processed data. A statement that the data "may be used for training" alone does not make clear whether derivative secondary use after training has also been permitted. In practice, dividing purposes of use into categories such as "training," "evaluation" and "improvement" and specifying for each whether third-party provision is permitted allows the people operating under the contract to understand the scope of the license without hesitation.

In contract negotiations, one sometimes hears the view that "under the Copyright Act, data use for AI training is freely permitted, so it cannot be restricted in detail by contract." Article 30-4 of the Copyright Act provides that, where the purpose is not to enjoy, or have others enjoy, the thoughts or sentiments expressed in a work, the work may be used to the extent necessary, and use for information analysis is included. However, the proviso to the same Article excludes cases that would unreasonably prejudice the interests of the copyright holder in light of the nature and purpose of the work and the manner of use.

This provision limits copyright for certain uses; whether contractual restrictions on use are valid is a separate question. There are also forms of use to which the Article does not apply at all, such as where a purpose of enjoyment coexists or where the right holder's interests would be unreasonably prejudiced. Where the data provider and the developer agree on conditions of use such as "limited to a specific training purpose," "no redistribution to third parties" and "prompt deletion of the data after termination of the agreement," that agreement becomes a contractual obligation binding both parties, unless it violates public policy or other mandatory provisions of law. Contractual restrictions on use cannot, however, be imposed directly on third parties who are not parties to the contract. Even for an act whose use is not restricted under the Copyright Act, whether it constitutes a breach of contract is judged separately as a matter of the contract's effect. It is important for the data provider to negotiate with an understanding of the difference between the Copyright Act's limitation provisions and conditions of use set by contract.

Responsibility for Handling Personal Information and Third-Party Rights

Data held by companies mixes a wide variety of information, including customer information, business partner information, materials purchased from outside sources and internal documents prepared by employees. Where personal information or other companies' copyrighted works are included, the risk of legal violations and the allocation of responsibility need to be agreed before the data is provided.

Under the Act on the Protection of Personal Information (APPI), providing personal data to an outside third party in principle requires the consent of the data subject. The Personal Information Protection Commission's guidelines take the position that, where personal data is provided in connection with entrusting all or part of its handling within the scope necessary to achieve the purpose of use, the recipient does not constitute a third party. Whether consent is required therefore turns on whether the data provision agreement is substantively an "entrustment" or a "provision" to an independent developer. If it is characterized as entrustment, the provider is obliged to exercise necessary and appropriate supervision over the entrustee, so if the developer uses the data beyond the scope of the entrusted work for its own independent model development, the matter cannot be handled within the entrustment framework alone. When entrusting to an overseas developer, it is also essential to check the standards under Article 28 of the same Act for provision to a third party in a foreign country. Confirm the developer's security control measures and include in the contract means for monitoring how the data is handled.

Where the data includes personal data, confirm before provision whether appropriate processing is needed and the legal basis. When creating and providing anonymously processed information, in addition to processing in accordance with statutory standards, carry out the public announcements and clear statements to recipients required at the time of creation and provision. As for pseudonymously processed information, third-party provision is in principle not permitted even with the data subject's consent, unless a statutory exception or entrustment or the like applies. Keep in mind that merely deleting names does not mean the data can be freely provided externally.

Advance measures are also needed where the data mixes in materials whose copyrights are held by third parties or data held under licenses from other companies. The provider either specifies in an exhibit to the contract the scope of data for which it has legitimate title or licenses and represents and warrants that it does not infringe others' rights, or clearly indicates the scope of third-party materials and the basis for using them (existing licenses or the Copyright Act's limitation provisions) and agrees which party will handle any additional rights clearance. Even if the parties agree on an allocation of responsibility between themselves, legal liability toward the right holders themselves is determined separately. When including representations and warranties, the clauses need to be structured with an eye to the cap on damages if a breach is discovered and to how models into which the data has already been fed for training will be treated.

In addition, when providing shared data with limited access under Article 2(7) of the Unfair Competition Prevention Act (technical or business information provided to specific persons on a business basis and accumulated and managed in a substantial volume by electromagnetic means, excluding trade secrets), confirm whether the management measures required for legal protection, such as markings indicating an intention to provide the data only to specific parties, can be maintained after the contract is concluded.

Dividing Rights Between the Data Provided and the Trained Model

The question "If we provide our data, who owns the finished trained model?" is examined by separating, for each component, whether rights exist and how they are allocated by law, and the contractual conditions of use. Even if the contract says nothing, there are parts whose ownership is determined by law. This is because the program, parameters and architecture of a trained model are technical outputs distinct from the data provided itself, and whether one has rights to use the data does not in itself determine ownership of the model.

When drafting the contract, distinguish the data provided from the program, parameters, know-how and so on that make up the model, confirm whether legal rights arise in each, and then set out the rights relationships. As indicated in the METI Contract Guidelines on Utilization of AI and Data (Japanese), information itself is not an object of ownership, and for parameters as well, confirm whether the conditions for exclusive rights such as copyright to arise are met. For parts that cannot be settled by the allocation of rights alone, agree concretely on the scope of usage rights and the conditions for disclosure and re-provision.

The main matters to be specified in practice are as follows.

  • Ownership of parts in which copyrights or rights to obtain patents arise, and the scope of transfer or license to the other party
  • The scope within which the provider may use the model (limited to internal use, or permitting provision of services to outside parties or resale)
  • Whether the developer may reuse or repurpose similar models or common technology for other customers
  • How to handle cases where the model produces outputs that directly reproduce features of the data provided

If the provider enters negotiations with the stance that "since the model was trained on our data, all rights in the model belong to us," it will collide with the business model of a developer that wants to roll out a common framework widely. From the developer's side, a realistic option is to propose a structure that separates contributions to the common model from the unique differential portion arising from the specific data provided, and grants the provider exclusive usage rights only to the differential portion.

Designing Quality Assurance and Disclaimer Clauses

How far the data provider guarantees the accuracy and completeness of the data provided is another point on which opinions tend to diverge in negotiations. For the developer, training on inaccurate data or data with many gaps directly affects the performance of the finished model. For the provider, however, accepting an unlimited quality guarantee could expose it to liability even for model malfunctions caused by defects in the data.

As a negotiating proposal from the provider, one option is to limit what the provider warrants to "data acquired and held on the basis of legitimate title, to the best of the provider's knowledge," and to include a clause that does not warrant (disclaims) the completeness or currency of the data or its fitness for a particular training purpose. This is a structure that takes the provider's position into account, and the balance should be adjusted based on the actual development purpose, the amount of consideration for the data and the difficulty of verification.

In addition, setting a cap on damages where the developer suffers losses caused by the data, based on a benchmark such as the contract price, makes risks easier for both parties to predict. However, it is necessary to confirm whether the cap applies in cases of willful misconduct or gross negligence and the limits imposed by public policy, and the cap clause does not extend to cases where a third party brings a claim directly. If the developer specifies as a premise of the contract that it will carry out the inspection and cleansing of the data it receives under its own responsibility, the division of roles in verification remains clear even if the scope of the provider's warranty is limited. That said, it should also be shared between the parties that carrying out cleansing does not necessarily produce the expected model performance.

Handling Data at the End of Provision and the Practice of Contract Negotiation

How data already provided and trained models are to be treated when the contract ends upon expiry of its term or is terminated midway is a matter that easily becomes a major dispute later if it is not decided at the time the contract is concluded.

For the data provided itself, set out obligations to return or delete it after termination, and specify whether copies and backups made by the developer are included in what must be deleted and whether a certificate of completion of deletion must be submitted.

For trained models, on the other hand, the end of the original data provision relationship does not necessarily mean that the influence reflected in the model's parameters through training can be technically removed. Whether deleting the original data alone can be regarded as erasing the training results inside the model differs depending on the training method, so the technical feasibility of retraining or suspending use, the method of verification and the allocation of costs need to be agreed in advance.

The contract should separately specify whether continued use of existing models is permitted after provision ends, whether only the use of the data for new training (additional training or retraining) is prohibited, or whether external provision of the model is to be discontinued. In cases where the model's output directly reproduces the original personal information or another company's creative expression, legal liability is not avoided even if the contract permitted continued use. Whether the consideration for the data is an ongoing usage fee or an initial lump-sum payment also changes how reasonable it is to permit use of the model at the end of the contract.

In contracts for training datasets, after confirming the scope of application of Article 30-4 of the Copyright Act, the conditions of use from development through the post-termination period are worked out concretely. It is important to put the main issues, such as the purpose of use, whether redistribution is permitted, responsibility for handling rights-protected materials, ownership of rights in the model and how to handle termination, into writing at the negotiation stage before the data is handed over. Once the data has been handed over and training has begun, changing the conditions afterward becomes difficult.

The basic issues of generative AI and copyright are explained in Checkpoints on Generative AI, Copyright and Training Data, and the perspective of personal information protection in companies in AI Services and Personal Data Protection: What Companies Should Check First. For the design of contract clauses on personal information and data management, you can consult our Data Privacy and Protection team, and for rights clearance for models and deliverables, our Intellectual Property team. If you would like comprehensive support from contract negotiation to building internal systems, please also make use of our Legal Outsourcing service.

Related articles

Articles connected to this topic.

Insight / 2026.07.10 Always Check Article Numbers and Sources Yourself: A Pitfall of Generative AI Insight / 2026.10.03 Internal Use and Copyright: What to Check When Sharing Articles, Preparing Training Materials, and Using AI Summaries Insight / 2026.10.02 How to Draft a Data Provision Agreement: Scope of Use, AI Training, and Treatment on Termination

Services connected to this topic

Legal outsourcing Ongoing legal team support for contract review and legal operations. Generative AI legal consulting Terms, privacy, copyright, AI governance, and internal AI use rules.
View AI Legal Lab articles