You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

November 23, 2023
On November 14, 2023, Thailand’s Personal Data Protection Committee (PDPC) published a draft notification on collection of personal data regarding criminal records. The draft notification aims to provide clarifications and prescribe further criteria for processing criminal record data under the Personal Data Protection Act (PDPA), which generally requires the processing of criminal records to be carried out under the control of the relevant official authority under the law or under a data protection measure implemented according to rules prescribed by the PDPC. After its eventual passage, the draft notification will have important implications for businesses’ recruitment and human resources activities in relation to individuals with criminal records. Key aspects of the draft notification include the following: “Personal data regarding a criminal record” and “criminal record data” denote personal data related to the investigations of criminal offenses, criminal prosecution, or criminal punishment that is official information or certified by the relevant supervisory authority, regardless of whether that action is connected to a final judgment. Under the draft notification, data controllers may process criminal record data for the purpose of a recruitment process, checking the qualifications of personnel, and considering the suitability of a person for a position if the processing activities are required by law or when a data controller obtains explicit consent from the data subject. Furthermore, the necessity of processing the criminal record data must be announced at the beginning of the recruitment process. Data controllers’ requests for explicit consent to collect a data subject’s criminal record data must also notify the data subject of the consequences of not providing consent or withdrawing consent. The draft notification sets the allowable retention period for criminal record data at a maximum of six months from the end of the processing activities specified above. After the retention period ends, the criminal
November 17, 2023
On October 3, 2023, Thailand’s Board of Investment (BOI) issued a new regulation clarifying the eligibility criteria for investment promotion under the BOI category “5.10 Development of software, platforms for digital services, or digital content.” To be eligible for BOI promotion under the digital activity category, projects must meet criteria related to local development, minimum investment amount, machinery and equipment, and development processes. These criteria for category 5.10 activities, along with the latest clarifications from the BOI, are detailed in the table below. Tax Incentives The BOI also clarified the method for calculating corporate income tax (CIT) exemptions. The CIT cap amount is calculated on an annual basis from the prescribed expenses incurred after applying for BOI promotion and occurring during the year for which the CIT exemption is claimed. The allowances include 100% of expenses for salaries for newly hired Thai IT personnel, technology-related training, and obtaining quality standards (such as ISO 29110). The revenue of projects that qualify for CIT exemption must be from sales or services directly related to software, platforms for digital services, or digital content developed as promoted by the BOI, including licensing fees, subscription fees, pay-per-use expenses, in-app purchase fees, usage fees, revenue sharing, advertising fees, and so on. For more details on BOI promotion for digital activities, or on any aspect of investment promotion in Thailand, please contact Athistha (Nop) Chitranukroh at [email protected] or +66 2056 5600, Napassorn Lertussavavivat at [email protected] or +66 2056 5662, or Thammapas Chanpanich at [email protected] or +66 2056 5561.
November 15, 2023
Four decisions from the Expert Committee under Thailand’s Personal Data Protection Act B.E. 2562 (2019) (PDPA) indicate that there will no longer be any relaxation of PDPA enforcement. The enforcement of Thailand’s seminal data protection law had been relaxed for more than a year when, on October 18, 2023, the Personal Data Protection Committee (PDPC) published the first decision made by the Expert Committee on the imposition of administrative measures against a company pursuant to authority granted to it under the Notification of the PDPC Re: Rules for the Consideration of the Imposition of Administrative Penalties by the Expert Committee B.E. 2565 (2022), which was one of the first subordinate regulations issued under the PDPA. Shortly thereafter, on October 19, October 25, and November 15, three additional Expert Committee decisions were published. These three decisions made by the Expert Committee are summarized below. October 18 Decision The complainant in this case lodged a complaint with the Expert Committee alleging that an insurance company contacted him to offer the company’s products without his consent. The complaint further claimed that when the complainant requested the company to disclose how his personal data had been acquired and asked the company to stop contacting him through any channel, the company did not take any action on the requests. The insurance company appeared to have obtained the personal data of the complainant from another source prior to the PDPA becoming fully effective (i.e., June 1, 2022). As the Expert Committee explained in its order, the company failed to comply with its obligations under the PDPA regarding the collection of personal data from another source, which requires consent as a legal basis; failed to comply with the grandfather provision by not publicizing opt-out procedures to enable the data subject to withdraw his consent easily; and
November 7, 2023
Under Thailand’s Royal Decree on Digital Platform Services, domestic and in-scope overseas digital platform operators that are required to notify the Electronic Transactions Development Agency (ETDA) of their operations must do so by November 18, 2023 (or by August 20, 2024, for small or low-impact platforms). This step is one of the essential requirements of the royal decree. Other key information on complying with the royal decree is as follows: The royal decree aims to regulate the operation of “digital platform services,” which refers to the provision of electronic intermediary services that create a connection between consumers, merchants or businesses, or other types of users in order to create an electronic transaction in whole or in part, regardless of whether a service fee is charged. The regulated digital platform services do not include digital platform services intended for offering the goods or services of a single digital platform service operator or an affiliated company that is an agent of the operator, irrespective of whether the goods or services are offered to third persons or to affiliated companies. The royal decree has extraterritorial effect, whereby overseas operators targeting the Thailand market are subject to the royal decree if their services are accessible in Thailand. Overseas operators are required to appoint a local coordinator in Thailand to coordinate with the ETDA. Compliance and Enforcement The ETDA released nine subordinate regulations under the royal decree; these took effect on August 21, 2023 (except for rules on platforms’ terms and conditions, which will take effect on January 3, 2024). Some important points on compliance and enforcement in the subordinate regulations, along with procedural guidance, are listed below. The ETDA has been emphasizing that both domestic and overseas digital platform operators need to notify the ETDA of their operations within the specified timeline (i.e.,