You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

November 7, 2025
Thailand and the United States signed a memorandum of understanding (MOU) titled “Cooperation to Diversify Global Critical Minerals Supply Chains and Promote Investments” on October 26, 2025, signaling a new strategic alignment aimed at developing Thailand’s mineral sector, particularly in rare earth elements (REEs). The MOU has implications for investments in technology, manufacturing, and other related sectors. This update outlines the key provisions of the MOU and the potential opportunities and legal navigating points for businesses. Objectives The primary driver of this agreement is the US initiative to diversify global supply chains for critical minerals and reduce reliance on current market leaders, particularly China. For Thailand, it represents a major opportunity to attract high-tech investment and develop its downstream processing industries. The cooperation is set to focus on five main areas: Technical knowledge: Exchange of technical expertise and international best practices to strengthen Thailand’s mining and processing sector. Joint cooperation: Establishing workshops, seminars, and scientific collaboration to boost innovation. Regulatory practice: Promoting good governance and streamlining regulatory and licensing procedures. Information sharing: Sharing data on potential projects and global market prices. Full-value chain: The MOU covers the entire mineral lifecycle, from exploration and extraction to processing, refining, and recycling. “First Opportunity to Invest” Clause The most debated provision within the MOU states that “participants expect to have the first opportunity to invest . . . in critical minerals assets that may be sold in Thailand.” Business implications: This clause is widely interpreted as granting US companies a first look or preferential access to investment opportunities in Thailand’s critical minerals sector. This could be a significant advantage for US-based or affiliated companies in mining, technology, and energy seeking to secure a foothold in a developing REE supply chain. Thai government position: Thai officials, including the prime minister, have publicly clarified
October 31, 2025
On September 29, 2025, Thailand’s Office of the Personal Data Protection Committee (PDPC Office) published its Regulations on the Review and Certification of Binding Corporate Rules B.E. 2568 (2025) (the Regulations). The Regulations provide clarity on the PDPC Office’s approach to reviewing and certifying binding corporate rules (BCRs) under Section 29 of the Personal Data Protection Act B.E. 2562 (2019) (PDPA), and aim to facilitate international data transfers within a group of undertakings or enterprises (a “corporate group”). In conjunction with this development, the PDPC Office also approved BCRs for two companies operating in Thailand on September 30, 2025. This milestone represents the first concrete progress since the PDPC’s Notification on Criteria for the Protection of Personal Data Sent or Transferred to a Foreign Country pursuant to Section 29 of the PDPA B.E. 2566 (2023) came into effect in March 2024. Some key features of the Regulations are set out below. Categorization of BCRs BCRs are classified into two types: (1) BCRs for Controllers (BCR-C) and (2) BCRs for Processors (BCR-P). The category must be clearly specified when submitting the BCRs to the PDPC Office. Documentation Requirement The applicant must prepare and submit the application (a standard template may be provided by the PDPC Office in the future) along with supporting documents for review and certification in the Thai language. If the supporting documents are in a foreign language, a certified Thai translation should be provided. The translation must be notarized by a notary public or qualified person. Supporting documents may include, among others, a binding instrument such as an intra-group agreement, or a list of entities subject to the BCRs. Expedited Process Requirement Organizations with existing BCR approvals under the EU or UK GDPR, or from countries announced by the PDPC under Section 28, may apply through an
October 26, 2025
AI-generated songs are now making waves in Vietnam on platforms like TikTok, with tracks such as “Say mot doi vi em” quickly gaining popularity and sparking widespread attention. This phenomenon raises a host of legal and ethical questions: Who is the author of these songs? Can they be protected by copyright? Who is responsible if there is an infringement? These questions are becoming increasingly urgent as AI music becomes more mainstream in Vietnam. Copyright Protection for AI-Generated Music in Vietnam Under current Vietnamese law, copyright protection is reserved for works that bear the mark of human creativity. The 2022 amendments to Vietnam’s Intellectual Property Law reaffirm that only works created by humans are eligible for copyright. In practice, if a human meaningfully contributes to the creative process—by providing prompts, making selections, editing, or arranging—their contribution may be protected. However, if a song is generated entirely by AI without significant human input, it is unlikely to qualify for copyright protection. When an AI-generated song does not qualify for copyright protection, the question arises as to whether the person who writes the prompts, edits, or compiles the work can still be considered the owner of an asset under the Vietnamese Civil Code. According to Article 105 of the Civil Code 2015, assets include objects, money, valuable papers, and property rights. While AI-generated music that is not protected by copyright is not considered money or valuable papers, it may be regarded as an object (in the form of a digital file or recording) or as a property right if it can be possessed, used, transferred, or exploited for value. Use of AI-Generated Works Without Copyright Protection If a song is not protected by copyright, does that mean anyone can use it freely? Not necessarily. The absence of copyright does not mean the
October 3, 2025
On September 26, 2025, the Contract Committee under Thailand’s Consumer Protection Board issued a regulation that aims to standardize contracts and enhance consumer protection within the beauty and wellness industry. The Notification on Prescribing the Beauty Service Business as a Contract-Controlled Business B.E. 2568 (2025), which takes effect on January 24, 2026, requires business operators to use a prescribed standard contract in Thai and adhere to strict mandatory provisions and prohibitions. These regulations apply to operators across all in-person and online service channels, including via digital platforms. “Beauty services business” is defined as the provision of services under an agreement allowing consumers to receive a series of treatments, either over a set number of sessions or within a set period. This includes massage, spa, other methods for cleanliness, beauty, or care of facial or body skin, and weight control and body shaping—including services offered electronically. The law excludes surgery, liposuction, and medical treatments performed by licensed practitioners. The notification establishes the following key requirements: Mandatory contract and formatting. All contracts with consumers must use the standard contract form, in Thai, with clear, readable text (minimum font size of 2 millimeters, no more than 11 characters per inch), and include all essential terms from the annexed form. Contract execution. Contracts must be made in duplicate, with one copy given to the consumer at signing. For agreements concluded through electronic channels, the process must comply with the Electronic Transactions Act and use the same required terms. Digital platforms. Business operators who provide services facilitated through a digital platform as an intermediary are ultimately responsible for ensuring the consumer receives a compliant contract. Prohibited clauses. The law prohibits clauses that limit or exclude liability for damages to life, body, health, mind, or property resulting from breach of contract or a wrongful act;