You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

September 17, 2026
Thailand’s Office of the Consumer Protection Board (OCPB) has released for public comment a draft bill to amend the Consumer Protection Act B.E. 2522 (1979), the country’s foundational consumer protection legislation. The draft amendment aims to modernize the nearly five-decade-old framework to address the rapid growth of digital commerce, online advertising, influencer marketing, and new business models. The public consultation period is open until October 10, 2026. Expanded Definitions Covering Digital Commerce The draft significantly broadens several core definitions to capture modern commercial activities: “Consumer” is expanded to include natural persons and nonprofit juristic persons who purchase or receive services, including those solicited by businesses and end users who do not directly pay for the goods or services. “Business operator” now explicitly covers advertising business operators and hired advertising persons, such as influencers and content creators. “Advertising media” is expanded to include digital platforms, social media, and social media user accounts. “Label” now encompasses electronic labels—symbols, codes, or other electronic formats displaying product information. Influencer and Advertising Disclosure Requirements In addition to these expanded definitions, “hired advertising person for selling goods or services” is a new definition covering influencers, content creators, live streamers, affiliate marketers, and virtual online media operators who receive monetary compensation or other benefits for advertising goods or services. Hired advertising persons—including influencers and content creators—must disclose to consumers that content is advertising and reveal their relationship with the business owner. Disclosure is required when the business owner employs the advertiser, pays or provides other benefits for the advertisement, or provides free or discounted products or services. These requirements apply where consumers would not otherwise know that the business has a connection to the person presenting the content. Labeling Requirements for Importers The draft introduces a clearer labeling obligation for importers of label-controlled goods, who must
September 11, 2026
Thailand’s National Broadcasting and Telecommunications Commission (NBTC) has published a new five-year master plan that will bring significant regulatory changes to the broadcasting and digital media sectors, including formal licensing requirements for internet-based audiovisual services. The Master Plan for Broadcasting and Television, 3rd Edition (B.E. 2569–2573/2026–2030) was published in the Government Gazette on September 1, 2026, and will affect OTT platforms, internet-based audiovisual service providers, and traditional broadcasters. Licensing Reform The NBTC will develop new licensing frameworks ahead of existing digital television license expirations, which are slated to occur between 2028 and 2030. This creates both uncertainty and opportunity for incumbents and new market entrants. New licensing criteria will also be developed for audiovisual services delivered over the internet, meaning previously unregulated internet-based providers may face licensing, fee, and content obligations for the first time. The plan also calls for a new law to govern converged communications services. OTT Regulation and Content Oversight The plan explicitly acknowledges and aims to lessen the regulatory asymmetry between traditional broadcasters—which are subject to licensing, fees, and content regulation—and internet-based services that currently face fewer obligations. The NBTC intends to develop regulatory frameworks to bring internet-based audiovisual services, including OTT platforms, streaming services, and user-generated content platforms, under content, consumer protection, and licensing requirements. Consumer Protection and Digital Rights The NBTC will strengthen its oversight of broadcasting, television, and telecommunications operators to ensure compliance with consumer protection and personal data protection requirements. This includes updating relevant notifications and orders and more strictly enforcing rules against practices that unfairly exploit consumers. These measures may layer NBTC-specific requirements on top of Thailand’s existing Personal Data Protection Act obligations. Stricter enforcement against practices that exploit consumers is a priority, with particular scrutiny on advertising practices. The NBTC will modernize complaint resolution processes, meaning service providers should
September 7, 2026
On September 4, 2026, Thailand’s prime minister convened the first meeting of the Data Center Business Policy Committee. The committee endorsed a draft policy framework for the data center industry and tasked four subcommittees with developing the standards that would sit beneath it, shifting away from fragmented, agency-by-agency approvals toward a unified national strategy aiming to maximize economic value while managing environmental and infrastructure concerns. Proposed Scope and Pillars of the National Data Center Policy Framework The proposed framework would cover all types of data centers, including internal or captive facilities operated within a company or its affiliates, rather than only commercial third-party providers. If adopted in this form, companies running private data centers purely for internal purposes would also become subject to regulatory oversight. Minimum safety and operational standards would be established, with uniform enforcement across all categories. The committee endorsed a draft policy framework with four key pillars: Industrial classification: Data centers exceeding 2 MW would be classified as industrial operations, which may require factory licenses and environmental impact assessments under the Factory Act. Resource pricing: Utility rates would be structured to reflect both direct and indirect costs, supporting green energy and green data center standards. Centralized screening: A centralized review would evaluate project suitability and resource allocation. Operators may be required to submit proposals through periodic “pitching” rounds, where projects are competitively assessed on their potential economic and strategic benefits to Thailand. Digital ecosystem: The framework would prioritize data sovereignty, tax incentives, and conditions promoting domestic digital businesses, AI, and cloud infrastructure. Multidimensional Evaluation Criteria and Subcommittees Four subcommittees will be established to develop standards responsible for the following dimensions: Economic: Criteria for assessing the economic viability of data center projects, for use in prioritizing data centers based on infrastructure readiness, demand type (including AI factories),
September 4, 2026
Foreign business restrictions on telecommunications, treasury center businesses, and intragroup support services were eased when Thailand published the Ministerial Regulation Prescribing Service Businesses Not Requiring Permission for Foreign Business Operations (No. 5) B.E. 2569 (2026) in the Government Gazette on August 28, 2026. The ministerial regulation expands the categories of service businesses that foreign investors may operate without a foreign business license (FBL) under the Foreign Business Act B.E. 2542 (1999) (FBA). Of particular relevance to the telecommunications, fintech, and technology sectors, the ministerial regulation exempts: Type 1 telecommunications licensees, which do not have their own networks; Treasury center businesses operated in accordance with Thailand’s exchange control regulations; and Certain intragroup administrative, human resources, and information technology management services. Telecommunications Services Foreign-owned businesses providing telecommunications services under a type 1 telecommunications license may now operate without obtaining an FBL. This may streamline market entry for qualifying telecommunications and digital infrastructure businesses. The exemption applies only to the FBA licensing requirement. Operators must continue to comply with applicable requirements under the Telecommunications Business Act and the regulations of the National Broadcasting and Telecommunications Commission, and the change does not affect foreign ownership restrictions applicable to type 2 or type 3 telecommunications businesses. Treasury Center Businesses The ministerial regulation also exempts qualifying treasury center businesses from the FBL requirement. This may facilitate centralized treasury functions in Thailand, including liquidity management, foreign exchange management, and intragroup funding arrangements. Treasury center operations remain subject to applicable requirements of the Bank of Thailand and other competent authorities. Intragroup Administrative, HR, and IT Services Certain administrative, human resources, and information technology management services provided between affiliated entities are also exempt, provided the relevant entities satisfy prescribed ownership or management criteria. The exemption is available where the service provider and recipient are related through specified ownership