You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

April 10, 2026
Thailand has introduced new regulatory guidance requiring digital platform operators to adopt structured, transparent, and fair fee practices. On March 16, 2026, the Electronic Transactions Development Agency (ETDA) published Announcement No. DPS 2/2569, titled “Guidelines for Transparency and Fairness in Digital Platform Service Fee Determination,” issued under the Royal Decree on Digital Platform Service Business Operations B.E. 2565 (2022). The guidelines establish a framework governing how digital platform operators should set, disclose, and adjust fees charged to users and related service providers such as logistics and payment providers. Although framed as best-practice guidance rather than legally binding rules with explicit penalties, the guidelines carry regulatory weight under the royal decree and represent a significant step toward structured governance of digital platform fee practices in Thailand. The guidelines establish various transparency principles and divide fees into two distinct categories—compulsory and additional—with specific governance principles for each. Transparency Principles The guidelines recommend that digital platform operators adopt several transparency measures to ensure that users can fully understand the costs of using a platform. Fee catalog. All fees should be consolidated into a single, accessible location, which should include the fee name, definition, scope of covered services, calculation methodology, rate, billing period, and calculation examples. Minimum service disclosure. Operators should disclose the minimum service that users can expect, such as baseline visibility, product listing capabilities, access to transaction data, and back-end dashboard access. Price structure disclosure. Operators should disclose the categories of costs underlying their fees, such as system maintenance, cybersecurity, and operational costs. While exact cost figures need not be made public, operators should be able to provide numerical data to regulators upon request. Clear fee formulas. Fee calculations should be simple and easy to understand—for example, percentage of net sales, cost per order, or cost per product listing. Operators should
April 10, 2026
As digital commerce continues to reshape consumer behavior in Thailand, the Office of the Consumer Protection Board (OCPB) has been taking steps to review and update key regulations for online platforms. The OCPB has had a particular focus on addressing the risks posed by e-marketplace businesses—from misleading product information to fraudulent online transactions. Some of the regulator’s current legislative efforts related to Thailand’s labeling regulations as well as potential changes to the country’s law on direct sales and marketing. Proposed Changes to Consumer Protection Labeling Regulations On February 24, 2026, the OCPB convened a public hearing to review the Notification of the Committee on Labels re: Specification of Goods as Controlled Label Goods B.E. 2565 (2022) and its annex issued under the Consumer Protection Act. The closed-door session, which started the OPCD’s process of seeking feedback on the proposed changes, brought together representatives from government agencies, business operators, and consumer groups. The OCPB explained that its review of the labeling regulations aims to address regulatory gaps arising from evolving commercial practices, particularly the expansion of e-commerce and cross-border transactions. Authorities highlighted recurring issues involving product information that is unclear, incomplete, or potentially misleading in digital sales channels. The proposed revisions are intended to improve consumers’ access to accurate and complete product information, ensure that label disclosures remain relevant amid the growth of e-commerce, and strengthen protections against deceptive or misleading digital advertising. The review is being undertaken pursuant to the Consumer Protection Act B.E. 2522 (1979). As part of the initiative, the OCPB signaled a potential update to the categories of “controlled label products” as well as enhanced disclosure obligations for business operators, with the broader aim of promoting greater transparency, reinforcing operator accountability, and aligning Thailand’s labeling framework with current market conditions. The OCPB secretary general emphasized that
April 9, 2026
As part of its ongoing public consultation process for the development of new practical guidelines under the Personal Data Protection Act B.E. 2562 (2019) (PDPA), Thailand’s Personal Data Protection Committee (PDPC) held a two‑day public hearing on April 1–2, 2026. The hearing followed an online questionnaire and stakeholder engagement activities conducted in March 2026 and reflects the PDPC’s continued efforts to develop guidance that aligns international regulatory standards with Thai operational realities. The public hearing provided a forum for participants from both the public and private sectors to exchange views with the PDPC on the proposed guidance so that it responds to the needs of the business community while supporting effective and balanced enforcement of the PDPA. The PDPC emphasized that the consultation process is part of a wider policy objective to build trust in the convenient, secure, and internationally aligned exchange of data. Structure of the Consultation Process According to the PDPC, the initiative to develop the draft PDPA guidelines is being implemented through three core phases: Review of international best practices. The PDPC has conducted a comparative review of data protection guidance and regulatory approaches in jurisdictions with internationally recognized standards, including Singapore, the United Kingdom, the European Union (EU), and Japan. These materials are intended to serve as a reference point for developing practical recommendations across key subject areas under the PDPA. Identification of practical issues and challenges. To ensure that the guidelines respond to real‑world compliance challenges in Thailand, the PDPC has gathered views from a broad range of stakeholders across the public sector, the private sector, and the general public. This phase included focus group discussions and questionnaires aimed at identifying areas to provide organizations with greater clarity and consistency on regulatory expectations. Preparation of draft guidelines. Insights from the comparative study and stakeholder
April 3, 2026
On March 16, 2026, Vietnam’s Ministry of Public Security released a draft version of a new Decree on the Prevention and Combating of Cybercrime and High-Tech Crime to replace the currently effective Decree 25/2014/ND-CP. In the draft, the ministry has proposed a comprehensive regulatory framework aimed at addressing violations occurring within the cybersecurity domain, including measures related to intellectual property. Acts of Online IP Infringement Article 9 of the draft decree notably introduces specific provisions addressing online intellectual property infringement, with detailed lists of acts considered to constitute infringement in the online environment. Copyright and related rights infringement includes: Uploading or sharing works, performances, sound recordings, video recordings, broadcasts, computer programs, software, research, documents, theses, or other intellectual creations on digital platforms without the consent of the rights holder. Unauthorized livestreaming of copyrighted television programs, sporting events, or artistic performances. Uploading, sharing, storing, transmitting, or providing links to infringing works or digital content via websites, social networks, applications, or digital platforms. Providing or using software, tools, devices, or access codes to circumvent technological protection measures or evade lawful control mechanisms implemented by rights holders. Using artificial intelligence (AI) tools to replicate the ideas or structure of another person’s work without significant new creativity or without proper attribution, thereby causing damage to the original author. Industrial property infringement includes: Manufacturing, trading, advertising, or distributing counterfeit goods bearing counterfeit trademarks, geographical indications, or industrial designs, as well as goods infringing industrial property rights through online platforms. Unauthorized registration, appropriation, or use of domain names, account names, or digital identifiers that create confusion regarding the rights holder or the origin of goods or services. Producing, using, or offering for sale products containing all or part of a patented invention via online platforms. Advertising or introducing products with technical features or characteristics identical