You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

May 6, 2026
Thailand has introduced new requirements for online social media platforms to verify the identity of paying advertisers before publishing their advertisements. On May 5, 2026, the Electronic Transactions Commission published the Notification on Measures for Prevention of Technology Crime for Online Social Media (No. 2) in the Government Gazette. The notification, which aims to prevent technology crimes such as fraud and scams, takes effect 180 days after publication (i.e., on November 1, 2026). Mandatory Advertiser Identity Verification Online social media service providers must verify the identity of every advertiser before publishing an advertisement. Verification remains valid for up to one year from the most recent verification date. The notification requires social media providers to use either of the following methods when verifying advertisers: Document-based verification: Examine government-issued identity documents (e.g., national ID, passport, or juristic person registration certificate), cross-check the connection between the advertiser and the identity documents (e.g., facial comparison with photo ID), and ensure that the identity documents are verifiable against reliable sources. Digital identity verification: Use an identity verification system with a level of assurance no lower than that prescribed by the Electronic Transactions Commission. Advertiser Data Collection and Retention Service providers must collect and retain certain data—including name, identification number, and contact details—from the start of the advertising service and for a minimum of 90 days after the end of the advertising service relationship. The same requirements apply where there is a third-party payer, such as an ad agency. Implications for Affected Businesses The notification raises two key areas of concern for affected businesses: Social media platforms must implement know-your-advertiser (KYA) onboarding as described above, including document upload and identity matching processes. The 180-day implementation window requires immediate technical and operational planning. The collection and retention of national ID cards, passport copies, and other personal
April 30, 2026
Vietnam’s Decree No. 134/2026/ND‑CP, which took effect on 9 April 2026, plays an important role in detailing and implementing Vietnam’s Intellectual Property (IP) Law in the context of rapid digital transformation and the growing application of artificial intelligence (AI). The new decree provides comprehensive guidance on the application of copyright and related‑rights regulations, addressing key issues such as authorship, ownership, statutory exceptions and limitations, registration procedures, and enforcement mechanisms. Through these measures, Decree 134 seeks to achieve an appropriate balance between safeguarding the legitimate interests of rightsholders and fostering innovation, research, and technological advancement, thereby strengthening the state’s framework for the effective management, protection, and exploitation of intellectual property in the digital and AI‑driven environment. Some notable aspects of Decree 134 are discussed below. Copyright for AI-Created Works Decree 134 provides important guidance on the determination of copyright and related rights in works created with the assistance of AI. Article 5a reaffirms the principle that human creativity remains central to copyright protection, clarifying that copyright or related rights arise only where a human makes a substantial and decisive intellectual contribution, exercises effective control over the creative outcome, and assumes responsibility for the content and its legality. At the same time, the provision confirms that AI is regarded solely as a technological tool rather than a rights‑holding subject, thus ensuring consistency with the fundamental concepts of authorship and ownership under the IP Law. By introducing requirements on transparency, proof of human contribution, and compliance with AI‑specific labelling and technical marking obligations, Decree 134 establishes a clear and enforceable legal framework for the responsible use of AI in creative activities. Lawful Use of Copyrighted Texts and Data Article 37a of Decree 134 sets out the specific conditions under which copyrighted texts and data may be lawfully used for scientific research, experimentation,
April 23, 2026
Vietnam has progressively positioned blockchain as a strategic technology within its broader digital transformation agenda over the past decade. From early policy orientations to more recent legislative developments, the regulatory approach has gradually shifted from high-level recognition to more concrete legal integration. Against this backdrop, a new draft decree regulating activities relating to product and goods identification, authentication, and traceability (the “Draft Decree”) marks a notable turning point. Rather than merely referencing blockchain as a policy priority, the Draft Decree incorporates blockchain directly into a nationwide regulatory system, positioning it as part of the underlying infrastructure for data governance and public administration in relation to the management, verification, and traceability of product-related data. Evolution of Vietnam’s Blockchain Legal Framework: The Draft Decree in Context Vietnam’s blockchain legal framework has developed in several distinct phases. The first phase, beginning around 2019, was characterized by high-level policy recognition in several resolutions of the Party Central Committee. Particularly, blockchain was identified as part of the broader category of digital technologies critical to industrial modernization and participation in the Fourth Industrial Revolution. These resolutions did not regulate blockchain directly, but established its strategic importance at the national level. The second phase (2023 to 2025) saw the introduction of national strategies and technology policies that more explicitly recognized blockchain as a priority technology. Those policies collectively signaled a clear policy commitment to developing blockchain infrastructure and applications. However, these instruments remained largely at a policy-level and did not establish binding regulatory frameworks. The third phase (from 2025) involves the gradual integration of blockchain into sectoral legislation. Laws such as the Law on Digital Technology Industry (2025), the Law on Personal Data Protection (2025), and the Law on Science, Technology, and Innovation (2025) have introduced concepts such as digital assets, crypto assets, and even specific
April 21, 2026
Thailand’s Personal Data Protection Committee (PDPC) has launched a public consultation period on a draft notification setting out criteria for data subject access requests (DSARs). The draft notification addresses practical uncertainties in handling DSARs by introducing standardized procedural requirements for data controllers. The consultation period runs from April 16 to May 15, 2026. The notification will enter into force 30 days from the date of its publication in the Government Gazette. Key Features of the Draft Notification The draft notification covers the following key areas: Scope of information subject to access. Data controllers must enable data subjects to access at least the following upon request: (1) personal data collected directly from them; (2) personal data obtained from other sources; and (3) the source of personal data obtained from other sources without consent. Information required under section 23 of the PDPA and information that must be recorded pursuant to section 39 of the PDPA—such as the categories of personal data collected and purposes of processing—must also be made available. Submission channels and formal requirements. Data controllers must provide at least in-person and postal channels for DSARs, while electronic or other channels are optional. Requests may be made either directly by the data subject or through an authorized representative, and must be signed and include sufficient identifying information, a preferred response method, and DSAR details. Identity verification documents (and proof of authority if the request is through a representative) are required, and additional documentation may be requested for verification or communication purposes. Data controllers may use different verification methods for DSARs submitted via electronic or other channels, provided this does not create undue obstacles to the exercise of data subject rights. Verification and response timelines. Data controllers must complete preliminary verification within seven business days of receiving a request. If a