You are using an outdated browser and your browsing experience will not be optimal. Please update to the latest version of Microsoft Edge, Google Chrome or Mozilla Firefox. Install Microsoft Edge

June 15, 2026

Synthetic Data in AI Model Training: Legal Challenges and Intellectual Property Risks

Dow Jones Risk Journal

The surge in AI development has led to a desperate demand for large, high-quality training data. However, real-world data can be expensive to collect, difficult to access, and often subject to strict privacy and regulatory constraints.

Synthetic data, which consists of artificially generated records that replicate the statistical properties of real-world data without reproducing specific individuals’ information, provides an appealing solution by generating artificial datasets at scale without relying on identifiable personal information. It combines speed, cost efficiency, and regulatory compliance, making it a sensible alternative for organizations seeking to reduce risks while maintaining data utility. When properly anonymized, synthetic datasets may fall outside the scope of laws such as the EU’s General Data Protection Regulation (GDPR) or Thailand’s Personal Data Protection Act (PDPA), reducing compliance burdens while still supporting high-quality model training.

However, relying on synthetic data without rigorous legal due diligence could be a strategic mistake. It replaces one set of known risks (scraping, direct privacy liability) with a new set of complex liabilities. The narrative that synthetic data is a “silver bullet” for privacy and IP compliance is dangerous and could be misleading.

While synthetic data addresses data scarcity, it also introduces new legal uncertainties. Legal counsel should anticipate downstream risks arising from compromised data sources. Models trained on unlawfully obtained data may need to be decommissioned, even if their outputs appear lawful.

What is synthetic data?

Synthetic data refers to artificially generated information created using AI techniques such as deep learning and generative models. Instead of copying real records, it reproduces the statistical patterns and relationships found in the original dataset.

Synthetic data generally falls into three categories:

  • Fully synthetic data – Entirely new data points generated from learned patterns. The model studies the structure of the original data and produces records that resemble real-world behavior without replicating any specific individual.
  • Partially synthetic data – Real datasets in which sensitive fields (names, ID numbers, contact details) are replaced with artificial values while nonsensitive attributes remain intact.
  • Hybrid synthetic data – A combination of real and synthetic records, often used where some genuine information must be retained for accuracy or operational purposes.

The appeal of synthetic data lies in its protection of privacy and its operational efficiency. Properly generated synthetic datasets exclude real personal identifiers and can often be used for development, testing, analytics, and model training without exposing the information of actual individuals. In highly regulated sectors such as healthcare and financial services, synthetic data allows organizations to work with large, realistic datasets while minimizing the legal and operational constraints associated with using real customer or patient information.

Synthetic data is often used in the following sectors:

  • Healthcare: Synthetic patient records and images for safe model development.
  • Finance: Simulated transactions for fraud detection and risk modeling.
  • Mobility and autonomous vehicles: Generated driving scenarios to train for rare or dangerous events.

Each of these sectors leverages synthetic data to accelerate AI innovation. It provides realistic, varied training examples without leaking sensitive details.

Intellectual Property considerations

Despite the clear benefits of using synthetic data, its use for AI training may still give rise to intellectual property risks. The main concerns relate to possible infringement and whether synthetic data can be protected by copyright.

Infringement Risks Arising from the Source Data

Although synthetic data can reduce privacy exposure, it does not eliminate IP risks. Every synthetic dataset starts with the same foundational step: an AI model must first access, copy, and analyze the original “source data.” If that source data is protected by copyright or contractual terms, training on it without permission may constitute infringement.

Some stakeholders adopt a more permissive view of AI training, characterizing it as a form of computational analysis that extracts abstract statistical patterns rather than protected expressive content, and therefore does not constitute infringement. However, this view reflects a policy-based interpretation rather than settled law.

Courts and regulators have increasingly indicated that using copyrighted works for AI training may amount to prima facie infringement, unless a specific legal exception applies. Developers often invoke defenses such as U.S. fair-use principles, but these are narrow, fact-dependent, and unsettled in the context of AI.

Recent U.S. cases, such as Bartz v. Anthropic and Thomson Reuters v. ROSS, have so far found fair use only where the underlying materials were lawfully acquired and the secondary use was genuinely transformative. Conversely, they have rejected fair use where the model was trained on pirated or unauthorized copies. In practice, this means that organic (real) data collected without permission still presents a significant copyright risk for model developers.

Copyrightability of Synthetic Data: Lack of Human Authorship

Even when synthetic data does not copy any specific protected work, it raises a different issue: copyright protection generally requires human authorship. Many copyright systems require a work to result from a human’s creative expression. Authorities in the U.S., U.K. and Thailand take a similar approach: the U.S. Copyright Office has repeatedly rejected registrations for fully AI-generated works on the basis that they lack human authorship. As a result, a fully synthetic dataset produced without meaningful human creative input may not be protected by copyright at all, meaning third parties could potentially reuse it freely. Nevertheless, when meaningful human judgment is involved in designing, selecting, or arranging synthetic samples, copyright may protect that creative selection or arrangement even if the individual records themselves are not protected.

Copyrightability of Synthetic Data: Originality and the Creativity Threshold

Aside from the issue of human authorship, synthetic data often fails the originality requirement. Modern copyright law does not protect works based solely on labor or investment (“sweat of the brow doctrine”). Courts require at least a minimal degree of creativity.

In the U.S., Feist Publications v. Rural Telephone Service Co. confirmed that originality requires independent creation plus a “modicum of creativity.” EU courts apply a similar test, requiring that a work reflect the author’s “own intellectual creation.”

For synthetic data producers, this creativity threshold is difficult to meet. Many synthetic outputs simply replicate statistical patterns without meaningful human creative contribution, leaving them ineligible for copyright protection. Developers should not assume that large or expensive synthetic datasets are automatically protected. To secure such copyright protection, it is necessary to clearly document the human creative decisions involved in designing or curating the synthetic data.

Compliance considerations

Synthetic data should not be presumed to fall outside privacy regulation. Under laws such as the EU’s General Data Protection Regulation and Thailand’s Personal Data Protection Act, information still qualifies as personal data if it relates directly or indirectly to an identifiable individual. Synthetic data may still fall within this scope when it is:

  • Generated from real individuals’ records,
  • Capable of being linked to a person when combined with other available information, or
  • Structured in a way that allows specific traits or behaviors of an individual to be inferred.

In these situations, regulators are likely to treat the synthetic dataset as containing personal data, meaning full compliance obligations still apply.

Ensuring true anonymization is technically challenging. Studies have repeatedly shown that even heavily anonymized datasets can be re-identified with the original individuals with high accuracy using only a few demographic attributes such as age, gender, and ZIP code. The same risks apply to synthetic datasets that replicate the structure of real-world data, especially in domains involving rare characteristics.

Therefore, anonymization cannot be treated as a single, conclusive action. As computational methods advance, datasets considered anonymous today may become identifiable tomorrow. Synthetic data remains a valuable tool, but organizations should deploy it with a realistic understanding of these evolving risks.

 

This article was originally published by Dow Jones Risk Journal in April 2026.

RELATED INSIGHTS​ 

July 17, 2025
On July 9, 2025, Thailand issued a notification that introduces comprehensive operational requirements for digital platform service providers operating as goods marketplaces, effective December 31, 2025 (i.e., 180 days after its publication in the Government Gazette). The regulation’s official name is Notification of the Electronic Transactions Committee Re: Other Actions for Digital Platform Service Operators in the Category of Marketplace for Goods with Specific Characteristics under Section 18(2) of the Royal Decree on the Operation of Digital Platform Service Businesses that are Subject to Prior Notification B.E. 2565 (2022), B.E. 2568 (2025). Scope of Application The notification applies exclusively to goods marketplace operators formally designated by the Electronic Transactions Development Agency (ETDA), which on the same day designated 19 platforms that had previously notified the ETDA of their operations. The goods requiring enhanced oversight by these operators are limited to those regulated by the Thai Food and Drug Administration (FDA) and the Thai Industrial Standards Institute (TISI). Development from Earlier Draft An earlier draft of the notification had included a requirement for offshore platforms to establish a local entity, but this requirement was removed from the final notification. Key Obligations Despite the removal of the local entity requirement, the notification imposes a range of additional obligations on designated goods marketplace operators: Transparency. Operators must implement robust transparency measures, including clear, accessible, and understandable disclosures to users in Thai. These disclosures must cover all relevant terms and conditions, comprehensive product information, and complaint management procedures. Operators must also submit an annual compliance report to the ETDA within 60 days after the end of their accounting period, including statistics on regulated goods. Business user registration and identity verification. Before permitting the sale or advertisement of regulated goods, operators must collect and verify business user information, including contact details, identification documents, registration
July 15, 2025
Thailand has established new safe harbor rules that require social media platforms to remove specified content within 24 hours of government notification. On July 5, 2025, the Notification of the Electronic Transactions Commission on Measures to Prevent Technological Crimes for Social Media Service Providers was issued and took effect. This followed a hearing in May 2025 where only a select group of social media and online communication platform operators were invited to attend and comment on draft rules that could exempt social media platform operators from joint liability under the amended Emergency Decree on Measures for the Prevention and Suppression of Technological Crimes in cases involving victims of technological crimes. Safe Harbor Rules The notification stipulates procedures that must be followed in order to receive the protection of the safe harbor rules. Upon being notified by the Division of Prevention and Suppression of Cybercrime, Office of the Permanent Secretary of the Ministry of Digital Economy and Society (MDES) of the presence of false or misleading information that may lead to the commission of a technological crime, social media service providers must immediately take down the specified content, with a maximum allowable turnaround time of 24 hours from the time of receiving the notification. Social media service providers are required to promptly report the outcome of each takedown to the MDES Division of Prevention and Suppression. This shift in Thailand’s regulatory approach to social media content moderation establishes clear government oversight mechanisms while providing platforms with liability protection for compliance. As the new rules took immediate effect, social media platforms need to ensure that they have adequate systems and processes in place to comply with the requirements.
July 11, 2025
Vietnam’s recent embrace of “regulatory sandboxes” reflects a deliberate policy choice to balance the need for robust oversight with an equally pressing imperative to catalyze innovation. A sandbox is a controlled, time-bound framework in which businesses may pilot emerging technologies, products, or business models under relaxed or tailor-made regulatory requirements, thereby allowing regulators to observe risks in real time while innovators validate commercial viability without bearing the full weight of the traditional compliance regime. By issuing sandbox regulations, the government of Vietnam is signaling its commitment to accelerating digital transformation, attracting investment, and developing a knowledge-based economy, all while safeguarding financial stability, consumer protection, and national security. This strategy is embodied in a suite of instruments that together establish sector-specific sandboxes: Decree No. 94/2025/ND-CP on the Regulatory Sandbox in the Banking Sector (Fintech Sandbox Decree), effective July 1, 2025. Law on Digital Technology Industry (DTI Law), effective January 1, 2026, and Law on Science, Technology and Innovation (STI Law), effective October 1, 2025. Resolution No. 222/2025/QH15 on International Financial Centers (IFC Resolution), effective September 1, 2025. In addition, a draft resolution on the pilot implementation of the crypto-asset market (Draft Crypto Pilot Resolution) is expected to introduce a dedicated sandbox for crypto-asset service providers later this year, further underscoring Vietnam’s holistic, forward-looking approach to regulating emerging technologies. Below is a brief summary of all the regulatory sandboxes, who they are open for, and what businesses are attracted. Fintech Sandbox Decree Under the Fintech Sandbox Decree, besides credit institutions and foreign bank branches, fintech companies operating in Vietnam can apply for a Certificate of Sandbox Participation issued by the State Bank of Vietnam to operate any of the following services in Vietnam: Credit scoring: A solution applicable to information technology systems of credit institutions, branches of foreign banks, and fintech
July 11, 2025
On June 10, 2025, Thailand’s Supreme Administrative Court accepted for consideration a pivotal lawsuit concerning the regulatory obligations of administrative agencies over internet-based television broadcasting services, commonly referred to as over-the-top (OTT) services. This court’s decision in the case may set important precedents for how OTT platforms are regulated, especially regarding consumer protections and advertising practices. Background A user of an OTT television application initiated legal action against the National Broadcasting and Telecommunications Commission (NBTC) and related officials, alleging that the lack of clear regulatory criteria and oversight allowed OTT operators to broadcast general television content while compelling users to view advertisements before and during programming. The plaintiff argued this constituted consumer exploitation and claimed that the responsible authorities neglected or delayed their statutory duties under the Act on the Organization to Assign Radio Frequencies and Regulate Broadcasting, Television, and Telecommunications Services B.E. 2553 (2010). Initially, the Central Administrative Court declined to accept the lawsuit. However, on appeal, the Supreme Administrative Court determined that the claim fell within its jurisdiction, noting that OTT television services—defined under section 4 of the governing act—are subject to the same regulatory framework as traditional television services, regardless of the transmission method (frequency, cable, internet, or other system). Implications for OTT Services The key implications for OTT services concern the following issues: Regulatory oversight: The court recognized that OTT television services are explicitly covered under Thailand’s broadcast regulatory regime. Regulatory agencies may be compelled to establish clear operational rules and oversight mechanisms for OTT providers. Consumer protections: The plaintiff’s claim that excessive or unavoidable in-program advertising constitutes consumer exploitation was acknowledged as a matter of public interest. This may prompt stricter advertising standards for OTT platforms. Licensing requirements: The case raises the prospect that OTT operators may be required to obtain licenses from the