Software > Software-News > GDPR-compliant AI: Securely anonymise personal data for AI models and LLMs!

GDPR-compliant AI: Securely anonymise personal data for AI models and LLMs!


Data masking and synthetic data enable secure AI in line with the GDPR!

Data protection for AI starts with the data used!

Artificial intelligence and large language models (LLMs) require large and meaningful datasets. However, corporate data often contains personal information such as names, addresses, contact details, account numbers, health information or other sensitive content. If such data is used without being checked for AI training, Retrieval-Augmented Generation (RAG), testing or analysis, this can give rise to significant data protection and compliance risks.

Data masking provides the technical foundation for protecting sensitive information whilst ensuring the data remains usable for AI applications. Personal data is identified, classified and protected using appropriate methods before being transferred to an AI model, an analytics platform or an external service.

Why must personal data be protected before being used by AI?

AI systems process information from databases, documents, emails, PDF files, images, cloud storage and numerous other sources. Personal information may be contained unnoticed, particularly in unstructured data sets.

An effective data protection strategy should therefore begin before any training or processing by an AI system takes place. Sensitive data is first automatically identified and classified. It can then be masked, pseudonymised, encrypted, replaced or removed in accordance with defined data protection rules.

This reduces the risk of personal information being disclosed in training data, prompts, vector databases, model outputs or test environments.

What methods protect data for AI and LLMs?

Depending on the type of data, its intended use and the level of protection required, various methods may be considered:

However, pseudonymised data is generally still considered to be personal data. Whether processing meets the requirements of the GDPR therefore also depends on the specific use case, the technical and organisational measures in place, and the possibility of re-identification.

IRI DarkShield detects and masks sensitive AI data!

With IRI DarkShield, personal and other sensitive information can be searched for, classified and masked within structured, semi-structured and unstructured data sets. These include, amongst others:

Search rules, regular expressions, dictionaries and machine learning methods help to identify relevant PII and PHI data. The information found can then be de-identified using rule-based methods. This makes DarkShield particularly suitable for datasets that are to be prepared for generative AI, RAG systems, LLMs or other machine learning applications.

Synthetic data for AI training that complies with data protection regulations!

If real production data is not to be used, IRI RowGen offers an additional option: the generation of synthetic test and training data.

In doing so, data formats, value ranges, dependencies, distributions and referential relationships can be specifically replicated. Organisations receive realistic datasets for development, quality assurance, analysis and AI projects without directly providing sensitive original information.

Data masking and synthetic data generation can also be combined. This allows suitable parts of existing data to be retained whilst protecting them, and missing or particularly sensitive information to be synthetically supplemented.

Is data quality maintained for AI models?

Appropriate masking techniques protect personal characteristics, whilst relevant structures, formats and statistical properties can largely be preserved. Which method is optimal depends on the training objective and the required data relationships.

Therefore, following any masking or data synthesis, it should be checked whether the data remains representative and whether the desired AI model delivers reliable results. Data protection and data quality must be considered together.

Data-protection-compliant AI as a controlled process!

A secure AI strategy does not begin with the model itself, but rather with the selection and preparation of the data. A controlled process comprises:

  1. Identifying data sources
  2. Recognising personal and sensitive data
  3. Classifying information
  4. Applying appropriate protection methods
  5. Documenting and verifying results
  6. Transferring only authorised data to AI systems

With IRI DarkShield for searching for and masking sensitive information, and IRI RowGen for generating synthetic data, organisations can integrate data protection into their AI processes at an early stage. This improves control, traceability and auditability, and supports the responsible use of AI.

Efficiency meets experience: For more than four decades, our software solutions have been supporting companies in data management and data protection – technologically leading, reliable in productive use and applicable across all industries.

In use since 1978: Numerous well-known companies, service providers, financial institutions and state and federal authorities are among our long-standing customers.

Maximum compatibility: Our software supports both classic mainframe platforms (Fujitsu BS2000/OSD, IBM z/OS, z/VSE, z/Linux) and modern open system environments such as Linux, UNIX derivatives and Windows.

Source: JET-Software GmbH
Press release from 16 Sep. 2026 about the software DarkShield
DarkShield
Demo version
request URL
Information
directly to the product website
Online demonstration
directly to the product website
Video appointment
request
Success story
request URL
Software exposé
request URL
Prices
directly to the product website
Customers
request URL
E-Mail-Contact