Data masking, synthetic test data and subsetting combine data protection with realistic testing!
Modern software must be developed quickly, tested reliably and continuously improved. To achieve this, development and quality assurance teams require realistic test data. However, the direct use of personal data or other sensitive production data poses significant data protection and security risks. Professional test data management establishes the necessary link between data quality, software quality and data protection.
Production-like test data replicates real data structures, formats and relationships as accurately as possible. This allows applications to be tested under realistic conditions, errors to be detected at an early stage and new functions to be reliably validated. At the same time, test and development environments must not allow any unnecessary conclusions to be drawn about customers, employees, patients or business partners.
How can production data be used securely for testing?
A key method is data masking. This involves specifically altering, replacing, encrypting or pseudonymising personal and confidential information. The original values are subsequently no longer recognisable, whilst the structure, format and business usability of the data are preserved.
Names, addresses, account details, telephone numbers or other sensitive information can thus be represented realistically without transferring the original values to less secure test systems. Masking helps organisations to adequately comply with data protection regulations such as the GDPR in software development, quality assurance and DevOps processes.
Data masking or synthetic test data?
Data masking and synthetic data generation take different approaches. In the case of masking, existing data sets serve as the basis. Sensitive content is altered, whilst the data structure and business context are retained.
Synthetic test data, on the other hand, is generated entirely from scratch. It is based on defined rules, formats, value ranges and distributions, and contains no real personal source data. In doing so, specific edge cases, boundary values, invalid inputs or rare scenarios can also be generated.
Which method is more suitable depends on the specific use case. Often, a combination of masked production data and synthetically generated data sets offers the greatest flexibility.
Reducing the size of relevant datasets through subsetting: For many tests, a complete copy of a production database is not required. Database subsetting reduces large datasets to a smaller, business-relevant subset. It is crucial that relationships between tables and records are preserved.
For example, a selected customer record must remain linked to the associated addresses, contracts, payments and transactions. Referentially correct subsets reduce storage requirements and provisioning times without compromising the validity of the tests.
Automated test data for CI/CD and DevOps: Manual processes quickly reach their limits when dealing with large volumes of data, complex system landscapes and short development cycles. Modern test data processes must therefore be automatable, repeatable and scalable.
Data profiling, classification, masking, transformation, subsetting and synthetic data generation can be integrated into standardised workflows. This enables the required test data to be provisioned as and when needed and updated regularly. Integration into CI/CD pipelines supports automated testing and accelerates the development of new software versions.
Integrated Test Data Management with IRI Voracity: IRI Voracity combines key functions for data management, data protection and test data provision within a unified platform. Sensitive data can be identified, classified, transformed and protected using reusable rules.
Among other things, the platform supports:
This enables organisations to centralise their test data processes and automate recurring tasks. The IRI Workbench provides an Eclipse-based development environment for this purpose, in which data rules can be created, managed and applied across various systems.
Improving data protection and software quality together!
Secure test data is an essential component of modern software development. It enables realistic testing without exposing sensitive production data to unnecessary risks. Data masking, synthetic data generation and database subsetting together create a robust foundation for data protection, test automation and reliable software quality.
IRI Voracity brings these processes together in an integrated platform. This results in repeatable and scalable test data processes that can be integrated into existing development environments and support organisations in the secure use of their data.
Online training course on test data management: The online training course ‘Test Data Management – Providing Test Data Securely and Purposefully’ demonstrates how secure and realistic test data is planned, generated and automatically provisioned in practice. Participants will learn how to use synthetic test data, data masking and database subsetting in a way that suits the specific test objective. Practical demonstrations using IRI Voracity, RowGen, FieldShield and DarkShield illustrate the technical implementation within development, testing, DevOps and CI/CD processes. The next session will take place on 12 and 13 January 2027.
Efficiency meets experience: For more than four decades, our software solutions have been supporting companies in data management and data protection – technologically leading, reliable in productive use and applicable across all industries.
In use since 1978: Numerous well-known companies, service providers, financial institutions and state and federal authorities are among our long-standing customers.
Maximum compatibility: Our software supports both classic mainframe platforms (Fujitsu BS2000/OSD, IBM z/OS, z/VSE, z/Linux) and modern open system environments such as Linux, UNIX derivatives and Windows.