Real IBANs in test data: the GDPR risk and what to use instead

A customer's IBAN in a staging database is personal data processed for a purpose nobody collected it for. This post reads the relevant GDPR articles with an engineer's eye and compares the practical alternatives: masking, pseudonymisation, synthetic data and generated IBANs.

By the Generate Random IBAN editorial team · · 8 min read

Most teams that handle payments have, at some point, restored a production dump into staging because it was the fastest way to get realistic data. If that dump holds customer or supplier IBANs next to names, the test data falls under the GDPR as squarely as the production data does, with weaker controls around it. This post goes through the articles that apply and then through what to put in the test database instead.

An IBAN that can be tied to a person is personal data, copying it is processing, and testing is not the purpose it was collected for. The GDPR asks you to process no more personal data than a purpose needs (Article 5(1)(c)), to build that limit into your systems by default (Article 25) and to secure what you do process (Article 32). A test environment full of real account numbers fails all three at once. Structurally valid generated IBANs take real account numbers out of the test environment altogether.

This is written from the engineering side. For how any of it applies to your organisation, talk to whoever owns data protection there.

Is an IBAN personal data?

Article 4(1) GDPR defines personal data as:

any information relating to an identified or identifiable natural person ('data subject'); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person

An IBAN sitting in a customers table next to first_name and email relates to an identified person. Even on its own, an IBAN is an identification number for an account that a bank can tie to its holder, so the person is identifiable indirectly. The one case where this does not apply is a business account held by a company: a legal person is not a data subject. In practice a production table mixes sole traders, private customers and companies, and you will not separate them before restoring the dump.

Recital 26 adds the test for when data stops being personal: it has to be "rendered anonymous in such a manner that the data subject is not or no longer identifiable", taking into account "all the means reasonably likely to be used" to identify the person. Replacing the IBAN column with a hash does not get there. A hash of a real IBAN is pseudonymised data, which the same recital says "should be considered to be information on an identifiable natural person", and the GDPR keeps applying to it.

Restoring a dump is processing, and testing is a new purpose

Article 4(2) defines processing as "any operation or set of operations which is performed on personal data", and its list includes "collection, recording, organisation, structuring, storage, adaptation or alteration, retrieval, consultation, use, disclosure by transmission". Taking a dump, copying it to another host, loading it, and letting QA read it are all on that list.

Article 5(1) then lists the principles every processing operation has to meet. Four of them are hard to satisfy with a production copy in staging:

Principle Article 5(1) text (shortened) Why staging struggles
Purpose limitation collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes The IBAN was collected to run a payment, not to test next quarter's release
Data minimisation adequate, relevant and limited to what is necessary in relation to the purposes A test of the payout flow needs IBANs with the right formats, not the real ones
Storage limitation kept in a form which permits identification of data subjects for no longer than is necessary Staging dumps are rarely deleted on a schedule, and old snapshots outlive the retention period of the source
Integrity and confidentiality processed in a manner that ensures appropriate security of the personal data Staging has broader access, fewer audits, and copies on laptops and in CI artefacts

Article 5(2) closes with accountability: the controller must "be able to demonstrate compliance", and that includes the staging copy.

What Articles 25 and 32 expect from the system itself

Two articles turn the principles into design requirements for the people building the software.

Article 25(1), data protection by design, asks the controller to "implement appropriate technical and organisational measures, such as pseudonymisation, which are designed to implement data-protection principles, such as data minimisation, in an effective manner". Article 25(2), data protection by default, is more concrete: "only personal data which are necessary for each specific purpose of the processing are processed", and that obligation "applies to the amount of personal data collected, the extent of their processing, the period of their storage and their accessibility".

Article 32(1) lists security measures "as appropriate", starting with "(a) the pseudonymisation and encryption of personal data" and ending with "(d) a process for regularly testing, assessing and evaluating the effectiveness of technical and organisational measures for ensuring the security of the processing". Article 32(2) says the risk assessment has to look in particular at "unauthorised disclosure of, or access to personal data transmitted, stored or otherwise processed".

Read together: the default state of a system that does not need real IBANs for its purpose is to not hold them. A test environment's purpose is to verify behaviour, and behaviour depends on the format of the IBAN, not on whose account it is.

Article 83 sets the ceilings for fines: up to 10 000 000 EUR or 2% of worldwide annual turnover for breaches of Articles 25 to 39 (Article 83(4)(a)), and up to 20 000 000 EUR or 4% for breaches of the Article 5 principles (Article 83(5)(a)). Supervisory authorities weigh many factors before setting an amount; the incident response and notifications after a staging leak are the more immediate cost.

Where test environments leak

Staging is riskier than production for ordinary reasons:

  • Access is wider: every developer, contractor and QA vendor can usually read the staging database, while production is locked down. A vendor with access is a processor under Article 28 and needs a contract that covers the data they can see.
  • Copies multiply: a dump on a laptop for a local reproduction, a fixture committed to the repository, a CI job that stores the database as an artefact, a screenshot of a customer record in a bug ticket, a log line that prints the request body with the IBAN in it.
  • Regions drift: staging often runs in whichever cloud region was cheapest, so a dump from EU production can end up outside the EU, which brings Chapter V (international transfers) into play.
  • Nothing expires: production has a retention policy; the staging snapshot from 2023 does not.

The alternatives compared

Approach Still personal data? Keeps IBANs valid? Effort
Production copy Yes Yes None, which is why it happens
Masking (overwrite digits with X or zeros) Depends on what is left; name and bank code may still identify No: a masked IBAN fails MOD-97, so every validation test fails Low
Pseudonymisation (hash or token, mapping kept) Yes (Article 4(5), Recital 26) No, unless the token is itself a valid IBAN Medium
Synthetic records, generated IBANs No: the data relates to nobody Yes, with the correct length, format and check digits for each country Medium, then zero

Masking is the common first attempt and the one that breaks most tests. An IBAN with XXXX in it does not pass your own validator, so you end up disabling validation in staging, which means staging no longer tests the thing you ship. Pseudonymisation keeps a link back to the person by definition: Article 4(5) requires the "additional information" needed to re-identify to be "kept separately", and as long as that mapping exists the data stays in scope.

Synthetic data is the approach that satisfies Recital 26. If the IBAN was never anyone's, there is nobody to identify. The requirement is that the synthetic IBANs behave like real ones in every check your code runs: correct country, correct length, correct BBAN layout, MOD-97 check digits that verify, and, for countries such as France, Italy or Belgium, national check digits that verify too. Test IBANs explains what "structurally valid" means and why a hand-written DE00 1234… is not.

Replacing real IBANs in a test database

A practical sequence that fits most schemas:

  1. Find every IBAN column. Search for columns named iban, account, bank_*, and also JSON blobs, event logs that store request payloads, and the audit table, which is the one usually forgotten.
  2. Keep the shape of the data. If production holds 60% German, 20% French and 20% Dutch IBANs, generate replacements in the same proportions, so that country-specific code paths still run.
  3. Generate per country, valid by construction. Use the generator for a handful, the export endpoint for up to 100 per country as CSV, JSON, SQL or plain text, the REST API for larger volumes, or the public MCP server if an agent is doing the seeding. Realistic mode (the default) uses real bank codes, so bank-code lookups and BIC derivation in your code still find a match.
  4. Replace consistently. If the same real IBAN appears in customers, mandates and payouts, it has to become the same generated IBAN in all three, or foreign-key-free joins break. Build the old-to-new map in memory during the anonymisation job and discard it at the end: a map you keep is a pseudonymisation table.
  5. Run the job on the production side. Anonymise before the data leaves the production boundary, in the same job that creates the dump. A dump that is "going to be anonymised later in staging" has already been transferred.
  6. Add a guard. Make the anonymisation job the only path that produces a staging dump, and alert on any restore that did not go through it. Comparing staging against production IBANs after the fact means moving production hashes around, which is more machinery than it is worth.

For a fresh fixture set rather than an anonymised dump, start here:

Each IBAN comes with its segments broken out, so you can seed bank_code and account_number columns from the same row.

What generated IBANs do not solve

A generated IBAN is not issued by any bank. It passes every structural check, which is what tests need, and it cannot be used for a payment, which is what makes it safe. Two limits:

  • It does not anonymise the rest of the row. Names, addresses, emails and transaction amounts need their own treatment; swapping the IBAN alone leaves an identifiable record with a fake account number.
  • It does not cover edge cases by itself. Keep a small, hand-checked regression set (letters in the BBAN, the 15-character Norwegian length, the 33-character Russian one), generated the same way, and review it like code.

Sources

  • Regulation (EU) 2016/679 (GDPR), EUR-Lex: Article 4(1), (2) and (5), Article 5, Article 25, Article 28, Article 32, Article 83(4) and (5), and Recital 26, quoted above from the Official Journal text.
  • SWIFT IBAN Registry, Release 102 (June 2026): the length and BBAN layout that a structurally valid test IBAN has to follow.
  • ISO 13616-1:2020 and ISO/IEC 7064: the IBAN structure and the MOD 97-10 check digits.

Need test IBANs?

Generate structurally valid IBANs for 91 countries, with correct check digits, ready to copy or download as CSV, JSON or SQL. They are not issued by any bank.

Open the IBAN generator