GBAF Logo
Global Banking & Finance Awards® 2026 Nominations open, free to enter Nominate now →
Why Synthetic Data Could Become Critical to Financial A - Technology news and analysis from Global Banking & Finance Review
Technology

Why Synthetic Data Could Become Critical to Financial A

Published by Barnali Pal Sinha

Posted on August 24, 2026

12 min read
Add as preferred source on Google

How privacy, scarcity, rare events and model governance are turning artificially generated data into a strategic infrastructure question for financial institutions.

Financial artificial intelligence has a data problem that is easy to underestimate. Banks, insurers, asset managers and payments firms increasingly want AI systems that can detect fraud, support underwriting, improve customer service, model liquidity and automate internal processes. Yet the data needed to train and test those systems is often the very data that institutions are least able to move, share or expose.

The scale of adoption is already significant. A joint Bank of England and Financial Conduct Authority survey found that 75% of responding UK financial firms were already using AI in 2024, with another 10% planning to do so within three years; firms expected the median number of AI use cases to rise from nine to 21. That growth increases demand not only for models, but for usable, governable and representative data.

Synthetic data is emerging as one possible answer. It is artificially generated information designed to reproduce selected statistical characteristics of real datasets without simply copying individual records. Used carefully, it can give financial institutions more room to train, test and stress AI systems while reducing dependence on raw customer data. Used carelessly, however, it can reproduce bias, create false confidence and even leak information about the real data from which it was generated.

Why finance has a particularly difficult data constraint

AI development works best when teams can experiment repeatedly with large, diverse datasets. Financial institutions operate under the opposite conditions. Customer, transaction, credit, insurance and investment data are highly sensitive. Access is fragmented across legal entities, jurisdictions and legacy platforms. Internal approvals can be slow, and data-sharing with vendors or external researchers can create privacy, confidentiality, operational and regulatory concerns.

The UK Information Commissioner’s Office describes synthetic data as data produced by synthesis algorithms to replicate the patterns and statistical properties of real information. It specifically identifies synthetic data as a privacy-enhancing technique that can be useful for AI training when organisations cannot access large real-world datasets.

For financial institutions, that matters because access friction is not a peripheral inconvenience. It can determine which models get built at all. A fraud model may need examples of scams that are extremely uncommon relative to legitimate payments. A credit model may need sufficient observations across different economic conditions and customer segments. A cyber model may need realistic attack patterns that organisations understandably do not want to reproduce on production systems. Synthetic data can be used to construct controlled training and testing environments around those gaps.

The strategic value is not simply privacy

The first attraction of synthetic data is usually privacy. The more important long-term value may be controllability. Real-world financial datasets reflect what has already happened. Synthetic data can be designed around what might happen, including conditions that are rare, extreme or operationally unsafe to recreate.

That creates several important applications. A bank could generate additional examples of unusual fraud patterns to improve testing without waiting years for enough real events. A payments firm could simulate transaction flows under new product rules before launch. An insurer could test how an AI model behaves in underrepresented claims scenarios. A lender could examine model sensitivity across carefully constructed customer profiles, while a trading or risk function could explore tail scenarios that are sparsely represented in historical records.

The Bank of England has explicitly noted that generative AI can support its work through the production of synthetic data, and its AI Consortium later highlighted synthetic data as a potential response to gaps in training-data adequacy. The direction is important: synthetic data is moving from an experimental privacy technique toward a broader component of AI development and validation.

This could make synthetic data especially valuable in financial services because many of the most consequential events are precisely the ones that appear least often in normal datasets: account takeover, sophisticated fraud, liquidity shocks, operational outages, defaults in unusual macroeconomic conditions or combinations of risk factors that have not occurred frequently enough to train a robust model.

Synthetic data could change how financial AI is tested

The more transformative use may be model testing rather than model training. Financial institutions already use scenario analysis and stress testing because historical averages are not enough for risk management. Synthetic data extends the same logic into AI development. Instead of asking only whether a model performs well on yesterday’s observed cases, institutions can ask how it behaves across thousands of controlled combinations of inputs.

That can support red-team exercises, fairness testing, boundary testing and pre-deployment validation. Teams can deliberately generate edge cases to examine whether a model fails in predictable ways. They can vary income, geography, transaction behaviour, device characteristics or other variables to test whether an AI system reacts consistently and whether unintended proxies are shaping outcomes.

The appeal is particularly strong for generative and agentic systems. As AI moves from prediction toward systems that retrieve information, call tools and initiate actions, institutions need realistic test environments in which failures do not affect real customers or live balances. Synthetic customer profiles, synthetic account histories and simulated transaction environments can create safer spaces for experimentation before an agent receives access to production data or execution rights.

Privacy benefits are real, but they are not automatic

Synthetic does not mean anonymous by definition. The ICO warns that if synthetic data closely mirrors the real data used to create it, there may still be a risk that information about individuals can be inferred. This produces a fundamental trade-off: the greater the fidelity to the source data, the more useful the synthetic dataset may be, but the harder it can be to guarantee that sensitive information has not been retained indirectly.

NIST reaches a similar conclusion from a more technical perspective. Its 2025 guidance on differential privacy notes that synthetic-data techniques that do not satisfy differential privacy generally offer only informal privacy guarantees, and that attacks on non-differentially private synthetic data have in some cases revealed information from the original datasets.

The implication is that institutions should not treat synthetic data as a shortcut around privacy governance. They still need to understand how the source data was obtained, how the generator was trained, whether individual records can be inferred, what privacy tests were applied, and whether the output remains sufficiently useful after safeguards are introduced. In higher-risk settings, techniques such as differential privacy may need to be combined with synthetic generation rather than assumed to be built into it.

Bias can be preserved, amplified or invented

A second problem is representativeness. Synthetic data is generated from assumptions, models and source datasets. If the source data contains historical bias or underrepresents an important group, the synthetic generator may faithfully reproduce that weakness. If the generator is badly specified, it may introduce patterns that never existed at all.

This matters because the EU AI Act places explicit emphasis on data governance for high-risk AI systems. Article 10 requires training, validation and testing datasets to be relevant, sufficiently representative and, to the best extent possible, complete and free of errors for their intended purpose. It also requires providers to examine potential bias and identify relevant data gaps.

Synthetic data may help fill those gaps, but generating more records does not automatically create better representation. A million artificial records built from a biased sample remain a biased million records. Financial institutions therefore need to distinguish between volume and informational coverage. The question is not how much data a model has seen, but whether the data reflects the decisions, populations and conditions the model will encounter in practice.

A new layer of model risk is emerging

As synthetic datasets become embedded in AI pipelines, they create a new dependency: the data generator itself becomes a model whose assumptions can shape downstream models. This creates a form of model-on-model risk. A credit model trained on synthetic borrowers may inherit distortions from the generator. A fraud model tested only against synthetic attacks may appear robust while remaining weak against real adversaries. A liquidity model may perform well in simulated conditions that fail to reproduce actual market interactions.

That means synthetic-data governance should sit inside model-risk management rather than outside it. Institutions need lineage from the original data through the synthesis process and into the final model. They need independent validation of the generator, utility testing against real holdout data, privacy testing, drift monitoring and clear documentation of which decisions can rely on synthetic evidence and which still require real-world validation.

The Bank of England’s earlier AI Public-Private Forum reached a similar practical conclusion: synthetic data can be useful when real data is insufficient, but models trained on it still need to be tested against actual data. That is an important constraint because it prevents synthetic data from becoming a self-contained substitute for reality.

Why this could reshape competition in financial AI

If synthetic data becomes reliable enough, it could change who is able to build sophisticated financial AI. Today, large incumbents often have an advantage because they possess enormous proprietary datasets. Smaller banks, fintechs and specialist vendors may have fewer observations, particularly for rare events. Synthetic generation cannot erase that advantage, because useful synthetic data still depends on credible source information and domain expertise. But it can reduce the marginal cost of experimentation and allow smaller datasets to support a wider range of controlled tests.

It could also make collaboration easier. Institutions are often reluctant to share raw customer or transaction data even when a joint dataset could improve fraud detection, anti-money-laundering research or system resilience. High-quality synthetic datasets may create a middle ground: participants can share statistically useful representations while reducing exposure of real records. The FCA has already included the use of synthetic data in financial services among the topics considered by its Innovation Advisory Group, reflecting the policy interest in this area.

The strategic contest may therefore move beyond who owns the most data. Competitive advantage could increasingly depend on who can turn sensitive data into safe, high-fidelity development assets, validate them rigorously and integrate them into governed AI pipelines. That is an infrastructure capability rather than a single model feature.

What banks should avoid

The biggest mistake would be to present synthetic data as “safe data” and stop there. Privacy, accuracy and fairness are separate questions. A dataset can be privacy-preserving but statistically poor. It can be highly realistic but discriminatory. It can be representative of normal conditions but useless for tail risk. It can also be generated securely and still produce a downstream model that fails once exposed to changing customer behaviour or adversarial fraud.

Financial institutions should also avoid validating synthetic data solely with aggregate similarity metrics. Two datasets can look similar at portfolio level while differing in the relationships that matter for a specific decision. For credit, the important question may be whether default relationships are preserved across customer groups. For fraud, it may be whether synthetic transactions reproduce sequence and network behaviour. For insurance, it may be whether tail dependencies remain realistic. Validation must be tied to the intended use.

The likely model: synthetic and real data together

The strongest case for synthetic data is therefore not replacement, but combination. Real data provides grounding. Synthetic data provides flexibility, coverage and safer experimentation. Differential privacy and other privacy-enhancing technologies can add protection. Human domain expertise defines which scenarios matter. Model governance determines whether the resulting system can be trusted.

This mixed approach is likely to become more important as financial AI moves into higher-stakes decisions. Institutions may train preliminary models on synthetic data, test them across generated edge cases, validate them against tightly controlled real datasets, and continue monitoring them on live outcomes. That architecture creates a clearer separation between broad experimentation and access to sensitive production information.

It also aligns with the broader regulatory direction. The question regulators are increasingly asking is not simply whether an institution used AI, but whether it can explain the data, assumptions, controls, testing and accountability behind the system. Synthetic data can support that discipline when it is governed properly because it allows firms to design specific test conditions and reproduce them. But it becomes another source of opacity if generation methods and limitations are not documented.

Conclusion

Synthetic data is unlikely to remove the financial industry’s dependence on real data. It may, however, change how much real data institutions need to expose during development, how they test rare scenarios, how they collaborate on sensitive problems and how quickly they can move from AI prototypes to controlled deployment.

Its importance will grow if three conditions are met: privacy protections become measurable rather than assumed; synthetic datasets are validated against the decisions they are intended to support; and financial institutions treat the generators themselves as governed models. If those conditions hold, synthetic data could become a critical layer of financial AI infrastructure—not because artificial data is inherently better than real data, but because it can make real-world financial intelligence safer to develop, broader to test and easier to govern.

References

1. Bank of England & FCA — Artificial intelligence in UK financial services – 2024

2. Bank of England — Financial Stability in Focus: Artificial intelligence in the financial system (April 2025)

3. Bank of England — Artificial Intelligence Consortium minutes, October 2025

4. Bank of England & FCA — AI Public-Private Forum: Final Report

5. FCA — Innovation Advisory Group

6. ICO — Privacy-enhancing technologies guidance

7. ICO — Guidance on AI and data protection: security and data minimisation

8. NIST — SP 800-226: Guidelines for Evaluating Differential Privacy Guarantees (2025)

9. European Union — Artificial Intelligence Act, Article 10: Data and data governance

10. OECD — Framework for the Classification of AI Systems

Related Articles

More from Technology

Explore more articles in the Technology category