TL;DR: Germany’s financial regulator has begun monitoring how AI is used by the banks and insurers it oversees, with fining powers that came into force on 29 July. BaFin will start with transparency duties — including whether customers are told when they are dealing with a chatbot — before extending to high-risk systems such as creditworthiness assessment from December 2027.

The framing is what carries beyond Germany. Rather than treating AI as a separate innovation question, BaFin has folded it into the same ICT risk regime that the EU’s operational resilience rules, DORA, already apply to other critical technology. Marina Marusenko, who manages risk at ING, summarised the stance: “BaFin treated AI as an ICT risk issue, not an innovation topic.”

Monitoring, not supervision

BaFin has been careful about the limits of its remit. Jens Obermöller, its director-general for cyber risks and technology, said the watchdog will review samples of applications used widely across firms in especially relevant areas, rather than examining every model at every institution. What the legislation requires, he said, is monitoring rather than supervision. Regulatory sandboxes and real-world testing environments are planned so banks can trial new applications.

That narrower scope does not reduce the evidential burden much. Any firm whose application gets selected must be able to explain how the system was validated, how it detects potential discrimination, and how performance is tracked after go-live. BaFin president Mark Branson framed the objective as fair access to financial services without anyone being discriminated against by AI.

What the guidance actually demands

The earlier guidance BaFin built on is unusually concrete about method. Conventional software disciplines — testing units, testing integrations, reviewing source code — stay relevant to validating AI, with depth scaled to how critical the business function being supported is. Lifecycle management sits at the centre. “A one-time approval at go-live was not sufficient,” Marusenko wrote, citing the need to keep watching for drift, data quality problems and version changes.

Third-party generative models complicate this further. A supplier can alter a model after approval without the bank controlling the change, or necessarily hearing about it, which makes regression testing and version pinning essential. BaFin also pushed firms towards adversarial work — simulated data poisoning, evasion attacks, and penetration testing aimed at AI-specific weaknesses.

Looking forward

For UK financial services, this is the nearest available template for what the FCA and PRA are likely to be measured against. Two points travel particularly well: ordinary software can quietly turn into an AI system the moment somebody connects it to an external model, and a firm unable to inventory where AI actually runs across its estate cannot know what needs testing.