Get all your news in one place.
100's of premium titles.
One app.
Start reading
inkl
inkl
Zoe Nauman

From MIT Lab to Regulated Reality: How Alexandre Berkovic Turned Research-Grade AI Into Banking Infrastructure

When banks fail an AI pilot, it is usually because the model's answer can't be reconstructed, challenged, or defended within an existing control framework.

U.S. financial institutions currently lose hundreds of billions of dollars annually to fraud and related activity. Nasdaq Verafin's 2026 Global Financial Crime Report estimated U.S. fraud losses at about $196 billion in 2025, including roughly $179 billion in bank fraud. That same year, the FBI's Internet Crime Complaint Center logged $20.9 billion in reported U.S. internet-crime losses, up from $16.6 billion in 2024.

That's where Alexandre Berkovic is making an impact. He specializes in machine learning and AI for financial crime and compliance, and works specifically with multi-agent systems that make regulated decisions auditable.

This includes onboarding, sanctions screening, and transaction alerting. Berkovic takes research-grade machine learning and makes it behave like an infrastructure that a supervised institution can actually run.

He says: "Improving a model's score is the easy part. The harder part is building methods with defined workflows and decision provenance. You also need to consider error handling, continuous evaluation, and of course, human oversight."

From Research-Grade Models to Regulated Systems

Berkovic's interest in data and systems began early. He reveals that from childhood he was drawn to coding and systems thinking: "I believed in the idea that data could be turned into reliable decisions."

He began to pursue his ideals by studying at Imperial College London, where he completed an MEng in Design Engineering with First Class Honors. This resulted in his thesis using generative adversarial networks to explore future art movements. Berkovic subsequently built audio-generation foundation models and founded Adorno AI.

Berkovic offers a direct account of that progression: "The technical demands of that project pushed me toward deeper formal training in machine learning and optimization."

To realize his goals, Berkovic then attended the Massachusetts Institute of Technology (MIT) as a Master of Business Analytics candidate in 2022–23, with work tied to MIT Sloan and the Operations Research Center. One project involved a latent text-to-image diffusion model designed to generate spectrograms from text prompts and then convert those spectrograms into sound.

Alexandre Jacquillat, who is the Associate Professor of Operations Research and Statistics at MIT Sloan and Berkovic's former instructor in 15.093, says: "15.093 is taken by some of the strongest quantitative students at MIT, and it is not an easy course in which to do well. Against that reference class, Alex stood in the top group of students I have taught."

That MIT stretch also opened commercial collaborations outside the classroom. Douglas Price, CEO of Pro Sound Effects, has known Berkovic since 2023 through the MIT community in Boston and its CAPD program, after an introduction from Leo Boussioux. Price first heard about Berkovic's data-science work through that network, then connected him with Pro Sound Effects product teams and gave him regular access to the company's audio data.

Pro Sound Effects holds millions of sound recordings and fully developed sound effects for media creation. The business problem was not a shortage of files. It was a shortage of data intelligence: how to make that library more useful for data science, and how clients actually wanted to receive and work with the material. Price says Berkovic quickly understood both the technical structure of the audio data and the commercial requirements, then designed a series of experiments that made millions of sound recordings more useful for the business and its clients.

Price describes him as a big-picture thinker who can break a problem down from both a technical and a business angle. In Price's account, Berkovic brought his own expertise together with outside resources, optimized the data for real-world application, and bridged the gap between what the company needed and what customers expected. "He communicates very well," Price adds, noting that Berkovic is "professional and personable at the same time," and that the working relationship grew into a friendship spanning Boston and Paris.

Berkovic began his early work in multi-agent systems and data science at Amazon, Schneider Electric, and Bear Robotics. He believes those roles reinforced a consistent pattern of moving between analytical methods, engineering environments, and practical deployment rather than treating research as separate from implementation.

That path led him to found Sphinx, an AI firm building multi-agent systems for financial crime and compliance work.

The company works with banks and fintechs, helping with onboarding, sanctions screening, and transaction alerting. The secret sauce? Its proprietary Interpretable Agentic Framework, which was devised by Berkovic. Clients include Alviere, Synctera, and Equals Money.

Berkovic says: "I think my work trajectory, which included rigorous data science training combined with repeated real-world deployment, is what led to me founding Sphinx.

"Our advanced machine learning systems, which we operate for financial institutions on regulated U.S. entities, have to be watertight when it comes to accuracy and auditability."

The Interpretable Agentic Framework

Sphinx's proprietary Interpretable Agentic Framework addresses a basic constraint of regulated AI deployment. Financial institutions and their regulators cannot rely comfortably on unexplained outputs when the underlying decision requires accountability.

Berkovic says: "Financial institutions and their regulators cannot accept black-box outputs."

That constraint shaped how he converted research-grade methods into production infrastructure. Rather than treating interpretability as an explanation bolted onto a single model after the fact, Sphinx's framework was built so regulated decisions leave a trail that analysts, reviewers, and auditors can reconstruct.

Chrisjan Wüst, Sphinx's co-founder and CTO, says he saw Berkovic's data-science instincts up close at Adorno, when Berkovic architected the core video-to-audio model and kept asking whether the research actually worked in practice.

Wüst adds: "Getting an LLM to do something impressive once isn't actually that hard anymore. Getting it to do the right thing consistently enough that a bank can rely on it is much harder."

A lot of Berkovic's work, he adds, has gone into that less flashy side of AI: evaluation, failure cases, messy real-world data, and making systems predictable enough for regulated environments.

Wüst says of his colleague's work with Sphinx: "Alex made the product architecture that made the company viable: agents that execute compliance procedures end to end, rather than tools that merely assist a reviewer.

He adds he has also seen Berkovic build and run mixed teams which include machine learning, NLP, optimization, and financial-crime specialists in the same shop: "In regulated AI, that mix is not optional. Model work, engineering, domain knowledge, and compliance have to move together. And Alex is incredibly talented at making that happen."

So why is Sphinx, and the system Berkovic has devised, making such an impact? Sphinx excels where a conventional black-box system can't. It can return an answer which doesn't leave analysts, reviewers, and auditors guessing how it got there.

The architecture has been undeniably successful and is operating in production inside publicly regulated U.S. financial institutions supervised by the OCC, the FDIC, and the Federal Reserve.

Berkovic says: "This is its deployment environment, and in turn that regulated deployment raises the standard for how systems are monitored and how their decisions are documented."

Deployment Method and the Hard Data Science Problems

Sphinx's deployment method is tied to the realities of existing compliance operations. Rather than starting with a traditional systems-integration project, clients provide credentials that let AI agents interact with the tools the compliance team already uses.

Jorgen Osio Norgaard, Chief Compliance and Risk Officer at Alviere and a Sphinx customer, says: "We had the first Sphinx agent running inside our environment in weeks, not the multi-quarter integration I had come to expect at Citi, JPMorgan, and WorldRemit."

That approach still leaves a difficult data science problem: understanding how experienced analysts actually make decisions.

Berkovic's account of the method is specific: "We systematically observe how experienced human analysts navigate cases, which systems they query, which decision criteria they apply, how they weigh conflicting evidence, and then encode those workflows into generative models and multi-agent systems."

He adds that the engineering issue is broad, which means they often have to deal with multiple challenges: "Achieving reliable, low-latency, and auditable behavior across heterogeneous U.S. and international institutions requires solving problems such as agent tooling, state management, error recovery, and continuous evaluation."

Another challenge is preserving the expertise experienced human analysts apply when cases contain incomplete or conflicting information. He states that problem without qualification: "Capturing tacit expert judgment in a form that remains accurate, stable, and fully auditable is a non-trivial data science problem."

The intended role of the agents is not to remove human judgment from the process. The stated approach is that repetitive analytical work can be absorbed at machine speed while human judgment remains focused on edge cases and situations requiring additional discretion.

Measured Results That Matter

Sphinx processes millions of alerts and handles hundreds of thousands of cases across multiple countries. It tracks several metrics, including false-positive reduction, straight-through processing, and onboarding speed, as well as screening precision, analyst time, and case-resolution time.

The following numbers cited by Berkovic come from Sphinx, not from independent public audits, and Berkovic is quite rightly proud of the results: "With Equals Money we achieved a 94 percent reduction in false positives and raised straight-through processing of onboarding cases from 42 percent to 87 percent, with zero engineering work required on the client side."

He adds: "Comparable quantitative improvements have been delivered to other U.S. institutions including Alviere and Synctera. We report sanctions-screening precision of 99.9 percent across our portfolio. Average onboarding time at Alviere is 2.4 minutes, and one other client is saving more than 200 analyst hours per month."

He goes on to frame the discipline behind those numbers: "These results are not marketing claims. We are measuring outcomes of careful feature engineering, model calibration, multi-agent orchestration, and continuous monitoring, Core data science disciplines applied to regulated financial crime workflows."

Across Equals Money, Alviere, and Synctera, the same framework supports onboarding and transaction alerting, and the company reports comparable deployments inside supervised U.S. institutions. Faster processing without an accountable decision trail would address only part of the deployment problem.

Others are quick to laud Berkovic's skills and say the numbers prove his system is making a big difference. Jenny Gai (Product Marketing, TRM Labs) was introduced to Alex through TRM's co-founders: "He brings clear technical authority to partner work. He has created deep out-of-the-box integrations and configurability that made it possible to align AI responses with an institution's risk thresholds."

Cyril Sayada, Partner at Sia, has known Berkovic for about three years as both a commercial partner and an independent validator. They have taken Sphinx and Sia into joint client conversations on AI in compliance.

Sphinx later hired Sia to stress-test its agents with a report built for banks, auditors, and regulators rather than a vanity write-up: "The moment that best captured his approach was when he commissioned an independent validation of Sphinx's agents," says Sayada.

"Most vendors ask for a document that makes them look good. Alex requested the opposite: a report that a bank's vendor risk team, an auditor, or a regulator could rigorously examine. He pressed us on sampling methodology and on whether the evidence would withstand scrutiny. That told me more about how he builds than any product demonstration could."

Berkovic's membership in the Association of Certified Anti-Money Laundering Specialists, known as ACAMS, places him inside the same anti-money-laundering discipline that Sphinx systems are built to support. The membership started in 2025 and reflects an ongoing connection to the regulatory and compliance side of financial-crime work, not only the technical side.

In 2024, he received two Élysée Palace invitations: a May gathering of French AI talent, then a September session preparing the AI Action Summit. Beyond those invitations, he has spoken at the Fintech Summit in San Francisco and Imagination In Action, plus a TEDx talk on managing AI agents and human judgment. In 2026, he served as a judge for both the MCP Hackathon and RoboHacks at Y Combinator.

He says: "If you build systems for AML teams, you cannot stay on the engineering side of the wall. ACAMS keeps me inside the same discipline the product has to serve. The Élysée sessions, the fintech stages, and the judging work do something similar. They force the work into rooms where regulators, operators, and builders argue about what actually holds up."

Values and Methodological Discipline

Berkovic identifies the values that guide the team and its data science work. He frames the standard in operational terms: "Quality above everything. In financial crime infrastructure, a 95 percent solution is unacceptable; residual error can translate into regulatory findings or material losses."

He then lays out the cultural expectations that follow from that standard: "Customer obsession without loss of technical focus. We want to maintain a clear methodological core while continuously incorporating the operational pain points of the institutions we serve.

"In addition, every data scientist and engineer is expected to drive projects from problem formulation through production monitoring.

We are also hardcore about our execution. We have sustained intensity when the work demands it, because reliability and throughput cannot be compromised. And we always place the customer first, team second, individual last."

This methodological discipline connects Berkovic's technical system to the working culture and to the practical requirements of regulated financial institutions.

The Practical Difference

Research-grade generative and multi-agent machine learning has not often moved from demonstration settings into production workflows inside regulated institutions. The challenge is not solved by producing a more capable model alone. It requires an architecture and deployment method that accounts for interpretability, operational reliability, error recovery, and human review.

Berkovic frames the scale of that deployment: "We are among the relatively few organizations that have taken generative and multi-agent machine learning systems in the financial-crime domain from laboratory demonstration to full production inside regulated banks, and we have done so with a native, low-integration architecture."

He adds: "We have processed millions of alerts and hundreds of thousands of cases across dozens of countries. The systems have helped clients avoid tens of millions of dollars in potential headcount costs and fraud losses."

In practice, the approach is simple. Clients hand over credentials. The agents work inside the tools analysts already use. Sphinx studies how those analysts actually move through a case, then turns the useful steps into agent workflows without wiping out the decision trail.

Berkovic says: "The system reduces hallucination and error rates by approximately 94 percent relative to comparable single-model systems." That figure is his reported comparison, not an independently established industry benchmark.

From Research Prototype to Supervised Infrastructure

Lab demos can skip the hard parts. Banks cannot. A regulated institution needs more than a strong model. It needs workflows you can evaluate, decisions you can trace, failures you can manage, and clear roles for the people who review the edge cases.

Berkovic connects those requirements to the larger goal of trust at scale, saying: "As generative models improve, the problem of reliably determining who is who and who is doing what online becomes more difficult, not less. Existing compliance systems were designed for an earlier technological era. We aim to build an intelligent data science layer that helps institutions and individuals maintain trust at global scale, preserving the ability of legitimate actors to operate while substantially raising the cost and difficulty of bad-actor activity."

That is the through-line from research prototypes to regulated production. Not one breakthrough. A discipline: build generative and multi-agent systems whose decisions can be rebuilt, whose errors can be contained, and that stay under human judgment rather than above it.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.