Building AI Without Fixing Your Client Data? Good Luck.
- Katie Tomlinson Broder

- Jul 29
- 4 min read
Why the least glamorous investment in your AI strategy is the most important.

Everyone is racing to build AI capabilities, and too many firms are doing it in the wrong order. They start with the AI application instead of the data foundation that makes it possible. The result is predictable: ambitious AI initiatives that never move beyond pilots.
To be clear, these are the right ambitions, and I've argued for many of these use cases in my recent writing. What rarely gets the same attention as the ambitions is the unglamorous data problem sitting beneath every one of them.
The reality is that few firms can confidently say their records accurately identify who their clients are or how they relate to one another. That may not sound like an AI problem, but it is. AI doesn't tolerate messy client records the way experienced humans do. It inherits the mess, scales it, and returns it with confidence. Before long, (already-hesitant) users stop trusting the outputs, and eventually, the technology itself.
AI Is Only as Good as the Client Record Underneath It
The challenge is that institutional relationships are genuinely complicated. A single client may exist as a parent company, multiple subsidiaries, dozens of funds, separate legal entities, and hundreds of contacts spread across a CRM, research platform, trading system, voting database, and market-data providers. Every system captures a different piece of the relationship but none tells the whole story.
Then there are the everyday issues every firm recognizes: duplicate records, inconsistent naming conventions, stale contacts, outdated hierarchies, incomplete ownership structures. Individually, they're manageable. Collectively, they become the foundation every AI application is expected to reason over.
Entity resolution is the discipline of determining which records describe the same real-world entity and how those entities relate to one another. Get it right and every downstream system inherits a coherent view of the client. Get it wrong and even the smartest model in the firm is reasoning about a person or an account that doesn't actually exist in the form it sees.
Why This Has Been So Difficult
For years, firms have treated entity resolution as a data-cleanup exercise. Teams periodically deduplicated records using deterministic rules and manual review. The process was labor-intensive, expensive, and almost immediately out of date. Clients merge, funds launch, organizations restructure, people change firms. The database begins drifting the moment the cleanup ends.
This isn't a new problem, and neither is the underlying science. For decades, researchers have shown that probabilistic record linkage consistently outperforms exact-match approaches, and that resolving entities jointly through their relationships can improve accuracy further. What held firms back wasn't the theory but the cost and complexity of implementing it.
Why It's Changing Now
That implementation barrier is finally beginning to fall. Modern entity-resolution platforms combine probabilistic matching, graph technology, machine learning, and increasingly large language models to maintain a continuously updated view of client identity. Instead of periodically cleaning a database, they resolve identity as new information arrives.
The opportunity extends beyond legal hierarchies. Modern graph-based approaches can also capture the relationships that actually drive institutional business: the analyst who keeps a mandate after changing firms, portfolio managers connected across multiple funds, recurring decision-makers, and the networks that relationship managers have traditionally carried only in their heads.
Early academic work on using large language models for entity resolution is particularly encouraging. While cost and scale remain real challenges, the results suggest that foundation models can resolve ambiguous records with surprisingly little task-specific training. A problem firms have lived with for decades suddenly has a credible path toward continuous improvement.
Why Financial Services Has the Most to Gain
Few industries depend on entity resolution as much as financial services.
The same capability that produces a trustworthy commercial view of a client relationship also underpins revenue attribution, broker votes, counterparty aggregation, beneficial ownership, and regulatory reporting. The commercial and compliance benefits come from the same underlying foundation. That's a rare investment: one that simultaneously improves revenue generation, operational efficiency, and risk management.
Start with the Foundation
Entity resolution is decidedly not exciting, but it is one of the few investments that makes almost every AI initiative more likely to succeed. The firms that treat clean, continuously resolved client data as their first AI project will be the ones whose copilots, agents, and client intelligence platforms deliver lasting value. Building AI without fixing your data will without a doubt limit the success of your ambitions.
As always, I'd love to hear where others believe AI can create the most value across financial services.
KTB
Author's Note: One of the most rewarding parts of my work over the past several months has been getting to know the founders building AI specifically for financial services. If you're trying to make sense of this rapidly evolving landscape, I'm always happy to share what I've learned, translate the technology into practical business use cases, and connect people with the teams I believe are building meaningful solutions.


