SPEAKER_00
Hello everyone. Thank you to the AI Engineering World's Fair team for providing this wonderful opportunity to share my research. It is truly an honor to be speaking alongside so many talented researchers and practitioners. My name is Varsha Shah. I am an Enterprise Technical Architect working at Tata Consultancy Services working for Microsoft. I am focused on artificial intelligence, enterprise compliance, finance governance, and intelligent automation. Today, I would like to share my research on AI-driven multi-document correlation for enterprise financial compliance and fraud detection. As organizations continue to digitalize their operations, they generate numerous amounts of data for the financial system across payroll, tax, procurement, and transaction systems. Ironically, while we have more data than ever before, compliance teams continue to struggle with hidden fraud patterns and regulatory risk. The reason is that most existing solutions analyze the documents independently. While many of the most critical risks only become visible when the information is connected across multiple systems. In this presentation, I'll introduce you to a framework that combines graph-based entity correlation, probabilistic risk modeling, and cross-jurisdictional normalization to uncover these hidden relationships and transform enterprise compliance from a reactive process into a proactive intelligence capability. The framework was evaluated using approximately 3 million financial records across four jurisdictions, demonstrating both strong detection performance and meaningful operational improvement. With that context, let's begin by looking at the compliance gap that organizations continue to face today.
SPEAKER_00
Let's begin by understanding the compliance gap that many organizations continue to face. Enterprise compliance has become significantly more complex over the last decade. Organizations now operate across multiple countries, regulatory frameworks, and financial systems, each with its own reporting standards and compliance requirements. At the same time, the volume of the enterprise has grown exponentially. Payroll records, tax filings, procurement transactions, and financial documents are generated every day, making manual reviews both time consuming and increasingly impractical. Adding to this challenge, fraud has evolved. Modern fraud rarely appears as an obvious error within a single document. Instead, it exploits subtle inconsistencies across multiple systems, patterns that often remain invisible when the records are reviewed independently. This creates a fundamental limitation. Traditional rule-based and document level NLP systems are designed to validate individual records, but they are not built to understand the relationships across documents. And that's the gap this research aims to address. Moving beyond isolated document analysis to uncover hidden risks through cross-document correlation.
SPEAKER_00
To better understand this limitation, let's look at why traditional document level analysis often fails to detect the most sophisticated fraud patterns. To understand why traditional approaches struggle, let's consider how most compliance systems operate nowadays.
SPEAKER_00
Multiple documents are analyzed together. For example, a payroll record may appear accurate. A vendor invoice may seem legitimate. A tax filing may be correctly submitted. But when these records are connected, they reveal inconsistencies that indicate fraud and compliance risk. The information already exists. What is missing is the ability to understand the relationships between these documents. That is why the research shifts the focus from document level validation to cross-document intelligence, enabling organizations to detect the risks that would otherwise remain hidden.
SPEAKER_00
So if the problem is understanding relationships rather than individual documents, what kind of architecture can solve this? Let me introduce you to the framework. Now that we have established the problem, let's look at the proposed framework. Rather than relying on a single model or algorithm, the solution is built on three complementary components that work together to transform raw enterprise data into actionable compliance intelligence.
SPEAKER_00
The first component is the entity correlation engine. The purpose of this is to connect related information across payroll, tax, procurement, and financial systems, creating a unified view of enterprise activity rather than isolated records. Once these relationships are established, the second component, the adaptive probabilistic risk model, evaluates the connected data to determine which patterns represent meaningful compliance risk factors. Instead of generating alerts based on single rules, it prioritizes cases using multiple risk signals. The cross-jurisdictional normalization layer provides the regulatory context by standardizing currency, tax structure, reporting standards, and compliance rules across different jurisdictions. This ensures that risks are evaluated consistently regardless of where the transaction originated.
SPEAKER_00
So individually, each component provides value. Together, they enable a framework that moves beyond document validation and towards enterprise-wide compliance intelligence. In the next few slides, I'll briefly explain how each of these components contributes to the process. So let's begin with the foundation of the framework, the entity correlation engine, which makes cross-document intelligence possible.
SPEAKER_00
The first component is a graph-based entity correlation engine. Its role is to connect entities across payroll, tax, procurement, and financial systems into a unified network. Rather than analyzing documents independently, it identifies relationships between employees, vendors, accounts, transactions, and regulatory files. And these connections reveal hidden patterns and structural anomalies that are often invisible to traditional document level analysis. In simple terms, this component answers one fundamental question: what is connected? It provides the relational foundation for the rest of the framework.
SPEAKER_00
Once relationships have been established, the next step is to determine which ones actually represent meaningful risk. That's the role of the adaptive probabilistic risk model. Instead of relying on static rules, it combines multiple risk indicators, such as anomaly strength and source reliability, with historic patterns to calculate confidence-based risk scores. This helps prioritize cases that require immediate attention while reducing unnecessary investigations. The important advantage is its ability to learn from audit outcomes, allowing the model to continuously improve its accuracy over time. Simply put, this component answers the question: what is most likely to be genuine compliance risk?
SPEAKER_00
Once we have identified high-risk cases, the final step is to ensure those risks are interpreted consistently across different regulatory environments. The final component is the cross-jurisdictional normalization layer. Global organizations operate across different countries, currencies, tax structures, and reporting standards. Without normalization, the same transaction can be interpreted differently depending on the jurisdiction.
SPEAKER_00
Together, the next step was to evaluate how the framework performs under real-world enterprise conditions. So before discussing the results, it's very important to briefly understand the evaluation environment. The framework was evaluated using approximately 3 million financial records collected over a five-year period across four different regulatory jurisdictions. The objective was to assess how the framework performed under realistic enterprise conditions where a large volume of interconnected financial data and varying regulatory requirements introduced significant complexity.
SPEAKER_00
With that foundation established, let's look at how the framework performed. So how effective was the framework in detecting hidden compliance risks? Let's look at the results.
SPEAKER_00
The framework achieved approximately 91 percent precision, meaning the vast majority of flagged cases were confirmed as genuine anomalies. It also achieved 87 percent recall, demonstrating its ability to identify most true fraud cases while minimizing missed detections. Together, these results produce an F1 score of 0.89, indicating a strong balance between precision and recall. More importantly, these results are achieved across four jurisdictions and large enterprise scale data, demonstrating that the framework performs consistently in complex real-world environments.
SPEAKER_00
The next question is equally important: How do these improvements translate into operational value for the compliance teams? Detection accuracy is one part of the story. The real value comes from improving the efficiency of compliance operations.
SPEAKER_00
Beyond detection accuracy, the framework delivers measurable operational benefits. One of the most significant outcomes was a 76 percent reduction in false positives, meaning investigators spend far less time reviewing cases that are ultimately legitimate. In addition, the framework reduces manual audit efforts by approximately 40 percent, allowing compliance teams to focus only on high-risk cases instead of routing document reviews. Ultimately, this translates into faster investigations, improved resource utilization, and greater confidence in compliance decisions. In other words, the framework doesn't just detect fraud more effectively. It helps compliance teams work more effectively.
SPEAKER_00
These operational improvements become even more meaningful when we compare the framework with traditional rule-based approaches.
SPEAKER_00
This comparison highlights the advantages of moving beyond traditional rule-based systems. Conventional approaches are effective at validating individual records, but they often struggle to detect complex fraud patterns that span multiple documents and systems. By incorporating entity correlation, probabilistic risk modeling, and regulatory normalization, the proposed framework consistently performs better than the baseline across key performance metrics. The key takeaway is simple. Connecting data across documents produces better detection, fewer false positives, and more actionable compliance intelligence than analyzing documents in isolation.
SPEAKER_00
One of the reasons the framework continues to improve over time is that it doesn't rely on static rules anymore. Instead, it continuously learns from completed audits and investigations and investigator feedback. This framework is designed to continuously improve through feedback. Every completed audit provides valuable information confirming fraud cases and strengthening future detection patterns. While false positives help refine the risk scoring and reduce unnecessary alerts. This creates a continuous learning cycle where the system becomes more accurate with each audit and investigation. As fraud patterns evolve and business environments change, the framework adapts rather than relying on manual rule updates. In simple terms, every completed audit helps make the next audit more efficient.
SPEAKER_00
This continuous learning capability enables a broader shift. From reactive compliance—issues identified after they occur—to identifying risk before they become audit findings.
SPEAKER_00
Beyond improving detection accuracy, this framework changes how organizations approach compliance. Traditionally, compliance has been reactive; issues were identified after audits, investigations, or regulatory reviews have occurred. This framework enables a more proactive approach through continuous monitoring, cross-document correlation, and adaptive learning. Instead of asking what went wrong, organizations can begin asking what is likely to go wrong next. That shift from reactive validation to predictive governance allows compliance teams to identify risks earlier, prioritize investigations more effectively, and make better-informed decisions. Ultimately, compliance becomes an ongoing intelligence function rather than a periodic review process.
SPEAKER_00
While the framework is research-driven, it has been designed with enterprise deployment in mind. Successful implementation depends on four key considerations. First is seamless integration with existing enterprise systems such as ERP, payroll, procurement, and tax platforms. Second is jurisdiction-specific configuration to ensure compliance with local regulations and reporting standards. Third is alignment with the audit framework, enabling investigators to focus on prioritized risk cases instead of manual document reviews. And finally, scalability. The evaluation demonstrates that the framework can effectively process millions of financial records, making it suitable for large enterprise environments. Together, these considerations help bridge the gap between research and practical enterprise adoption.
SPEAKER_00
As I conclude, I would like to leave you with four key takeaways. First, many of today's most significant compliance and fraud risks exist between the documents, not within them. Cross-document analysis and cross-document intelligence matters. Second, combining graph-based entity correlation, probabilistic risk modeling, and cross-jurisdictional normalization creates a more comprehensive and scalable approach to enterprise compliance. Third, the results demonstrate that this particular framework not only improves detection accuracy but also reduces false positives and significantly lowers manual audit efforts, delivering measurable operational value. Finally, by continuously learning from audit outcomes, the framework enables organizations to move beyond reactive compliance toward predictive intelligence-driven risk management. I believe this represents an important step toward the future of enterprise financial governance where AI is not only helping organizations detect risks but also anticipates and prevents risks going forward.
SPEAKER_00
Thank you once again to the AI Engineering World Fair team for this wonderful opportunity, and thank you all for your time and attention. Please feel free to reach out to me with your questions and your enthusiasm via the email below or connect with me on LinkedIn. I am happy to help if you are building such amazing systems for compliance using AI. I want to transform them, and thank you all again. See you next time. but they are not built to understand the relationship across the documents. And that's the gap this research aims to address. Moving beyond this isolated document analysis to uncover the hidden risk through the cross-document correlation.
SPEAKER_00
To better understand this limitation, let's look at why traditional document level analysis often fail to detect the most sophisticated fraud patterns.
SPEAKER_00
To understand why traditional approaches struggle, let's consider how most compliance system operates nowadays. contracting contracting contracting contracting ! contracting multiple documents are analyzed together. For example, a payroll record may appear accurate to us. A vendor invoice may seem legitimate to us. A tax filing may be correctly submitted. But when these records are connected, they are revealing the inconsistencies and indicate the fraud and compliance risk. The information already exists. What missing is the ability to understand the relationship between these documents. That is why the research shifts the focus
SPEAKER_00
from document level validation to cross document intelligence, enabling the organizations to detect the risk that would otherwise remain hidden. So if the problem is understanding relationships rather than individual documents, what kind of architecture can solve this? Let me introduce you to the framework. Now that we have established the problem, let's look at the proposed framework. Rather than relying on a single model or algorithm, the solution is built on three complementary components that work together to transform the raw enterprise data into the actionable compliance intelligence. The first component is the entity correlation engine.
SPEAKER_00
The purpose of this is to connect the related information across the payroll, tax, procurement, financial systems, creating a unified view of enterprise activity rather than isolated records. Once these relationships are established, the second component that is adaptive probabilistic risk model, which is which evaluates the connected data to determine which pattern represent meaningful compliance. the risk factors. Instead of generating alerts based on single rule, it prioritizes the cases using the multiple risk signals here. The cross jurisdictional normalization layer provides the regulatory context by standardizing the currency, tax structure, reporting standards,
SPEAKER_00
and compliance rules across the different jurisdictions. This ensures that the risk are evaluated consistently regardless of where the transaction originated. So individually, each component provides a value. Together, they enable the framework that move beyond the document validation and towards the enterprise-wide compliance intelligence. In the next few slides, I'll briefly explain you how each of these components contribute to the process. So let's begin with the foundation of the framework, the entity correlation engine, which makes the cross document intelligence possible. So the first component is graph-based entity correlation engine.
SPEAKER_00
The role is to connect the entities across the payroll, tax, procurement, financial system into unified network. Rather than analyzing the documents independently, it identifies the relationship between the employee, vendors, account, transaction, regulatory files. And these connections reveals the hidden pattern and structural anomalies that are often invisible to the traditional document level analysis we were doing. In simple terms, this component answers one fundamental question, what is connected? It provides the relational fundamental for the rest of the framework.
SPEAKER_00
Once those relationships are established, the next step is determining which one actually represents the meaningful risk.
SPEAKER_00
Once relationships have been established, the next step is to determine which one actually represents meaningful risk. the next step is to determine which one is the fundamental risk. That's the role of the adaptive probabilistic risk model. Instead of relying on static rules, it combines the multiple risk indicators, such as anomaly, anomaly strengths or the source reliability, historic patterns has happened to calculate confidence-based risk score. This helps prioritize the cases that require immediate attention while reducing the unnecessary investigations.
SPEAKER_00
The important advantage is its ability to learn from the audit outcomes, allowing the models to continuously improve its accuracy over the time. Simply put, this component answers the question that what is most likely to be genuine compliance risk. Once we have identified the high risk cases, the final step is to ensure those risks are interpreted consistently across the different regulatory environments. The final component is cross judicional normalization layer. Global organizations operate across different countries, currencies, tax structure and reporting standards. Without normalization, the same transaction can be interpreted differently depending on the judiction.
SPEAKER_00
name name name name name name together the next step was the evaluated how the framework performs under real world enterprise conditions. So before discussing the results let's it's very important to briefly understand the evaluation of the environment. The framework was evaluated using the approximately 3 million of the financial records collected over five years of period over the four different regulatory jurisdictions. The objective was to assess how the framework performed under the realistic enterprise condition and where a large volume of interconnected financial data and varying
regulatory requirement introduced a significant complexity here. With that foundation established let's look at how the framework performed. So how effective was the framework in detecting hidden compliance risk. Let's look at the results. So here now look at the detection performances. The framework achieved approximately 91 percent of precision meaning the vast majority of the flag cases were confirmed as a genuine anomaly here. It also achieved 87 percent recalls demonstrating its ability to identify most true fraud cases while minimizing the misdetections here. So together
these results produce a F1 score that is 0.89 indicating a strong balance between the precision and recall here. More importantly these results are achieved across four jurisdictions and large enterprise scale data demonstrating that the framework performs consistently in the complex real world environments. The next question is equally important. How do these improvements translate into operational value for the compliance teams? Detection accuracy is one part of the story. The real value comes from improving the efficiency of the compliance operations.
SPEAKER_00
The real value comes from the performance of the compliance operations. Beyond detection accuracy the framework delivers measurable operational benefits. One of the most significant outcome was 76 percent reduction in false positive meaning um investigators spend uh far less than reviewing the cases that they uh they are ultimately uh legitimate. So in addition the framework reduces the manual audit efforts by approximately 40 percent allowing compliance teams to be focused on the only high risk cases instead of uh routing the document reviews.
SPEAKER_00
Ultimately uh this translates into faster investigations improved uh resource utilization and uh greater confidence in compliance decision. In other words the uh the framework doesn't just detect fraud more effectively. It helps compliance teams work more effectively. So these operational improvements become even more meaningful when we compare uh uh the uh the framework with the traditional rule-based approaches.
SPEAKER_00
So these are the advantages of the rule-based approaches. This comparison highlights the advantage of moving beyond the traditional rule-based system. Conventional approaches are effective at validating the individual records but they often struggle to detect the complex fraud patterns uh that spans multiple uh documents and systems. So by incorporating entity correlation, probabilistic risk modeling, and regulatory normalization, the proposed framework consistently um performs the baseline across the key performance metrics. The key takeaway is simple. Connecting data across documents produces better detection, fewer false positive,
SPEAKER_00
and more actionable uh compliance intelligence than analyzing the documents is uh in in isolations. One of the reason the framework continues to improve over the time is that uh doesn't rely on the static rules anymore. Instead it continuously learn from the completed audit and investigate uh and the investigator feedbacks here. So this framework is designed to continuously improve through the feedback. Every uh completed audit provides the valuable information uh confirm the fraud case strengthen the future detection pattern. While false positive helps refining the risk scoring and reduce the
SPEAKER_00
unnecessary alerts. This creates a continuous learning cycle where the system becomes more accurate with each audit and investigation. As fraud patterns evolve and business uh environments changes the framework adopts rather than relying on manual rule updates. In simple terms every completed audit helps make the next audit more efficient. So this is how the process of reporting is to improve the assessment. This continuous learning uh capability enables a broader shift here. From reactive to compliance issue after they occur to identifying the risk before they become audit uh findings here.
SPEAKER_00
contracting contracting contracting contracting contracting contracting contracting contracting contracting contracting beyond improving detection accuracy this framework changes how organization approach compliances traditionally the compliance has been reactive issue were identified after audits investigation or regulatory reviews have happened this framework enables more proactive approach through continuous monitoring cross document correlation adaptive learnings instead of asking what went wrong organizations can begin asking what is likely to go wrong next that shifts from reactive validation to predictive governance allowing the compliance teams to identify the
SPEAKER_00
risk earlier prioritize the investigation more effectively and make better informed decision ultimately the compliance becomes an ongoing intelligence functions rather than a periodic review process while the framework is research driven it has been designed with the enterprise deployment in mind successful implementation depends on four key consideration first one is seamless integration with existing enterprise systems such as erp payroll procurement tax tax platforms second the judicial specific configuration to ensure the compliance with local regulations and reporting the compliance and reporting standards third one is alignment with the audit framework enabling the
SPEAKER_00
investigator to focus on prioritized risk-old cases instead of manual document reviews and finally scalability the evaluation demonstrate that the framework can effectively process millions of financial records making it suitable for the large enterprise environments so together these consider these considerations help bridge the gap between the research and the practical enterprise adoption as i conclude i would like to leave you with four key takeaways first is many of today's most significant compliance and fraud risk exist between the documents not within them the analysis and cross-document analysis and cross-document intelligence
SPEAKER_00
and cross-document intelligence second is combining the graph based entity correlation probabilistic risk modeling and cross-geridictional normalization creates some more comprehensive and scalable approach to enterprise compliance third is the result demonstrated that this particular framework is not only improves the detection accuracy but also reduces the false positive and significantly lowers the manual audit efforts delivering the measurable operational value. Finally by continuous learning from the audit outcomes and the framework enables the organizations to move beyond reactive compliance
SPEAKER_00
toward predictive intelligence driven risk management here. I believe this represents an important step toward the future of enterprise financial governance where AI is not only helping the organizations to detect risk but also anticipates and prevent this risk going forward.
SPEAKER_00
So thank you once again to the AI engineering world first team for this wonderful opportunity and thank you all for your time and attention. Please feel free to reach out to me with your questions and your enthusiasm on this email below or connect me on the LinkedIn. I am happy to help you if you are building with such kind of amazing systems for compliances using the AI I wanted to transform them and thank you all again see you next time.