By the time this presentation ends, someone somewhere will make life-altering decisions based on something generated entirely by AI.
For the last 30 years, digital infrastructure is based on these unwritten rules. If you see a face or you hear a voice, you trust someone behind it. If you see a signature, you trust someone has approved it. Every business transaction, every government workflow is based on this foundation of trust. So trust is an invisible layer. On top of it, all digital infrastructure is based. But generative AI trashed it completely. An AI agent can clone a voice. A synthetic face can bypass identity check. AI can impersonate people at a speed no criminal organization has been able to do before.
So in the era of generative AI, seeing is no longer believing, neither is hearing. The cost of deception has collapsed, and the speed of deception has exploded.
Hi, my name is Rachna Shirvaswa. I work for California Department of Financial Protection and Innovation. Our mission is to protect 39 million Californians and their financial identities. Let me tell you what we do. When a massive fraud attack happens in the state of California, we identify the fraud, we examine the fraud, and we run enforcement. Let me take you behind the scenes of what we actually do.
Fraud comes, we investigate the fraud. When the investigation happens, we collect all the evidence for the fraud. When we have all the evidence, we take that evidence to the court. And in the court, the defense attorney has only one job: to attack the system that we have built.
So the question arises: how should we build a system which is credible? What are the attributes we need to have in a system which can appear in the court? And appear in the court means the system should be defendable. Defendable means we should be able to explain the system at every step of the process. We should be able to reproduce the issue at every step of the process. Because everything, every data that we produce, is going to appear in the court. So how do you build that kind of system?
Since we deal with financial data, we have to ensure that the data is secure, for sure. We have to ensure that the data is evolving and learning. Why? Because the fraud space is changing every day. We have to ensure that the data is available to appear in the court and reproducible. So what I'm going to show you now is the system that we built and the lessons we have learned the hard way. So the first and foremost decision that we have taken was to build our solution offline. So you might be wondering, it's extreme. Why do you want to create a solution offline when we have a secure cloud?
So before we show you what we built, let's talk about what the industry does in the normal way. So the gold standard of security is encryption. Encryption means you encrypt the data at rest or encrypt the data in transit. But the thing is, for a machine learning model to work, the data has to be decrypted and available in plain text in the model memory. And if the data is already decrypted and present in plain text, it is prone to prompt injection attack.
Secondly, let's talk about private endpoints. Private endpoints are actually an isolated environment given to you from your cloud provider. So the thing is that according to the CLOUD Act, the federal government can access your data, which is in the cloud, and they don't even have to tell you that your data has been accessed. Let's talk about certification. FedRAMP compliant, ISO certified, SOC 2 compliant. All these compliance standards are just paper. And we have seen these highly compliant organizations fail over and over again. So the only way you can make your data secure and available in the court to present as evidence is to build the solution offline.
So like everyone else, we thought, what a big deal. Let's just download an open source model, create an isolated environment, spin up some GPU, add some system prompt, add some guardrails, and then we are running. Push the live data into it. We did the same as everyone else, and the system collapsed in two hours. And our first instinct was, we took a free model from online. The model is not good for our use case. But the issue was, we were treating the machine learning model as a magic box instead of as a data pipeline. So whatever garbage came into the system, we were expecting the model to clean up and do all the processing.
So to solve this problem, we introduced Kafka for data ingestion, Spark for data processing, and a large language model for reasoning. Three different tools to solve three different problems. We use Kafka for these three things. First, the data that enters into our system does not come at a consistent speed. When a fraud attack happens, it attacks all over the state. So we get a high spike of traffic into our system. So we needed a tool that can buffer that traffic so that the downstream component can access the data at a consistent speed. And Kafka did that.
Secondly, we wanted all the events that occur in the system to come and be stored in a sequential order. When did you open the account? When did the transaction happen? The order of events in the fraud is extremely important. Kafka helped us solve that problem. But the main problem that Kafka helped us solve is reproducibility. In Kafka, it lets you move the checkpoint to the point at which the decision was made, and it helped you replay the event. And this replayability of the event is the proof that we present in the courtroom.
So now Kafka did solve our data storage problem, but the data that enters into the system is really messy. We get 10 different formats of bank statement. We get audio file. We get screenshot, fax statement. And if you dump all this data to model memory, then no wonder the model hallucinates. So we introduced Spark because Spark let us process this data behind the scenes in clusters of CPU instead of working on GPU. And Spark let us clean the data, and then when we sent the clean data to the model memory, to our utter surprise, the same model started giving us such good results, it found the connection between the points which we never expected.
So we learned our first lesson. The lesson is most of the data problems in AI are data engineering problems wearing AI masks. So I recommend you consider using data engineering tools to solve those problems. Now our model memory is filled with very condensed clean data, but then we found our next issue. The issue was the memory was filled with a lot of sensitive information: your credit card number, your bank account number, your social security number.
So we introduced something called a cryptographic vault, which is basically a SHA-256 cryptographic algorithm with a hardware security module. So in the simple sense, it means that as we get the data into the system, the cryptographic vault converts the data into a cryptographic hash. The key that is used to do the conversion is physically attached to our server rack. So if tomorrow somebody is able to get our data, they have to walk into our office physically, break the server rack, get the key, to actually make sense of the data. Then we learned our second lesson. If the stakes are really high, trust hardware over software.
Now we have a model which is working really well. All the PII is redacted. We thought, let's just run a quick round of load testing and release this product. The problem is, in the cloud, the scaling is unlimited. You get a high spike of traffic, it spins up more servers, it distributes the load to the servers, and you are done.
But when you are creating an isolated environment, you have limited GPU, limited compute. So we have to step back and figure out what is the actual problem here. The problem was we were using one state-of-the-art machine learning model to do every single processing. The same model was doing the summarization, entity extraction, fraud ring detection. So we were actually making a neurosurgeon take the blood pressure of every single patient.
To solve this, we introduced a triage nurse, or semantic router. And the goal of semantic router is, as we get the data, it analyzes the data and forwards the request to the smallest possible model which is capable of processing that request. And by making this simple architecture change, we found that more than 80% of our tasks could be easily done by the smallest, fastest, cheapest model. And by not adding any new GPU, we could process three times more traffic, and the cost of processing each request reduced by nearly 70%.
Now we have done the unit testing, we have done the load testing, and then we found the hardest problem of all. The problem was, we built this highly secure system, but the system was not learning. The system did not know what is happening in the threat space. So the question arises: how do we make our system learn without creating a security hole? Usually in industry, people solve this problem by configuring a software firewall. But any configuration can be misconfigured. And once you have misconfigured it, your highly secure system will be a highly exploited one.
We decided not to trust the configuration, and we took help from physics. We introduced a one-way data diode, which is basically a fiber optics cable physically cut into half. The first half of the cable is connected to the internet to receive the data. The second half of the cable is connected to our solution. The first half of the cable has a laser transmitter, which receives the data from the internet and transmits the data to the second half. The second half has a laser receiver to receive the data from the first half, but there is no laser transmitter from our end to the outside world. So it is physically impossible for data to leak from the system. And this is how we ensure 100% guarantee of the security of the system.
The data diode solves the problem of directionality of the data. But anything entering into the data is considered to be unsafe until proven. So anything from outside first lands into a quarantine zone, where we run a Spark job that runs a validation on each input data. And then, if all the validation is successful, the data goes to the production layer.
In the production layer, we also save the data into Apache Iceberg. Apache Iceberg is a time-traveled, queryable, immutable data store. We enter the data in this data store because two years from now, if one of our results from our system goes to the court, we cannot go to the court and tell them, this system's result is produced by AI and we don't know anything about it. At that moment, we travel the Apache Iceberg, go to the point at which the decision was made, and get the state of the system at that moment from the database. And that state of the system appears as proof in the court.
And that is how we build a solution which is secure, which is evolving, and which is defensible in the court. This is the whole architecture end-to-end. If you see this, only at the very end of the system, when a user logs into the system, we authorize the user into multi-factor authentication, and only then is the data reverted. Data is not opened till the very end at the browser where the user is actually evaluating the threat case. So don't look at this solution as a fraud detection solution. This is the architecture of the future. Soon, the same architecture will be used to predict healthcare, banking statements, legal, and other domains.
In the end, I just want to say one thing. Build solutions that can be trusted. And remember, trust is not our policy. Trust is a physical property of the system. You have to build the trust from day one into your hardware, into your physics, into your architecture, or it's not there. It's that simple. So years from now, nobody will remember the models you trained. Nobody will remember the benchmark you received. People will only remember whether the solution that you have built can be trusted when it is needed the most. Thank you so much. Thank you so much. We have to ensure that the data is evolving and learning. Why? Because the fraud space is changing every day.
We have to ensure that the data is available to appear in the court and reproducible. So what I'm going to show you now, the system that we built, and the lessons we have learned the hard way. So the first and foremost decision that we have taken was to build our solution offline. So you might be wondering, like, it's extreme. Like, why are you, why do you want to create a solution offline when we have a secure cloud? So before we show you what we build, let's talk about what do industry do in the normal way. So the gold standard of security is encryption. Encryption means you encrypt the data at rest or encrypt the data on the transit.
That means the thing is, but for machine learning model to work, the data has to be decrypted and available in the plain text in the model memory. And if the data is already decrypted and present in the plain text, it is prone to prompt injection attack. Secondly, let's talk about private endpoints. Private endpoints are actually an isolated environment given to you from your cloud provider.
So the question, the thing is that according to the cloud act, federal government can access your data, which is in the cloud, and they don't even have to tell you that your data has been accessed. Let's talk about certification. FedRAM compliant, I also certified SOC 2 compliant. All these compliance are just paper. And we have seen these highly compliant organizations fail over and over again. So only way you can make your data secure to be available in the court to present as evidence as evidence is to build the solution offline.
So like everyone else, we thought, like, what a big deal. Let's just download an open source model, create an isolated environment, spin up some GPU, add some system prompt, and add some guardrails, and then we are running, push the live data into it. We did the same like everyone else. And the model, the system collapsed in two hours. And our first instinct was, you know, we took the model, free model from online. Model is not good for our use case. But the issue was, we were treating machine learning model as a magic box, instead of as a data pipeline. So anything that garbage comes into the system, we were expecting the model to clean up and do all the processing.
So to solve this problem, we introduced Kafka for data ingestion, Spark for data processing, and large language model for reasoning. Three different tools to solve three different problems. We use Kafka for these three things. First, the data that enter into our system does not have come at a consistent speed. When a fraud attack happen, it attack all over the state. So we get a high spike of traffic into our system. So we needed a tool that can buffer that traffic, so that the downstream component can access the data at a consistent speed. And Kafka did that. Secondly, we wanted all the events that occur in the system to come and store in a sequential order.
When did you open the account? When did the transaction happen? The order of events in the fraud is extremely important. Kafka helped us solve that problem. But the main problem that Kafka helped us solve is reproducibility. In Kafka, it let you move the checkpoint to the point at which the decision was made, and it helped you replay the event. And this replayability of the event is the proof that we present in the courtroom. So now Kafka did solve our data storage problem, but the data that enter into the system is really messy. We get 10 different format of bank statement. We get audio file. We get screenshot, fax statement.
And if you dump all this data to model memory, then no wonder model hallucinate. So we introduced Spark because Spark let us process this data behind the scene in the clusters of CPU, instead of working on GPU. And Spark let us clean the data, and then when we process, send the clean data to the model memory, to our utter surprise, the same model started giving us so good results, it found the connection between the points which we never thought expected.
So we learn our first lesson. The lesson is most of the data problem in AI are data engineering problem wearing AI mask. So I recommend you to consider using data engineering tool to solve those problems. Now our model memory is filled with a very condensed clean data, but then we found our next issue. The issue was the memory was filled with a lot of sensitive information, your credit card number, your bank account number, your social security number. So we introduced something called cryptographic vault, which is basically a SHA-256 cryptographic algorithm with hardware security module.
So in the simple sense it means that as we get the data into the system, cryptographic vault, which converts the data into cryptographic hash. is the key that is used to do the conversion is physically attached to our server rack. So if tomorrow somebody is able to get our data, they have to walk into our office physically, break the server rack, get the key to actually make sense of the data. So we have to do the data. So we have to do the data. Then we learn our second lesson. If the stakes are really high, trust hardware over software.
Now we have a model which is working really well. All the PII is redacted. We thought, let's just run a quick round of load testing and release this product. The problem is in the cloud, the scaling is unlimited. You get a high spike of traffic. It spins up more server. It distributes the load to the server and you are done. But when you are creating an isolated environment, you have limited GPU, limited compute, limited and you have to do it. So we have to step back and figure out what is the actual problem here. The problem was we were using one state of the art machine learning model to do every single processing.
The same model was doing the summarization, entity extraction, fraud ring detection.
So we were actually making a neurosurgeon take the blood pressure of every single patient. To solve this, we introduce triage nurse or semantic router. And the goal of semantic router is, as we get the data, it analyzes the data and forward the request to the smallest possible model which is capable of processing that request. And by making this simple architecture change, we found that more than 80% of our tasks could be easily done by the smallest, fastest, cheapest model. And by not adding any new GPU, we could process three times more traffic and cost of processing each request reduced to nearly 70%.
Now we have done the unit testing, we have done the load testing, and then we found the hardest problem of all. The problem was, we built this highly secure system, but the system was not learning. System did not know what is happening in the thread space. So the question arises, how do we make our system learn without creating a security hole?
System is not learning. Usually in industry, people solve this problem by configuring software firewall.
But any configuration can be misconfigured. And once you have misconfigured, your highly secure system will be highly exploited one. We decided to not trust the configuration and we took help from physics. We introduced one-way data diode, which is basically a fiber optics cable physically cut into half. The first half of the cable is connected to the internet to receive the data. Second half of the cable is connected to our solution. First half of the cable has laser transmitter, which receives the data from the internet and transmit the data to the second half. Second half has a laser receiver to receive the data from the first half.
Second half, but there is no laser transmitter from our end to the outside world. So it is physically impossible for data to leak from the system. And this is how we ensure 100% guarantee of the security of the system.
Data diode solves the problem of directionality of the data. But anything entered into the data is considered to be unsafe till proven. So anything from outside first lands into quarantine zone. where we run a Spark job that runs a validation on each input data. And then all the validation is successful. Then the data goes to the production layer. In the production layer. We also save the data into Apache Iceberg. Apache Iceberg is a time-traveled, queryable, immutable data store.
We enter the data in this data store because two years from now, one of our results from our system goes to the court. We cannot present, we cannot go to the court and tell them, you know what, this system's result is produced by AI and we don't know anything about it. At that moment, we travel the Apache Iceberg, go to the point at which the decision was made and get the state of the system at that moment from the database. And that state of the system appear as a proof in the court. And that is how we build a solution which is secure, which is evolving and which is defensible in the court.
This is the whole architecture end-to-end. If you see this only at the very end of the system, when a user logs into the system, we authorize the user into multi-factor authentication and only then the data is reverted. Data is not opened till the very end at the browser where the user is actually evaluating the threat case.
So don't look at this solution as a fraud detection solution. This is the architecture of the future. Soon, the same architecture will be used to predict the healthcare, banking statements, legal and other domains. In the end, I just want to say one thing. Build solution that can be trusted. And remember, trust is not our policy. Trust is a physical property of the system. You have to build the trust from day one into your hardware, into your physics, into your architecture, or it's not there. It's that simple. So years from now, nobody will remember the models you trained. Nobody will remember the benchmark you received.
People will only remember the solution that you have built can be trusted when it needed. The most. Thank you so much. Thank you so much.