Hello. Thank you for coming. I'm going to talk to you about agentic sites, how we call it, building hyper-personalized websites. I'm not going to just talk about it. I'm going to show you what we're building. I've been working on this project for a bit now and will try to show you what is possible today with AI. I work at Adobe. I'm a principal scientist at a product that not many people know, Adobe Experience Manager, content management. We run a lot of websites, properties for big brands, and my background is in open source, contributing to a lot of foundations and projects.
What are agentic sites and how are we building this thing? We are looking for sites that are looking at what intent the user browsing has, what is the user doing, what is the user trying to achieve. And the end goal is to personalize these pages for the current user browsing so that eventually this drives higher engagement or conversions, whatever the marketing teams want to achieve. And these pages are personalized in real time based on the user that is accessing the site and what the user is doing.
The stack we're using is AM Edge Delivery. This is the part of the product we have where all the content is on the edge. And then we have a backend service that powers this experience with different LLM providers, LLM services. We use Cerebras for fast inference, or we can also use Bedrock and a bunch of others. I'll be showing Cerebras today, and you will see the reason why.
The engine that is personalized in this is in the rich content and blocks. Different blocks on the site are customized depending on what the user persona is. We don't want the whole site to be generated. If you talk to marketing people, they have very strict brand guidelines. You don't want to just come up with or have some hallucinations there. So what is personalized is different sections of the site, and we use the whole site as a corpus. We build a RAG from the whole site. So what is generated is grounded on the existing site.
We try to solve the problem where one size fits all. We want hyper-personalized experiences. Also, we want to help our customers do more automatic authoring, so they do not have to create thousands of different variations of the site, but use AI for this and then do these multiple layers of personalization.
Some examples of what we're doing or showing in the demo are instant persona adaptation, query generation when the user searches for something on the site, where the page with the results is customized for them, and also something like recommendations, where after you browse the site for a period of time, we can create a page that recommends something based on what we think you are looking for.
For marketeers, they can define this strategy in natural language, and they can use analytics to drive the loop of personalization and what is the end goal, and how this goes back again to adapt the personalization to improve that whole cycle. Everybody's talking about loops in this conference, so that's one of the loops there.
How does the architecture look? It's a dynamic front end with some blocks, what I mentioned before, and with Edge Delivery Services, you compose these blocks, and they are updated in real time with the AI. In the back end, we do the evaluation of the models and the providers, and one thing we realized is that this is very dependent on the site. So we have a bunch of prompts, and we run them across a huge variety of models and providers, and then we look at the accuracy, we look at the speed, but this is going to depend highly on what type of site, how big the site is, what different area the site is targeting, what type of commerce it is, and so on. So we run this evaluation continuously.
We use PromFu. Anybody heard about PromFu? Okay, some people. PromFu allows you to evaluate models and prompts against multiple model providers, and you can do local models and any of the features that we have to do with a bunch of OpenAI-compatible providers, and a lot of them, basically. We look for two things. Why? Accuracy. That's typically what people look for. But also we want the speed because we don't want the site generation to take more than one or two seconds, right? Because this is already proven, that the faster the site, the more conversions it generates, or the better the experience it is for the user.
What I mentioned is different sites may have different requirements, so you may have to run this evaluation of models depending on the site. So we have 15 prompts for this example site, and at the top you can see with Cerebras on the Gemma 4 model that was announced last week, we can get an average latency of 1.1 seconds generating a page. You can compare that to the second one, which is 4.6 seconds, right? So the difference is huge. And that's why we use Cerebras for this use case. And you can see that different providers, different models have different speeds.
And here is a... Let me... I can show you the whole thing here. Not this one, this one, right? So at the bottom we have others. Sometimes maybe some of them may be good. They don't need to be perfect, but they're good enough if they're fast enough. So that's going to be the kind of decisions that you need to make on whether the model is good enough for your use case or not. So we're looking, yeah, average 1.1 seconds, and then the next ones are going from 4 seconds higher. And you don't need a huge LLM to do this sort of work because you are generating text, you are deciding where to put blocks and how to organize the website. You don't need lots of information for that.
So this browsing and the queries are being recorded. These are the metrics or the data we gather from the user. And this is fed into the LLM to personalize the site. And then in this example, we personalize the hero card, the products, the blog feeds, and the navigation based on the persona. Also, some of the buttons, like our call to action navigation, we can also personalize those.
We create, and I'll show you, the For You page, which is a recommendation. And this is an interesting one because this you could pre-generate, right? As the user browses your site, you gather the signals and you could keep generating this. So in this case, you wouldn't need such a big speed. But that's interesting because if a user wants to buy something, you could just say, okay, For You, I will recommend these three products or something like that. And then they can see this recommendation. And if they go there, that could be prefetched for them. And obviously, you have to keep updating it as the user navigates around the site and so on. So that's also something to consider on the cost, cost of doing multiple generations, multiple LLM calls.
When the user runs a query, a dynamic personalized page is shown to them. When these queries are also grouped into personas or intent types, what is this person trying to do in the site? Is he trying to buy something? Is he trying to just get information? So you can get marketers to decide what type of groups, how many groups you want to have, how you want to deal with customers. And AI will choose the blocks and the suggestions for those groups of people.
And we can adapt the different blocks, the sequence of the blocks, and media. You could also do media. One of the things we consider is there was some model announced today or yesterday, the Nano Banana Light. So you could even generate images very fast on the fly, obviously not as fast as text, but that's also something that would be, I don't know, something marketing people would want, generated images. That depends a lot on the quality, if it's on brand.
And this site, in this example, we have a product site, and then we have guides, experiences, blogs, and the whole response of the LLM is grounded there. And there are comparisons. We can do comparisons between products that are tailored, and the product pages can be tailored for the user.
Okay. This is a bit of the stack. I'm not going to spend too much time here, but in the browser you have some layers. You have the browser where the signals get gathered from the user, and then we have the backend. We run some of these things in Google. So the backend is basically just calling the LLM and doing some reasoning using the RAG that is built on the site to do the generation. And you have, obviously, the vector database, the inference machinery, and Adobe Experience Manager is serving the pages and the static content.
So let me show you, because I think this is... We call this audience of one, because the idea in marketing, they always dream of being able to personalize things for each individual. So we call it audience of one. So I have this site. This is a site that is absolutely generated. The example site is a coffee machinery site. So I can go and read some stories, and I can go and look at some products. Let's go and look at some products. Let's go and click here. Okay. So I'm browsing around the site, and I have this debugging tool thing, which, let me just go here, I think.
Let's see. So down there are the signals that the browsing is giving us. I don't know if you can see it much, because I cannot see it much. So the user is bucketed into the exploring category. We have the pages that they have visited, and then we have how much time is spent on each page. All of this data is now available for the LLM. So if I go here, I already have a For You page that was generated for me based on my browsing. And you will notice that it's slightly different than everything else.
But if I go here and I run a query, like I want, I'm looking for a coffee machine to prepare coffee while camping, then you're going to see some things like the text is customized: camping shouldn't mean compromising on your whatever routine. So you're going to see things like the coffee tips for camping, the coffee machine that is being recommended, the Arco Viaggio, or the Nano, which are good for a camping trip, right?
So you saw how fast this was. I'm going to run it here, something similar that I have here. And I can run it in the debug mode here. And you will see, let's make this bigger. Total time, 1.64 seconds to generate the page. So this includes a round trip to the LLM. This is using Cerebras Gemma 4, so the Gemma model from Google running on Cerebras on their very fast chips. We get 2,300 tokens per second, which is not bad, I would say. And if I run it again, probably something like that, the LLM time is one second, and again, 2,200 tokens per second. So this is something that we only dreamed about before.
On this example site, we have some other options because we've been showing this to customers. So we have the ability to change the different models, temperature, tokens, and so on. And we can show and try the different models and see how they behave. Besides automatic tests with PromFu, then we can manually come and tweak things and see how that works.
And we also have of one labs. So we built this tool that generates an agentic site for any site we want. So if somebody wants to have a demo for a customer, come here and enter the URL. In less than an hour, you have an agentic site. I did this last week with the AI engineering site, and I got this site that is just a search box and a few things. Let me open it here, the full page. Yeah. Okay. So I could say, Europe AI conferences. So these suggestions are also AI generated, and I get a page that is more focused on these European conferences.
If I go back, I can search for anything the same way I did with the article. So as a specific... There was one that was generating a good comparison side to side. Let me see if this one... Okay, here, this one. I went on this generated page with a very good comparison. If I'm looking at two conferences and I need to decide, if I figure out that the user wants to do that, this is great because that gives them a side-by-side comparison on the fly.
Now, I think this is cool already, but then I have this idea that probably a bunch of people are talking about. Is the web the future still, and so on? Nobody knows. But we can also do something with this, with this audience of one, these generative sites. So imagine you have your personal assistant and you ask a query through, in this case, through Google, and you say, I want to buy, I don't remember what the query said, it was something like, I want to buy a machine, and I get this on my Google TV, right? This is absolutely personalized to my query. Okay. Okay. No, go back.
This is absolutely personalized to my query. So I'm there in my living room. I don't need a phone. I don't need a computer. I don't need anything. Just my voice and something that will show me something that is absolutely personalized to me. Okay. So that one. So what I was trying to show, and hopefully you remember from this session, is that this is now possible. It's only going to get better from here on. It's only going to get cheaper. It's only going to get faster. And you will be able to have huge personalization options for sites and for other things.
And you can do this with intent-driven personalization. So what is my user trying to do? What does my user want to buy? These sorts of questions. And you can assemble a page just for them. And you can also do this with multiple models. And eventually it's just going to be faster and faster, right? So that's it. Thank you for coming, and I hope you got the idea. Thanks. providers, and then we look at the accuracy, we look at the speed, but this is going to depend highly on what type of site, like how big is the site, how, I don't know, what different, what different area is the site targeting, what type of commerce it is, and so on. So we run this evaluation continuously.
We use PromFu. Anybody heard about PromFu? Okay, some people. So PromFu allows you to evaluate models and prompts against multiple models providers, and you can do local models and any of the features that we have to do with a bunch of open AI compatible providers, and a lot of them, basically. We look for two things. Why? Accuracy. That's typically what people look for. But also we want the speed because we don't want the site generation to take more than one or two seconds, right? Because people, this is already proven that people want the faster the site, the more conversions it generates,
or the better the experience it is for the user. What I mentioned is different sites may have different requirements, so you may have to run this evaluation of models depending on the site.
So we have 15 prompts for this example site, and we have at the top, you can see with Cerebras on the GEMA 4 model that was announced last week, we can get an average latency of 1.1 seconds generating a page. You can compare that to the second one, which is 4.6 seconds, right? So the difference is huge. And that's why we use Cerebras for this use case. And you can see that different providers, different models have different speeds. And here is a... Let me... I can show you the whole thing here. Not this one, this one, right? So at the bottom we have other... Sometimes maybe some of them may be good.
They don't need to be perfect, but they're good enough if they're fast enough. So that's going to be the kind of decisions that you need to make on whether the model is good enough for your use case or not. So we're looking... Yeah, we're looking... Yeah, average 1.1 seconds, and then the next ones are going from 4 seconds higher. And you don't need a huge LLM to do this sort of work because you are generating text, you are deciding where to put blogs and how to organize the website. You don't need lots of information for that. So this browsing and the queries are being recorded. So these are the metrics or the
the data we gather from the user. And this is filled into the LLM to personalize the site. And then in this example, we personalize the hero card, the products, the blog feeds, and the navigation based on the persona. Also, what are some of the buttons like our call to action navigation, you can also... We can also personalize those. We create and I'll show you the for you page, which is a recommendation. And this is an interesting one because this you could pre-generate, right? As the user browses your site, you gather the signals and you could keep generating this. So in this case, you wouldn't need such a big speed. But that's interesting
because it will be if a user wants to buy something, you could just say, okay, for you, I will recommend these three products or something like that. Yeah. And then they can see this recommendation. And if they go there, that could be prefetch for them. And obviously, you have to keep updating it as the user navigates around the site and so on. So that's also something to consider on the cost, cost of doing multiple generations, multiple LLM calls.
When the user runs a query, a dynamic personalized page is shown to them. When these queries are also grouped into personas or intent types, so what is this guy trying to do in the site? He's trying to buy something. He's trying to just get information. So you can get marketers to decide what type of groups, how many groups you want to have, how you want to deal with customers. And AI will choose the blogs and the suggestions for those groups of people.
And we can adapt, yes, the different blogs, the sequence of the blogs, and media. You could also do media. One of the things we consider is there was some model announced today or yesterday, the Nano Banana Light. So you could even generate images images very fast on the fly, obviously not as fast as tests, but that's also something that would be, I don't know if it's not something like marketing people would want to have generated images. That depends on the quality a lot, if it's on brand.
And this site, in this example, we have a product site, and then we have guides, experiences, blogs, and the whole response of the LLM is grounded there. And there's comparisons. We can do comparisons between products that are tailored, and the product pages can be tailored for the user. Okay. This is a bit of the stack. I'm not going to spend too much time here, but the browser, you have some layers. You have the browser where the signals get from the user, and then we have the backend. We can have the backend. We run some of these things in Google. Some of these are
the backend. So the backend is basically just calling the LLM and doing some reasoning using the rack that is built on the site to do the generation. And you have, obviously, you have to have the vector database, the inference machinery, and the Adobe Experience Manager is doing the serving the pages and the static content. So let me show you, because I think this is, so we call this audience of one, because the idea of a marketing, the, they always dream on being able to personalize things for each individual. So we call it, yeah, audience of one. So I have this site. This is a site that is absolutely generated.
Example site is a coffee machinery. So I can go and read some stories, and I can go and look at some products. Let's go and look at some products. Let's go and look at some products. Let's go and click here. Okay. So I'm browsing around the site, and I have this debugging tool thing, which, let me just go here, I think.
Let's see. So down there is the signals that the, that the browsing is giving us. So I don't know if you can see it much, because I cannot see it much. Then, so the user is bucketed into the exploring category. We have the pages that they have, have visited, and then we have how much time is spending on each page. All of this data is now available for the LLM. So if I go here, I already have a For You page that was generated for me, and, uh, based on my browser. And you will notice that it's slightly different than everything else. But if I go here and I run a query, uh, like I want, I'm looking for a coffee machine to, uh, prepare coffee while camping.
And then you're going to see some things like the text is customized camping. And then you're going to see some things like the text is customized camping shouldn't mean compromising on your, uh, whatever routine. So, you're going to see some things like the coffee tips for camping, uh, the coffee tips for camping, uh, machine, uh, that are being recommended, the Arco Viaggio, uh, and, uh, or the nano, which are, um, good for, um, for, um, for the, um, for the, um, for a camping trip, right? So you saw how fast this was. I'm going to run it here, uh, something similar that I have here. And I can run it on the debug mode here. And you will see, let's make this bigger.
Total time, 164 seconds to generate the page. So this includes a round trip to the LLM. This is using Cerebras Gemma 4. So the Gemma model from Google running on Cerebras on their, uh, very fast chips, uh, we get 2,300 tokens per second, which is not bad, I would say. And if I run it again, uh, probably something like that, uh, the LLM time is one second and again, 2,200 tokens per second. So this is something that we only dreamed about before. On the, on this side, example side, we have some other options. Uh, so because we've, we've been showing this to customers, so we have the ability to change the different models, temperature,
temperature tokens, and so on. And we can, uh, we can show and try the different models and see how they behave. Besides automatic tests with PromFu, then we can manually come and tweak things and see, and see how that, how that works. And, uh, we also have, uh, of one labs. So we have, we build this tool that generates an agentic site for any site we want. So if somebody wants to have a, uh, demo for a customer, come here and enter the URL in less than an hour, you have an agentic site. I did this last week with the AI engineering site, and I got this site that is just a search box and a few things.
Uh, let me open it here. The full page. Yeah. Okay. So I could say, uh, Europe AI conferences. So these suggestions are also AI generated and I get a page that is, uh, more focused on, it should be more focused on, on the, on this European conferences. If I go back, did I go, I can search for anything the same way I did with, with the article. So I, as a Pacific, there was someone that was generating a good comparison side to side. Let me see if this one. Okay. Here, this one, I went on this generated, uh, page with, uh, very good comparison.
If I'm looking at two conferences and I need to decide if I figure out that the user wants to do that, this is great because that gives them a side by side comparison on the fly.
Now, this, this is, I think this is cool already, but then we have, uh, I have this idea that probably the, um, a bunch of people are, we are talking about is the web there, is, is the web, the future is still and so on. Nobody knows, but we can also do something with this, uh, with this audience of one, these generative sites. So imagine you have, uh, you have your personal assistant and you ask a query through, in this case, through Google and you say, I want to buy, I don't remember what the query said, it was something like, I want to buy, uh, a machine and I get this on my Google TV. Right? So this is absolutely personalized to my query. Okay. Okay. No, go back.
This is absolutely personalized to my query. So I'm there in my living room. I don't need a phone. I don't need a computer. I don't need anything. Just my voice and something that will, uh, kind of show me something that is absolutely personalized to, to me.
Okay. Okay. So that one. So, what I was trying to show, and hopefully you remember from this session, is that this is now possible. It's only going to get better from here on. It's only going to get cheaper. It's only going to get faster. And you will be able to have, uh, huge personalization options for sites and for other things. And you can do this, uh, with intent driven. So what is the, what is my user trying to do? What does my user want to buy? These sort of questions. And you can, uh, assemble a page just for them. And you can also do this with, uh, multiple models. And, and eventually it's just going to be faster and faster. Right? So that's it.
Um, thank you for coming. And I hope you, you got the idea. Thanks.