AI Engineer

Training Taste — Thais Castello Branco, Taste Labs

1827 summary words 8 min summary Watch video

Start with the signal

8 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: AI-generated design quality will not be fixed through stronger base models alone; applications need measurable, inference-time systems for context, brand adherence, creativity, and verification to prevent generic output.
  • Why it matters: This is a concrete quality-control architecture for creative agents: decompose subjective quality into measurable attributes, use targeted classifiers and retrieved constraints, then block or revise outputs that fail contextual and brand-fit checks.
  • Best use: Use it as product and evaluation input for any agent that generates customer-facing artifacts, especially presentations, websites, marketing assets, or interfaces.

Executive Summary

Thais Castello Branco argues that “AI slop” is not merely a subjective complaint but a production failure that can be operationalized. In her framing, slop has three recurring properties: repetition, lack of fit to the user or situation, and low apparent intent. As generation becomes nearly free, the scarce and valuable capability shifts from production to judgment—discerning what is appropriate, distinctive, and faithful to a particular context.

Taste Labs approaches the problem at two layers. At the frontier-model layer, it helps labs evaluate model failures and build post-training data or environments to address them. At the application layer, where off-the-shelf models tend toward stylistic averages, it focuses on supplying context, enforcing preferences, producing controlled novelty, and verifying outputs before they ship.

The company’s research claim is that slop can be detected quantitatively. After analyzing more than two million websites across ten years and comparing human-made with synthetically generated sites, Taste Labs extracted structured design features—such as color, typography, layout, and audience—and trained narrow “probes” to identify them. Combined probe signals reportedly predicted slop very well and outperformed common LLM-as-judge approaches.

Its product direction is a practical control plane for creative output. A proposed creativity API would create deliberate, domain-appropriate divergence rather than relying on higher model temperature. Its first public product, a Brand API, converts a brand URL into structured constraints that an agent can follow and that a verifier can score. The central operational lesson is to replace one-shot creative generation with retrieval, constraints, targeted evaluators, and explicit gates.

Key Takeaways

  • Claim: Creative-agent quality should be treated as a judgment and verification problem, not simply a model-capability problem. | Evidence: Taste Labs distinguishes deterministic or near-consensus design components—color palette selection, contrast, and alignment—from aesthetic choices where experts naturally disagree; it works both with frontier labs on post-training/environments and with applications on context, judgment, and verification. | Implication: For customer-facing generative systems, define and test the objective subproblems separately while using context-specific criteria and human or curated data where consensus is weak. | Caveat: The speaker does not claim subjective quality can be reduced entirely to one universal score; some aesthetic disagreement remains inherent.
  • Claim: Slop is usefully characterized by repetition, poor contextual fit, and low intent. | Evidence: Castello Branco contrasts a pet-shop website and a finance-firm website that converge on the same visual design: even if each is superficially acceptable, the shared pattern signals that the system did not understand their differing contexts. She attributes low intent partly to quick one-shot prompting and to systems that fail to elicit the user’s real goal. | Implication: Quality evaluation should include diversity across requests, appropriateness to audience and use case, and whether the system has captured enough user intent—not just visual polish or task completion.
  • Claim: Targeted feature classifiers can detect design slop more effectively than asking a general LLM to judge quality. | Evidence: Taste Labs analyzed over 2 million websites from the past decade, synthetically generated comparison sites, mined structured features including colors, typography, layout, and audience, and trained individual “probes” or baby classifiers. It reports that combinations of these probes predicted slop with high accuracy and performed better than most LLM-as-judge methods. | Implication: Use specialized, composable evaluators for observable failure modes instead of relying exclusively on a monolithic model-based quality judge; require validation before treating the reported result as production-grade evidence. | Caveat: No benchmark definition, quantitative accuracy, dataset composition, or independent validation is provided in the talk.
  • Claim: Inference-time controls are at least as important as model-layer improvements because user context and intent are resolved during the interaction. | Evidence: The speaker argues that better base models are necessary but insufficient: applications still use off-the-shelf models that collapse toward average styles, while the end-user back-and-forth that establishes intent happens at inference time. | Implication: Agent architecture should reserve first-class stages for brief construction, context retrieval, constraint application, output evaluation, and revision rather than treating the base model as the sole quality lever.
  • Claim: Creativity should be controlled divergence from domain norms, not random variation. | Evidence: Taste Labs’ proposed “creativity API” is described as an inspiration system that makes agents generate outputs outside the generic mean. Castello Branco rejects simply increasing sampling temperature: a strong pitch deck should preserve category expectations while intentionally breaking selected rules. | Implication: For creative workflows, explicitly encode which conventions are mandatory, which dimensions may vary, and how much divergence is appropriate; randomness alone will increase incoherence as often as originality. | Caveat: The talk provides no implementation details or evaluation results for the creativity API.
  • Claim: Brand extraction and adherence offer a high-leverage near-term way to raise generative design quality. | Evidence: Taste Labs’ Brand API, in beta with design partners, takes a brand URL and extracts structured components for an agent to follow and for a system to judge against. In its demonstration, Claude Design’s default slide deck for the General Intelligence Company of New York was less faithful than a deck generated with the extracted brand information. | Implication: Organizations should treat brand systems as retrievable, machine-readable specifications and use adherence scoring as a release gate for generated collateral. | Caveat: The example is qualitative and does not establish how robust brand extraction is across sites, assets, ambiguous brands, or legal usage contexts.
  • Claim: For users without established brands, retrieval of coherent prebuilt style systems may outperform generating a style from scratch. | Evidence: Taste Labs is creating an index of pre-created brand systems so a user seeking, for example, a “dreamy” look can retrieve a cohesive system rather than induce a style ad hoc at generation time. | Implication: A style-library or reference-retrieval layer can make consumer creative tools more reliable, while reducing generic model defaults and avoiding the need for users to articulate a full design brief.

Detailed Brief

The measurement thesis behind Taste Labs

  • Claims: The speaker believes subjective domains become more tractable when decomposed into components with different levels of objectivity rather than evaluated as a single fuzzy concept.; Internet design was already becoming more homogeneous before generative AI, likely because trends spread quickly; AI accelerates and generalizes that convergence across unrelated contexts.; The initial goal is not to recreate peak human craft or settle the nature of taste, but to raise a currently low baseline of generated quality.
  • Evidence: The research spans approximately ten years of web design history and compares historical sites with synthetically generated websites.; The design-feature schema mentioned includes colors, typography, layout, and audience.
  • Caveats: Homogeneity is not inherently a quality failure: shared conventions can improve usability. The relevant failure is unthinking convergence that ignores the specific brand, audience, and task.; The transcript does not disclose methodology for labeling human versus AI sites or defining ground-truth slop.
  • Implications: A useful quality program should distinguish desirable standardization from inappropriate sameness.; Evaluation datasets should cover multiple verticals and contexts so that a model is not rewarded for applying one fashionable visual pattern everywhere.

Product pattern: structured generation plus structured verification

  • Claims: Taste Labs frames its Brand API as both a generation aid and a judging mechanism: the same extracted components that guide an agent can be used to identify where its output departs from the source brand.; The company sees existing brand work as high-value latent training and control data that organizations have already paid designers to create.; Its proposed systems divide the anti-slop problem into distinct interventions: creative divergence for repetition, brand conditioning for fit, and classifiers or gates for intent and quality control.
  • Evidence: The Brand API is described as the first public release and is in beta with unnamed design partners.; The demonstration compares an original General Intelligence Company of New York brand, Claude Design’s default reproduction, and a higher-fidelity result produced with extracted brand data.
  • Caveats: The presentation is a founder talk and product narrative, so claims about quality improvement and classifier performance should be treated as directional pending published metrics and customer results.
  • Implications: The strongest implementation pattern is not a better prompt alone but a closed loop: retrieve or derive specifications, generate, score against those specifications, and regenerate or route exceptions to review.

Notable Concepts & Terms

  • AI slop: The speaker’s label for generic generative output, operationalized through repetition, poor fit to context, and low intent.
  • Fit: Whether an output feels correct for a particular person, brand, moment, audience, and use case; it is the key counterweight to generic cross-context templates.
  • Inference-time quality control: Application-layer context gathering, conditioning, evaluation, and revision performed when a user interacts with an agent, rather than only through model training.
  • Probes: Taste Labs’ narrow classifiers for individual structured design characteristics; their combined signals are used to predict slop.
  • Creativity API: A proposed system for deliberate, domain-aware novelty that preserves key conventions while selectively departing from norms.
  • Brand API: Taste Labs’ beta product that extracts a brand URL into structured, agent-usable components and enables adherence checks.
  • LLM-as-judge: The general approach of having an LLM rate output quality; Taste Labs claims its targeted probe approach outperformed it for slop detection.
  • Out of distribution: In this talk, purposeful departure from the model’s generic average output—not arbitrary randomness—to create output with distinctiveness and craft.

Operator Notes / Why Ken Should Care

  • Add a creative-output release gate to any agent that produces external assets: score brand adherence, context fit, repetition against recent outputs, and basic design constraints before delivery.
  • Convert existing brand guidelines, websites, component libraries, approved collateral, and style references into a versioned machine-readable context package rather than relying on prompts such as “match our brand.”
  • For new or consumer users, offer curated style-system retrieval and a short intent-elicitation flow before generation; avoid one-shot generation as the default interaction.
  • Separate evaluators by failure mode—contrast, layout, typography, duplication, brand adherence, audience fit—and combine their scores; benchmark this stack against an LLM judge on real acceptance outcomes.
  • Demand published methodology and held-out performance data before adopting any slop-detection vendor or using its scores as an automated rejection criterion.

Source/Metadata

  • Title: Training Taste — Thais Castello Branco, Taste Labs
  • Transcript words: 6154
  • Duration seconds: 906
  • Timestamp note: No timestamps or chapters were present. The supplied transcript contains a substantial duplicated second pass of the talk.
Full transcript 3123 words · 27 min read
0:13

Okay, amazing. It's great to meet everyone. I'm Thais. I'm the founder of Taste Labs. For those of you who don't know us, we came out of Stout a few weeks ago, and our whole mission is how do we end AI Slop? It's my personal enemy. And so we really believe that to solve this problem of Slop, we have to decode subjective domains. There's been so much effort being put into getting models and agents amazing at things like coding and math, and it's time that we put all that same effort into making them great at things like design and writing. And so design is this first pillar that we're starting with, and it's been incredibly exciting. We work primarily in two ways. So we work a lot with the Frontier Labs on how do we evaluate their models, understand where they're breaking, understand what could be better about them, and then construct the right either post training data or our environments to fix that problem. And part of this is how do you take something as fuzzy and large as design and break it down to a level that you can identify what is best solved through each method? What are elements of design that are almost, once you boil down the problem, become so specific that they almost become deterministic. So for example, if you're trying to train a model to be good at selecting color palettes or have contrast or alignment, those are things that if you define the problem in the context in a specific enough way, you can get to an answer that's pretty objective or that at least most experts would agree to. But maybe other things like aesthetics, you naturally will see this expert disagreement, and so then you want to lean on things that are closer to data. So anyway, we spend a lot of time thinking about all those problems. But on the other side is also, without even touching the model layer, right? How do we actually help agents and app layer companies to produce better things? And there's a lot that goes into that, right? You have these different sets of problems at the application layer because you're using an off-the-shelf model that tends to collapse in terms of style, tends to collapse to the mean. So how do we force that creativity back to the system? How do we avoid these patterns of slop, which we'll talk about a lot today? How do you understand user preferences or brand preferences so that you can maintain adherence to that style? So there's lots of things that actually need to be solved as context or judgment or verification at the app layer, which is why we kind of work across both. But maybe I'll start with more of a philosophical question of how do you define something that is great? Like how do you define greatness? And for something like math, it's easier, right? Because there's one objective answer and great is the same as correct. But then for something like writing or design, it's much harder, right? Like how do you define what's a great tweet or what's a great art piece or what's a great website? I don't know what's the last time that you interacted with a poem or walked into a coffee shop and for some reason it hit different and it felt very special. But probably it's a combination of things that it felt very unique. It felt almost a little different. It called your attention. It felt like it was made with a lot of care and attention to detail and craft and it almost had this sense of authenticity. And I think that's a lot of what AI is missing today. It's how do we take things that are not necessarily average, right? How do we produce things that are purposely out of distribution? And slop is the opposite of that, right? I think it is hard to define what is great sometimes, but I think it's pretty easy to define what is slop in the sense that most people would agree. I think the sense of repetition, of soullessness, is something that all of us feel right now when using AI. And I think it's quite magical, by the way, that AI has gotten to a point that any human on the planet that is not even a designer, that is not an engineer, can click a button and suddenly make an entire PowerPoint or make a website or make a web app. That's pretty cool. But it comes with consequences, right? It comes with consequences of suddenly now the cost of generation is basically going to zero. But the average person hasn't necessarily honed their taste. Think about the amount of effort and work that a designer puts in throughout their life to build up their taste, right? There's all this process of getting exposed to many things and learning to spot patterns and learning to develop a point of view and doing things in a courageous way that maybe are a little bit against the norm. Learning what not to do and how to have restraint and that's very hard. The average person doesn't necessarily have the time or the skills to go and develop taste in everything, in design. And so I think it would be a bad case scenario for us to just say, okay, the way to fix slop is for everyone to have taste, because I don't think that's necessarily realistic. I think how do we understand this better so that we can make, even for the average person, the ability to create something great and to understand maybe their own taste easier. So that's a lot of what we're focusing on. So yeah, I think this phenomenon of slop, by the way, is not new. If you were on the internet as social media emerged, you probably saw a lot of slop before that. But I do think that AI has been this accelerating force, right? Instead of being able to create things very easily with a click of a button and the thoughtlessness around it. And there's these three characteristics that I would say repeat in slop. So repetition, so you start seeing the same thing many, many, many times. The second is lack of fit, which I actually think is very related. So fit is this ability for something to feel correct for a specific context, right? For a specific moment in time, for a specific person. But suddenly if you have repetition, and let's say one person asked for a website for their pet shop and the other one asked for a website for their finance firm, and somehow those designs converge and look the same, that's quite odd, right? If you were actually crafting that with care, you wouldn't converge necessarily on those things. And so this lack of fit and lack of understanding of context is a huge problem that leads to slop. And the third is maybe low intent, which is probably a mix of you're going to have a bunch of people prompting really quickly and maybe just wanting to one-shot something. But I think there's actually this intent interpretation piece that's missing in the systems that we're building. How can you help your user, right? How can you help them better understand the intent that they have so that you can add more color and add more context onto what you're trying to create? Okay, and I'm a big believer, by the way, that in order to fix something, first you have to measure it and you first have to understand it. I think that's exactly why we're so focused on how do we turn these domains into something a bit more verifiable, so that we can attach a measure to it. So you'll go on a little bit of a research journey with me here now, but we basically wanted to figure out can we measure slop? Can we actually measure this quantitatively and spot this? And what does that look like? So we analyzed over 2 million websites from the past 10 years, kind of way back machine style to try to understand all the trends across design. How is the internet changing? How is design changing over time? And two things were interesting. And we also, by the way, then kind of synthetically generated a set of design websites so we could compare. How do human-made sites compare to AI generated ones? And there were a few things that were interesting. So one was that you already kind of saw a bit of a collapse on the internet before even AI. So you saw the internet becoming more homogenous, using more similar color palettes, using more similar layouts, which is probably a function of trends spreading more quickly. But with AI, I think you saw this repetition happening a lot more and being almost more identified regardless of context. So even in completely different buckets, you saw patterns that were very similar. So we built this—I call this probes—but basically we did two things. So we did this pattern mining on all this data to understand what are features that we can extract from all these sites? What are all these characteristics that we can make more objective, right? Colors, typography, layout, audience. How can we distill this down into things that become almost structured? And then how do we train up these probes? So think of these as baby classifiers. How do we train the ability to spot this one characteristic? And for all these slop sites, we started identifying what are the probes that basically mean this site is very likely to be AI slop. And especially when you start combining them and you see the frequency of multiple of these happening at once, it became very likely that you could actually measure and predict slop. And we saw a super high ability to do that prediction, which was really cool to see. This performed better, by the way, than most LLM as a judge methods of asking an LLM to judge if that is great human quality versus AI generated slop. So that was pretty cool to see. And I think it shows this pattern that we see in AI is an actual quantitative thing that we can see in slop, which I find really cool. But obviously we don't want to stop there, right? We don't want to just measure slop. We want to also solve it. And so there's a few—I mentioned this before—but as the cost of production basically goes to zero, I think the thing that becomes expensive and matters more than ever is judgment. I don't even want to use the word taste here, it's judgment. I think it's this ability to discern what's right, this ability to break down a problem so that you can actually understand it and create solutions for it. And so yes, there's this side of judgment that is human judgment that I actually think is more valuable than ever. But there's also this side of how do we build the right tools and systems to fix pieces of this problem, right? So how do we fight slop, my enemy?

0:19

And by the way, I think there's a lot of conversation going around how do you fight slop at the model layer? How do we make models better? How do we make models have a higher bar? Which, don't get me wrong, it has to be solved and we're working very hard to solve that too. But I actually think this problem of inference time is equally, if not even more important. Because that's actually when you interact with the end user. And this back and forth of how do you understand this context and intent happens at the moment of inference time. So I don't think that we can ignore and just make models better and not solve this or the rise of slop will keep existing. So maybe breaking down a few of those pieces and a few of the ways that we've thought about solving this or a few solutions that we've built to solve this. But I think, for example, for something like repetition, one of the things that we're working on is—I've nicknamed it, I don't know if that's going to be the official name—but the creativity API. How can we create a system that almost becomes an inspiration machine for your agent? So that it can produce something that's actually out of distribution instead of something that is in that same average and mean that we're seeing happen with the slop sites. So this is one of the ways that practically, if we can intentionally produce something that's out of distribution, you can improve this overall quality. And by the way, I don't think that this can be something just like randomness. It's not just about turning up a temperature of a model and fingers crossed, hoping for the best. I think it's much more like how do we understand even what are rules or expectations in specific domains? Like let's say you asked for a slide deck for a picture of your startup. What does a good pitch deck look like? And then how do you almost intentionally break rules to create things that are more creative, right? Because usually creativity isn't randomness, isn't doing something that completely feels off for that situation. It's like you intentionally maybe diverge on a couple of things while maintaining adherence to expectations of that category for others.

0:23

So that's one of the things we're working on. The second one on this problem of fit, I think it's interesting, but brands, as probably a lot of you who are designers know, take so much effort to create great brands. Great brands are the work of dozens of designers putting in a lot of craft and thought and care. And so we've almost already pre-done the work of defining what is great for that specific company and then we're not using it well. So this brand adherence actually I think is a huge problem and one of the things that can very more easily raise that bar of quality. So I'll touch on an example on this one specifically and then same with intent and judgment. I think the baby classifiers was a good example. How we can actually use this to even become a gate for slop and not let your agent ship slop. But so the brand API is the first product that we're releasing to the public. This is already in beta testing with a bunch of our design partners and essentially what it does is it can take a brand URL and extract this into very specific components that are good for an agent to follow. So basically how do we turn something as fuzzy as a brand into something so structured that it becomes easy for your agent to follow that but also for you to judge against it, right? Because I think the piece that we can't forget here is this judgment and verification. So yes this goes and helps your agent to produce something better, but how can we also add a way for you to judge okay is the agent actually staying on track? Is it actually performing well to adhere to this brand or where is it failing? So this is the first flow I would say that we are seeing that is really helping to improve quality. And what's cool is of course we're talking here about an example of a brand that already exists but let's say you have an agent or you have an app and the person that is using your app actually doesn't have a brand, let's say they're an average consumer. Can we actually—one of the things that we're creating is basically a repository, an index of brands, of pre-created brand systems so that if they want something that feels dreamy why not retrieve a dreamy brand system that already has been thought out to be cohesive instead of doing a generative approach the moment of which might end up not so great or might end up again in those pillars of slop.

0:31

And I want to show you a real example of this in action. So there's this company that I think is awesome called the General Intelligence Company of New York. They have a sick website you guys should check it out. But basically if you ask Claude Design to create a slide deck in their branding, the middle one is basically what it comes up with. So the one on the left is the original brand. This is the default and if you use this extraction actually in the process it creates something that's way more high fidelity with the original and even in the details I would say it feels right. So this is just to show an example of it in action. But yeah, I think all of us would agree that human taste and the peak of human craft is always going to be deeply valuable and that right now I think the challenge is we are almost not even earning the right to debate this. How can we have models reach this pinnacle of taste? I don't think it's about that at all. It's how do we first just raise the bar? The bar is kind of really on the ground and so I think all of this work that we're putting into how do we decompose a problem and how do we measure it is exactly so that we can at least improve this bar of quality and I think we have to start with that.

0:36

That's it. Thank you very much for the time. This is awesome. effort into making them great at things like design and writing. And so design is this first pillar that we're starting with, and it's been incredibly exciting. We work primarily in two ways. So we work a lot with the Frontier Labs on how do we evaluate their models, understand where they're breaking, understand what could be better about them, and then construct the right either post training data or our environments to basically fix that problem. And part of this is like how do you take something as fuzzy and large as design and break it down to a level that you can identify

1:10

what is best solved through each method? What are elements of design that are almost like once you kind of boil down the problem, become so specific that they almost become deterministic. So for example, if you're trying to train a model to be good at selecting color palettes or have contrast or alignment, those are things that if you define the problem in the context in a specific enough way, you can get to an answer that's like pretty objective or that at least most experts would agree to. But maybe other things like aesthetics, you naturally will see this expert disagreement,

1:38

and so then you want to lean on to things that are closer to data. So anyway, we spend a lot of time thinking about all those problems. But on the other side is also without even touching the model layer, right? How do we actually help agents and app layer companies to produce better things? And there's a lot that goes into that, right? You have these different sets of problems at the application layer because you're using an off-the-shelf model that tends to collapse in terms of style, tends to collapse to the mean. So how do we force that creativity back to the system? How do we avoid

2:05

these patterns of slop, which we'll talk about a lot today? How do you understand like user preferences or brand preferences preference so that you can maintain adherence to that style? So there's lots of things that actually need to be solved as context or judgment or verification at the app layer, which is why we kind of work across both. But maybe I'll start with more of a philosophical question of like how do you define something that is great? Like how do you define greatness? And for something like math, it's easier, right? Because there's kind of one objective answer and great is the same as

2:39

correct. But then for something like writing or design, it's much harder, right? Like how do you define what's like a great tweet or what's a great art piece or what's a great website? I don't know what's the last time that you interacted with a poem or walked into a coffee shop and for some reason it kind of like hit different and it felt very special. But probably it's a combination of things that it felt very unique. It felt almost a little different. It kind of called your attention. It felt like it was made with a lot of care and attention to detail and craft and it almost had this sense of like

3:10

authenticity. And I think that's a lot of what AI is missing today. It's like how do we take things that are not necessarily average, right? How do we produce things that are purposely like out of distribution? And slop is kind of the opposite of that, right? I think it is hard to define what is great sometimes, but I think it's pretty easy to define what is slop in the sense that most people would agree. I think the sense of like repetition, of kind of soullessness, is something that all of us feel right now when using AI. And I think it's quite magical, by the way, that AI has gotten to a point that any human on the

3:38

planet that is not even a designer, that is not an engineer, can click a button and suddenly make an entire PowerPoint or make a website or make a web app. That's pretty cool. But it comes with consequences, right? It comes with consequences of suddenly now the cost of generation is basically going to zero. But the average person hasn't necessarily honed their taste. Like, think about the amount of effort and work that a designer puts in throughout their life to like build up their taste, right? Like there's all this process of like getting exposed to many things and learning to like spot

4:09

patterns and learning to develop a point of view and like kind of do things in a courageous way that maybe are a little bit against the norm. Learning what not to do and how to like have restraint and that's very hard. Like the average person doesn't necessarily have the time or the skills to go and develop taste in everything, let's say in design. And so I think it would be a bad case scenario for us to just like be like, okay, the way to fix slop is for everyone to have taste, because I don't think that's necessarily realistic. I think how do we how can we understand this better so that we can make

4:38

even for the average person the ability to create something great and to understand maybe their own taste easy, more easy. So that's that's a lot of what we're we're focusing on. So yeah, I think this phenomenon of slop by the way is not new. If you were in the internet as social media emerged, you probably saw a lot of slop before that. But I do think that AI has been this kind of like accelerating force, right? Instead of like being able to create things very easily with a click of a button and the like thoughtlessness around it. And there's kind of these three characteristics that I would say repeat

5:09

and slop. So a repetition, so you start seeing the same thing many, many, many times. The second is lack of fit, which I actually think is very related. So fit is kind of this ability for something to feel correct for specific context, right? For specific moment in time, for a specific person. But suddenly if you have repetition, and let's say one person asked for a website for their pet shop and the other one asked for a website for their finance firm, and somehow those designs converge and look the same, that's quite odd, right? Like if that wasn't, if you were actually crafting that with care, that wouldn't,

5:41

you wouldn't converge necessarily on those things. And so this lack of fit and lack of understanding of context is actually a huge problem that like leads to slop. And the third is maybe low intent, which is probably a mix of, yeah, you're gonna have a bunch of people prompting really quickly and maybe just wanting to one-shot something. But I think there's actually this like intent interpretation piece that's missing in the systems that we're building. Like how can you help your user, right? Like how can you help them better understand the intent that they have so that you can add more color and add

6:08

more context onto what you're trying to create? Okay, and I'm a big believer, by the way, that you, in order to fix something, first have to measure it and you first have to understand it. I think that's exactly why we're so focused on like how do we turn these domains into something a bit more verifiable, so that we can attach a measure to it. So you'll go on a little bit of a research journey with me here now, but we basically wanted to figure out can we measure slop? Like can we actually measure this quantitatively and spot this? And what does that like look like? So we analyzed over 2 million websites

6:43

from the past like 10 years, kind of like way back machine style to try to understand all the trends across like design. How is the internet changing? How is like design changing over time? And two things were interesting. And we also, by the way, then kind of synthetically generated a set of design websites so we could kind of like compare. Like how does human-made sites compare to AI generated ones? And there were a few things that were interesting. So one was that you already kind of saw a bit of like a collapse on the internet before even AI. So you saw kind of the internet becoming more homogenous,

7:18

using more similar color palettes, using more similar layouts, which is probably a function of more, I would say there's a kind of trend spreading more quickly, let's say. But with AI, I think you saw this repetition happening a lot more and being almost more like identified kind of regardless of context. So even in completely different buckets, you saw patterns that were very similar. So we built this, I call this probes, but basically we did two things. So we did this like pattern mining on all this data to understand like what are features that we can extract from all these sites? What are all these

7:48

characteristics that we can make more objective, right? Colors, typography, layout, audience. How can we like distill this down into things that become almost like structured? And then how do we train up these like probes? So think of these as like baby classifiers. Like how do we train the ability to spot this one characteristic? And for all these slop sites, we identified, we started identifying like what are the probes that basically mean this site is very likely to be AI slop. And especially when you start combining them and you see the frequency of multiple of these happening at once, it became

8:20

very likely that you could actually like measure and predict slop. And we saw a super high, basically, ability to do that prediction, which was really cool to see. This performed better, by the way, than like most LLM as a judge methods of like asking an LLM to like judge if that is a great human quality versus like AI generated slop. So that was pretty cool to see. And I think it kind of shows this pattern that we see in AI really being an actual quantitative thing that we can see in slop, which I find really cool. But obviously we don't want to stop there, right? We don't want to just

8:51

measure slop. We want to also solve it. And so there's a few, I think I mentioned this before, but like the, as the cost of production basically goes to zero, I think the thing that becomes expensive and matters more than ever is judgment. I don't even want to use the word taste here, is judgment. I think it's this ability to discern what's right, is this ability to break down a problem so that you can actually understand it and create solutions for it. And so yes, there's this side of judgment that is human judgment that I actually think is more valuable than ever. But there's also this side of like, how do we build the right tools and systems to like

9:22

fix pieces of this problem, right? So yeah, how do we fight slop, my enemy? And by the way, I think there's a lot of conversation going around, how do you fight slop at the model layer? Like how do we make models better? How do we make models have a higher bar? Which, don't get me wrong, it has to be solved and we're working very hard to solve that too. But I actually think this problem of inference time is equally, if not even more important. Because that's actually when you interact with the end user. And this kind of back and forth of how do you understand this context and intent

9:55

happens at the moment of inference time. So I don't think that we can ignore and just make models better and not solve this so the rise stop will keep existing. So maybe breaking down a few of those pieces and kind of a few of the ways that we've thought about solving this or a few solutions that we built to solve this. But I think, for example, for something like repetition, one of the things that we're working on is, I've nicknamed it, I don't know if that's going to be the official name, but like the creativity API. How can we create a system that almost becomes an inspiration machine

10:21

for your agent? So that it can produce something that's actually out of distribution instead of something that is in that same average and kind of mean that we're seeing happen with like the slop sites. So this is one of the ways that practically, if we can intentionally produce something that's out of distribution, you can improve this like overall quality. And by the way, I don't think that this can be something just like randomness. It's not just about like turning up a temperature of a model and kind of fingers crossed, hoping for the best. I think it's much more like how do we understand

10:50

even like what are rules or expectations in specific domains? Like let's say that you asked for a slide deck for for the picture of your startup. Like what is it what does a good pitch deck look like? And then how do you almost like intentionally break rules to create things that are more creative, right? Because usually creativity isn't like randomness, isn't doing something that completely feels off for that situation. It's like you intentionally maybe diverge on a couple of things while maintaining kind of um adherence to to expectations of that category, let's say for others.

11:20

So that's one of the things we're working on. The second one on this problem of fit, I think um it's interesting, but brands as probably a lot of you who are designers know, take so much effort to create great brands. Like great brands are the work of dozens of designers uh putting in a lot of like craft and thought and care. Um and so we've almost like already pre-done the work of defining what is great for that specific company and then we're not using it well. So this like brand adherence actually think is a huge problem and one of the things that can very more easily let's say like raise that bar

11:51

of quality. So I'll touch on an example on this one specifically and then same with like intent and judgment. I think the baby classifiers was a good example. Um like how it how we can actually like use this to even become a gate for slop and not let your agent uh ship slop. But so the brand API is the first product that we're releasing to to the public. This is already in in beta testing with a bunch of uh our design partners and essentially what it does is it can take let's say a brand URL and extract this into like very specific components that are good for an agent to follow. So basically how do we turn

12:23

something as fuzzy as a brand into something so structured that it becomes easy to uh for your agent to follow that but also for you to judge against it right because I think the piece that we can't forget here is this judgment and verification. So yes this goes and helps your agent to produce something better uh but how can we also add a way for you to judge okay is the agent actually staying on track? Is it actually performing well to adhere to this brand or how is it failing or where is it failing? So this is the first flow I would say that we we are seeing that is really helping to improve

12:52

quality. Um and what's cool is of course we're talking here about an example of a brand that already exists but let's say you have an agent or you have an app and uh the person that is using your app actually doesn't have a brand let's say they're an average consumer can we actually one of the things that we're creating is basically like a repository like an index of brands uh of pre almost like pre-created brand systems so that if they want something that feels dreamy why not retrieve a dreamy brand system that already has been thought out to be cohesive instead of doing like a generative

13:20

approach the moment of that might end up not so great or might end up again in those pillars of slop. And I want to show you a real example of this in action so um there's this company that I think is awesome called the General Intelligence Company of New York they have a sick website you guys should check it out um but basically if you ask Claude Design to create a slide deck uh in their branding the the middle one is basically what it comes up with so the one on the left is is the original brand uh this is kind of the the default and if you kind of use this extraction actually in the process it

13:49

creates something that's way more high fidelity with the original um and that even like in the details I would say like feels right so this is just to show an example of it in in action um but yeah I think we I think all of us would agree that like human human taste and kind of the peak of human craft is always going to be like deeply valuable and that right now I think the challenge is we are almost even not earning the right to debate this like how can we have uh models like reach this like pinnacle of taste I don't think it's about that at all it's like how do we first just like raise the bar like the bar is

14:26

kind of really I would say on the ground and so I think all of this work that we're putting into like how do we decompose a problem and how do we measure it is exactly so that we can at least like improve this bar of quality and I think we have to start with that that's it uh thank you very much for for the time uh this is this is awesome thank you

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note