Open Reader

The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI

completed 14:09 Aug 21, 2026 Watch on YouTube

Current Status

completed

Video ID

7GMKdpLsxwU

RAG / Chat

Enabled
The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI
Description

Most people can spot a vibe coded app in two seconds and cannot say why. Hassan El Mghari names it: the purple gradient background, italics in the header, a scroll to explore prompt nobody asked for, all caps pills with wide letter spacing, too many emoji. He reckons you could list thirty such tells, and naming them is what lets you tell an agent to avoid them. That is most of what Hallmark does, the design skill he shipped six weeks ago to more than 10,000 users. It codifies the patterns as gates and hands the model a library of themes, on the through line that strong inspiration produces much better work. El Mghari leads developer experience at Together AI and is not a designer. He has shipped roughly ten apps a year for five years, a few reaching millions of users, and credits design for that reach. He starts an app in Codex or Claude Code, then iterates with a smaller open source model, and made the case by putting up two landing pages, one from GLM 5.2 and one from Opus 4.8, and asking the room to guess which was which. Almost nobody could. The rest is practical: keep a vault of screenshots you admire and paste them in, the one habit he says to take away; record a voice note and let the prompt run three paragraphs; send one or two features per prompt instead of seven; write accumulated preferences into a skill file or AGENTS.md. Whatever the agent hands back is a base, never the finished app. Speaker info: - https://x.com/nutlope - https://www.linkedin.com/in/nutlope/ - https://nutlope.com Timestamps: 0:00 - Ten apps a year, and why design is the edge 1:41 - The apps: logos, comics, subtitles, a cloud agent 3:08 - The tells that give away a vibe coded app 4:11 - Hallmark: slop gates and a theme library 6:53 - Iterating with a smaller open source model 7:30 - Guess which page came from which model 10:50 - Give the agent references and screenshots 12:18 - One or two features per prompt

Summary

Generated by gpt-5.6-terra

At-a-Glance

  • Verdict: Skim
  • Core thesis: AI-generated app UIs become materially more credible when builders codify anti-slop preferences, supply strong visual references, use detailed staged prompts, and treat the first agent output as an iteration base rather than a finished product.
  • Why it matters: As AI coding makes implementation cheap, generic visual patterns increasingly signal low quality; design direction and iteration are becoming a practical differentiation layer for agent-built products.
  • Best use: Use this as a concise operating checklist for improving the presentation quality of agent-generated prototypes, especially by building reusable design-context files and an inspiration library.

Executive Summary

Hassan El Mghari argues that the major failure mode of vibe-coded apps is not that they fail to function, but that they visibly look machine-generated: purple gradients, italic headlines, oversized all-caps pills, gradient logos, arbitrary emoji, stock decorative graphics, and uneven spacing. His practical claim is that spending an additional 10–20% of effort on UI and UX creates disproportionate product differentiation, even for simple one-page AI applications.

His solution is to make design taste operational rather than intuitive. He built Hallmark, a design skill that encodes both prohibited “AI slop” patterns and a set of visual themes. More generally, he recommends maintaining an agents.md-style preference file that records recurring corrections—such as logo treatment, layout choices, and elements to avoid—so agents receive persistent design guidance instead of rediscovering it on every project.

The highest-leverage tactic is reference-driven generation. El Mghari saves screenshots of products he likes in an inspiration vault, then provides multiple references and explicit instructions such as combining aspects of Duolingo with other applications. He says detailed, two-to-three-paragraph prompts, sometimes dictated as voice notes, produce better outcomes than sparse prompts; however, complex builds should still be divided into small feature-level tasks rather than bundled into one enormous request.

He also makes a cost-and-speed argument for model routing: use strong models to establish a base, then use sufficiently capable smaller models for frequent UI iteration. He presents GLM 5.2 as an example that, in his comparisons, generated landing pages competitive with Opus 4.8 while being faster and roughly one-fifth the cost. The talk is useful as a tactical design-conditioning framework, though its evidence is primarily personal demonstrations rather than controlled evaluation.

Key Takeaways

  • Claim: Generic AI UI patterns have become recognizable quality signals, so eliminating them can create a meaningful competitive advantage for otherwise simple AI apps. | Evidence: El Mghari identifies recurring tells: purple-gradient backgrounds, italicized headers, “scroll to explore” prompts, spaced all-caps pills, gradient logos, excessive emojis, arbitrary graphics, and spacing or padding defects. He says he has built roughly 10 apps annually for five years and attributes much of the adoption of several projects to design and UX. | Implication: Ken should treat recognizable AI-generated visual conventions as explicit quality-control failures in agent-built front ends, not merely subjective aesthetic preferences. | Caveat: Avoiding familiar AI patterns does not by itself produce excellent product design; the speaker frames it as a better starting point that still needs iteration.
  • Claim: Design taste can be translated into reusable agent instructions instead of relying on ad hoc prompting each time. | Evidence: Hallmark codifies “AI slop gates” that tell models what not to use, alongside a library of themes that supplies positive visual direction. El Mghari recommends accumulating comparable preferences in an agents.md or Markdown file as repeated UI issues emerge, such as bad logo generation or undesirable layouts. | Implication: A persistent design policy file can function as part of an agent control plane: it reduces repeated corrections and makes output style more consistent across projects and operators.
  • Claim: Visual references and screenshots are the speaker’s most important input for improving agent-generated UI. | Evidence: He says he no longer builds without supplying substantial inspiration, maintains an “inspiration vault” of apps and sites he likes, and asks models to synthesize multiple references—for example, a mix of Duolingo and other apps. He reports that outputs are consistently much better when screenshots are included. | Implication: For agent workflows that generate customer-facing interfaces, Ken should make curated reference sets a standard project artifact rather than expecting models to infer taste from a text-only specification. | Caveat: Reference-driven design can drift into imitation if teams do not separately specify product constraints, brand distinctions, and what should not be copied.
  • Claim: Prompt detail improves initial UI quality, but implementation should be decomposed into small, well-specified iteration tasks. | Evidence: El Mghari typically uses two-to-three-paragraph prompts and sometimes records one-to-three-minute voice notes describing the user, interaction model, components, visual direction, and inspirations. Yet when building multiple features, he recommends one or two features per prompt, often queuing separate follow-up tasks after a large initial app-generation request. | Implication: The appropriate agent workflow is a broad, context-rich initialization followed by bounded change requests—not repeated one-shot attempts or a single prompt containing every feature and refinement.
  • Claim: Smaller, fast models can be the better default for high-frequency UI iteration, provided they clear a quality threshold. | Evidence: In audience comparisons, he says a GLM 5.2 landing-page output was difficult to distinguish from an Opus 4.8 output; he states GLM 5.2 was faster and the Opus result cost five times more. He also cites Cursor’s Composer 2.5 as an example of a fast post-trained open-source model that feels effective for iterative work. | Implication: Ken can route expensive frontier models toward initial architecture or difficult reasoning, while assigning rapid visual refinements, state changes, and polish passes to lower-cost models with tight feedback loops. | Caveat: The comparison is anecdotal and presentation-based, not a standardized benchmark; his recommendation depends on the model being strong enough for the relevant design and coding task.
  • Claim: One-shot generation should be treated as a prototype, not a shippable endpoint. | Evidence: He shows an image-playground app initially generated in one prompt with visible AI tells, then describes reaching a substantially cleaner result through one or two follow-up prompts that improved the logo, loading states, animations, and spacing. Hallmark examples are similarly presented as stronger bases rather than final designs. | Implication: Product quality gates for agent-generated applications should require at least one intentional refinement pass focused on interaction states, visual hierarchy, branding, and layout consistency before publication.

Detailed Brief

Practical examples and scope of the speaker’s approach

  • Claims: El Mghari’s examples are predominantly lightweight, visually polished, one-page AI applications rather than complex enterprise software.; His objective is not elite design originality; it is making applications feel substantially more intentional and polished than unconditioned AI output.
  • Evidence: Examples include a logo-creator app used by about 85,000 people, a video-subtitle generator used by about 8,000 people, an AI chat application for open-source models, a landing-page variation generator, Make Comics, and a cloud coding agent that creates pull requests from a GitHub repository in a sandbox.; He says Hallmark had more than 10,000 users approximately a month and a half after launch.
  • Caveats: Usage figures are self-reported examples and do not establish that design alone caused adoption.; The methods are most directly demonstrated on landing pages and smaller consumer-style apps, not dense operational software with accessibility, compliance, and complex workflow requirements.
  • Implications: The framework is immediately applicable to prototypes, demos, growth pages, and agent-built internal tools, but should be supplemented with product-design and accessibility review for production systems.

Notable Concepts & Terms

  • AI slop gates: A negative design specification: an explicit list of common AI-generated visual patterns that the model should avoid.
  • Hallmark: El Mghari’s design skill that combines anti-slop constraints with reusable visual themes to condition website generation.
  • agents.md: A persistent project instruction file for retaining learned UI preferences, recurring corrections, and style constraints across agent sessions.
  • Inspiration vault: A curated collection of screenshots and product references used as visual context for new application builds.
  • Reference-driven generation: Providing screenshots and named design inspirations so an agent has concrete visual targets instead of relying on abstract aesthetic language.
  • Iteration model routing: Using a lower-cost, faster but capable model for repeated refinements after an initial build, rather than paying frontier-model costs for every change.
  • GLM 5.2: An open-source model El Mghari presents as unusually capable at design-oriented generation and useful for fast UI iteration.

Operator Notes / Why Ken Should Care

  • Create a version-controlled UI-design context file for agent projects: prohibited patterns, brand tokens, typography and logo rules, component preferences, and examples of successful interaction states.
  • Establish a shared, rights-safe screenshot/reference library tagged by product type, visual tone, layout, onboarding pattern, and interaction pattern.
  • Add a mandatory agent-generated UI review checklist before external release: visual hierarchy, spacing, loading/empty/error states, logo quality, animation restraint, mobile behavior, and accessibility.
  • Test a two-tier model-routing policy: premium model for initial build or difficult changes, fast economical model for bounded refinement tasks; measure time-to-acceptable-UI and total cost rather than model preference alone.
  • Avoid shipping raw one-shot UI output, particularly for demos or public landing pages, even when functionality is complete.

Source/Metadata

  • Title: The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI
  • Transcript words: 2789
  • Duration seconds: 849
  • Timestamp note: No usable timestamps or chapters were present in the provided transcript.

Transcript

2712 words en Processed in 77.3s

. All right. Hello, everybody. Welcome. My name is Hassan. I lead the developer experience team over at Together AI, and I'm super excited to be here today to talk to you about how to make AI apps look good or stop letting your agents ship ugly UIs. I'm especially passionate about this project because a big part of my job and my personal life, I love building a lot of these AI apps. And I'll go to the next slide. I've been building about 10 apps a year for the last five years, and I've been lucky enough that some of these apps have gotten millions of users who have tried them out. And I think the number one reason for that is honestly just design and UX. And so that's what I want to talk to you about today, how I approach design and UX in my apps as someone who's not a designer. But before we move on, really quick, I work at Together AI. We're an AI-native cloud platform. We help you do three things. We help you run open source AI models, chat models like GLM 5.2, image models like Nano Banana, audio models, vision models. We have all the different modalities and all the big open source models on our platform through our inference API. We let you fine-tune models on your own data. And then we also have a GPU cluster product where you can reserve H100s and B200s to train your own models or do your own inference. So getting into some demos, I just want to start with a couple apps that I've built so you get a sense of the type of stuff that I build. They're usually very simple one-page apps where I try to index on the design and some animations and things like that. So this is a logo creator app that I built that got about 85,000 people who used it. This is one called Make Comics where it'll create a comic book from scratch starring you as the superhero. And this one, we actually have in person at the Together AI booth at this conference, where you can go and there's a little iPad and you can take a picture and get a real comic book printed out that you can take home. Other stuff, this is generating subtitles from videos. I upload a lot of videos to YouTube and to Twitter, and so I needed something like this. And I looked into generating subtitles with open source models. This one had about 8,000 people who checked it out. AI chat app for open source models. This one's fairly straightforward. This one is a website that I built to help you build five or six different variations of whatever landing page you want to build, and then you can choose one of them. And then the last one I'll show off really quickly is an AI cloud agent where you can ask it to do things and give it a GitHub repo, and it'll spin up a sandbox and actually create a PR for you from scratch. And so this is just an example of some of the stuff that I build and that I put out. And I think the big takeaway here is, we live in a world now where more and more AI apps are slop. You can look at it, and within a second or two, you can tell that this looks AI-generated. And so I think just doing a little bit of extra effort, a little 10 to 20% after, focused on the UI, is a really, really big competitive advantage for these things. So like I said, vibe-coded apps all look the same. They have the same tells. They have the same purple gradient background in every single one. They have italics in headers. They always have the scroll to explore for some reason. They have these pills that are all caps and with spaced-out letter spacing. They have these gradient logos. They use a bunch of emojis. And so it's the same kind of stuff, sometimes some spacing and padding issues. But the point is, if you really think about it and you look at all of these AI-generated websites, you can make a list of 20 or 30 different things that are like, okay, this is what AI slop is, right? All these random graphics. And so I'm going to talk about two different ways that I've tried to overcome this. And the first one I'm going to start with is this design skill that I built called Hallmark. And Hallmark basically takes all of these AI slop patterns and codifies them and tells AI models, hey, don't use these. Don't do a purple gradient. Don't use italics in the title. And all of these AI slop gates, is what I call them, or slop patterns. So that's one thing it does. And then the second thing it does, which I think is really important when you're building stuff, is it gives your AI model a lot of different themes. And so I built a bunch of these different themes. We feed them into Hallmark. And so when you ask Hallmark to create a website for you, it'll use these as context. And that's going to be a theme throughout this talk as well, that if you give AI models really, really good inspiration, they tend to perform very, very well. And so I launched this about a month and a half ago. I have a little over 10,000 people who have tried it out so far. These are some examples of, I've mostly indexed it on landing pages, but landing pages using Hallmark. And so this is build a landing page for an indie podcast, and it gives you that. And there are different skills and stuff that I'm not going to get too deep into. But the main things I just wanted to cover are the AI slop gates and the themes. And those two things really do a lot. And so this is the announcement tweet, and I got a bunch of really great feedback, and we're still iterating a lot to try to make it really, really good. But I found this to be one way that I try to avoid AI slop websites. And so these are some more examples of Hallmark-generated pages. This one's build a landing page for an invoicing app. And you can see they don't have a lot of the same tells that you'd expect from an AI-generated website. And these also are just one shot. This is one single prompt, and you get this whole website. And really, a lot of the magic also is iterating on these things over and over and over again. So this is another one. And then before and after, this is a pretty good example. This is build a page for a learning app for kids. And on the left is without Hallmark, and on the right is with Hallmark. You can see it's just a little bit cleaner, it's a little bit nicer. It has more animations, which I didn't record a video to show. But it just looks a little bit nicer in general. This is another one as well. On the left, the classic AI-generated page with the gradients and everything like that. And on the right is with Hallmark. You can look at the one on the right and still say, well, that's not a perfect landing page. But it does give you a much better base to start out from, right? And then you iterate your way to something that really, really looks incredible. And then I've been indexing on this thing, but app iteration is also very important. And I think specifically, a lot of people undervalue the importance of using smaller, faster models for these app iteration cycles. A lot of the apps that I build now, I'll start in Codex or Cloud Code to build a base. And then for iterating, I'll use usually a smaller open source model like GLM 5.2. GLM 5.2 is amazing. Who here has used it? A show of hands. Okay, a few people. So this is a model that came out two weeks ago that, in my opinion, was the first open source model to actually be very, very good at design. And actually, we're going to play a little game here where one of these landing pages was generated by GLM 5.2, which is a cheaper open source model, and the other one was generated with Opus 4.8. Raise your hand if you think GLM 5.2 is A. Okay, four hands. That was the GLM 5.2 one. Right, so they're almost indistinguishable in certain ways where the one on the left is generated with GLM 5.2. It was created way faster as well because it's a smaller model, so it's inherently a lot faster. And the one on the right cost five times as much and was a lot slower as well. So anyway, this is another one where the left and the right, the left one arguably is even more AI-generated, and that's the one that Opus actually created. Right, so anyway, when I say use a cheaper model to iterate with, it still needs to be sufficiently good, and GLM 5.2 is one of those models I feel is really, really incredible for this kind of stuff. So this is another example where it's very hard to tell the difference. And for iteration, so this is a really simple website I one-shot with just that one prompt. This is using GLM 5.2. And you can see this does have a lot of the AI tells. It's impressive that it actually works. It's an image playground where you send it a prompt and you get an image. But just with a little bit of iteration on the left, I'm not going to go through all of it, you get to a website that looks roughly similar, but it just looks a lot nicer, and it has a much better logo and way better loading states and animations and just better spacing overall with just one or two follow-up prompts. So I think iteration is extremely, extremely important. So the final part of my talk is just takeaways for how I approach building apps that look good as someone who's not really a designer as well. A lot of other people speaking on this track are incredibly talented designers, and you should take a look at their thing. My objective is not necessarily to make the best-designed app in the world, it's just to make apps that look and feel really good, or at least much, much better than purely AI-generated slop. So my main takeaways for this: one is getting familiar with a lot of these AI tells. We went through some of the patterns of AI slop. A lot of the time, I sit down with people and I show them a website and they're like, oh, that's AI-generated in two seconds. But they can't tell me why. They're like, yeah, it just looks AI-generated, I have no idea what it is. And so I think it's worth understanding that, well, yeah, it's purple gradient, and it's this logo, and it's this thing. And when you understand that stuff, I think it becomes a lot easier to bake that into the apps that you build, or ask the AI models to, hey, don't do this or don't do that. So that's one tip I have. Another one that's tangential is saving your preferences in some sort of skill or markdown file. You can do it at agents.md, you can do whatever you want. But the big thing here, I think, is as you build stuff, you start to build an intuition for what you want, what you don't want. You generate a website, and every single time the logo looks like crap and you have to regenerate the logo, you can start to build out an agents.md that has a bunch of this stuff: avoid this. For the logo, do them this way. For this, do it that way. And as I've gone, I've built up a substantial agents.md that helps me do this. And to a degree, this is what Hallmark is as well. Right? If you don't have your own, you can use something like Hallmark. This is maybe the most important one. If you take one thing away from this talk, please give your agents references and screenshots. I don't build anything nowadays without giving AI models a lot of inspiration. And I have this inspiration vault for anytime I look at a website or an app that I really like, or that I'm like, wow, this looks amazing. I'll just save it. I'll save it somewhere. And now I have a really large collection of them. And so when I'm building a new app, I usually will be like, well, actually, I think I want this to look like a mix of Duolingo and a mix of this app and a mix of this app. And I'll paste in a ton of screenshots for the AI model to look at. And the output will always be way, way, way better. So always try to use references and screenshots. Longer and more specific prompts are always better. These days I will just record voice notes. And so I'll start a voice note for one, two, three minutes, and I'll just rant and be like, well, I want to build an app that looks like this. And here's how you should use it. And here's the type of user that's going to use it. And we have a text box here, and we have an image upload thing. And here's how it should look. And here's some inspiration, right? And so usually, all my prompts now are two to three paragraphs. So they're a lot longer prompts, and they usually do a lot better. This is an example of a much longer prompt. I'm not going to go through all of it, but it just produces something that's just a little bit better. The thing on the bottom was produced with this prompt. Break things down into steps. I just talked about making your prompts longer. But also, if you're building seven different features, you probably don't want to put them all in the same prompt. Right, you probably want one or two features per prompt where you talk about them extensively. And usually I'll send off, the start of the app will be a huge prompt with a bunch of screenshots of inspiration. And then immediately I'll queue up a bunch of other messages of like, oh, we'll do this feature and do this feature and do this feature, and I'll let it go. So that's something that I've seen helps. Cool, this is the one I talked about. Iterating with a cheaper open source model can be better than using closed source models for everything. For iteration tasks, I found they're a lot better. I think Cursor did a great job proving this with Composer 2.5, which is an open source model that they post-trained. And it feels magical using it for iteration, because it's so, so, so fast. And so you'll see that in some of these open source models. Cool. And then the last thing is I see people trying to one-shot apps and then being like, well, that's it, I'm done. I'm going to post this as is. And it's really, really important to understand that whatever your agent creates is just the base, right? And it's on you to give it additional context, give it additional inspiration, and go back and forth and make the apps look and feel really good, keep them simple, and iterate that way. Cool, that's all I got. Thank you so much for coming. well, that's it, I'm done. You know, I'm gonna post this as is. And it's really, really important to understand that like whatever your agent creates is just the base, right? And it's on you to kind of like give it additional contact, give it additional inspiration, and go back and forth and make the apps look and feel really good, keep them simple, and kind of iterate that way. Cool, that's all I got. Thank you so much for coming.