Bounded Autonomy: Between Free Will and Determinism — Angus J. McLean, Oliver
Description
Angus McLean spent time building a complex agent application to generate his CV. Four letters beat it: HTML. He puts the improvement at 100x. The talk is from Oliver's AI Director, where agents generate around 4,000 creative assets a day for 200 plus brands, assets you have probably seen and had no idea were AI. The core argument: models are naturally verbose and tend toward complexity, and so are the developers working with them. His counter is to strip back. Replace internet access with curated documentation, ask how little context you can use and still complete the task, and never automate a job you cannot do yourself. Speaker info: - https://uk.linkedin.com/in/angusjmclean
Summary
Generated by claude-sonnet-4-530-second take
Angus McLean, AI director at Oliver (a GenAI ad agency generating 4,000 assets daily for 200+ brands), argues that practitioners building agents are over-engineering and moving too fast. His core thesis: LLMs are "closed boxes" doing semantic math, not learning systems, so treat them like flexible databases with severe constraints. Context windows—not model intelligence—drive most agentic advances, meaning developers should impose deliberate limits, use simpler solutions (his CV builder: raw HTML beat complex pipelines 10-100x), and recognize that AI is fundamentally translation between representation formats. The talk is grounded in real production experience scaling risky, client-facing agents under time/cost pressure.
Key takes
- LLMs haven't fundamentally changed since the 1990s—recent gains come from brute force (massive compute, larger context windows) rather than architectural breakthroughs. Most "innovations" are band-aids masking core limitations like catastrophic forgetting and data inefficiency. Implication: Stop chasing shiny tools; fundamental constraints remain.
- Context windows are soft constraints that shape model behavior more than guardrails do—feeding curated documentation instead of internet access dramatically improves results because models are "very susceptible to SEO" and promotional content. Implication: Control what you exclude from context, not just what you include.
- The challenge has flipped from "lack of context" to "too much context"—abundance kills scrappiness. McLean advocates asking "how little context can I use and still get the task done?" rather than maxing out windows. Implication: Self-imposed constraints force better design and token economics.
- AI is fundamentally translation between representation formats—text→image, audio→video, unstructured→structured. Knowledge production is compression/summarization. This means the structure of data is "a property of the observer," not inherent. Implication: Use multiple representation structures (markdown for hierarchy, graphs for relationships, folders for fast retrieval, timelines when relevant) depending on task needs.
- Workflows beat autonomous agents in production—capitalism naturally breaks work into "small, easily repeatable chunks" (Adam Smith's pin factory). Oliver finds structured multi-step workflows more effective than freeform agents. Implication: Don't automate jobs you can't do yourself; understand the task before delegating it.
- Models naturally trend toward complexity and verbosity—they'll suggest the most complicated solution. McLean's HTML CV example: four simple letters (literally "HTML") outperformed a complex pipeline by 10-100x. Implication: Simplicity-first testing shortens feedback loops with reality.
Useful details
- Oliver's scale: 3,000 staff, 46 countries, 4,000 AI-generated assets/day for 200+ brands with media spend ranging from £20k to millions. This creates a fast feedback loop for measuring what works in the wild.
- Agency structure shift: Accounts now 20% (down from 50%), Creative 20%, Strategy 20%. Creative and Strategy roles increasingly use agents for ideation, copywriting, content production, audience insight, competitor analysis, performance optimization.
- Concrete failures: Models fail at trend identification (if it's too new, they don't recognize it) and spotting promotional vs. consumer content when given internet access.
- Context window evolution: GPT-2 had 512 tokens; Gemini 3.5 Pro vastly larger. Yet "total knowledge doubles every 12 hours," so windows will never be enough.
- Example use case: Social media intelligence—50,000 tweets clustered and organized into strategic insights/standout examples for creative/strategy teams, generated "almost instantly" vs. manual reports.
- Campaign personalization: Generating localized ads for different cities (NYC vs. Miami) via deep research on ideal personas/territories.
- Historical computing parallels: Space War built with 4,000 words; Crash Bandicoot devs hacked PS2 memory constraints for performance gains. Constraints breed creativity.
- Practical tools mentioned: TF-IDF for cluster labeling (older approach), markdown, knowledge graphs, clustering, folder systems, timelines, MCP for model handoffs in long-running agents.
Caveats / counterpoints
- McLean explicitly says this is not prescriptive—"These are my ideas, and I don't expect you to follow them if you're not that interested." He frames it as "conventional wisdom for unconventional times" based on his specific context (ad agency production).
- No acknowledgment of frontier model improvements beyond compute—dismissing recent advances as "band-aids" may understate architectural gains (e.g., reasoning models, multimodal fusion). His "closed box" framing is a mental model, not a technical argument.
- Limited discussion of when autonomous agents do work—he mentions workflows are "often" better but doesn't define boundary conditions or task types where autonomy succeeds.
- "Don't automate a job unless you can do it yourself" is good advice but may bottleneck delegation at scale. It's unclear how this applies to greenfield problems or tasks no human has done before.
- Unclear whether "keep it simple" is survivorship bias—his HTML example worked, but we don't know how many simple solutions failed before he found it.
Ken relevance
High relevance for Ken's agent systems and AI ops:
- Constraint-driven design aligns with Ken's emphasis on forcing functions and avoiding over-engineering. The "how little context can I use?" question is directly applicable to token budgets and cost control in production agents.
- Multiple representation structures (markdown, graphs, folders, timelines) maps to Ken's work on structured outputs, memory systems, and retrieval optimization. McLean's "structure is a property of the observer" reframes how to think about data pipelines.
- Workflows > autonomous agents confirms Ken's intuition about bounded autonomy and orchestration. Oliver's production experience validates structured pipelines for client-facing/high-risk work.
- Context curation over guardrails is a tactical insight: controlling what models don't see (e.g., blocking SEO spam, promotional content) may matter more than prompt engineering what they do see.
Moderate relevance for content/business opportunities:
- Oliver's scale (4,000 assets/day with media spend) is a case study in productionizing GenAI with real ROI measurement. Their feedback loop model could inform Ken's thinking on content iteration.
Low direct relevance for investing/GTM unless evaluating tooling companies—McLean's skepticism about "tools come and go" and "band-aid solutions" suggests durability is in workflows/systems, not point solutions.
Watch verdict
Skim. The transcript delivers the key arguments clearly: treat LLMs as constrained databases, impose self-limits, use simple solutions, structure data for task needs, prefer workflows over autonomy. McLean's production grounding (4,000 assets/day, real media spend) adds credibility. However, the talk lacks deep technical examples or novel frameworks—it's more philosophical reframing than tactical playbook. Ken can extract the core principles (constraint-driven design, context curation, representation plurality) from this summary without watching. Full watch only if interested in hearing ad agency war stories or his delivery style.
Transcript
Thanks so much for coming everyone. Happy Friday. My talk today is called Bounded Autonomy Between Free Will and Determinism. It's about changing the way we think about our interactions with large language models. This talk could have been called a number of different things. I toyed with the idea between automation and customization, between oversight and agency, and between possibilities and constraints. But what it really is is conventional wisdom for unconventional times, and it's based off some of my personal experience in designing agents within industry. It's for people actively experimenting with agents. It's both for beginners who are overwhelmed by the pace of change and experts stuck in a rut in need of a new perspective. What it is not going to be is overly technical, definitive, or prescriptive. These are my ideas, and I don't expect you to follow them if you're not that interested. Great. I'm Angus. I'm an AI director at Oliver. We're a start-up. We've been in the advertising industry for a few years, and then we switched into almost fully Gen AI now. We've got 3,000 staff across 46 countries. You probably haven't heard of us, but I guarantee you've definitely seen our work. You probably didn't realize it. You probably didn't notice. I didn't when I joined the company because you didn't know it was AI when you saw it. This is some of our work that we've done for Johnny Walker. We generate around 4,000 assets a day for more than 200 brands, many of which you've probably interacted with today, maybe even this morning or this afternoon. Unlike other Gen AI content agencies, we actually put quite a lot of media spend behind these assets, anything ranging from 20 grand to a few million, and that enables us to measure these assets' performance in the wild. This gives us quite a good feedback loop. We've got huge volumes of data, and it allows us for much faster iteration and a deeper understanding of what actually works. Right. Who knows what an ad agency looks like? Anyone? Probably not. I didn't think so. Yeah, you probably think it's maybe a bit like Mad Men, but essentially there's three parts of an ad agency. There's the accounts department, there's the creative department, and the strategy department. Accounts typically made up 50% of the agency before. Now we have 20% creative and 20% strat. Accounts manages the client relationships and keeps the project on track. Creative turns the ideas into compelling ads and content. That's historically been the core of the agency. And then strategy is layered on creative. So that's how do we get to those images? How do we define the insight, the audience, and the direction behind the work? Creative and strategy have previously been knowledge work, but they're now increasingly agentic. We don't just do image generation. We do ideation. We do copywriting. We do content production. And those are all done with forms of agents. I'm on the strategy side. I've been for a long time. And we do audience insight, trends analysis, competitor analysis, and performance optimization, all with different agents because we're predominantly customer facing. So we operate in quite a high, fast-paced and high-risk environment. And when we scale these images, it can be just as useful. It can be just as negative for a brand if there's poor reception of these images as if you did a good job. Why do we use agents? We use them primarily for speed and secondarily for scale. Agents allow us to move faster and be more reactive. For creative teams, this allows us to generate content at speed, and this is especially important for iteration and testing. For strategy teams, agents allow us to scale our research so we can get much closer to the consumer, and we can eventually convert them more easily down the line. So typically something we might do is campaign personalization or territory personalization. So that will be we'll have ideal personas, and then we'll try and deep research each of those personas or each of those territories. A good example is we do advertising in multiple cities around the world. We might want to localize that for New York. We might want to localize that for Miami. And in general, the goal is to do much more with much less and create more effective advertising. So my first piece of advice here today is slow down. I think AI is moving very fast. There's a blink-and-you'll-miss-it mentality. Everything seems to be coming and going very quickly. There's a lot of change. I've seen a lot of tools come and go in the last few years. But if you did blink and you did miss it, was it really that important? Probably not. The actual core of LLMs hasn't changed. Let's say at least since the 1990s. Emma Strubel would argue even further back. But no matter how advanced LLMs seem, today's large language models don't actually understand the data they're presented with. And we know this because they've still got several clear limitations. First one is data efficiency. Humans learn from very few examples, whereas models need massive data sets in order to come up with relatively simple conclusions. It's especially evident when it comes to learning. Models don't continuously learn without forgetting in the way that humans do. It's a closed box, as we'll see in a second. One could argue that recent gains come more from less material breakthroughs and more from brute force model improvements. You just heard in the last presentation the amount of compute required to generate those images, 400 marathons. That's maybe too much to be using. This is a famous diagram. It's a bitter lesson. What we end up doing is we end up creating band-aids to go around these model constraints. Most of our tools are band-aids, and you could argue even the way that models themselves are being trained is a form of band-aid. One of the ways to spot these band-aids is that they're temporary. They're a quick fix. They're not a long-term solution. They're quite superficial. They often mask the symptoms rather than fixing the problem, and they're often inadequate, and they don't fully fix the issues. This is how I like to think of LLMs. It's very simplistic, but it's just a closed box with knowledge inside. I prefer to think of it as a flexible database capable of doing semantic math than anything else. It's fully closed. I don't expect any form of emergence or the model to actually learn anything. This is quite evident when we're trying to locate trends. Especially in advertising when we use trends, you probably come up against it not recognizing certain models. But for us, the biggest problem is trend identification. If it's really new, the model won't recognize it. Most recent advances in agentic capabilities have been largely due to increased context window. This is especially true with long-running agents because the longer context windows allows them to do longer-running tasks. You've got things it needs to do that. History of actions, tool output, structure over time, goals and plans. So it's fully closed. I don't expect any form of emergence or the model to actually learn anything. This is quite evident when we're trying to... Especially in advertising when we use trends. So you probably come up against it not recognizing certain models. But for us, the biggest problem is trend identification. If it's really new, the model won't recognize it. Most recent advances in agentic capabilities have been largely due to increased context window. This is especially true with long-running agents because the longer context windows allows them to do longer-running tasks. You've got things it needs to do that: history of actions, tool output, structure over time, goals and plans. It has to be organized. It has to know what it's doing. It has to be able to store and retrieve information in the short term. So without large context, the system forgets mid-task. It can't work on long complex multi-step workflows. And you may have heard of this famous incident of it deleting a lot of emails because the context ran out. This is a model size context difference. GPT2 versus Gemini 3.5 Pro. Does anyone remember the 512 context windows? They're very small. Context windows keep getting larger, but they'll never be enough. The total knowledge in the world keeps doubling about every 12 hours. So we're always going to want more. However big they get, it's not going to be enough. I think that's the clear problem. So how does that work from a developer's perspective? Actually, I find context constraints the models as much as guardrails. So context is a soft constraint rather than a hard constraint. So you can feed stuff in and you can move stuff around and shape indirectly how the model is performing. So a practical everyday example of this is not giving the model access to the internet. And instead giving it high quality documentation, you'll get much better results. And typically when we do this, we find that models are really bad at spotting promotional content. So when we're doing advertising, we're looking up competitors. It'll soak up all of the information they wrote themselves rather than consumer information, which is ultimately what we're after. And they're very susceptible to SEO. I think as soon as you take the knowledge out of the model and you give it another tool, it's quite limited by the way that it uses that tool. So in the past, this is originally what we used to do. So when we had really small context windows, we used to do TFIDF for cluster labeling. So you would take the most frequent words within the cluster. You cluster the text corpus, you'd do the most frequent words within the cluster, and then you'd label that. You'd also do top K and that sort of thing. Whereas now, context assembly is much more dynamic, right? And we've gone from lack of context to too much context, right? So now I think we need to think more about what we can exclude from that. And the challenge is no longer getting context in, but to a certain extent, keeping the noise out, right? And constraints actually create creativity, right? Abundance stops you being scrappy. So if suddenly progress stops tomorrow, how would you make the most of what you have? I think it's important to set self-imposed constraints. Everyone knows that you shouldn't use the full context window, but maybe how little of the context window can you use and still get the task done is a more interesting question. I'm not a token billionaire yet. I know there are some people probably in this room. This parallels with early computing, lack of compute, very difficult times. That's from the Model Railroad Club. That was from yesterday. But great things come out of constraints and limitations, right? So if you can build Space War with only 4,000 words, that's a pretty big deal. And I don't know if you've ever seen the developers of Crash Bandicoot, but they talk a lot about how they were using the memory function in the PS2 to get massive improvements, right? So if they can do it, I imagine you probably can. Other things to try. Maybe using an old, this is just experimental. I'm not suggesting you do this in production. Use an older, smaller version of a model or harness. So if you're using this, this can help you understand, connect with the model maybe a bit more. Building your own harness, building your own memory and compaction. And also I think pre-processing and archiving stuff is really important. So how do file systems work? How do you negotiate knowledge graphs? That sort of thing. And this will overall improve your ability to prompt and control the model. You'll be much closer to it. And you'll have better fundamentals and better best practices. And in general, you just have an improved understanding of the data that you're working with. You never know what's going to come in handy later. This is Rosenblatt actually pulling out wires in the perceptron. And that later formed the basis of what became Dropout. The next one is keep it simple. Have you ever built anything and then realized that the model could just do it better on its own? Probably. Got a lot of nods there. Happened to me recently. I was trying to do my CV. What would win? This was my very complex CV application. Did not work. Well, it did work, but it didn't work as well as four simple letters. Those four simple letters were HTML. Yeah, I was pretty blown away by this. So that was over a 10x improvement. I would say it's probably a 100x improvement. So I think just because you have the power of the gods, that doesn't mean you should use it. Models are naturally verbose and they tend towards complexity. They are going to suggest the most complicated solution. So don't waste your time. Don't waste tokens and don't make extra work for yourself. Your ideas will collide with reality pretty fast. So what matters the most is building a simple version that works. And I think when we're talking about using agents, we talk about shortening that feedback loop. But maybe you should shorten your feedback loop with reality when you're building products. This is, I think, probably the most interesting bit. So AI at its core is just translation. This is the Attention is All You Need paper. It was done on English to French initially. And then I think more interesting is the idea of being able to translate text into images, images into audio, audio into video. Right? And I think there's an argument that knowledge production is just in itself summarization. So what matters the most is building a simple version that works. And I think when we're talking about using agents, we talk about shortening that feedback loop. But maybe you should shorten your feedback loop with reality when you're building products. This is, I think, probably the most interesting bit. So AI at its core is just translation. So this is the attention is all you need paper. It was done on English to French initially. And then I think more interesting is the idea of being able to translate text into images, images into audio, audio into video. Right? And I think there's an argument that knowledge production is just in itself summarization. Right? Like in this talk, I'm trying to summarize my experiences of the last few years. We're compacting that down into knowledge or good knowledge. Right? So different types of data can be converted into common internal representations and then transformed into something else. You've got an unstructured input and you turn it into an unstructured output. This is what MCP is in the model handoffs for long running agents. This is all a very similar idea. Right? You've got something structured on one side, something very unstructured on the other. And I think these two things can coexist. And if you can manipulate something through a representation space, then that means that the structure of that input is a very similar idea. Data. It's not an inherent property of that object. Right? So it's actually more of a property of the representation of the observer. It's what do I want to see this piece of content in what format? Right? And sometimes, maybe the easiest way of thinking about it is with slides. What can you turn into a diagram? And what do you want written? What do you want voiced over? And I think it's super interesting. So this has a lot of relevant implications when you're building. This means, ideally, you should use multiple representation structures. So you might want to use markdown for human readable hierarchy and authoring. You might want to use graph relationships and references. You might want to use clustering as well if you're dealing with large bodies of text or more unstructured bodies of text. And you might want to use folders for stuff you need to retrieve really fast. You could also use timelines if that's relevant to your task. And finally, there were a few more practical bits of advice. But finally, I think it's so important just to have fun and experiment, right? I think so much of what we do is for work. But actually, a lot of this stuff you can only learn through thoughtful play and experimentation. I go to a lot of hackathons now, more than I ever did before. And I just don't see... I'm not allowed to experiment with this stuff in the way that I would in my role. So a lot of that then helps me even more. I've got one minute left. It's really interesting. I will say one more thing. So I was going to talk about how agents fit into the workplace and whether you should use more structured agents or less structured agents. And I will just say, so I think if you look back to Adam Smith and his pin factory, capitalism naturally breaks tasks down into small, easily repeatable chunks. And often we have found workflows to be far more effective. I was going to say that I would... One of the other things I was going to say is don't automate a job unless you can do it yourself. This is my old job. This was social media intelligence. And what you can see here is a practical report. This was done by an NLM. It clustered the data. It did everything. And this is typically the sort of information that we would want from an advertising agency. Because we're an AI engineer, I've done it for OpenClaw just on the last year sort of data. This is all the sort of insight that we would get. This is 50,000 tweets clustered and then organized into essentially strategies. So we can create this almost instantly for our creative teams or our strategic teams. And then it's got standout tweets and it's breaking it all down for you. And I think that's my time there. So thank you very much. Thank you very much. So this is the attention is all you need paper. It was done on English to French initially. And then I think more interesting is the idea of being able to translate text into images, images into audio, audio into video. Right? And I think there's an argument that knowledge production is just in itself summarization. Right? Like in this talk, I'm trying to summarize my experiences of the last few years. We're compacting that down into like knowledge or like good knowledge. Right? So different types of data can be converted into common internal representations and then transformed into something else. Like you've got an unstructured input and you turn it into an unstructured output. This is like, I think, I guess what MCP is in the model handoffs for long running agents. This is all like a very similar idea. Right? You've got something structured on one side, something very unstructured on the other. And I think these two things can sort of coexist. And if you can manipulate something through a representation space, then that means that the structure of that input is a very similar idea. Data. It's not an inherent property of that object. Right? So it's actually more of a property of the representation of the observer. It's like, what do I want to see this piece of content in what format? Right? And sometimes, I guess, maybe the easiest way of thinking about it is with slides. What can you turn into a diagram? And what do you want written? What do you want voiced over? And I think it's super interesting. So this has a lot of relevant implications when you're building. This means, ideally, you should use multiple representation structures. So you might want to use markdown for human readable hierarchy and authoring. You might want to use graph relationships and references. You might want to use clustering as well if you're dealing with large bodies of text or maybe more unstructured bodies of text. And you might want to use folders for stuff you need to retrieve really, really fast. You could also use timelines if that's relevant to your task. And finally, there were a few more practical bits of advice. But finally, I think it's so important just to have fun and experiment, right? I think so much of what we do is for work. But actually, a lot of this stuff you can only learn through thoughtful play and experimentation. I go to a lot of hackathons now, more than I ever did before. And I just don't see... I'm not allowed to experiment with this stuff in the way that I would in my role. So a lot of that then helps me even more. I've got one minute left. It's really interesting. I will say one more thing. So I was going to talk about how agents fit into the workplace and whether you should use more structured agents or less structured agents. And I will just say, so I think if you look back to Adam Smith and his pin factory, like capitalism naturally breaks tasks down into small, easily repeatable chunks. And often we have found workflows to be far more effective. I was going to say that I would... Like one of the other things I was going to say is don't automate a job unless you can do it yourself. This is my old job. This was social media intelligence. And what you can see here is like a practical report. This was done by an NLM. It clustered the data. It did everything. And this is typically the sort of information that we would want from an advertising agency. Because we're an AI engineer, I've done it for OpenClaw just on the last year sort of data. This is all the sort of insight that we would get. This is 50,000 tweets clustered and then organized into essentially strategies. So we can create this almost instantly for our creative teams or our strategic teams. And then it's got standout tweets and it's breaking it all down for you. And I think that's my time there. So thank you very much. Thank you very much.