How to Keep Shipping When You Walk Away from Your Desk — Zack Proser, WorkOS
Description
Simon Willison fires up four parallel agents and is wiped out by 11am. That is the problem Zack Proser is solving: not that the tools are too slow but that human attention is still the hard constraint. His loop: voice brief at 184 words per minute, agent dispatched to an isolated git worktree, laptop closed, progress checked from a phone on LTE miles away via remote control. The talk covers four layers that make this sustainable: signal agents that read Slack and Linear on a loop so you never open them yourself, verification gates from lint and build up to browser click through and critic passes, a weekly agent run over your JSONL conversation history to surface inefficiencies and generate missing skills, and an Oura ring connected via MCP so Claude can tell you that you did not sleep last night. You can ignore it. But at least you thought about it. Speaker info: - https://linkedin.com/in/zackproser - https://github.com/zackproser
Summary
Generated by claude-sonnet-4-5At-a-Glance
- Verdict: Watch fully
- Core thesis: AI coding tools enable unprecedented productivity but risk rapid burnout; sustainable developer workflows now require signal layers, voice-first flows, remote control coding, and self-improving systems to balance agent scale with human limits
- Why it matters: Ken is building/investing in AI agent systems and workflows—this presents a tested framework for sustainable AI-assisted development that addresses the critical bottleneck (human attention/burnout) before it becomes endemic
- Best use: Practical playbook for AI-augmented dev workflow design; watch for signal layer architecture, verification gate patterns, and the philosophical shift from 'agents as productivity multiplier' to 'agents plus human judgment as system'
Executive Summary
Zach Proser, WorkOS applied AI team member, argues that despite AI coding tools delivering unprecedented output, developers are hitting burnout by 11am due to relentless context switching and inability to scale attention alongside infinitely scalable agents. He demonstrates this with a Slack bot bug fix where Claude Code autonomously verified its own work via MCP connections to Slack and Linear, completing the entire loop without human intervention—which was 'incredible but also scared' him because there's no ceiling to how much can be piled on.
The solution framework has four pillars: (1) Signal layers—letting agents filter Slack/Linear/etc. to surface only high-priority asks while maintaining focus, preventing distraction traps; (2) Voice-first flows—dictating at 184 words/min vs typing at 90, enabling parallel work across multiple Cursor/Codecs windows; (3) Remote control—starting deep focus sessions at desk then walking away while agents continue work, accessible via phone (Claude Code remote control flag), exploiting diffuse mode for creative solutions without stopping productivity; (4) Self-improving systems—agents reviewing their own conversation JSONLs weekly to identify missing skills/tools that caused friction, auto-building new MCP skills to tighten loops.
Critical caveats: Proser stresses 'speed requires safety'—verification gates are essential (lint/build/tests, browser click-through verification, constitutional AI where another agent checks compliance). He warns the default path is 'burnout turbo' if mindless scaling continues. For early-career devs, his rule is 'don't use AI to do something you don't already know how to do'—build foundational battle scars first so you can catch hallucinations immediately. For chunky cross-stack features, he proposes git worktrees for true parallelism, agent teams with clear prompts, and continuous integration where feedback loops back into the system constantly.
Operational implications: Proser runs agents on cron jobs overnight, reviews output in morning, merges small percentage; plans Linear ticket system with 'agent ready' tag and 15-minute polling loop 24/7. He uses Claude Hooks to extract struggle patterns at session end, archives them weekly for meta-analysis. He connects his Oura ring via MCP so Claude literally tells him 'you didn't sleep, we're tackling part of this today and the rest tomorrow.' The philosophical shift is agents scale infinitely, humans don't—judgment, taste, and knowing something truly solves business needs remain human, while agents handle minutiae with verification criteria.
Key Takeaways
- Claim: AI coding tools create paradox: highest productivity ever but developers are 'completely fried by 11am' from context switching and adrenaline dumping | Evidence: Simon Willison fires up 4 parallel agents, wiped out by 11am; Proser's own Slack bot bug fix completed autonomously via Claude Code + MCP (Slack/Linear) but he realized 'there's no ceiling' to how much could be stacked, which 'scared' him | Caveat: No quantified burnout metrics provided; anecdotal observations from Proser and Willison; assumes all developers experience similar attention degradation | Implication: Ken's agent systems must architect for human attention as the hard constraint, not agent capability—sustainable workflows require deliberate attention management or operators will churn out despite shipping faster | Timestamp: 00:30 / 02:15
- Claim: Signal layers (agents filtering comms) prevent context-switch cost without disconnecting from team/tickets | Evidence: Proser gives Claude Code read access to Slack to surface @mentions/DMs/high-priority asks plus Linear via MCP to find real tickets, creating 'just enough facade' to maintain focus; avoids '80% guaranteed' distraction from manually combing Slack | Caveat: Risk of missing nuanced human asks that agents misclassify; requires trust that agent triage is accurate; no discussion of false negatives (missed urgent items) | Implication: Ken should explore MCP-based triage layers for customer/team comms in his systems; trade-off is control vs attention preservation—critical for multi-agent orchestration where human is reviewing/directing, not executing | Timestamp: 05:40
- Claim: Voice-first coding enables 184 words/min vs 90 typing, unlocking parallel workflows where multiple agents are running before traditional dev finishes first prompt | Evidence: Proser: 'imagine developer speaking across three Cursor windows or Codecs and into Claude across multiple tabs at 184 wpm—they're off and running while traditional dev is still typing first prompt'; compounds over years | Caveat: Assumes voice recognition accuracy sufficient for technical prompts; no mention of noisy environments or transcription errors; may not suit all cognitive styles (some think better by typing) | Implication: Voice interfaces for Ken's agent systems could dramatically increase operator throughput for parallel task queueing; consider voice as input modality for high-level orchestration/feedback loops rather than low-level code editing | Timestamp: 07:50
- Claim: Remote control (Claude Code) + diffuse mode principle allows walking away from desk while agents continue work, exploiting creative breakthroughs away from screen without stopping productivity | Evidence: Pass --remote-control flag to Claude Code session; start at desk with deep focus, queue agents, then access session from phone miles away on trail; when 'genius idea' hits during walk, fire it back and 'come back to it having already been applied'; Proser filmed 32-min proof of reviewing PRs from phone in woods | Caveat: Requires trust in verification gates (see next takeaway); limited to Claude Code ecosystem currently; assumes reliable mobile connectivity; no discussion of security implications of remote sessions | Implication: Ken should evaluate remote orchestration for his agent workflows—enables humans to operate in review/direction mode rather than execution mode; paradigm shift from 'diffuse mode = stop working' to 'diffuse mode = different working mode' | Timestamp: 09:20 / 11:00
- Claim: Verification gates are mandatory for safe remote/autonomous work: (1) lint/build/unit tests, (2) browser click-through verification, (3) constitutional AI where second agent verifies compliance with criteria | Evidence: Gate 1: hooks every time agents verify at code level; Gate 2: tell Claude Code '--Chrome' to verify login not broken via browser; Gate 3: constitution of requirements + second agent gives feedback if not met (Anthropic constitutional AI pattern) | Caveat: No specifics on how to define effective constitutions; browser verification limited by what can be tested in UI; constitutional AI gate adds latency and token cost; human still ultimate verifier for 'business needs actually met' | Implication: Ken must architect multi-tier verification for agent systems; constitutional AI pattern (agent-reviews-agent) is key for autonomous loops; trade-off is safety/reliability vs speed—critical when agents run overnight/unsupervised | Timestamp: 12:30
- Claim: Self-improving systems: agents review their own conversation JSONLs weekly to identify friction points, missing skills/tools, then auto-build new MCP skills to tighten loops for next week | Evidence: All Claude Code conversations saved locally in JSONL; scheduled pass at end of week where agent looks for 'patterns where you had to spend thinking tokens to get something right or eliminate ambiguity,' identifies skill deltas; Claude Code has built-in skill to build/evaluate/improve skills from natural language | Caveat: JSONL files 'not really meant for AI consumption'—questioner notes they're long/junky; Proser suggests intermediary step (hooks at session end to save key struggle bits to separate data store like Obsidian/markdown) or just point Claude at raw JSONLs; no metrics on actual skill improvement over time | Implication: Ken should implement feedback loops where agent systems learn from their own execution history—turns every session into training data for workflow improvement; key is structuring the 'struggle signal' so agent can learn what tools/context would prevent future friction | Timestamp: 15:00 / 16:20
- Claim: For early-career devs, rule is 'don't use AI to do something you don't already know how to do'—build battle scars first so you can catch hallucinations immediately | Evidence: Proser ships TypeScript/RAG/AWS deployments with AI because he 'used to do that the hard way for many years'; has scar tissue to 'immediately catch Claude and say no, that's insane'; advises questioner to 'still code some stuff by hand, figure out what's painful, build confidence in skills' before AI-accelerating them | Caveat: Risks creating two-tier developer ecosystem (those with foundation vs those without); unclear how long 'build foundation' phase should last before AI acceleration; advice may become outdated if models improve to point where they teach fundamentals reliably | Implication: Ken's GTM for agent tools should consider operator skill level—senior operators can leverage more autonomy, junior operators need guardrails/learning modes; potential market segmentation by experience level | Timestamp: Q&A 18:00
- Claim: For chunky cross-stack features, solution is git worktrees for true parallelism, agent teams with clear prompts, and continuous integration where human feedback loops back constantly | Evidence: Questioner notes discrete bugs work great but chunky features 'grind to a halt'; Proser proposes worktrees so agents don't block each other, agent teams with defined roles, 'continuously getting to point where I'm constantly testing app, they're building against spec, I'm flowing back my rage into system and it's fixing as we go' | Caveat: No concrete implementation details or success stories for this pattern; acknowledges 'everyone's struggling with that'; assumes models will improve to make this more reliable; git worktrees add complexity/merge conflicts | Implication: Ken should explore orchestration patterns for multi-agent teams on complex tasks—key is decomposition, clear agent roles, and tight feedback loops; this is bleeding edge and still being figured out across industry | Timestamp: Q&A 21:30
- Claim: Night shift agents on cron jobs generate content/code overnight; human reviews in morning, 'freaks out, screams at it, eventually merges small percentage' | Evidence: Proser runs overnight cron jobs for content generation, reviews/merges small % in morning; plans Linear system with 'agent ready' tag and 15-min polling loop 24/7 churning through personal/company work | Caveat: Low merge rate suggests output quality still requires heavy human filtering; no specifics on what % merges or how to improve it; 24/7 loop assumes tasks are sufficiently decomposed and verifiable | Implication: Ken's agent systems should support asynchronous operation with review queues—humans become reviewers/approvers rather than builders; key is surfacing confidence scores or verification results so human knows what to scrutinize | Timestamp: Q&A 19:00
- Claim: Connecting biometric data (Oura ring) via MCP allows agents to factor developer state into task planning: 'you didn't sleep, we're tackling part of this today, rest tomorrow' | Evidence: Proser connected Oura ring to Claude via MCP GitHub projects; Claude will literally tell him 'you didn't sleep last night, so we're going to tackle the first part of this and not the rest'; Proser admits he ignores it but 'at least I thought about taking a break' | Caveat: Mostly humorous example; unclear if this meaningfully improves outcomes vs being gimmick; privacy implications of giving AI access to health data; Proser overrides it anyway | Implication: Ken could explore contextual awareness (calendar, workload, biometrics) for agent pacing/task selection—agents that respect operator limits might reduce burnout; novel surface area for differentiation in agent UX | Timestamp: 17:40
Detailed Brief
The Burnout Paradox: AI Tools Enable Nuclear Productivity but Ancient Nervous Systems Can't Keep Up
- Claims: Despite shipping more than ever with AI coding tools, developers are exhausted by 11am from context switching and adrenaline dumping; Agents can scale infinitely (loop until criteria met), but human attention is 'still in meatspace' and 'degrades under load'—it's the hard constraint; Simon Willison's observation: fires up 4 parallel agents, wiped out by 11am; suggests everyone must find personal limits individually; Default path is 'burnout turbo, super fast and easier than ever, enabled by LLMs'; intentional path requires preserving self while directing agents to do minutiae
- Evidence: Proser's Slack bot bug fix: Claude Code autonomously fixed sentence-case acronym mangling via MCP connections to Slack/Linear, verified own work, completed loop without human intervention—'felt incredible' but 'scared me because there's no ceiling'; Proser describes 'window hell' of constant context switching; giving Claude read/write Slack access created signal layer that prevented distraction; Show of hands at talk: significant audience felt 'completely fried despite getting more done than ever'; Proser: 'Why not fix 150 bugs every day? Agents can do that, but I can't sit on top of all that, ensure quality, and show up next day for another 8-hour session without being completely destroyed'
- Caveats: No hard data on burnout rates, productivity metrics, or attention degradation—mostly anecdotal; Assumes all developers experience similar limits; some may handle higher cognitive load; Unclear whether issue is tool design, operator discipline, or inherent human-AI mismatch; 'Finding personal limits' is vague guidance—no framework for how to systematically determine them
- Implications: Ken's agent systems must treat human attention as scarce resource, not agents as bottleneck—invert traditional scaling assumptions; Sustainable AI-augmented workflows require deliberate throttling, verification gates, and time away from desk—not just more parallelism; Market opportunity for tools that manage human cognitive load rather than maximize agent throughput; GTM/product messaging should address burnout risk explicitly rather than pure productivity gains
Four-Pillar Framework for Sustainable AI-Augmented Development
- Claims: Signal layers: agents filter comms (Slack/Linear/email) to surface high-priority asks, preventing context-switch cost without disconnecting from team; Voice-first flows: dictating at 184 wpm vs typing at 90 wpm enables parallel workflows where multiple agents are running before traditional dev finishes first prompt; Remote control: start deep focus session at desk, then walk away while agents continue work accessible via phone (Claude Code --remote-control flag); exploits diffuse mode for creative solutions without stopping productivity; Self-improving systems: agents review own conversation JSONLs weekly to identify friction/missing skills, auto-build new MCP skills to tighten loops for next week
- Evidence: Signal layers: Proser gives Claude Code Slack read access to scan @mentions/DMs, Linear via MCP for tickets—'just enough facade to maintain focus' and avoid '80% guaranteed' distraction from manual combing; Voice-first: Proser regularly hits 184 wpm; 'imagine speaking across three Cursor windows/Codecs and Claude tabs—they're off running while traditional dev is typing first prompt'; compounds over years; Remote control: Proser filmed 32-min video proving PR reviews from phone in woods; start session on dev machine with file system access, access from phone miles away on trail, fire ideas back that get applied before returning; Self-improving: Claude Code conversations in local JSONL files; scheduled end-of-week pass looks for 'patterns where you spent thinking tokens to eliminate ambiguity or struggled,' identifies skill deltas, uses built-in skill to build/evaluate/improve new skills
- Caveats: Signal layers risk missing nuanced human asks agents misclassify; requires trust in agent triage accuracy; Voice-first assumes sufficient recognition accuracy, quiet environment; may not suit all cognitive styles; Remote control currently Claude Code-specific; requires mobile connectivity and trust in verification gates; security implications not discussed; Self-improving systems: JSONL files 'long/junky, not meant for AI consumption' per questioner; Proser suggests intermediary step (hooks to save struggle bits to Obsidian/markdown) or just point at raw files; no metrics on actual improvement
- Implications: Ken should explore MCP-based triage layers for comms in his systems—critical for multi-agent orchestration where human is reviewing/directing; Voice interfaces could dramatically increase operator throughput for parallel task queueing—consider voice for high-level orchestration rather than low-level code; Remote orchestration enables humans in review/direction mode rather than execution mode—paradigm shift from 'diffuse mode = stop working' to 'diffuse mode = different working mode'; Feedback loops where agents learn from execution history turn every session into training data for workflow improvement—key is structuring 'struggle signal'
Verification Gates and Constitutional AI Patterns for Safe Autonomous Operation
- Claims: Speed requires safety: multi-tier verification gates are mandatory when agents run remotely/autonomously; Gate 1 (minimal): lint, build, unit tests with hooks every time to verify at code level; Gate 2 (browser): tell Claude Code '--chrome' to click through and verify login not broken, functionality intact; Gate 3 (constitutional AI): define constitution of requirements, second agent verifies compliance and gives feedback if not met (Anthropic pattern)
- Evidence: Proser's bug fix example had Claude Code verify own work via Slack channel interaction and outcome inspection before declaring 'definitively fixed'; Passing --chrome flag to Claude Code gives it browser access for click-through verification; Constitutional AI reference to Anthropic's approach where constitution defines 'what you must do' and another agent checks correctness; Proser uses hooks to trigger verification at session end, PR merge, or explicit 'this is done' trigger
- Caveats: No specifics on how to define effective constitutions or verification criteria beyond basic examples; Browser verification limited by what can be tested in UI—backend/database changes may need other gates; Constitutional AI gate adds latency and token cost—trade-off between safety and speed; Human remains ultimate verifier for 'business needs actually met'—judgment, taste, quality still human domain per Proser
- Implications: Ken must architect multi-tier verification for agent systems—single-layer testing insufficient for autonomous loops; Constitutional AI pattern (agent-reviews-agent) is key for overnight/unsupervised operation—second agent acts as quality gate; Trade-off is safety/reliability vs speed—critical design decision when agents run 24/7; Verification criteria should be part of task definition, not afterthought—'verify your own work' as explicit instruction; Hooks/triggers for verification (end of session, PR merge) enable continuous validation without manual oversight
Chunky Features, Agent Teams, and Git Worktrees for Complex Multi-Surface Tasks
- Claims: Discrete bugs/UI fixes work great with parallel agents, but chunky cross-stack features 'grind to a halt' per questioner; Solution: git worktrees for true parallelism (agents don't block each other), agent teams with clearly defined prompts/roles, continuous integration where human feedback loops back constantly; Verification gates and unit tests become 'even more important' for large features; As models improve, more complete harnesses will make this more reliable—still bleeding edge
- Evidence: Questioner describes workflow: small bugs work great, many parallel tasks, but chunky backend/database/frontend features halt parallelization; Proser's proposal: git worktrees so each agent has own branch, agent teams with defined roles (backend agent, frontend agent, etc.), CI constantly testing, human 'flowing back insults and rage into system and it's fixing as we go'; Acknowledges 'everyone's struggling with that' and 'I think that's going to keep evolving'; References need for 'really clearly defined prompts' and 'building against a spec' with continuous integration
- Caveats: No concrete implementation details or success stories for this pattern—purely proposed solution; Git worktrees add merge complexity and potential conflicts between agent branches; Agent teams require orchestration layer to coordinate—who assigns roles, resolves conflicts, integrates work?; Assumes human has capacity to constantly provide feedback and 'rage' into system—doesn't this replicate burnout problem?; Relies on future model improvements ('as models get better') rather than current capability
- Implications: Ken should explore orchestration patterns for multi-agent teams on complex tasks—this is current frontier, no clear best practices yet; Key design challenge is decomposition: how to break chunky feature into parallelizable sub-tasks with clear interfaces; Human role shifts to architect/integrator/feedback provider rather than executor—but feedback intensity may still cause burnout; Git worktrees pattern suggests each agent should have isolated workspace to avoid blocking—architectural implication for agent runtime design; Continuous integration becomes critical coordination mechanism—tests define truth, agents iterate against them
Asynchronous Operation: Night Shift Agents, Review Queues, and Linear Ticket Polling
- Claims: Proser runs agents on cron jobs overnight for content generation, reviews in morning, 'freaks out, screams at it, merges small percentage'; Planned Linear system: tag tickets 'agent ready,' 15-min polling loop 24/7 churning through personal and company work; Human role becomes reviewer/approver rather than builder; Low merge rate currently but goal is better verification systems to increase it
- Evidence: Questioner asks about night shift for agents; Proser confirms he uses cron jobs for content generation; Proser describes morning routine: review overnight output, provide feedback, merge subset; Planned architecture: 'Linear tickets for everything, subtasks for bugs/features, mark with agent ready tag, loop every 15 minutes all day/night'; Goal is 'churning through' with tags like 'agent ready' to indicate tasks are sufficiently specified
- Caveats: Current merge rate is 'small percentage'—suggests output quality requires heavy human filtering; No specifics on what % merges, how to measure quality, or how to improve acceptance rate; 24/7 loop assumes tasks are decomposable and verifiable without human in loop—big assumption for complex work; Risk of generating large volume of low-quality output that takes longer to review than to build from scratch; 'Agent ready' tagging requires human judgment upfront—moves bottleneck to task specification rather than execution
- Implications: Ken's agent systems should support asynchronous operation with review queues—separation of execution and approval; Key is surfacing confidence scores, verification results, or test outcomes so human knows what to scrutinize vs rubber-stamp; Trade-off between volume and quality: more autonomous agents generate more output but require more review; Task specification becomes new bottleneck—need frameworks for decomposing work into 'agent ready' units; Polling frequency (15 min) suggests near-real-time but not instant—balance between responsiveness and resource cost
Early-Career Developer Considerations and Foundational Skill Development
- Claims: Rule: 'Don't use AI to do something you don't already know how to do'—build battle scars first so you can catch hallucinations immediately; Proser ships TypeScript/RAG/AWS with AI because he did it 'the hard way for many years' and has scar tissue to 'immediately catch Claude and say no, that's insane'; Still recommends going deep on fundamentals, coding by hand, figuring out what's painful—then AI-accelerate once confident; Advantage of LLMs for learning: 'in the past I didn't know the names of things I didn't know, so I couldn't ask; now you can ask and go faster and deeper'
- Evidence: Questioner asks about skill deficit: 'early in career, skill development very important, learned by doing deep work and overcoming hurdles, but this flow has made it harder to learn'; Proser's response: 'Don't use AI for what you don't know how to do already... I can immediately catch Claude because I've spent time building that up'; Recommends: 'Still code some stuff by hand, build those skills, figure out what's painful; once you develop confidence, then it's okay to ship faster with LLMs'; Frames LLMs as learning accelerator: 'If you're more honest with yourself about what you don't know, you can go faster; can test me, where am I missing this, my mental model is murky'
- Caveats: Advice assumes there's objective way to know when 'battle scars' are sufficient—no clear threshold; Risks creating two-tier ecosystem: experienced devs who AI-accelerate vs junior devs who grind fundamentals without AI; Unclear how long 'build foundation' phase should last—6 months? 2 years? Project-dependent?; Advice may become outdated if models improve to reliably teach fundamentals and provide accurate feedback; Tension between 'still go deep' and 'use AI to go faster/deeper on learning'—which is it?
- Implications: Ken's GTM for agent tools should segment by operator skill level—senior operators get more autonomy, juniors get guardrails/learning modes; Potential product differentiation: 'training wheels mode' for junior devs where agents explain reasoning, cite sources, walk through decisions vs 'high autonomy mode' for experts; Market positioning should address skill development explicitly—frame as 'complement to learning, not replacement' to avoid stigma; Long-term risk: if junior devs don't build foundation with AI crutch, industry may face skill gap crisis in 5-10 years; Counter-argument: AI may redefine what 'foundational skills' means—judgment/taste/verification may matter more than syntax/API memorization
Biometric Integration and Contextual Awareness for Agent Pacing
- Claims: Connected Oura ring to Claude via MCP; Claude factors sleep/recovery data into task planning; Example: 'You didn't sleep last night, we're tackling first part today, rest tomorrow'; Proser admits he overrides it ('the hell with you, you're a machine, do what I want') but 'at least I thought about taking a break'; Frames this as part of 'holistic work view'—not just conversations, tickets, skills, but condition of body, optimal focus times, sleep quality
- Evidence: Proser: 'I love my Oura ring, one of first things I did was connect it via MCP; there's a couple GitHub projects that enable you to do that'; Direct quote of Claude response: 'You didn't sleep last night, so we're going to tackle the first part of this and not the rest, you're going to do it tomorrow'; Acknowledges he ignores it but notes 'I tell him the hell with you... but at least I thought about taking a break'; Frames as part of broader context: 'Not just conversations with colleagues, tickets, skills, but condition of your body, times you're able to focus best, how much sleep you're getting'
- Caveats: Mostly humorous/experimental example—Proser overrides agent's advice, so unclear if it meaningfully improves outcomes; No evidence this actually prevents burnout or improves productivity vs being gimmick; Privacy implications: giving AI access to health data, sleep patterns, biometrics—what's stored, who has access?; Risk of paternalistic AI: agent deciding you can't work vs respecting operator autonomy; Assumes biometric data is accurate proxy for cognitive capacity—oversimplifies human variability
- Implications: Ken could explore contextual awareness (calendar, workload, biometrics) for agent pacing/task selection—novel UX surface area; Agents that respect operator limits and suggest breaks might reduce burnout—differentiation in agent design; Privacy/data governance critical if integrating biometrics—need clear policies on what's tracked, retained, shared; Could extend to other context: time of day, meeting load, recent output velocity, error rates—multi-modal operator state model; Counter-risk: operator may resent AI telling them what they can/can't do—balance between helpful nudge and annoying nanny
Notable Concepts & Terms
- Signal layers: Agent-mediated interface to comms/tickets (Slack, Linear) that surfaces high-priority asks while filtering noise, preventing context-switch cost without disconnecting from team—Proser's term for attention-preserving facades
- Remote control (Claude Code): Pass --remote-control flag to Claude Code session; start on dev machine with file system access, then access same session from phone/different network miles away—enables coding from trails/parks while agents run on dev machine
- Diffuse mode principle: Two modes of thinking: focus mode (hardcore in IDE, clear blueprint, ideal for execution but increases blind spots/inhibitions) vs diffuse mode (walking away, playing with dog, shower—lowers inhibitions, enables creative solutions); previously meant stopping work, now means different work mode via remote control
- Constitutional AI (verification pattern): Define constitution/requirements for agent work, second agent verifies compliance and gives feedback if not met—Anthropic's approach adapted for developer verification gates where agent-reviews-agent output
- Voice-first flows: Dictating prompts at 184 wpm vs typing at 90 wpm, enabling parallel workflows where developer speaks across multiple Cursor/Codecs/Claude windows simultaneously, agents running before traditional dev finishes first typed prompt
- Git worktrees: Git feature allowing multiple working directories for same repo; Proser proposes this for agent parallelism so each agent works in isolated branch without blocking others—architectural pattern for multi-agent teams on chunky features
- Claude Hooks: Triggers in Claude Code that fire at session end, PR merge, or explicit completion; Proser uses to extract struggle patterns, save key bits to separate data store (Obsidian/markdown), enable self-improving system by capturing what could be more efficient
- MCP (Model Context Protocol): Connection protocol that gives Claude Code read/write access to external systems (Slack, Linear, Oura ring, etc.)—enables signal layers, biometric integration, ticket reading; Proser's examples all use MCP connections
- Agent ready (Linear tag): Proser's planned workflow: tag Linear tickets that are sufficiently specified/decomposed for autonomous agent work; 15-min polling loop picks up tagged tickets, agents work on them 24/7—moves bottleneck to task specification
- Night shift agents: Agents running on cron jobs overnight to generate content/code; human reviews in morning, provides feedback, merges subset—asynchronous operation pattern where human becomes reviewer/approver rather than builder
Operator Notes / Why Ken Should Care
- Ken: This is a tested playbook for sustainable AI-augmented development, not aspirational—Proser is shipping this at WorkOS today with specific tools (Claude Code, MCP, Linear, Slack)
- Critical inversion: stop treating agents as productivity multiplier and start treating human attention as scarce resource—this reframes entire agent system design around throttling/pacing rather than maximizing throughput
- Verification gates (lint/build/browser/constitutional AI) are non-negotiable for autonomous operation—single-layer testing insufficient when agents run unsupervised; multi-tier validation is key architectural requirement
- Voice-first interfaces could be massive throughput unlock for Ken's systems—184 wpm dictation enables parallel agent queueing in ways typing can't; consider voice as orchestration layer
- Remote control coding pattern (start at desk, walk away while agents work, access from phone) exploits diffuse mode without stopping work—paradigm shift from 'step away = stop working' to 'step away = review/direction mode'
- Self-improving systems via conversation JSONL analysis: every session becomes training data for workflow improvement if you structure 'struggle signal'—agents identify missing skills/tools that caused friction, auto-build next week
- Chunky features are still hard: git worktrees + agent teams + continuous integration proposed but bleeding edge with no proven patterns yet—opportunity for Ken to pioneer orchestration patterns for complex multi-surface tasks
- For Ken's GTM: segment by operator skill level—senior operators get high autonomy, juniors get learning modes with guardrails; 'don't use AI for what you don't know' rule suggests training wheels product tier
- Asynchronous operation (night shift agents, review queues, 24/7 Linear polling) shifts human to approver role—key is surfacing confidence/verification results so operator knows what to scrutinize vs rubber-stamp
- Biometric integration (Oura ring) is experimental but hints at contextual awareness opportunity—agents that respect operator limits/suggest pacing could differentiate; privacy/governance critical if pursuing this
- Early-career dev tension: advice to 'build foundation first' may create two-tier ecosystem; long-term risk of skill gap if juniors lean too hard on AI without foundational battle scars to catch hallucinations
- Trade-off between speed and safety pervades entire talk—more autonomy = more throughput but requires more verification; Ken's systems must make this trade-off explicit and configurable per use case
- Key philosophical shift: agents scale infinitely, humans don't—judgment, taste, knowing something truly solves business need remains human domain; agents handle minutiae with verification criteria, humans direct and approve
Watch Map
- 00:00: Intro: WorkOS context, show of hands on AI coding burnout
- 00:30: Concrete story: Slack bot bug fix with autonomous Claude Code + MCP loop
- 02:15: Burnout paradox: nuclear tools but ancient nervous system, 'wiped by 11am'
- 03:30: Agents not bottleneck, we are: attention degrades under load
- 04:20: Four-pillar framework introduction
- 05:40: Signal layers: Claude reads Slack/Linear to surface high-priority asks, prevent distraction
- 07:50: Voice-first flows: 184 wpm dictation enables parallel workflows
- 09:20: Remote control: Claude Code --remote-control flag, code from phone on trail
- 11:00: Diffuse mode principle: creative breakthroughs when walking away, now without stopping work
- 12:30: Verification gates: lint/build/browser/constitutional AI patterns
- 13:50: Daily workflow: deep focus morning session, queue agents, walk away with phone access
- 15:00: Self-improving systems: agents review own JSONL conversations weekly, build missing skills
- 16:20: Claude Hooks for struggle pattern extraction, Obsidian/markdown archival
- 17:40: Oura ring MCP integration: Claude factors sleep into task planning (humorous but real)
- 18:00: Q&A: Early-career dev skill development, 'don't use AI for what you don't know' rule
- 19:00: Q&A: Night shift agents on cron, Linear 'agent ready' tag, 15-min polling loop
- 20:10: Q&A: Voice vs reading agent output, OpenAI advanced voice mode for 2-hour walks
- 21:30: Q&A: Chunky cross-stack features—git worktrees, agent teams, continuous integration (bleeding edge)
- 22:50: Closing: Build one signal layer, add verification gates, use margin for life away from desk
Source/Metadata
- Title: How to Keep Shipping When You Walk Away from Your Desk — Zack Proser, WorkOS
- Transcript words: 6317
- Duration seconds: 1517
- Timestamp note: Timestamps manually inserted based on content flow since transcript did not include chapter markers; approximate based on 1517-second duration (25:17 video length)
Transcript
Music Hey everyone, I'm Zach. I work at WorkOS. Thanks for coming. WorkOS provides drop-in APIs that allow you to take your software and go upmarket and sell larger deals to enterprises. What I'm going to talk about now is the way that I'm finding to try and maintain balance with all the insane new tools that we're getting every day. Show of hands, if anyone is AI coding with agents lately and feels a little bit like this: despite getting more done than ever before, you're completely fried at the end of the day, adrenaline dumping constantly. This has been my experience and I've noticed that some of the worst of it is the context switching was always super expensive for me, and now it's worse than ever before. So I think a lot of us are feeling this way. The tools are insanely powerful and our skills are more in demand than ever before and yet we're exhausted by 11 a.m. So I'll make this concrete with a recent story. I'm on the applied AI team at WorkOS. I was building a Slack bot that democratized uniform blogging for everybody so that anybody, even if they've never written a blog post before, can come into a Slack channel, make a simple request and get a uniform blog post that does everything the correct way, right? And there was a bug that one of my colleagues reported that said, hey, we need to use sentence case and the sentence case pass right now is mangling some of our acronyms like SCIM and SSO. So normally I'd be living in that kind of window hell that I just showed you and instead this time I made a very minor change that ended up being super impactful. So I gave Claude Code, which is my current preferred aperture for working, the ability to read and write to Slack as well and it already had my Linear ticket access. And so I told it, you need to fix this and then you also need to verify your own work and don't stop until you've done that. And so roughly this is what it looked like. If my terminal's on the left, I ran Claude and I said, fix the sentence case enforcer, it's mangling acronyms. And so it went and did that and because it had an MCP connection to Slack, it fired it into this blog post channel. And then the blog bot that it's working on, sitting in that code base on my working directory, picked it up and ran all the way through and got to the step that was relevant. And then Claude verified the access and the outcome and then said, okay, now I have definitively fixed this bug. So when I came back to it, I came back to a completed loop that had been fixed and I didn't have to tell it, hey, this part's still broken. So that felt incredible. And I think that that's an important part of what we're going to all build into our toolkit and we already are. But it also scared me because I realized that there's nothing, there's no ceiling to this, right? So the tools are nuclear now and our nervous system is still relatively ancient. And so the thing that I'm thinking about recently is how do I find my own kind of developer balance in this world, right? So why not just stack that same process? Why not use that harness and just fix 150 bugs every day at work? Well, the agents can do that, especially if you give them enough context, if you give them the right verification criteria, if you give them the tools to verify their own work. But at the end of the day, I can't actually sit on top of all of that and make sure that the quality is there and also show up the next day for another eight hour work session and not be completely destroyed. So what I think I'm finding and I think a lot of other folks have been talking to you recently are finding is that the agents are not the bottleneck now. And I think that's going to increasingly be the case, but we are. So agents can scale infinitely, especially now that they're made available as of last night via Claude API. You can give them verification criteria and the tools that they need in order to match that criteria. But our attention is still, in meat space, if you will, and it still degrades under load. It's still the hard constraint, essentially. And this is something that's not just me and not just you, everyone that just raised their hand at the beginning of this talk feeling it. As Simon Wilson said recently, last week, he fires up four parallel agents and he's wiped out by 11 a.m. And so he suggests that all of us finding our own individual balance and our personal limits is now something that we all need to do on our own. I'm definitely seeing this also on the applied AI team. So here's a couple of tips and tricks or things that I've used recently and found success with and so I'll share them. Some of them I'm sure you've already seen. So we essentially need to bring the human developer into balance with this new way of working. Because now that we have these hypercharged tools, it's faster than ever to burn yourself out, especially if you just scale linearly in terms of what you're taking on at work and what you're outputting. And so the way that I'm breaking it down in my mind is that agents are going to do better in terms of infinitely scaling, looping infinitely until they have reached the criteria you've given them. But we still have judgment, taste, knowing that something is actually solved, knowing that the criterion is actually met in terms of human needs and business needs. And so this is an early breakdown of the stack that I'm seeing. So the first I'm calling signal layers for lack of a better word and I'll develop that a little bit in a second. The second is voice first flows. I've been doing voice first coding now for about a year and a half and it's been life changing. And then remote control, which is becoming more and more recent. Right now it's a Claude Code specific thing. And then the system improving itself. So changing just minimal things about the way that you work and the way that you store your own message history, can enable incredibly powerful passes now that you have agents that can rip through all that material in seconds and find patterns for you on a loop without you even needing to remember to do it. So let's take a quick look at what each of these means. So calling back to the bug that I showed you that I fixed and had Claude do it. My problem is that if I were to go and comb through Slack myself, it's 80% guaranteed that I'm going to get distracted by some other thread. I'm going to find something else or somebody's going to have a new ask for me and that's going to pull me off task. And so instead I had Claude Code be able to read my Slack and do it on a loop so that it can see are there at mentions, are there DMs, are there actually high priority asks that need to be actioned. Meanwhile, it's always had access to my Linear via MCP and so it can duplicate asks and find the real tickets. So calling back to the bug that I showed you that I fixed and had Claude do it. My problem is that if I were to go and comb through Slack myself, it's 80% guaranteed that I'm going to get distracted by some other thread. I'm going to find something else or somebody's going to have a new ask for me and that's going to pull me off task. And so instead I had Claude code be able to read my Slack and do it on a loop so that it can see are there @mentions, are there DMs, are there actually high priority asks that need to be actioned. Meanwhile, it's always had access to my Linear via MCP and so it can duplicate asks and find the real tickets. And this is just enough of a facade for me to allow me to continue to focus and maintain my attention on the things that are key for me to be able to do as the human developer. And I think that there's a ton of tools that are coming out that we're all seeing here too that are going to make this a bespoke experience wherever you want to work. But it's about managing the now extra insane levels of traffic and pinging and noise that we're all going to deal with. The second is voice first flows. Highly recommend it if you haven't tried this yet. I'm a person that loved to type. I've grown up typing since I was three years old. I think at my best I was hitting 90 words per minute in a non-standard way with my giant sausage fingers. But now with voice first tools, it's significantly faster. I regularly hit 184 words per minute on a given day. And what that enables is not just speaking into one thing quickly and having it done faster. It enables parallel workflows. So imagine if at the top I'm a developer who's speaking across three different Cursor windows or into Codecs and then also into Claude across multiple tabs. And because it's 184 words per minute, they're now off and running while a traditional developer is still typing in their first prompt. And I think that if you consider that small things grow quickly in terms of software, how does this compound over the course of a year, two, three years of working? And this has been really key for me because as I get more comfortable with voice flows, it also enables what I'm going to show next, which is spending less and less time at your actual desk while still getting work done. The next is remote control. Right now, this is a Claude code specific thing, but I expect that's going to rapidly change. So before we go look at exactly what that means in the Claude ecosystem, we'll touch on the diffuse mode principle. I'm sure everyone has heard about this before. The basic idea is that there's two modes of thinking. If you're hardcore focused in your IDE and you're typing and you're searching for symbols, you're likely in focus mode. You likely have a very clear blueprint in your mind of what you're trying to build and how you want to do it. And that focus is excellent for getting something over the line and building something exactly the way you want. But it's also ideal for having blind spots and missing what you need, creative solutions to things and increasing your inhibitions. But paradoxically, when you get up and walk away and you start playing with your dog or walking your kids to the park or taking a shower, it's like there's a flash of insight and you get the full-form solution. And the thesis, the diffuse mode principle, is that your subconscious is always churning on these hard problems. And so as soon as you walk away and open the aperture and lower your inhibitions, diffuse mode allows you to see more creative solutions quickly. And the key thing I want to stress here is that we've always had this and there's been hundreds of books written about it and we've all talked about it for decades. But it used to mean that diffuse mode and walking away from your desk meant stopping work. And that is no longer the case, especially with things like remote control. So what remote control means in the Claude code ecosystem example. If I'm starting a Claude code session, I pass the remote control flag or I run remote control or I have in my config enable remote control, then I can start at my desk, it's running on my dev machine, it has access to the file system, etc. But then as soon as I pull up Claude on my phone on a different network, miles away on the trail, I can still see that session and I can still talk to it and send messages and poke it. And that's incredibly powerful because I get my best ideas and solutions as soon as I've walked away from the desk. And so what this enables now is, and what I'm going to propose that enables is starting your day in focus mode, getting everything loaded up that you need to do that's super important, and then getting your agents churning, making sure work is proceeding the way that you want, and then leaving the desk, reducing your RSI and the number of physical injuries you're getting from sitting in the same position all day, and going and taking a walk, but still being super productive. And I have done a ton of experiments with this, and I even filmed a 32 minute film last year proving that you can do this and review PRs from your phone in the woods. So it's possible. And the really nice thing about it is that when you talk to that session through your phone, that session on Claude code is still running on your machine, still has access to do whatever it needs. If you have some genius idea of this is the design that's going to nail it, then you just fire that back. You don't have to remember to do it when you get back to your desk, you're going to come back to it having already been applied. So of course, in order to do this, I often say that speed requires safety. And so there's levels of doing this and having verification. And now with new tools coming online, not only Chrome use, but also computer use for agents, this is going to get more sophisticated. Gate one is the minimal lint and build and unit test. Let the agents with hooks every single time verify their own work at the code level to make sure nothing's broken. Gate two is when you tell Claude code that you must verify your own work with the browser, click through it and ensure that you haven't broken login, for example. Three is closer to constitutional AI in the way that Anthropic conceives of it, where they're talking about there's a constitution of what you must do and another agent will come and verify that you did that correctly, otherwise give you feedback that you need to action. And so taking all this together, how does this actually change a working software developer's day? So as I said before, I propose a deep focus session in the beginning of the day, you might go through all the backlog tasks in GitHub, the software development lifecycle chores that you need to queue up, click through it and ensure that you haven't broken login, for example. Three is closer to constitutional AI in the way that Anthropic conceives of it, where they're talking about there's a constitution of what you must do and another agent will come and verify that you did that correctly, otherwise give you feedback that you need to action. And so taking all this together, how does this actually change a working software developer's day? So I said before, I propose a deep focus session in the beginning of the day, you might go through all the backlog tasks in GitHub, the software development lifecycle chores that you need to queue up, fire those all into codecs, start working on the features you really care about in an IDE perhaps or in Cloud Code, and then essentially you walk away after getting them going on the work tracks that you've identified for that day, because you have access to them on your phone. So even when you're out wandering around on the edge or on LTE, you can fire messages back to them, and you can start reviewing the PRs as they come through on your phone. And now, because agents are quite reasonable to use with Opus 4.6, it's quite reasonable to leave a natural language comment on a PR in GitHub Mobile at Cloud or at Cursor Agent or at Vercelbot, and say this needs to change. And most of the time, it's going to get it right. And so this enables this complete loop where you're actually spending less and less time away from your desk, less and less time at your desk, you're spending less and less time injuring yourself, your wrists, your hands, and you're getting oxygen and getting better ideas when you're out walking around, but you're still in the loop and you're still directing work forward and making progress. So I'll just quickly show that a key learning that I found in doing this as an experiment is that if you just use the tools on your own, in the beginning of the week, Monday feels amazing, I can rip through a ton of work, Tuesday feels the same, now I got a bunch of random asks in the middle of the week that threw me off course and then by Friday I'm completely wasted and I don't remember what the hell I did. And I know that work shipped, but it's all disorganized. As my illustrious colleague who's with us here, Nick, reminded me, all of Cloud Code's conversations are saved locally in JSON-L files. And what that enables you to do is to start working smarter and have a scheduled pass where your agent goes back and reviews your own conversations with it at the end of every week, at the end of every day if you want, and say, look for the patterns where you had to do a significant amount of spending thinking tokens to get something right, or you and I had to go back and forth and eliminate ambiguity in order to get a task done correctly and figure out the skills that are missing. What's the delta if you had these tools, this MCP server, or these skills? How could we tighten that loop so that doesn't happen next week? And then this is a way in which just by working regularly with your own tools, your entire system or your entire harness can start to get smarter. There is a built-in skill in Cloud Code now to not only build its own skills, but evaluate skills, improve skills, and take natural language prompt and just create bespoke skills that you need. So I highly recommend doing that and then tightening that loop so that you can still get your work done, still deliver what you need to at work, but spend less and less time at your desk. And so what that starts to look like is you're working, you're still paying attention to your main preferred aperture, whatever it is. Maybe it's Z, maybe it's cursor, maybe it's Cloud Code in the terminal, but the patterns are being built up because you're not trashing all of the context that you're building up while working. And so you're treating all of those sessions as gold, which they are, because a single pass with Opus 4.6 can reveal a ton of skills that if we had this next week, I can do this way more efficiently, way more quickly, in a way more reliable manner. Last thing I'll just share for giggles. I love my Uber ring, and one of the first things I did was connect it via MCP. There's a couple of GitHub projects that enable you to do that, and I gave it to Claude. And so when I'm arguing with Claude about a project, there are times he will literally come back and say, you didn't sleep last night, and so we're going to tackle the first part of this, and we're not going to do the rest of it, and you're going to do it tomorrow. And I tell him, the hell with you, you're a machine, do what I want, and I just do it anyway. But at least I thought about taking a break. And so this is a fun thing too, but I do think there is something to this as well, where you start to actually look at your work holistically, not just in terms of the conversations you're having with which colleagues, your tickets and everything, the skills that you have, but then also the condition of your body, what times are you able to focus the best, how much sleep are you getting. And I think this is super important because if we just do this mindlessly, the default path is going to be burnout, but now burnout turbo, super fast and easier than ever, enabled by LLMs, right? Whereas the intentional path is a little bit more like how do I preserve myself, still do my best work, and direct agents to do the minutia for me while I'm still responsible for the quality and the review and actually shipping. So if you find any of this interesting, I would recommend trying to build one signal layer. It can be as simple as just plugging in Slack or linear or whatever you find to be the highest cost context switch for you into your preferred pane of glass that you're working with. Add some verification gates you don't have. For Cloud Code, it's as simple as passing dash dash Chrome now and giving it access to its own browser. And then with the margin that you get back, use that to go to a picnic in the park alone, right? Or go on a walk or play with your dog or whatever the case may be. So the tools are nuclear now. Our nervous systems are still ancient. And so what I'm thinking about these days is trying to find some developer balance. Hope that was helpful. Thank you so much. Any questions? Yes. Yeah. So I guess I'm somewhat early in my career. Uh huh. And that means skill development is also very important. Yep. And also doing deep work or at least I learned to program by doing a lot of deep work, getting into running into issues. Yep. And overcoming those hurdles. Totally. And I'm also super on board with this new flow of working, but it's felt like a skill deficit for me. Like it's made it harder to learn. Hope that was helpful. Thank you so much. Any questions? Yes. Yeah. So I guess I'm somewhat early in my career. Uh huh. And that means skill development is also very important. Yep. And also doing deep work. I learned to program by doing a lot of deep work, getting into running into issues. Yep. And overcoming those hurdles. Totally. And I'm also super on board with this new flow of working, but it's almost felt like a skill deficit for me. Like it's made it harder to learn. [SPEAKER_03] So have you found a balance for that where you can still push forward in skill, but while getting benefits of this approach? [SPEAKER_03] Yeah. Excellent question. [SPEAKER_03] So to repeat in case it's not recorded, the question is if I'm early in my career, this is all skill advancement, but then how do I actually do the hard skill development? And my fear is that this could take it away from me. How do I manage that or maintain it? The way I think about that is the best piece of advice I saw for that was don't use AI to do something that you don't know how to do already. I'm shipping TypeScript systems and RAG systems and doing AWS deployments because I used to do that the hard way for many years. So I have that battle, those battle scars and scar tissue of doing it. [SPEAKER_03] And I can immediately catch Claude, for example, and say no, that's insane. [SPEAKER_03] We're not doing that. The second it says something that's a hallucination or is not a good idea because I've spent that time building that up. [SPEAKER_03] I think you should absolutely still do that. I think you should go deep on those things. And I think it's even possible with LLMs and AI to go deeper, faster, and to say like, test me, where am I missing this? My mental model here is still murky. So I highly recommend doing that. Build some of those skills, still code some stuff by hand, figure out what's painful about it. But once you start to develop those skills and you have confidence in them, then it's okay to, if you've shipped a ton of Ruby apps, it might be okay to start shipping Ruby apps faster with LLMs and Claude. If my focus is skilling up and making sure that I'm on a solid foundation, I would recommend that. And I would also say don't get discouraged because in the past when I was coming up and learning, I didn't know the names of the things that I didn't know. So I couldn't ask. Now you can ask and go faster and deeper on it. It's almost like if you're more honest with yourself about what you don't know, you can go faster. And I still think there's a super bright future for that. So yeah, great question. Yeah, so you talked about getting Claude to look at your own chat history with all those JSONL files. I've tried stuff like that and the problem I found is that those JSONL files are not really meant for AI consumption. They get really long. There's a lot of junk in there. Do you just point Claude straight at it? Or do you have some kind of intermediary step where something passes that into a more amenable format? Yeah, that's a great question. I mean, I have just pointed at it before. The question is how do you handle Claude going back and reviewing and doing aggregate analysis on JSONL files if they're super gross and not meant for AI consumption? I have had success just pointing it at it. But the other thing you can do is use hooks so that at the end of every coding session, you can say save the key bits that we talked about and especially highlight where we struggled or where we spent a lot of extra time and put them in a separate data store. It could be Obsidian, could be a flat file of markdown for that week, could be just a simple archive. And then you run your analysis at the end of the week on that. Would it be using AI or just something deterministic? [SPEAKER_04] It would be used with Claude Hooks. [SPEAKER_04] So at the end of each one of these sessions, or when I say that this is done or we merge the PR, that would be the trigger. [SPEAKER_04] Yeah, but how do you determine which other bits to save? [SPEAKER_04] You could just tell it in the prompt basically. You could say look specifically for things that could make more efficient in the future or indications of struggle. [SPEAKER_04] It would be an AI prompt. [SPEAKER_04] And say you're specifically looking for things that we could make more efficient in the future or indications of struggle and return. [SPEAKER_04] You also make use of nighttime, like a night shift for your agent? I do. The question is do you make use of night shift for agent? I've been experimenting with Claude. [SPEAKER_02] I do that with cron jobs and I have some content generation for me. [SPEAKER_02] Then I wake up in the morning, review it and freak out and scream at it and then eventually merge a small percentage of them. [SPEAKER_02] I'd like to get to the point where with better systems and verification that's churning through. [SPEAKER_02] I think the thing I'm going to end up settling on is linear tickets for everything, subtasks for bugs and for feature requests. [SPEAKER_02] And then marking tickets with a tag that says agent ready and then having a loop that's literally going every 15 minutes all day long and all night long. Churning through personal and company work too. [SPEAKER_04] Does your agent then speak back to you or do you read it? It's very verbose, and obviously reading is faster than speaking. Yeah, great question. The question is does your agent speak back to you in terms of voice, in terms of speed and everything. I actually do all of that. I have, most of the time with Claude Code, I'm speaking to it and then it's writing back to me and I'm reading it. But I do a lot of work in OpenAI's advanced voice mode. That's one of my favorite things about ChatGPT. It's one of the only reasons I would use ChatGPT over Claude and I'll go for a walk for two hours and I'll brain dump and talk back and forth and sharpen an idea and then say at the end, okay, now make this a succinct transcript or architecture that I can paste. That's stuck at GPT 4.1, right? I can talk to it, and even that is quite intelligent enough to have a conversation of arguable quality with it before. [SPEAKER_01] But there's also the voice space is moving so quickly that even last night I found one that's an open source project that's basically Whisper flow. It's one of the only reasons I would use ChatGPT over Claude and I'll go for a walk for two hours and I'll brain dump and talk back and forth and sharpen an idea and then say at the end of that, okay, now make this a succinct transcript or architecture that I can paste. That's stuck at 4.1, right? This GPT 4.1 or something. I think I can talk to it, even that is quite intelligent enough to have a conversation of arguable quality with it before. [SPEAKER_01] But there's also the voice space moving so quickly that even last night I found one that it's an open source ghost pepper that's basically whisper flow. It's local only and doesn't use an API. I think there's a tremendous amount of ability. I have OpenClaw on Twilio. So at night I can ask for the call as opposed to it reading to me and then I can talk to it and say this is what I want you to do tomorrow. I find general voices super efficient. Thank you. Yes, sir. Do you do every kind of work with that? So I've had similar workflows and I found it works really nicely when you're doing small bugs or UI fixes and stuff like that and I can do many things in parallel. But then when I have a more chunky feature, something that needs to touch the backend, database, frontend, that changes the way the application works or a factor. It feels like it grinds all of that to a halt and I'm not able to parallelize anymore. [SPEAKER_00] Yep. [SPEAKER_00] Are you able to work on that sort of more chunky task yourself? [SPEAKER_00] Yeah, I think it's a great question. I think everyone's struggling with that. The question is these flows work really well for discrete bite-sized tasks and I agree with that. And then how do you do bigger chunkier ones that are going to touch the entire stack. It's going to be an entire new feature for a distributed cloud system. Where my mind goes for that is git work trees, first of all, so the agents can run in truly parallel without stopping each other's work. And then agent teams with really clearly defined prompts. And then again, it makes the verification gates and the unit tests even more important. [SPEAKER_02] And then continuously getting to the point where I'm constantly testing the application. They're constantly building against a spec. [SPEAKER_02] I'm flowing back my feedback and then into the system and then it's fixing it as we go. [SPEAKER_02] I think that's going to keep evolving though. And I imagine that as models get better, there's going to be more and more complete harnesses to make that more reliable. [SPEAKER_02] That's a great question. All right. Thank you so much. And say like you're specifically looking on the hunt for things that we could make more efficient in the future or indications of struggle and return. Yep. You also make use of nighttime, like a night shift for your agent? I do, I have, the question is do you make use of night shift for agent? I've been experimenting with OpenClaw. So I do that with cron jobs and I have like some, doing some content for me. Then I wake up in the morning, kind of review it and freak out and scream at it and then eventually merge like a small percentage of them. I'd like to get to the point where with like better systems and verification that's like churning through. I think the thing I'm going to end up settling on is going to be linear tickets for everything, subtasks for bugs and for feature requests. And then marking tickets with a tag that says agent ready and then having a loop that's literally going every 15 minutes all day long and all night long. Churning through personal and earmuffs company work too. Hi. Do you let your agent then speak back to you or do you read it? It's very verbose what you do because obviously reading is faster than speaking. Yeah, great question. The question is do you let your agent speak back to you in terms of voice, in terms of speed and everything. I actually do all of that. I have, if I'm working most of the time with Cloud Code, I'm speaking to it and then it's writing back to me and I'm reading it. But I do a lot of work just in like OpenAI's advanced voice mode. That's one of my favorite things about that. ChatGPT. It's one of the only reasons I would use ChatGPT over Cloud and I'll go for a walk for two hours and I'll like brain dump and talk back and forth and sharpen an idea and then say at the end of that, okay, now make this a succinct transcript or architecture that I can paste. That's stuck at 4.1, right? This GPT 4.1 or something. I think it's, I can talk to, even that is quite, you know, intelligent enough to have like a, I've had, you know, conversations of arguable quality with it before. But I, and there's also the voice space is moving so quickly that even last night I found one that, that it's an open source ghost pepper that's basically whisper flow. It's local only and doesn't use an API. I think there's like a tremendous amount of ability. I have OpenClaw on Twilio. So at night I can ask for the call as opposed to it reading to me and then I can talk to it and say this is what I want you to do tomorrow. Yeah, I find, I find general voices super efficient. Yeah. Thank you. Yep. Yes, sir. Do you do every kind of work with that? So I've had similar workflows and I found it works really nicely when you're doing like small bugs or UI fixes and stuff like that and I can, you know, do many things in parallel. But then when I have a more chunky feature, you know, something that needs to touch the backend, database frontend, that changes like the way the application work or a factor. It feels like it grinds all of that to a halt and I'm not able to parallelize anymore. Yep. Are you able to work on that sort of like more chunky task yourself? Yeah, I think, I think it's a great question. I think everyone's struggling with that. The question is, you know, these flows work really, really well for discrete bite sized tasks and I agree with that. And then how do you do bigger chunkier ones that are like, it's going to touch the entire stack. It's going to, you know, entire new feature for like a distributed cloud system. Where my mind goes for that is, is get work trees, first of all, so the agents can run in truly in parallel without stopping each other's work. And then agent teams with really clearly defined prompts. And then again, it makes the verification gates and the unit tests and even more important. And then kind of, you know, continuous integration getting to the point where I'm constantly testing the application. They're constantly building against a spec. I'm flowing back my insults and rage and, you know, into the system and then it's like fixing it as we go. I think that's going to keep evolving though. And I imagine that as models get better, there's going to be more and more complete harnesses to make that more reliable. Yeah, that's a great question. All right. Thank you so much.