Your Finance Agent's Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI
Description
The slowest part of shipping a production finance agent is not the model or the GPUs, it is you, the developer in the loop. Ramana Siddanth Emani's point is that the same agent harnesses you use to build products can automate your own developer loop. Coding agents can multiply how much you ship; run an army of them across separate git worktrees and they clear tasks in parallel, with skills making sure each one uses the right patterns. The tasks come from where they already live, QA reports, Jira tickets, GitHub pull requests, and a sub agent pulls the traces and logs, writes and runs end to end tests, builds, and reports back, needing your context only at a few steps. Point this at your bug queue and a month later you have shipped far more, having stepped further out of the loop as the agents improve, while keeping a human as the final verifier. At Auditoria, where the work is finance, that means agents talking to agents and reconciling source data, so you spend your time verifying rather than grinding. Speaker info: - https://x.com/siddanth2486 - https://www.linkedin.com/in/siddanth-emani Timestamps: 0:00 - Your bottleneck is you 1:05 - From bugs to pilots to production 2:37 - Automating the developer loop 3:03 - Coding agents that multiply output 3:39 - Skills for the right patterns 4:22 - Sub agents and where tasks come from 5:18 - Pulling traces, testing, reporting back 7:56 - Auditoria in the finance sector 9:04 - Stepping out of the loop safely 11:35 - Turning customer patterns into features 13:04 - Keep a human as verifier
Summary
Generated by gpt-5.6-terraAt-a-Glance
- Verdict: Skim
- Core thesis: Production agent reliability is constrained less by model choice than by developer-harness velocity: parallelize work with sub-agents and isolated Git worktrees, encode organizational workflows as skills, connect operational systems through tools, and keep humans as accountable verifiers rather than execution bottlenecks.
- Why it matters: This is a relevant control-plane pattern for scaling agent-assisted software operations, especially where bug investigation, testing, deployment, and audit requirements span many systems.
- Best use: Use it as a concise architecture prompt for designing an autonomous bug-fix pipeline and deciding where human approval must remain; do not expect a concrete implementation or security design.
Executive Summary
Ramana Siddanth Emani argues that impressive agent demos fail after reaching production not primarily because teams lack a stronger model, GPU, or framework, but because their development and remediation loops are too slow. Since models, infrastructure, and frameworks can be swapped over time, the operational differentiator is a harness that rapidly turns production failures into investigated, tested, reviewed, and deployed fixes.
The proposed harness rests on four primitives: sub-agents working in parallel, isolated Git worktrees to prevent those agents from colliding, reusable organizational “skills” that encode approved workflows, and tool connectivity—especially via MCP—to retrieve logs, traces, tickets, data, and other context. A central, low-friction interface should let a human oversee this activity without continuously moving among dashboards, Jira, GitHub, Kubernetes, and coding tools.
His target workflow assigns each QA-reported Jira bug to a separate worktree. An agent then parses requirements, performs root-cause analysis, gathers logs and traces, follows TDD, runs local and end-to-end tests, opens a PR, and proceeds through deployment environments. The speaker's position is that humans should primarily supervise the process and validate the outcome in staging rather than manually coordinate each intermediate handoff.
For finance, the talk adds an important constraint: automation does not eliminate accountability. Human auditors and controllers remain responsible for review and SOX sign-off; an agent cannot be the party blamed for an incident. The longer-term vision is a recursive improvement loop in which completed bug-fix runs are analyzed for bottlenecks, allowing the harness itself to improve over days or weeks. This is directionally useful, but the talk remains conceptual and does not specify permissions, evaluation gates, rollback controls, or audit-evidence design.
Key Takeaways
- Claim: The main production bottleneck for agent systems is developer-loop velocity, not merely model capability. | Evidence: The speaker contrasts easily swappable future models, GPUs, and frameworks with the immediate need to fix production bugs as new customer data and failure cases appear. | Implication: Prioritize the operational path from incident signal to safe fix over repeatedly replatforming around the newest model. | Caveat: This is a strategic assertion rather than a measured comparison; the talk provides no reliability or productivity data showing that harness design dominates model quality in a particular environment.
- Claim: Parallel sub-agents should receive independent tasks in isolated Git worktrees so they can make progress without code conflicts. | Evidence: The speaker describes Git worktrees as isolated folders where each agent writes code, suggesting that QA Jira tickets can each be assigned to a separate worktree; he estimates that a 48 GB MacBook could support roughly 50 active worktrees/sub-agents. | Implication: Treat worktree isolation as a basic concurrency primitive for coding agents, and design task decomposition so agents are not editing the same surfaces. | Caveat: The 50-worktree estimate is presented as an illustrative claim, not a benchmark, and actual concurrency will depend on model execution, repository size, tests, and local resource consumption.
- Claim: Organizational skills and system connections are what turn general-purpose agents into production-capable operators. | Evidence: Skills are described as an organization's “secret recipes” that guide agents through correct workflows; MCP-connected tools can expose logging systems, authentication gateways, databases, traces, tickets, and other third-party systems. | Implication: Build versioned, narrow, approved skills around recurring incident workflows, then expose only least-privilege tools and data required for each role and task. | Caveat: Broad access to client data and production systems materially increases authorization, data-governance, and blast-radius risk, none of which the talk operationalizes.
- Claim: A bug-fix lifecycle can be largely automated end to end, with human involvement concentrated in oversight and final validation. | Evidence: The example agent parses bug requirements, conducts root-cause analysis, retrieves traces and logs, creates an isolated worktree, applies TDD, runs test scripts and local end-to-end tests, opens a PR, and moves builds through development and staging before QA retests. | Implication: Automate the repeatable middle of the remediation pipeline, while explicitly defining approval checkpoints based on risk rather than assuming all intermediate steps are safe to delegate. | Caveat: The claim that humans are needed only at the beginning and final staging validation underplays material approval points such as production deployment authority, security review, merge controls, and ambiguous incident triage.
- Claim: The human should remain a verifier and accountable owner, but should not be the throughput ceiling for an agent fleet. | Evidence: In the finance context, the speaker notes that a human auditor reviews code and a controller signs off for SOX compliance; agent-to-agent review does not preserve accountability if production failures occur. | Implication: For regulated workflows, separate autonomous execution from accountable approval, and make the approval artifact auditable rather than relying on informal supervision. | Caveat: Human accountability alone is insufficient for compliance; the system would also need durable evidence of agent actions, approvals, test results, data access, and deployment provenance.
- Claim: Harnesses can improve recursively by treating each completed production-fix run as data for finding and eliminating process bottlenecks. | Evidence: The speaker proposes letting a loop run for one or two days across five or six bug tickets, then asking the agent to analyze bottlenecks and progressively remove them; he also proposes background aggregation of recurring customer-session patterns, called “dreaming,” to improve the system. | Implication: Instrument agent runs and mine them for recurring failures, but put proposed harness changes through an offline evaluation and change-control process before allowing them to alter production behavior. | Caveat: Self-modifying operational loops can propagate flawed assumptions or unsafe workflow changes unless improvements are evaluated, versioned, and approved before adoption.
Detailed Brief
Control-plane UX and orchestration model
- Claims: The speaker sees orchestration overhead—not only coding effort—as the next limiting factor once teams operate many sub-agents.; A minimal interface can reduce the context switching involved in supervising a software-delivery lifecycle.
- Evidence: His proposed single-pane view combines production-agent status, Kubernetes services and pods, logs, Jira tickets, GitHub PRs, and a coding session.; He frames the benefit colloquially as reducing the number of “neck rotations” between multiple monitors and application windows.
- Caveats: A consolidated dashboard can simplify supervision, but it can also conceal essential risk distinctions if it treats incidents, code changes, credentials, and deployment approvals as equivalent queue items.
- Implications: The useful design objective is not merely one pane of glass; it is a role-aware control plane that surfaces exceptions, provenance, approvals, and blocked decisions without forcing operators to inspect every agent action.
Goal-driven autonomy
- Claims: The speaker recommends combining durable goals with recurring loops so agents can investigate objectives asynchronously rather than requiring continuous interactive steering.; A stated example is a report data-discrepancy goal where source data does not match agent-generated output.
- Evidence: He describes setting the goal, allowing the agent to investigate via connected systems, and monitoring or interacting from a phone rather than keeping a laptop open.
- Caveats: The transcript does not define completion criteria, escalation conditions, budget limits, or safeguards for an agent that cannot resolve the discrepancy.
- Implications: Persistent goals are appropriate for bounded, observable operational objectives, but they require explicit stop conditions and escalation policies before they can be trusted unattended.
Notable Concepts & Terms
- Developer harness: The operational layer around coding agents that supplies task intake, context, tools, execution isolation, testing, review, deployment, and supervision.
- Git worktrees: Isolated working directories tied to a repository; the speaker uses them as the concurrency boundary that lets multiple coding agents work without directly colliding.
- Skills: Reusable organization-specific procedures or recipes that constrain agents toward approved ways of diagnosing and resolving recurring production issues.
- MCP tools: Tool and server connections that can give agents access to external operational context such as logs, authentication systems, tickets, and databases.
- Human as verifier, not throughput ceiling: The proposed governance model: agents execute high-volume intermediate work while humans retain supervision, acceptance, and accountability.
- Recursive self-improvement: Using the records of completed agent tasks and production failures to identify harness bottlenecks and iteratively improve the automation loop.
- Dreaming: The speaker's term for background aggregation and compaction of recurring customer-session patterns into data that can inform system improvements.
- SOX sign-off: The finance-specific accountability constraint invoked to argue that autonomous agents cannot replace responsible human auditors and controllers.
Operator Notes / Why Ken Should Care
- Prototype a bounded Jira-to-PR remediation lane using one worktree per ticket, but restrict it initially to low-risk services and require deterministic test evidence before PR creation.
- Create a skill registry for recurring incident types; require each skill to specify allowed tools, required evidence, success criteria, owner, and version history.
- Define risk-tiered approval gates for merge, staging, production, and data access rather than adopting the speaker's broad “human only at start and end” model.
- Instrument every agent run with task inputs, retrieved evidence, tool calls, code diff, test outputs, reviewer decision, deployment status, and rollback outcome to support both learning and audit.
- Treat recursive harness improvement as a governed change-management workflow: evaluate proposed workflow changes offline before granting the system authority to adopt them.
- Design the orchestration console around exceptions and approvals, not simply dashboard consolidation; ensure operators can rapidly see agent ownership, permissions, blocked states, and blast radius.
Source/Metadata
- Title: Your Finance Agent's Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI
- Transcript words: 3482
- Duration seconds: 822
- Timestamp note: No timestamps or chapters were present. The transcript substantially repeats the main presentation content, so the unique substantive material is shorter than the stated word count.
Transcript
Hello, everyone. Welcome to this session about your finance agent's bottleneck is you. Sorry for the rude title. I don't mean to call the audience here the bottlenecks, but I'm here to talk about the harnesses that you guys are developing and using, these internal harnesses to build your production agents. My name is Siddhant Timani, and I'm a data scientist at Auditorial AI, and we build production agents for finance. So if you're a CFO in the audience, I would love to speak to you after the session. This talk is in between the harness engineering track and AI for finance. So this talk is mostly about identifying the bottlenecks within your developer harnesses. And if you're a developer yourself, how do you be 10x productive with the agent harnesses that you're using? All of us have seen beautiful demos in this AI Engineers World Fair. But once these demos are promoted to pilots, and you start onboarding new customers, the agents have never seen this future data. So all of us know production bugs are very high, and production gods build by the air. So that's a hard fact. Writing code is very easy. So shipping beautiful demos and showing them to a lot of people is very easy nowadays. So what is the problem? And why do these demos fail in production? Is it the model? Do you need a better model? Fable 5, perhaps? Or do you need faster GPUs? Or do you need a better framework? Maybe. Or your RALF loops are not working properly. So what is the answer? If we wait three and a half months, we are awarded with a new model in the market. So we can easily swap models. If we wait perhaps one year, we have new chips, we have faster GPUs. And again, writing code is easy. So we have new frameworks every day. So you can swap your framework every now and then. So how do we, in real time, fix these production bugs? The answer is your developer velocity. The model capability increases very exponentially. And the developers have to spend a lot of time every day to automate your developer loop. I'm talking about four primitives here. All of you need to think about loops. And at the end of the session, I hope you can 10x your production code. First, we have sub-agents. Nowadays, whatever harness you're using, you can spawn new sub-agents. You can have an army of them. And Git worktrees are your best friend. So think of worktrees as isolated folders. And inside these folders, the agent writes whatever code it's generating. So you want these worktrees to be in parallel. So the sub-agents are doing independent tasks and are not fighting over the same thing. Second, we have skills. These are your organization's secret recipes. So make sure you have a lot of skills. Because these skills, once you start giving them to your agents, the agents will always make sure to use the correct and proper workflows to solve whatever production bug you're facing. And of course, all of us have seen a lot of MCP tools being shipped into the market right now. Everybody says the agent can connect to whatever MCP tool and whatever third-party server there is. And your client data can live in any system you want. At the end of the day, if you have a lot of sub-agents, you have a lot of work to orchestrate. So minimal UX is the key here. Let's look at the sub-agents. With you as the orchestrator, let's say with 48 GB of RAM on your MacBook, you can have 50 active worktrees. That is 50 active sub-agents working independently on different tasks. So where do these tasks come from? Let's say the production software you're going to ship has a lot of bugs that your QA is reporting. So all the Jira tickets can be thought of in separate different worktrees. So different worktrees are handled by a separate agent, and these agents can spawn multiple sub-agents to solve that particular task. You don't want to queue up your tasks, because the agent will do that a lot better than you. Let's look at an example harness. What if the QA reports a lot of bug tickets, and somehow magically there is an agent which parses the requirements, does a root cause analysis, pulls all the traces, pulls all the logs, puts all this in a separate worktree, does the TDD, implements the fix? Because it's in your local system, you have to do test scripts, local end-to-end testing. You create a PR, you submit the PR to your team for review. And after review, you merge it into your master branch, let's say. After merging it, obviously, you have to build a Docker image, deploy it into your development environment, test it. Again, build an image for your stage environment, test it, deploy it to stage. And then you go back to the QA saying, here you go, you can test it around. So I would like to ask a question to the audience. At what points do you think the human contact is required in these steps one to nine? I would say the human is only required at steps one and nine, because in between steps, the agent can do a lot better work. There needs to be a human to see what work the agent is doing, and then there needs to be a human at the end to validate after the work is being shipped to stage. And obviously, we need minimal UX because humans love minimal UX. So in the image, if you squint your eyes and see, the image shows the production agent software that you're building, the project dashboards, which show all your Kubernetes services, pods, examples, all the logs, system logs, all your JIRA tickets, all your GitHub PRs, and maybe a cloud code session at the bottom. So this is basically a Mac OS widget. And you don't need to open multiple windows to do all of this work. A developer does a variety of things in their software lifecycle. So you can use just this one widget to do a lot of things. You can see from the graph also, the number of neck rotations to ship one change reduces drastically. And I imagine all of you have two to three monitors on your table, and you just keep rotating your neck, orchestrating these agents. So Auditoria works in finance. So there's a lot of regulation and policies happening in finance right now. So what does it look like for orchestrating a team of sub-agents in the finance sector? If we take AI out of the picture, usually what happens is you have a human auditor who reviews the code, and you have a controller who signs off under your SOX compliance. And reviewing agent to agent, it doesn't really keep the accountability. If something goes wrong in production, you can't say, Claude is doing this. Something is wrong. But let's say you have all these sub-agents, and you're using these harnesses to fix bugs in real time. What is the bottleneck? It becomes human attention because you yourself have to orchestrate all these different tasks. And moving fast and breaking things in the finance sector is a lot different. So let's look at part two, which is removing yourself from the loop. Till now, I've been saying a human is required to see what the agent is doing, and at the end also to validate what the agent has done. But with the self-improvement of the agent and model capabilities these days, we get Fable 5 and Mythos 5 and GPT 5.6 also. So what does it look like when you have this recursive self-improvement in your internal developer harnesses? So all your production failures become input. Let's say you keep automating these developer harnesses every day, and you ask the agent to upgrade itself, essentially. So you do a task. You let the loop run, let's say, one or two days. You solve five to six bug tickets. And you just tell the agent to analyze all the bottlenecks in this process, make a list of them, and somehow slowly keep removing these bottlenecks every day. At the end of one month, let's say, you have a really nice self-automated loop where you just type in one sentence and just say, fix this bug for me. And the agent goes off, connects to all your database systems, fetches all the logs, traces tickets, and migrates it to the Jira-to-QA pipeline. And you can just book a vacation, maybe, or work from home. And what does it look like internally? And what happens when you steer less and ship more? Nowadays, how many of you know you can give goals to your agents? You can just set a goal and forget about it. Anybody? Nice. So what if you combine goals and loops? You can just set a goal saying there is some data discrepancy in this report. And in the production bug, the source data is not matching with what the agent has generated. So you can just set a goal to fix this, look into this, set a loop. You can even close your laptop because you can do it from your phone nowadays. And if you look at the last but one point, which is dreaming, let's say a lot of customers are using your production software. So you can do it from your agent. And they're doing the same type of patterns. And they're facing the same type of problems. So you let the agent dream like humans dream in the background so that it collects all the sessions that your customers are using and compacts them into a set of data points which your system can use and basically upgrade yourself. So with a combination of all these features, basically, you can essentially remove yourself out of the loop. But as I said before, the developers do a lot of variety of things in their software development lifecycle. And sitting behind a desk from nine to five and just writing code is not valid anymore. So just an overview of what I've covered till now in the session. You can have a team of sub-agents working in parallel worktrees. You can have skills, your organizational secret recipes, your customers' recipes. You can give all of these to an agent. Your agent can connect to whatever third-party server there is. It can be a logging system. It can be an authentication gateway. And you just compress all of this into one pane of glass because minimal UX is the key. And you can set goals and loops for autonomy. If you think this particular work can be done by the agent a lot better, you can just ship it to the agent. Always have the human as a verifier, but not the throughput ceiling, because human attention is very limited. So thank you for your time. And thank you for your service. I hope you learned something from this session. Thank you. And the developers have to spend a lot of time every day to automate your developer loop. So I'm talking about four primitives here. All of you need to think about loops. And at the end of the session, I hope you can 10x your production code. So first, we have sub-agents. Nowadays, whatever harness you're using, you can spawn new sub-agents. You can have an army of them. And Git worktrees are your best friend. So think of worktrees as isolated folders. And inside these folders, the agent writes whatever code it's generating. So you want these worktrees to be in parallel. So the sub-agents are doing independent tasks and are not fighting over the same thing. Second, we have skills. These are your organization's secret recipes. So make sure you have a lot of skills. Because these skills, once you start giving it to your agents, the agents will always make sure to use the correct and proper workflows to solve whatever production bug you're facing. And of course, all of us have seen a lot of MCP tools being shipped into the market right now. Everybody says the agent can connect to whatever MCP tool and whatever third-party server there is. And your client data can live in any system you want. And at the end of the day, if you have a lot of sub-agents, you have a lot of work to orchestrate. So minimal UX is the key here. Let's look at the sub-agents. With you as the orchestrator, you can have, let's say, with 48 GB of RAM on your MacBook, you can have 50 active worktrees. That is 50 active sub-agents working independently on different tasks. So where do these tasks come from? So let's say the production software you're going to ship has a lot of bugs that your QA is reporting. So all the Jira tickets can be thought of in a separate different worktrees. So different worktrees are handled by a separate agent, and these agents can spawn multiple sub-agents to solve that particular task. You don't want to queue up your tasks because the agent will do that a lot better than you. Let's look at an example harness. What if the QA reports a lot of bug tickets, and somehow magically there is an agent which passes the requirements, does a root cause analysis, pulls all the traces, pulls all the logs, puts all this in a separate work tree, does the TDD, implements the fix. Because it's in your local system, you have to do test scripts, local end-to-end testing. You create a PR, you submit the PR to your team for review. And after review, you merge it into your master branch, let's say. After merging it, obviously, you have to build a Docker image, deploy it into your development environment, test it. Again, ship build an image to your stage environment, test it, deploy it to stage. And then you go back to the QA saying, here you go, you can test it around. So I would like to ask a question to the audience. At what points do you think the human contact is required in this steps one to nine? So I would say the human is only required at steps one and nine, because in between steps, the agent can do a lot better work. There needs to be a human to see what work the agent is doing, and then there needs to be a human at the end to validate after the work is being shipped to stage. And obviously, we need minimal UX because humans love minimal UX. So in the image, if you squint your eyes and see, the image shows the production agent, software that you're building, the project dashboards, which shows all your Kubernetes services, pods, examples, all the logs, system logs, all your JIRA tickets, all your GitHub PRs, and maybe a cloud code session at the bottom. So this is basically a Mac OS widget. And you don't need to open multiple windows to do all of this work. A developer does a variety of things in their software lifecycle. So you can use just this one widget to do a lot of things. So you can see from the graph also, the number of neck rotations to ship one change reduces a lot drastically. And I imagine all of you have like two to three monitors on your table, and you just keep rotating your neck, orchestrating these agents. So auditoria works in finance. So there's a lot of regulation and policies happening in finance right now. So what does it look like for orchestrating a team of sub-agents in the finance sector? If we take AI out of the picture, usually what happens is you have a human auditor which reviews the code, and you have a controller which signs off under your SOX compliance. And reviewing agent to agent, it doesn't really keep the accountability. If something goes wrong in production, you can't say, Claude is doing this. Something is wrong. So but let's say you have all these sub-agents, and you're using these harnesses to fix bugs in real time. What is the bottleneck? It becomes a human attention because you yourself have to orchestrate all these different tasks. And moving fast and breaking things in sector, in the finance sector, is a lot different. So let's look at part two, which is removing yourself from the loop. Till now, I've been saying a human is required to see what the agent is doing, and at the end also to validate what the agent has done. But with the self-improvement of the agent and model capabilities these days, we get Fable 5 and Mythos 5 and GPT 5.6 also. So what does it look like when you have this recursive self-improvement in your internal developer harnesses? So all your production failures become input. So let's say you keep automating these developer harnesses every day, and you ask the agent to upgrade itself, essentially. So you do a task. You let the loop run, let's say, one or two days. You solve five to six bug tickets. And you just tell the agent to analyze all the bottlenecks in this process, make a list of them, and somehow slowly keep removing these bottlenecks every day. At the end of one month, let's say, you have a really nice self-automated loop where you just type in one sentence and just say, fix this bug for me. And the agent goes off, connects to all your database systems, fetches all the logs, traces tickets, and migrates it to the Jira to QA pipeline. And you can just book a vacation, maybe, or work from home. And what does it look like internally? And what happens when you steer less and ship more? Nowadays, how many of you know you can give goals to your agents? You can just set a goal and forget about it. Anybody? Nice. So what if you combine goals and loops? You can just set a goal saying there is some data discrepancy in this report. And in the production bug, the source data is not matching with what the agent has generated. So you can just set a goal to fix this, look into this, set a loop. You can even close your laptop because you can do it from your phone nowadays. And if you look at the last but one point, which is dreaming, let's say a lot of customers are using your production software. So you can do it from your agent. And they're doing the same type of patterns. And they're facing the same type of problems. So you let the agent dream like humans dream in the background so that it collects all the sessions that your customers are using and compacts it into a set of data points which your system can use and basically upgrade yourself. So with a combination of all these features, basically, you can essentially remove yourself out of the loop. But as I said before, the developers do a lot of variety of things in their software development lifecycle. And sitting behind a desk from nine to five and just writing code is not valid anymore. So just an overview of what I've covered till now in the session. You can have a team of sub-agents working in parallel work trees. You can have skills, your organizational secret recipes, your customers' recipes. You can give all of these to an agent. Your agent can connect to whatever third-party server there is. It can be a logging system. It can be an authentication gateway. And you just compress all of this into one pane of glass because minimal UX is the key. And you can set goals and loops for autonomy. If you think this particular work can be done by the agent a lot better, you can just ship it to the agent. Always have the human as a verifier, but not the throughput ceiling because human attention is very limited. So thank you for your time. And thank you for your service. I hope you learned something from this session. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you.