AI Engineer

Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon

3761 summary words 17 min summary Watch video

Start with the signal

17 min read

Summary

At-a-Glance

  • Verdict: Watch fully
  • Core thesis: RL is not needed for operational success in bounded ETL failure remediation; structured state representation, deterministic rules for facts, and external safety constraints deliver ~99.85% MTTR reduction—ablation shows RL matched (not beat) a simple deterministic policy, but provides inspectable learned decision surfaces for future context-rich incident histories.
  • Why it matters: Ken is building agent systems; this is a rare honest case study showing when NOT to use RL, how to separate learned policy from safety constraints, and how to measure agent reliability with repeated-seed ablations instead of one-shot demos.
  • Best use: Study the architecture diagram, safety layer design, and ablation results (slide ~7-9 minutes); use as template for Ken's own agent systems where action space is small, state is compact, and deterministic baselines should be tried first.

Executive Summary

Anna-Marie Benzon presents a capstone project that uses reinforcement learning (tabular Q-learning) to automate ETL pipeline failure remediation in AWS Glue environments. The baseline problem: manual incident recovery takes ~2.5 working days due to handoffs, incomplete context, and risk of unsafe fixes. The system architecture uses EventBridge to catch job-failed events, Lambda to run the agent, CloudWatch + Glue Data Catalog for read-only evidence gathering, and a three-layer intelligence stack: (1) deterministic anomaly rules for observable facts (schema drift, null rate spikes, type changes), (2) Q-learning policy for action selection (retry, coerce schema, rollback, quarantine, escalate, or log), and (3) external safety override that escalates critical anomalies if the policy proposes passive actions. Every decision is audited; escalation is treated as a valid outcome, not failure.

The critical finding from ablation studies across 30 seeds: the RL policy achieved identical success rates to a hand-coded deterministic policy (0 percentage point difference, ±0.19pp confidence interval). Deterministic action selection beat random by 15.63 points, and the safety override reduced non-escalation by 15.03 points (intentional—escalates when autonomy is inappropriate). Mean resolution time for successful cases was ~5.24 minutes vs. 216,000 seconds (2.5 days) baseline, a 99.85% MTTR reduction. However, reliability came from structured state representation, sensible decision logic, and external safety constraints—not from RL sophistication. RL's value is the inspectable learned decision surface for future richer incident histories, not immediate performance gains over deterministic rules.

The benchmark uses sanitized synthetic data (schemas, logs, incident scenarios) to preserve experimental reproducibility without exposing client context. Rule-based anomaly detector achieved precision=1.0, recall=0.8, F1=0.889 (conservative but correct). Success rate across benchmark was 74.63% ±1.51pp; non-escalation rate 88.63% ±0.89pp. Validation boundary is explicit: synthetic scenarios only, agent responds after failure (not predictive), real incident diversity may exceed current state space, some remediations are simulated, and production deployment would require shadow mode validation, strict approval gates, version policies, and rollback support.

Five engineering takeaways: (1) Use deterministic logic for directly measurable facts. (2) Use learning only where contextual action selection adds real value. (3) Place safety constraints outside the learned policy so updates cannot redefine authority. (4) Treat escalation and validation as first-class outcomes. (5) Evaluate across repeated seeds and compare against simple baselines—a single favorable run is a demo, not evidence. The goal is not to eliminate human judgment but to stop spending it on the same recognizable failure at 2 AM. Code, benchmark, and reproducibility instructions are publicly available on GitHub. Speaker explicitly seeks feedback on state representation, reward design, and safety boundaries for operational agents.

Key Takeaways

  • Claim: The RL-guided system achieved identical success rates to a deterministic policy in the controlled benchmark (0 percentage point difference, ±0.19pp confidence interval). | Evidence: Ablation study across 30 seeds (seed 42-71) showed RL policy matched hand-coded deterministic policy; deterministic beat random by 15.63 points, proving structured logic—not RL—drove success. Reported success rate 74.63% ±1.51pp, non-escalation 88.63% ±0.89pp. | Caveat: Results are from synthetic benchmark scenarios; real production incident diversity may exceed current state space, and online learning in production would require strict approval gates, shadow mode validation, and continuous monitoring before granting execution authority. | Implication: For Ken's agent systems: RL is not a magic bullet for operational tasks with small state/action spaces—start with deterministic baselines, use RL only when contextual action selection adds measurable value, and run repeated-seed ablations to separate architecture gains from learned policy gains. | Timestamp: timestamp unavailable
  • Claim: Manual ETL incident recovery baseline is ~2.5 working days (216,000 seconds); the RL-guided system reduced mean resolution time to ~5.24 minutes for successful cases (~99.85% MTTR reduction). | Evidence: Manual baseline modeled from client capstone context (queuing, investigation, approval). System architecture: EventBridge → Lambda → CloudWatch logs + Glue Data Catalog → RL policy → safety override → Glue API execution → S3 audit logs. Measured across benchmark scenarios. | Caveat: MTTR reduction applies only within the controlled benchmark scope; 'successful cases' are 74.63% of incidents—remaining 25.37% require escalation or fail validation. Production MTTR gains depend on real incident distribution, shadow mode validation, and operational team trust. | Implication: For Ken: even deterministic automation of routine failures delivers 2+ order-of-magnitude latency gains; the expensive part of incidents is not diagnosis but handoffs and risk management. Agent value comes from compressing the fast path, not handling every edge case. | Timestamp: timestamp unavailable
  • Claim: The safety layer sits outside the learned policy and overrides passive actions (e.g., 'log only') when anomaly severity is critical, converting them to escalations; escalation is treated as a valid capability, not agent failure. | Evidence: Safety override reduced non-escalation rate by 15.03 points (intentional design). Example: policy proposes 'log' for critical anomaly → override escalates. Ablation shows guarded system escalates more when autonomy is inappropriate. Every proposal, override, execution result, and validation outcome is written to audit record. | Caveat: If success is measured only by non-escalation rate, the optimization target is wrong—forcing autonomy on high-risk cases trades latency for safety. Safety layer logic is deterministic and must be updated manually; it is not part of the learned policy update loop. | Implication: For Ken's agent systems: build safety constraints as a separate layer outside the policy update loop so model updates cannot silently redefine their own authority. Design escalation as a first-class outcome; the ability to say 'I should not do this automatically' is a capability, not a deficiency. | Timestamp: timestamp unavailable
  • Claim: The intelligence layer deliberately separates three concerns: (1) deterministic anomaly rules for observable facts, (2) Q-learning policy for contextual action selection, (3) safety override outside learned policy. | Evidence: Anomaly rules: schema profiler extracts structure/types/null rates, drift detector compares to baseline, data quality analyzer checks completeness/validity/consistency, error classifier maps log patterns to failure families, risk scorer outputs operational risk level. Q-learning receives compact state (failure category, risk level, retry count, drift severity, data quality) → selects from 6 actions (retry, coerce, rollback, quarantine, escalate, log). Tabular Q-table is cheap to evaluate and every decision is inspectable. | Caveat: Deterministic rules assume observable data conditions can be captured explicitly; speaker notes 'with richer incident history, some classifiers could become learned components' but cautions 'ML-ready is not the same as ML-required—simplest reliable component should own each decision.' Current state space may not capture all real production incident diversity. | Implication: For Ken: use explicit rules for facts that can be measured directly (schema changes, type mismatches, null rate thresholds); use learning only where action outcomes vary meaningfully by context and manual preference encoding becomes difficult. The design thesis is 'rules for facts, learning for bounded choices, guardrails for authority.' | Timestamp: timestamp unavailable
  • Claim: The project includes a sanitized public benchmark with generalized synthetic schemas, logs, and incident scenarios to preserve experimental reproducibility without exposing client context; rule-based anomaly detector achieved precision=1.0, recall=0.8, F1=0.889. | Evidence: Capstone uses client-provided synthetic data; public GitHub repo uses newly generalized synthetic data with no client documents, infrastructure identifiers, or business-specific values. Ran 4 controlled experiment groups across 30 seeds (42-71) with 95% confidence intervals. Detector was conservative: flagged anomalies were correct but missed some positive cases. | Caveat: Perfect precision does not mean perfect detection—for operations, the distinction matters. Synthetic benchmark may not represent full production incident distribution. Production validation is 'the next evaluation boundary'; next step is shadow mode deployment on representative incident traces where recommendations can be compared with human decisions before granting execution authority. | Implication: For Ken: structure agent benchmarks for independent review and reproducibility; report aggregates with confidence intervals across repeated seeds, not single runs. Be explicit about validation boundaries—controlled benchmark results are credible visibility demonstrations, not production claims. Treat shadow mode as the bridge between synthetic eval and live deployment. | Timestamp: timestamp unavailable

Detailed Brief

Problem Context and Engineering Objective

  • Claims: Cloud ETL failures rarely arrive as clean, well-labeled exceptions—common cases include late/unavailable sources, schema drift, datetime incompatibilities, null rate spikes, type changes, and runtime errors not in runbooks; Manual recovery workflow: inspect logs → diagnose → repair → rerun → validate; latency comes from handoffs, incomplete context, and need to avoid unsafe fixes; Client capstone evaluation modeled manual recovery baseline at ~2.5 working days (216,000 seconds) representing normal queuing, investigation, and approval; Engineering objective: compress that loop for routine, recognizable failures while escalating cases that are uncertain, novel, or high risk
  • Evidence: Opening scenario: engineer spends all day checking logs, schema, upstream data; now past midnight asking 'what changed?'—failure is small but inspection/diagnosis/safe response/rerun/validation is expensive; AWS architecture: Glue ETL job emits job-failed event → EventBridge triggers Lambda → Lambda gathers evidence from CloudWatch (logs) + Glue Data Catalog (schema metadata) → constructs state for RL engine → policy proposes bounded response → safety layer checks → Glue API executes → S3 stores audit logs/quarantined outputs → job is rerun and validated; Operational loop: monitor → diagnose → score → decide → check safety → act → verify recovery
  • Caveats: Manual baseline is modeled, not measured from real production incidents; 2.5 days represents 'normal' case, not worst-case escalation; ETL failures in real production may have higher diversity, political/business context, or cascading dependencies not captured in synthetic benchmark
  • Implications: For Ken: agent value proposition is not 'replace human judgment' but 'stop spending human judgment on the same recognizable failure at 2 AM'—compress the fast path, escalate the novel/high-risk cases; Architecture lesson: separate read-only evidence gathering (CloudWatch logs, schema metadata) from write actions (Glue API execution), and audit every proposal/override/result for compliance and debugging

Intelligence Layer Design: Deterministic Rules, Q-Learning Policy, Safety Override

  • Claims: Intelligence layer separates three concerns: (1) deterministic anomaly rules for observable facts, (2) Q-learning policy for contextual action selection, (3) safety override outside learned policy; Anomaly rules: schema profiler (structure/types/nesting/null rates), drift detector (compare to baseline, identify additions/removals/type changes), data quality analyzer (completeness/validity/consistency), error classifier (map log patterns to failure families), risk scorer (operational risk level); Q-learning policy: receives compact state (failure category, risk level, retry count, drift severity, data quality condition) → selects from 6 actions (retry, coerce, rollback, quarantine, escalate, log); Tabular Q-learning chosen because state/action spaces are small, Q-table is cheap to evaluate, and every decision is directly inspectable: 'for this state, these were action values, this action won'; Each incident modeled as single-step contextual decision, not long-horizon control task—deliberate formulation to choose one safe operational response from bounded action set; Safety layer evaluates policy proposal against anomaly severity and operational constraints; passive actions overridden for critical conditions, high-risk/unknown cases escalated; every proposal/override/execution result/validation outcome written to audit record; Escalation included in action space—not agent giving up, but system correctly recognizing boundary of evidence or authority; 'for operational agent, ability to say I should not do this automatically is a capability'
  • Evidence: Example failure path: Glue job-failed event → log classifier detects datetime format incompatibility (0.9 confidence) → policy proposes schema coercion → safety override does not fire (not critical anomaly) → executor discovers coercion unavailable for this case → system records proposed action, reports execution unavailable, sends incident for manual review; Design thesis: 'rules for facts, learning for bounded choices, guardrails for authority'; Speaker notes: 'With richer incident history, some classifiers could become learned components, but ML-ready is not the same as ML-required—simplest reliable component should own each decision'; Ablation result: RL policy matched deterministic policy (0pp difference ±0.19pp CI); deterministic beat random by 15.63pp; safety override reduced non-escalation by 15.03pp (intentional—escalates when autonomy inappropriate)
  • Caveats: Deterministic rules assume observable conditions can be explicitly captured; production may surface conditions not in current rule set; Tabular Q-learning does not scale to high-dimensional state/action spaces; speaker chose it because spaces are small and inspectability was priority; Safety layer logic is deterministic and separate from policy update loop—must be maintained manually, not learned; Example shows two distinct controls: policy safety (is action safe in principle?) and implementation capability (is action available in environment?)—agent must represent both explicitly
  • Implications: For Ken's agent systems: start with explicit rules for directly measurable facts (schema changes, type mismatches, thresholds), use learning only where action outcomes vary meaningfully by context; Inspectability is a first-order design requirement for operational agents—every decision should be traceable to state/action values, not a black box; Escalation is not failure; build it into action space and reward model—optimization target is 'safe resolution when appropriate, escalation when not' rather than 'maximize autonomy'; Separate policy proposal from safety enforcement so policy updates cannot silently redefine their own authority

Benchmark Design, Ablation Results, and Validation Boundary

  • Claims: Sanitized public benchmark built to preserve experimental logic without exposing client context: generalized synthetic schemas, records, logs, incident scenarios; no client documents/infrastructure IDs/business values; Capstone uses client-provided synthetic data; public repo uses newly generalized synthetic data for independent reviewability; Ran 4 controlled experiment groups, repeated robustness evaluation across 30 seeds (42-71); reported aggregates include 95% confidence intervals; Rule-based anomaly detector: precision=1.0, recall=0.8, F1=0.889—conservative but correct; 'perfect precision does not mean perfect detection, distinction matters for operations'; RL-guided workflow: mean resolution time ~5.24 minutes for successful cases; success rate 74.63% ±1.51pp; non-escalation rate 88.63% ±0.89pp; Compared to manual baseline (216,000 sec), minute-scale result is ~99.85% MTTR reduction within benchmark scope; Ablation: RL policy matched deterministic policy (0pp ±0.19pp); deterministic beat random by 15.63pp; safety override reduced non-escalation by 15.03pp (intentional); Reliability came primarily from structured state, sensible decision logic, external safety constraints—not from RL alone; 'useful engineering result'; RL provides inspectable learned decision surface rather than immediate success rate advantage; value becomes significant as incident histories get richer, action outcomes vary by context, manual preference encoding becomes difficult; Validation boundary: synthetic scenarios only, agent responds after failure (not predictive), real incident diversity may exceed state space, some remediations simulated/bounded, online learning would require strict approval gates/version policies/rollback support/continuous monitoring; Next step: shadow mode deployment on representative incident traces where recommendations compared with human decisions before agent receives execution authority
  • Evidence: Speaker: 'Where did reliability come from? Primarily from structured state, sensible decision logic, external safety constraints, not from RL alone. That is a useful engineering result.'; Speaker: 'In current benchmark, RL provides inspectable learned decision surface rather than immediate success rate advantage. Its value becomes more significant as incident histories become richer.'; Speaker: 'A single favorable run is a demo, not evidence'—emphasizes repeated-seed evaluation and simple baseline comparisons; Five takeaways for engineering teams: (1) use deterministic logic for directly measurable facts, (2) use learning only where contextual action selection adds real value, (3) place safety constraints outside learned policy, (4) treat escalation/validation as first-class outcomes, (5) evaluate across repeated seeds and compare against simple baselines
  • Caveats: Benchmark uses synthetic data; production incident distribution may differ; Success rate 74.63% means 25.37% of incidents require escalation or fail validation—agent handles fast path, not all cases; MTTR reduction is within benchmark scope; production gains depend on real incident distribution, team trust, shadow mode validation; Deterministic baseline already matched RL policy in current benchmark—RL value is future-looking (richer histories, varied contexts) not immediate; Online learning in production would require additional safety infrastructure not yet demonstrated
  • Implications: For Ken: structure benchmarks for reproducibility with sanitized data, repeated seeds, confidence intervals; make validation boundaries explicit; Ablation studies are critical—compare learned policy against deterministic baseline and random baseline to isolate where value comes from; Report both success rate and non-escalation rate; if escalation rate goes up with safety overrides, that is correct behavior, not failure; Shadow mode is the bridge between synthetic eval and live deployment—run recommendations alongside human decisions, measure agreement, identify drift/novel cases before granting execution authority; Honest engineering result: RL did not beat deterministic policy in this case, but provides inspectable learned surface for future context-rich scenarios—this is more valuable for Ken than overstated claims

Notable Concepts & Terms

  • Bounded remediation action: Agent action space is explicitly limited to 6 safe operations (retry, coerce schema, rollback, quarantine, escalate, log); system cannot invent arbitrary fixes or make unbounded changes—design constraint for operational trust.
  • Safety override / guardrails outside learned policy: External deterministic layer that evaluates policy proposals and overrides passive actions (e.g., 'log only') when anomaly severity is critical, converting to escalation; sits outside RL update loop so policy cannot redefine its own authority.
  • Escalation as first-class outcome: Agent action space includes 'escalate' as valid choice; system is designed to recognize boundary of evidence/authority and defer to humans for uncertain/novel/high-risk cases—not treated as failure but as correct operational behavior.
  • Tabular Q-learning for single-step contextual decision: Each incident modeled as one decision (not multi-step control task); Q-table maps compact state (failure category, risk level, retry count, drift severity, data quality) to 6 actions. Chosen for small state/action space and full inspectability—every decision traceable to Q-values.
  • Ablation study with repeated seeds and confidence intervals: Experimental method: run 4 groups across 30 seeds (42-71), compare RL policy vs. deterministic baseline vs. random baseline, report aggregates with 95% CI. Core finding: RL matched deterministic (0pp ±0.19pp), proving reliability came from architecture not RL sophistication.
  • Shadow mode deployment: Next validation step: run agent recommendations alongside human decisions in production without granting execution authority; measure agreement, identify drift/novel cases, build operational trust before live deployment.
  • MTTR (Mean Time To Resolution): Manual baseline ~2.5 working days (216,000 sec) vs. agent ~5.24 minutes for successful cases = ~99.85% reduction within benchmark scope. Key caveat: applies only to 74.63% of incidents that agent can handle; remaining 25.37% escalated.
  • ML-ready vs. ML-required: Speaker's design principle: just because a component could be learned does not mean it should be—use simplest reliable component for each decision. Deterministic rules for directly observable facts, learning only where contextual action selection adds measurable value.

Operator Notes / Why Ken Should Care

  • Critical for Ken's agent systems: this is a rare honest case study showing when NOT to use RL—ablation proved RL matched (not beat) deterministic policy in small state/action space. Reliability came from structured state, sensible decision logic, and external safety constraints, not RL sophistication. RL's value is inspectable learned decision surface for future richer incident histories, not immediate performance gains.
  • Safety architecture: build safety constraints as separate layer outside policy update loop so model updates cannot silently redefine their own authority. Design escalation as first-class outcome—agent that says 'I should not do this automatically' is more valuable than agent that optimizes for non-escalation at any cost.
  • Benchmarking discipline: run repeated seeds (not single runs), report confidence intervals, compare against simple deterministic baseline and random baseline to isolate where value comes from. Make validation boundaries explicit—controlled benchmark results are not production claims. Use shadow mode to bridge synthetic eval and live deployment.
  • For Ken's content/investing: this talk is a counterexample to AI hype—speaker shows RL did not outperform deterministic rules, but correctly positions RL as inspectable learned surface for future context-rich cases. Honesty about validation boundaries and next steps (shadow mode) is rare and valuable.
  • Operational agent design thesis: 'rules for facts, learning for bounded choices, guardrails for authority'—use explicit logic for directly measurable conditions (schema drift, type changes, null rate thresholds), use learning only where action outcomes vary meaningfully by context, and enforce safety outside the learned policy. Every decision must be auditable.
  • For Ken's GTM/agent workflow: the value proposition is not 'replace human judgment' but 'stop spending human judgment on the same recognizable failure at 2 AM'—compress the fast path for routine incidents, escalate novel/high-risk cases. Manual MTTR was 2.5 days due to handoffs/incomplete context/risk management, not inherent technical difficulty. Agent reduces handoff latency, not cognitive complexity.
  • Code/benchmark available on GitHub (speaker requests feedback on state representation, reward design, safety boundaries); public repo uses sanitized synthetic data for reproducibility without exposing client context. Next step is shadow mode on representative incident traces—compare recommendations with human decisions before granting execution authority.

Watch Map

  • timestamp unavailable: Opening scenario: engineer at midnight debugging ETL failure—latency is inspection/diagnosis/safe response, not fix itself
  • timestamp unavailable: AWS architecture diagram: EventBridge → Lambda → CloudWatch + Glue Data Catalog → RL policy → safety override → Glue API → S3 audit logs
  • timestamp unavailable: Intelligence layer: deterministic anomaly rules (schema profiler, drift detector, data quality analyzer, error classifier, risk scorer) → Q-learning policy (6 actions: retry, coerce, rollback, quarantine, escalate, log) → external safety override
  • timestamp unavailable: Example failure path: datetime format incompatibility → policy proposes coercion → executor discovers coercion unavailable → records proposal, reports unavailable, escalates (shows policy safety vs. implementation capability)
  • timestamp unavailable: Sanitized public benchmark: generalized synthetic schemas/logs/incidents, 4 experiment groups, 30 seeds (42-71), 95% CI; anomaly detector precision=1.0, recall=0.8, F1=0.889
  • timestamp unavailable: Results: RL-guided MTTR ~5.24 min vs. manual 216,000 sec = 99.85% reduction; success rate 74.63% ±1.51pp, non-escalation 88.63% ±0.89pp
  • timestamp unavailable: Ablation: RL matched deterministic (0pp ±0.19pp), deterministic beat random by 15.63pp, safety override reduced non-escalation by 15.03pp (intentional)—reliability from structured state/logic/safety, not RL
  • timestamp unavailable: Validation boundary: synthetic scenarios, agent responds after failure (not predictive), real diversity may exceed state space, some remediations simulated, production would require shadow mode/approval gates/rollback support
  • timestamp unavailable: Five takeaways: (1) deterministic logic for measurable facts, (2) learning only where contextual action adds value, (3) safety outside learned policy, (4) escalation/validation as first-class outcomes, (5) repeated seeds + simple baselines
  • timestamp unavailable: Closing: goal is not eliminate human judgment but stop spending it on same recognizable failure at 2 AM; code/benchmark on GitHub, speaker requests feedback on state representation/reward design/safety boundaries

Source/Metadata

  • Title: Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon
  • Transcript words: 1905
  • Duration seconds: 880
  • Timestamp note: Timestamps were not present in the transcript; watch_map notes reference conceptual segments and slide content rather than specific minute marks.
Full transcript 1825 words · 8 min read
0:00

SPEAKER_00

Imagine you are this engineer. A production data job failed hours ago. The dashboard went stale. You have spent all day checking the logs, the schema, and the upstream data. And now it is past midnight. The same question keeps coming back. What changed? The failure itself may be small, but the expensive part is everything around it. Inspection, diagnosis, choosing a safe response, rerunning the job, and confirming that we did not make the data worse.

0:34

SPEAKER_00

Hi, I'm Anna-Marie Benzon. In this talk, I will show an RL-guided system that selects bounded remediation action for ETL failures. The central question is not simply whether an agent can act, but whether it can act usefully, explainably, and within boundaries that an operations team would actually trust. Cloud ETL failures are rarely arrived as one clean, well-labeled exception. We see late or unavailable sources, schema drift, datetime incompatibilities, null rate spikes, type changes, and runtime errors that do not match anything in the runbook.

1:14

SPEAKER_00

The usual response is a human workflow. Inspect the logs, form a diagnosis, attempt a repair, rerun the job, and validate the output. Each step is reasonable. The latency comes from handoffs, incomplete contacts, and the need to avoid an unsafe fix. In the capstone evaluation, the manual recovery baseline was modeled at roughly 2.5 working days. This represents an incident moving through normal queuing, investigation, and approval. So the engineering objective is specific. Compress that loop for routine, recognizable failures while escalating the cases that are uncertain, novel, or high risk.

1:59

SPEAKER_00

This diagram shows the end-to-end AWS architecture from my capstone. An existing AWS Glue ETL job emits a job-failed event. Amazon EventBridge catches that event and triggers the Lambda function that runs the agent. Lambda gathers evidence from two read-only sources. CloudWatch provides the error logs while the Glue data catalog provides the current schema metadata. The system uses those signals to classify the failure, assess the data quality and operational risk, and construct the state pass to the RL decision engine. The policy then proposes a bounded response.

2:39

SPEAKER_00

The safety layer checks that proposal before the executor can use the Glue API to re-trigger the job or apply an approved remediation. Amazon S3 stores agent artifacts, audit logs, and quarantined outputs. Finally, the job is rerun and validated. So this is a close operational look. Monitor, diagnose, score, decide, check safety, act, and verify recovery. The capstone implementation uses synthetic data provided by the client. The public repository preserves this pattern through a sanitized, generalized deployment template. The intelligence layer deliberately separates three concerns. Deterministic anomaly rules establish observable facts.

3:31

SPEAKER_00

A field disappeared, a type change, or the null rate crossed a threshold. The Q-learning policy handles contextual action selection. Given the current incident state, should the system retry, coerce the schema, roll back, quarantine, escalate, or simply log the event? Then, a safety override sits outside the learned policy. For example, if the anomaly is critical and the policy proposes a passive action, such as logging, the override converts that choice into an escalation. This separation is the design thesis of the project. Rules for facts, learning for bounded choices, and guardrails for authority.

4:10

SPEAKER_00

Before selecting an action, the system has to establish what actually happened. The schema profiler extracts structure, types, nesting, and null rate statistics. The drift detector compares the current profiler with a baseline and identifies additions, removals, and type changes. The data quality analyzer checks completeness, validity, and consistency. The error classifier maps log patterns into failure families. And the risk scorer turns those signals into an operational risk level. These components are deterministic by design. For directly observable data conditions, an explicit rule is easier to validate, explain, and audit than an opaque inference.

4:57

SPEAKER_00

With richer and representative incident history, some classifiers could become learned components. But ML-ready is not the same as ML-required. The simplest reliable components should own each decision. The policy receives a compact state: failure category, risk level, retry count, drift severity, and data quality condition. It then selects from six actions. Retry, coerce, rollback, quarantine, escalate, or log. I use tabular Q-learning because the state and action spaces are small. The Q-table is cheap to evaluate, and every decision can be inspected directly. For this state, these were action values, and this action won.

5:40

SPEAKER_00

Technically, each incident is modeled as single-step contextual decision implemented with tabular Q-learning, rather than as a long horizontal control task. That formulation is deliberate. The system needs to choose one safe operational response from a bounded action set. The value of the learned policy here is not sophistication for its own sake. It is a structured way to learn action preferences from outcomes while retaining a decision surface that an engineer can inspect. The learned policy does not have final authority. It proposes an action. The safety layer evaluates the proposal against the anomaly severity and the system's operational constraints.

6:24

SPEAKER_00

Passive actions are overridden for critical conditions, and high-risk or unknown cases are escalated. Every proposal, override, execution result, and validation outcome is written to an audit record. Notice that escalation is included in the action space. That's not the agent giving up. It is the system correctly recognizing the boundary of its evidence or authority. For an operational agent, the ability to say, I should not do this automatically, is a capability. If success is measured only by non-escalation, the optimization target is wrong. Here is one failure path. The agent receives a glue-style job failure event.

7:09

SPEAKER_00

The log classifier detects a datetime format incompatibility with 0.9 confidence. Based on the encoded state, the policy proposes schema coercion. The safety override does not fire because this is not classified as a critical anomaly. But the executor then discovers that automatic coercion is not available for this specific case. The system does not pretend that the fix happened. It records the proposed action, reports that execution was unavailable, and sends the incident for manual review. This example shows two distinct controls. Policy safety and implementation capability. An action can be safe in principle and still be unavailable in the current environment.

7:55

SPEAKER_00

A robust agent must represent both conditions explicitly. To make the work independently reviewable without exposing the client context, I build a sanitized public benchmark around the generalized AWS Lambda-style architecture. The capstone implementation uses client-provided synthetic data. The public repository uses newly generalized synthetic schemas, records, logs, and incident scenarios. It contains no client documents, infrastructure identifiers, or business-specific values. I ran four controlled experiment groups and repeated the robustness evaluation across 30 seeds from 42 through 71. The reported aggregates include 95% confidence intervals.

8:35

SPEAKER_00

This preserves the system design and experimental logic in a form that other engineers can inspect and rerun while maintaining the confidentiality boundary. On the controlled benchmark, the rule-based anomaly detector achieved precision of 1, recall of 0.8, and an F1 score of 0.889. That means the detector was conservative. The anomalies it flagged were correct in this benchmark, but it still missed some positive cases. For operations, that distinction matters. Perfect precision does not mean perfect detection. For cases where the RL-guided workflow resolved the incident successfully, mean resolution time was about 5.24 minutes.

9:23

SPEAKER_00

Across the 30 runs, the simulated success rate was 74.63%, plus or minus 1.51 percentage points. The non-escalation rate was 88.63%, plus or minus 0.89 points. The chart compares that minute-scale result with a modeled manual baseline of 2.5 working days, which is 216,000 seconds. Within the benchmark, that is approximately a 99.85% reduction in MTTR. These figures quantify performance within the controlled benchmark. Within that scope, they show that the architecture can automate the fast path for known failure conditions. Production validation is the next evaluation boundary. The ablation results are, in my view, the most useful part of the project.

10:18

SPEAKER_00

The RL policy matched the equivalent deterministic policy, a difference of 0 percentage points within a 0.19 point confidence interval. On this compact state space, the learned policy maintained the same success level as the hand-defined policy. By contrast, deterministic action selection beat random selection by 15.63 points, and enabling the safety override reduced non-escalation by about 15.03 points. That decrease is intentional. The guarded system escalates more often when autonomy would be inappropriate. So, where did the reliability come from? Primarily from structured state, sensible decision logic, and external safety constraints, not from RL alone.

11:04

SPEAKER_00

That is a useful engineering result. In the current benchmark, RL provides an inspectable learned decision surface rather than an immediate success rate advantage. Its value becomes more significant as incident histories become richer. Action outcomes vary by context. And manually maintaining every preference becomes difficult. This slide defines the current validation boundary. The results come from synthetic scenarios. The agent responds after a failure signal. It does not predict a failure before it happens. Real incident diversity may exceed the current state space. Some remediation actions are simulated or deliberately bounded.

11:52

SPEAKER_00

And online learning in a production environment would require strict approval gates, version policies, rollback support, and continuous monitoring. The result is credible visibility demonstration of the system design with a clear path toward production validation. The next step is a shadow mode deployment on representative incident traces, where recommendations can be compared with human decisions before the agent receives execution authority. There are five takeaways I would leave with an engineering team. First, use deterministic logic for facts that can be measured directly. Second, use learning only where contextual action selection adds real value.

12:36

SPEAKER_00

Third, place safety constraints outside the learned policy. So a policy update cannot silently redefine its own authority. Fourth, treat escalation and post-action validation as first-class outcomes, not exception paths. And fifth, evaluate across repeated seeds and compare against simple baselines. A single favorable run is a demo, not evidence. A practical self-healing system does not need the largest possible model. It needs a clear state, bounded action, reproducible evaluation, observable decisions, and the discipline to stop when uncertainty exceeds its authority. This brings us back to the engineer in the opening video. The goal is not to eliminate human judgment.

13:25

SPEAKER_00

It is to stop spending that judgment on the same recognizable failure at 2 in the morning. Before, the response is manual log inspection, schema tracing, delayed dashboards, and recovery process measured in working days. After, the routine path becomes event-triggered diagnosis and RL-guided but safety-constrained action. Explicit validation and recovery measured in minutes when the case is supported. The unusual or high-risk failures still go to the humans. That is the point. Human attention is reserved for incidents where context, trade-offs, or authority genuinely require it.

14:06

SPEAKER_00

The code, synthetic benchmark, experiment scripts, and reproducibility instructions are available in the GitHub repository on screen. If you work on agent reliability, data quality, or protection incident automation, I would especially value your feedback on state representation, reward design, and safety boundary. Thank you for watching. If you work on agent reliability, data quality, or protection incident automation, I would especially value your feedback on state representation, reward design, and safety boundary. Thank you for watching name

Reading tools

Type to find a passage

Appearance
Ask this transcript

Add a note