AI-SRE / AgentOps
Measure token cost per incident, latency, quality and false positives through the internal Aurora pilot.
Build the team capability behind Infra Fox’s real wedge: AI-SRE / AgentOps, safe-agent deployment and GenAI FinOps—grounded in AWS, measured on internal systems and proven before client rollout.
General AWS/DevOps remains the base-load. This curriculum develops only the capabilities needed to earn credibility in the Agentic Ops wedge.
Measure token cost per incident, latency, quality and false positives through the internal Aurora pilot.
Turn the OpenClaw work into a repeatable safe-agent deployment assessment.
Extend existing cloud cost credibility into model and agent unit economics.
Broad GenAI app building, new verticals and unrelated AI requests that do not strengthen the wedge.
Teach them in a spiral through the pilot. They are gates to safe judgment, not a six-week classroom prerequisite.
LLMs, tokens, context windows, prompts, tools, agents, workflows, state, memory and human approval boundaries.
IAM, VPC, Lambda/ECS, Step Functions, Bedrock, AgentCore, CloudWatch, KMS and Secrets Manager.
Signals, failure modes, retries, timeouts, idempotency, queues, escalation, SLI/SLO and incident response.
Structured events, logs, metrics, OTEL traces, model/tool spans, latency, quality, false positives and attribution.
Prompt injection, tool abuse, excessive agency, data leakage, identity, sandboxing, guardrails and auditability.
Tokens, model routing, cache, tool-call cost, unit economics, budgets, anomaly detection and cost per outcome.
The whole team reaches L2 across the map. Builders reach L3. The pod lead and Suresh reach L4 in the wedge; L5 follows real client evidence.
Describe the problem and place the capability on the Agentic Ops map.
Explain architecture, data flow, trust boundaries, cost and failure modes.
Build and troubleshoot with docs/AI; produce evidence and a runbook.
Design production trade-offs, lead incidents and review others’ work.
Advise clients, define standards, package offers and mentor the team.
Each module uses Problem → Concepts → Build → Explain → Verify → Break → Reflect. Nothing is complete until its gate is passed.
Trace user intent → model → tool → AWS resource → observation → approval → outcome.
Agents vs workflows · autonomy · state · tool contracts · human-in-the-loop
Diagram the proposed AI-SRE loop and threat boundaries before writing code.
Given an unsafe tool call, identify which boundary failed.
Select models using quality, latency, context and cost—not popularity.
Bedrock · inference · token accounting · model routing · quotas · privacy
Create a small evaluation harness comparing two models on ops tasks.
Diagnose throttling, runaway tokens and degraded answer quality.
Know when Step Functions is safer than an open-ended agent loop.
State machines · retries · timeouts · idempotency · compensation · approvals
Orchestrate a read-only diagnostic workflow with explicit approval gates.
Recover from duplicate execution and partial tool failure.
Build bounded, inspectable agent behavior with explicit state and exits.
Graphs · nodes · edges · state · memory · tool calling · termination
Implement symptom → evidence → hypothesis → recommendation with no write access.
Stop a looping agent and prove the termination safeguard.
Deploy an agent with durable identity, isolation, observability and lifecycle ownership.
Runtime · identity · gateway/tools · memory · policy · environment separation
Deploy the internal Aurora pilot in a sandboxed AWS environment.
Repair missing permissions without broadening access unnecessarily.
See one request across model, tool, infrastructure and business outcome.
Traces · spans · correlation · CloudWatch · latency · quality · false positives
Instrument model calls, tool calls and approval waits; create an ops dashboard.
Find a latency regression using traces rather than guesses.
Turn LLM usage into an operational unit metric suitable for pricing and governance.
Attribution · unit cost · budgets · anomaly detection · showback · optimization
Create the reusable Agent Cost Governance dashboard.
Investigate a 3× cost spike and distinguish usage, model and retry causes.
Design least-agency systems that resist hostile input and contain tool misuse.
Prompt injection · Bedrock Guardrails · IAM · data boundaries · egress · audit
Convert OpenClaw findings into an assessment checklist and test sandbox.
Respond to a prompt-injection attempt that reaches a privileged tool.
Ground answers in owned runbooks while enforcing authorization and citations.
Knowledge Bases · chunking · retrieval · metadata · access control · evaluation
Create an assistant over approved Infra Fox runbooks and internal documents.
Diagnose a confident answer sourced from stale or unauthorized content.
Translate internal evidence into one safe, repeatable client engagement.
Discovery · maturity · risk · architecture · economics · roadmap · boundaries
Run the audit with one warm design partner after internal metrics exist.
Handle a no-access or low-data client environment without inventing certainty.
The pilot is the learning environment, evidence generator and future demo asset. It runs against Infra Fox’s own environment before any client system.
Do not skip to a client pilot. First establish the internal loop and cost-governance dashboard; only then offer one AI-SRE Readiness Audit to a warm design partner.
Because agentic systems fail in new ways, this specialization deliberately gives more weight to troubleshooting and safety review.
Every project or module ends with evidence, an architecture review, a failure exercise and a reusable artifact.
Can the engineer draw model, tools, data, identities, networks, state and approvals?
Are timeouts, retries, idempotency, fallbacks, failure domains and escalation explicit?
Are prompt injection, excessive agency, secrets, data exposure and audit risks tested?
Are quality, latency, false-positive and cost claims supported by reproducible measurements?
Is token cost tied to an incident or business outcome, with budgets and ownership?
Can someone else deploy, observe, stop, recover and roll back using the runbook?
Can the submitter explain, verify and modify every AI-generated artifact?
Are constraints, uncertainty, trade-offs and recommendations communicated honestly?
Own modules 03–08 and the Aurora pilot. One builder, one reviewer; rotate after the first validated milestone.
Complete foundation modules, shadow milestones, reproduce incidents and join weekly brown-bags.
Own positioning, architecture gates, design-partner decision and strategic distillation into public content.
Capture decisions, protect client data, follow approval gates and teach one verified lesson back to the team.
Whole team · baseline now
Target: L1–L2 vocabulary and AWS AI mapPilot builder · during pilot
Target: supporting ML delivery and operations depthSuresh + one lead · after pilot begins
Target: architecture and delivery depth backed by pilot evidenceInternal evidence—builds, incidents, reviews and teach-backs—sets the L-level. A certificate is supporting evidence only.
Baseline the team · complete modules 01–03 · start Aurora pilot · publish the already-earned OpenClaw security article · first brown-bag.
Gate: architecture approved and one read-only diagnostic flow running.Public content is the final stage of knowledge capture—not the starting point. Real numbers, no client names and no AI slop remain non-negotiable.
Aurora pilot findings and OpenClaw assessment work · every 4–6 weeks.
Token-cost dashboards and unit economics in the proven case-study format · every 4–6 weeks.
Continue useful AWS FinOps and security content from existing client work at the current cadence.