Agent Observability & Tracing: Debug Production LLM Systems
Published 9/2026
Created by Frank Robotics Lab
MP4 | Video: h264, 1920×1080 | Audio: AAC, 44.1 KHz, 2 Ch
Level: Intermediate | Genre: eLearning | Language: English | Duration: 71 Lectures ( 7h 35m ) | Size: 4 GB
Instrument a real agent with OpenTelemetry, catch silent failures, judge live traffic, debug a 3-bug incident.
What you’ll learn
Instrument a real tool-calling agent with OpenTelemetry: spans, attributes, events, nesting and context propagation
Stand up Arize Phoenix and Langfuse yourself with Docker, and compare both on one identical trace
Catch silent failures that never raise: empty results, hidden retries, graceful fallbacks and confidently wrong answers
Attribute latency and cost per span, per feature, per user and per tenant, and see why run-level averages hide regressions
Build an LLM-as-judge on sampled live traffic, wire its verdicts back into the trace, and catch the judge being wrong
Detect prompt injection and quality drift from trace patterns rather than from the prompt text
Debug an unfamiliar trace under time pressure, correlate it to a user complaint, and diff a good run against a bad one
Work a multi-cause incident end to end: three seeded bugs, one red herring, a verified fix and an honest postmortem
Price observability at real scale and decide what to sample using arithmetic instead of instinct
Roll this out on a team: ownership, code review for instrumentation, and a runbook built from a real incident
Requirements
Comfortable Python: functions, classes, dictionaries, reading a traceback
Any OpenAI-compatible API key. Total spend across the whole course is a few dollars
Docker and Docker Compose installed, for the two tracing backends
No prior OpenTelemetry or observability experience needed. It is built from the first span
Description
This course contains the use of artificial intelligence.
Your agent returned a confident, well formed, completely wrong answer. Every span in the trace says status OK. Your eval suite passed this morning. Nothing is red, and something is badly broken.
An eval score is a photograph. Production is a video. This course is seventy one lectures about closing that gap, and every number in it was measured on a real system rather than typed because it sounded right.
What you actually build. A real tool-calling agent (search, fetch, calculate, summarise) instrumented with OpenTelemetry from scratch. Two real self-hosted backends, Arize Phoenix and Langfuse, compared on one identical trace. An LLM-as-judge running on sampled live traffic. A review queue. A runbook. Around twenty five small diagnostic tools, each written to answer one question the traces raised.
The failures are seeded and real. A search tool that always succeeds and always returns nothing, so the model rephrases six times and burns its budget. A fetch that serves the right document with stale numbers, where the agent does flawless arithmetic and lands eighty nine percent low. A graceful fallback that hides the truth. A prompt injection arriving through a poisoned document rather than a user prompt. None of them raise an error.
Then a capstone incident. Three bugs of three different kinds, no hints. One is a red herring: the loudest thing in the latency data, and not your problem. You learn to abandon it properly, by measuring what it leaves unexplained. The third bug has no error, no slow span, no anomaly at all, and nothing in any dashboard this course builds will flag it. It is visible only by adding up numbers the system already gives you and noticing they disagree.
Including the parts that did not work. When I scored every detector built in this course against that capstone, seven of thirteen contributed nothing. The LLM judge did not merely miss the bug, it returned grounded equals true and was right to. That result is in the course rather than hidden, because your coverage is always narrower than it feels.
Measured, not asserted. Storing every span for a million agent runs a day costs four dollars and thirty two cents a month, which makes sampling traces to save money the wrong instinct. A judge flipped its verdict ten times out of ten purely because the evidence arrived in a different order. Every figure quoted on screen is read off a measurement ledger that ships with the course, so you can check me.
Ten sections, seventy one lectures, roughly seven and a half hours, built on one small system you can run end to end on a laptop.
Who this course is for
Engineers already shipping something with an LLM in it, who have been told it is broken and had no idea how to find out why
Anyone whose agent passes its evals and still misbehaves in production
Backend and platform engineers who know normal APM but have not traced a non-deterministic tool-calling loop
ML and AI engineers who write evals and want the live half of the picture
Not for beginners looking for a first Python course, and not a prompt engineering course
AGTAOBERFRTRADFERDEDBUEGFRAPRODEU

you must be registered member to see linkes Register Now