Skip to content Skip to sidebar Skip to footer

Ultimate LLMOps Bootcamp – Production LLM Systems

Ultimate LLMOps Bootcamp – Production LLM Systems

Published 9/2026
MP4 | Video: h264, 1920×1080 | Audio: AAC, 44.1 KHz, 2 Ch
Language: English | Duration: 12h 30m | Size: 1.87 GB

Build, evaluate, deploy, operate real LLM systems- RAG, evals, fine-tuning, K8s, guardrail, agents. 8 GB laptop, no GPU.

What you’ll learn
Explain how tokens, prefill, decode, attention and the KV cache drive your latency, memory and cost
Choose an open model on measured quality, licence, format and memory — not on benchmark headlines
Serve models behind an OpenAI-compatible API and compare Ollama, llama.cpp and vLLM in practice
Build a RAG pipeline and diagnose chunking, retrieval, grounding and prompt-injection failures
Create a golden set and an eval baseline that turn “it looks better” into a ship-or-reject decision
Run a laptop-scale LoRA fine-tune and decide honestly whether the result deserves to ship
Package, version, sign and verify model artifacts with an OCI supply-chain workflow
Deploy and canary an LLM on Kubernetes, then promote or roll back on quality evidence
Measure TTFT, ITL, queue depth, traces and token cost without trusting a misleading average
Autoscale on a queue rather than CPU, and manage model versions with GitOps
Put a gateway in front: virtual keys, budgets, caching, fallbacks, guardrails and CI eval gates
Operate a bounded tool-calling agent over MCP, with a read-only tool boundary and an audit trail

Requirements
A laptop with at least 8 GB RAM and around 40 GB free disk for local models and containers
Comfort with the command line, Git, containers and Docker Compose
Helpful but not required: basic Kubernetes (Deployments, Services, kubectl) for the later modules
No machine learning or data science background is needed — the course builds the model theory it uses
No GPU, no cloud account, no paid API key and no Docker Desktop subscription required

Description

Most LLM courses stop where production work starts.
You get a prompt that returns a good answer, and then the hard questions begin. Is this change actually better, or does it just read better? What blocks a release? How do you package a model so someone else can verify exactly what shipped? When the assistant gets slow, or expensive, or wrong, how do you find out before your users tell you?

This bootcamp answers those questions by building one system and operating it the whole way through.

You build one assistant, called OpsMate, and take it through the complete LLMOps lifecycle.
You start with what a model actually does — tokens, prefill and decode, the KV cache, quantization — and why those decide your latency, memory and bill. You serve it behind an OpenAI-compatible API. You give it your own documents with RAG. You build a golden set and record a baseline, so “better” becomes a number instead of an opinion. You fine-tune with LoRA and then let the evaluation gate decide whether it ships. You package and sign the model as a versioned artifact, deploy it to Kubernetes, canary it, and promote it on quality evidence rather than on a green deploy. You add observability, autoscaling, GitOps, a gateway with budgets and guardrails, and finally a bounded tool-calling agent.

The labs are built around evidence, not around everything working.
You will watch a weak prompt lose an A/B test. You will watch a fine-tuned model fail its quality gate and get blocked — which is the gate doing its job. You will hit the promotion lie, where the tag moves and the bytes do not. You will apply a bad liveness probe on purpose and watch it turn a slow model load into a restart loop. Every one of those is a real failure from building this course, and each one teaches a decision you will have to make for real.

Everything required runs on an average 8 GB laptop.
No GPU. No cloud account. No paid API. No Docker Desktop subscription. The whole course uses open models and local infrastructure, and every lab tells you its resource path up front. Kubernetes shows up later as one deployment environment, taught on a local cluster — this is a Production LLMOps course, not a Kubernetes administration course.

What you walk away with
A working system you built yourself, and a repeatable method for the question that actually matters in this job:is this model change ready to promote, and what is my evidence?

Who this course is for
DevOps, SRE and platform engineers moving into LLM and GenAI operations
Backend and application engineers who need to take an LLM feature past a local demo
MLOps engineers extending from predictive models into open-model LLM systems
Technical leads who need a defensible model-release and operating framework
Engineers who want real production practice without first buying a GPU or a cloud subscription
Not for: learners after prompt-writing tips, a cloud certification, or a Kubernetes introduction

UMATREORIBOTARFDGECAPOEPRODUCTEN

you must be registered member to see linkes Register Now

Leave a comment