Marsof Academy
Course 4 · Unit 4

Engineering LLM applications in production

Part of Applications with LLMs

  • 12 lessons
  • ≈ 40 h of study
  • Level: intermediate

The daily work of an AI engineer. Versioned prompts, structured outputs, tools, evals, guardrails, cost, observability, RAG and agents in production, and an end-to-end support assistant.

Topics covered

  • Prompt engineering
  • Prompt templates
  • Structured outputs
  • JSON Schema
  • Function calling
  • Idempotency
  • Evals
  • LLM-as-judge
  • Bootstrap
  • Guardrails
  • Prompt injection
  • PII
  • Prompt caching
  • Model routing
  • Streaming
  • Traces
  • LLMOps
  • Citations
  • Reindexing
  • Agents in production
  • Checkpoints
  • Fine-tuning vs RAG
  • Multimodal
  • Voice

Lessons in this unit

  1. Serious prompt engineering — templates, few-shot and versions 65 min
    A production prompt isn't a clever sentence, it's code. You'll learn its anatomy, when examples help, how to separate instructions from data and how to version it so nobody breaks it on a Friday afternoon.
  2. Structured outputs — JSON you can trust 65 min
    Your code needs data, not prose. You'll learn the three ways to ask a model for JSON, what each one guarantees, how to extract and validate it, and how to build a repair loop that doesn't spin forever.
  3. Function calling in depth — tools that hold up in production 90 min
    You already know how to validate arguments. Here you make the leap to production with good tool design, parallel calls, retryable errors and those that aren't, idempotency so you don't charge twice and call budgets.
  4. Evals for LLM applications — knowing whether your change improves anything 95 min
    "It seems to work better" isn't a metric. You'll build a golden set with automatic checks, understand an LLM judge's biases, set up a regression gate for CI and learn to decide with a confidence interval whether one prompt is better than another.
  5. Guardrails and security — layered defence 95 min
    Sooner or later, your application will read hostile text. You'll learn what direct and indirect injection and jailbreaks look like, and build a layered defence with an input filter, marked data, leak detection, a link allowlist and pseudonymisation of personal data.
  6. Cost and latency — the bill and the clock 65 min
    Every token costs money and time. You'll learn to compute the cost and latency of a request and of a conversation, and to bring them down with prompt caching, semantic caching, batching, model cascades and streaming, knowing what each technique risks.
  7. LLMOps observability — seeing what happens inside 90 min
    When a user says "the assistant gave me a nonsense answer", you need to be able to find that request and see every step. You'll learn to instrument with traces and spans, to compute latency and cost per feature, to log without violating privacy and to close the loop between production and evals.
  8. RAG in production — freshness, permissions and citations that hold up 70 min
    Building a RAG that works in a demo is easy. Keeping it correct when documents change, each user can see different things and answers must cite real sources is the real work. You evaluate chunking strategies, plan incremental reindexing and verify citations sentence by sentence.
  9. Agents in production — budgets, humans and resumption 75 min
    An agent that works on your laptop can spend 400 dollars in one night, hang waiting for a confirmation or repeat a payment when it restarts. You learn to give it step, cost and time budgets, human approval points and checkpoints so it can resume without repeating effects.
  10. Prompting, RAG or fine-tuning? Deciding with judgment 70 min
    "We should fine-tune" is one of the most expensive sentences you can hear on an AI team. You learn a framework for deciding between prompting, RAG and fine-tuning, how to prepare a clean SFT dataset (normalized, deduplicated and leak-free) and how to do the cost math for LoRA.
  11. Multimodal and voice in applications 65 min
    Images, scanned documents and voice open up new products and new problems. You learn how much an image costs in tokens and how to fit it to a budget, when to use OCR or a vision model, and how to design a voice assistant that responds quickly by streaming sentences to speech synthesis as soon as they're ready.
  12. Capstone project — an end-to-end support assistant 120 min
    You bring the whole level together in a real system. Input guardrails, pseudonymization, an order tool, RAG with verification, handoff to a human, a budget and traces in one pipeline, plus a release gate that decides with data whether the new version goes to production.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.