A guided video lesson

Closed-Loop Multimodal Evals: From Routing to Marketplace Feedback

A source-grounded lesson on building a production multimodal image-enhancement system with selective routing, bounded editing, layered QA, offline human alignment, online drift correction, and marketplace feedback.

Open the complete source

Chapter pages

Choose where to begin

  1. 01Scope: a production multimodal eval problem0:19 - 0:56
  2. 02Why food imagery matters at marketplace scale0:56 - 2:01
  3. 03The merchant-side quality problem2:01 - 2:34
  4. 04Threading the needle: authenticity, faithfulness, and diversity2:34 - 3:18
  5. 05Long-tail input quality and concrete defects3:18 - 3:43
  6. 06Design goals for the agents3:43 - 4:23
  7. 07Why agentic flexibility needs guardrails4:23 - 5:06
  8. 08The representative end-to-end example5:06 - 5:23
  9. 09Core orchestration: understand, route, edit, and gate5:23 - 6:19
  10. 10Logging as the foundation for diagnosis6:19 - 7:13
  11. 11What the router consumes and decides7:13 - 7:45
  12. 12Evaluating routing with binary and multi-branch metrics7:45 - 8:35
  13. 13Human-aligned offline tuning and the recall guardrail8:35 - 9:45
  14. 14Failure mode: sending a good image for enhancement9:45 - 10:14
  15. 15Failure mode: a recall miss and cross-modal hallucination risk10:14 - 10:45
  16. 16From static offline models to online drift correction10:45 - 12:10
  17. 17Prompt optimization, versioning, and safe closed-loop deployment12:10 - 13:21
  18. 18The bounded enhancement-and-QA loop13:21 - 14:13
  19. 19Worked example: sweet potato fries and pass at K14:13 - 14:49
  20. 20Pairwise comparison: deciding what counts as better14:49 - 15:47
  21. 21Failure modes: faithfulness, completeness, and reward hacking15:47 - 17:20
  22. 22Multimodal uncertainty should block publication17:20 - 17:42
  23. 23Publish-ready QA as a redundant final defense17:42 - 18:35
  24. 24Several feedback loops, not just the model loop18:35 - 19:19
  25. 25The diagnoser: a higher-level feedback abstraction19:19 - 19:47
  26. 26Dogfooding, human feedback, and regression replay19:47 - 20:29
  27. 27Production outcomes and segment-specific tuning20:29 - 21:36