A guided video lesson
Agent Harnesses: From General Loops to Evidence-Driven Custom Systems
A source-grounded lesson on how agent harnesses assemble context, run tools, specialize workflows, and use private evals plus observability to improve the model, context, or harness.
Chapter pages
Choose where to begin
- 01Why harnesses and evals are the decision to examine0:00 - 1:00
- 02Owning intelligence: model, context, and harness1:00 - 2:46
- 03The general agent loop and its middleware levers2:46 - 5:22
- 04From general harness to domain-specific cognitive flow5:22 - 7:36
- 05When distribution shift justifies a custom harness7:36 - 9:08
- 06Private evals, observability, and the learning loop9:08 - 10:32
- 07Harbor turns agent tasks into comparable benchmarks10:32 - 13:06
- 08Observability exposes the path to an agent failure13:06 - 14:41
- 09The data flywheel for continuously improving agents14:41 - 17:31
- 10LangSmith Engine automates trace curation and fixes17:31 - 19:20
- 11Dogfooding Engine and benchmarking harness behavior19:20 - 20:48
- 12The unresolved spectrum from off-the-shelf to fully custom20:48 - 23:57