This chapter explains what to monitor after an agent configuration is deployed to production and how to use the results for the next adjustment. The evaluation target is not limited to routing or image-quality checks. It also includes the marketplace's overall quality and health, as well as user behavior.
The presenters explain that, after a configuration is deployed to production, the team tracks marketplace quality and health. Offline routing correctness and an enhancement result that passes QA do not by themselves show whether the system helped in production. The team also needs to see whether the system leads to desirable outcomes in an environment with real users and merchants.
One example of such an outcome is conversion. In this talk, it is described as a behavior path in which a user adds an item to the cart and then completes an order. Thus, the team tracks not only whether an image looks better, but also what happens along the path from adding to the cart to completing the order. However, the presentation gives no improvement percentage or experimental design for proving an effect. We should therefore understand conversion as an example of a production behavior metric, not as proof that image enhancement increased orders.
図:本番でコンバージョンを測定し、その結果を設定の調整と展開に戻すマーケットプレイスのループです。
Caption: A marketplace loop in which production conversion is measured and the result is returned to configuration tuning and deployment.
The slide shows production conversion measurement entering a marketplace feedback loop and leading to configuration tuning and deployment. It is not a separate funnel diagram showing the path from adding to the cart to completing an order. This part of the presentation appears around 20:42 in the video.
全体平均だけでなく、セグメントごとに切り分ける
Slice results by segment, not only by the overall average
マーケットプレイスの利用状況は、すべて同じではありません。そこで、指標を一つの全体平均にまとめるだけでなく、セグメント(segment)ごとに切り分けて調べます。この作業は、英語では slicing and dicing と呼ばれます。地域、端末の種類、料理の種類などで結果を分けると、全体の数字だけでは見えない違いを確認できます。
Marketplace usage is not uniform. Therefore, the team does not only combine metrics into one overall average; it also examines them by segment. In English, this work is called slicing and dicing. Separating results by geography, device type, or dish type can reveal differences that an overall number hides.
地域ごとに、画像改善後の品質や行動の結果を見る
端末の種類ごとに、結果の違いを見る
料理の種類ごとに、結果の違いを見る
Examine image-quality and behavioral outcomes by geography
Caption: A flow that measures production conversion, separates results by dimensions such as geography and dish type, tunes each segment, and returns the best configuration to production.
The slide explicitly names geography and dish type as dimensions for separating results. It also shows segment-specific tuning inside a feedback loop. It does not show a device-type breakdown or a detailed metric dashboard. This part of the presentation appears around 21:07 in the video.
全体の改善と、一部の悪化を分けて考える
Separate overall improvement from deterioration in one segment
Even when the overall average improves, results may get worse for a particular geography, device, or dish type. Conversely, a small overall change may hide a large improvement in one segment. This is why the team examines segment-level results, finds the parts that are improving or lagging, and makes adjustments suited to each one.
This connects to the earlier goal of optimizing the marketplace globally. If one store or user group improves while the change harms other stores, the system has not succeeded overall. Segment-level checking helps find this kind of imbalance behind the average. The point here is not a report that regression definitely occurred in a particular segment; it is a way to guide tuning decisions.
The presentation gives no conversion values, experimental design, segment thresholds, or details about the other quality and health metrics. Conversion is presented as one example of a metric to track. The final applause and thanks add no technical claim.
The key point of this chapter is that deployment is not the end. Measure behavior and marketplace health in production, then examine the results by segment. Use that information to tune the configuration while balancing overall improvement with safe outcomes for each segment.