A few things you can build:
1️⃣ Annotation queues: Review and label one (or many) runs at a time
2️⃣ Experiment comparisons: Compare outputs and metrics across experiments side by side
3️⃣ Trace reviews: Build trace or thread history views to inspect app behavior