tl;dr -- jev is a very useful paradigm which has been hiding in plain sight.
also, pour one out for log probs disappearing from so many new closed inference apis... 😢
prompting llms with max_tokens: 1 w/ log probs for fast, parallelized judgments may be an old pre-structured outputs trick gaining new life, but i'll admit i'm rather surprised at how flexible and useful this pattern appears. there's a lot more depth here than i expected -- the demos are quite impressive!
back in the golden days of early llms circa 2023, i recall experimenting with fine tuning oai's ada to output either a 1 or a 0 token, and inspecting the log probs to roughly gauge confidence. at zapier we were exploring this around field mapping -- matching output fields from a prior step to input fields in another.
seems like training a base model specifically to return calibrated probabilities takes that quite a bit further. it's so cheap that you can ask tons of speculative questions and just throw away the answers. add a hierarchy of follow-up questions and you can really tune the plinko machine. quantity truly has a quality of its own.
with short questions and shared context, the rough math works out to ~60 field judgments per penny with fine-tuned ada versus potentially ~11k with jev batching -- amazing!
what other interesting paradigms are lurking in plain sight?