i forgot that logprobs used to be universally exposed from LLM inference APIs! we're entering some market maturity phase of LLMs where it's viable to slice off interesting parts of the LLM stack and package them in new ways that are useful
tl;dr -- jev is a very useful paradigm which has been hiding in plain sight.
also, pour one out for log probs disappearing from so many new closed inference apis... 😢
prompting llms with max_tokens: 1 w/ log probs for fast, parallelized judgments may be an old pre-structured outputs trick gaining new life, but i'll admit i'm rather surprised at how flexible and useful this pattern appears. there's a lot more depth here than i expected -- the demos are quite impressive!
back in the golden days of early llms circa 2023, i recall experimenting with fine tuning oai's ada to output either a 1 or a 0 token, and inspecting the log probs to roughly gauge confidence. at zapier we were exploring this around field mapping -- matching output fields from a prior step to input fields in another.
seems like training a base model specifically to return calibrated probabilities takes that quite a bit further. it's so cheap that you can ask tons of speculative questions and just throw away the answers. add a hierarchy of follow-up questions and you can really tune the plinko machine. quantity truly has a quality of its own.
with short questions and shared context, the rough math works out to ~60 field judgments per penny with fine-tuned ada versus potentially ~11k with jev batching -- amazing!
what other interesting paradigms are lurking in plain sight?