I’ve been experimenting with Jev and WebMCP 😁 - Jev selects the next tool and fills enumerable inputs. If the same call needs free text or a number, it falls back to LLM.
If that same call also needs free text, a number, or another non-enumerable value, an LLM fills only those remaining fields.
What's also interesting is that based on the confidence score given by Jev we can make decisions in harness - if Jev is 97% sure we might trigger the tool - but if the confidence is, let's say, 60% we might want to do additional checks!
Evals next? 🤔
1/2 🧵