RL hold-out-1 experiment im curious about to study how model specialization relates to model capacity and data allocation for how many domains to specialize for
a common way to get a model to specialize in N domains is Multi-Teacher-On-Policy Distillation
train N specialized teachers, do OPD with routing for a set of prompts to teach a model all those abilities
for a given domain like data viz -> if we remove X% of the teachers (ie. Don’t specialize on those skills), does our performance on data viz increase?
how is this affected by which domains get left out? does it suffer if similar domains are removed but benefit if very different domains are?
we still see vertical focused specialized models like GPT-Cyber
clearly looks like allocating a lot of data and parameters for a given vertical boosts perf in that vertical more than a general purpose mixture
points towards a future where the specialist models always win for high value domains