💡This new paper introduces SFGA, a statistics-first gating architecture for cost-aware SFT data procurement.
Instead of sending every case to an LLM judge, SFGA starts with low-cost blind measurements across three intrinsic quality axes:
🔹 Diversity
🔹 Utility
🔹 Redundancy
Each is evaluated together with confidence intervals.
Only when evidence is weak, borderline, or conflicting does the system escalate to an adjudicative debate between a buy-advocate and a reject-advocate, resolved by a presiding verdict.
And even after escalation, SFGA does not blindly trust the LLM judge. It audits the process through advocate swapping, explicitly measuring negativity skew and positional bias.
The contribution is not a new estimator or debate algorithm, but the architecture itself:
✅ Cheap statistics first
✅ Confidence gating
✅ Selective escalation
✅ Bias auditing
For AI data markets, this direction matters: LLM judges should be used where they add the most value — on cases that are genuinely ambiguous, costly, and worth debating.
Paper: