The latest models stay faithful to source information, but their findings and responses miss key details experts include. MLCR-AA scores each answer on completeness, accuracy, and concision. Accuracy is the easier of the two deciding dimensions, and completeness is where models separate: Claude Fable 5 reaches 73.8% completeness at 90.1% accuracy, while GPT-5.6 Terra (max) records 93.7% accuracy and 33.9% completeness. For medical record review, an answer that is accurate but incomplete can still be unsuitable.