> They had to use AI to analyze the transcripts--specifically, the same model responsible for some of the bad behavior!
Yeah this is pretty concerning; anecdotally and IME, using a different model to review materials/give feedback is often more fruitful, too (though ofc I understand it's not possible in this specific use case)
(Claudes will gloss over Claude problems; OpenAI models will gloss over OpenAI model problems; OpenAI models giving feedback on Claude material or vice versa, or some other permutation, often results in better feedback)
顯示更多