가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Shashwat Goel
@ShashwatGoel7
Training Decision Assistants Past work: Training AI Co-scientists, ΔBelief-RL, Measuring Long Horizon Execution
가입 June 2020
2.5K 팔로잉 중    4.2K 팬
I went digging into OpenAI consumer training opt-out policy wordings, and I wonder if there's a loophole that allows training on hidden CoT reasoning which makes the guarantee vacuous. From OpenAI's consumer terms, opt-out apply to Input and Output (together called Content), defined as: - "Your Content. You may provide input to the Services (“Input”), and receive output from the Services based on the Input (“Output”)". ... - "Opt out. If you do not want us to use your Content to train our models, you can opt out by..." Notice, we do not receive hidden CoT from OpenAI, so its unclear if it counts as "Output" in this definition. This is important, because hidden CoT mostly contains a lightly processed version of everything useful for training, including user input and the output returned to users. Some evidence that made me particularly suspicious: other parts of the terms imply Hidden CoT is not considered Content, which then means opt-out would not apply to it: - "You are responsible for Content, including ensuring that it does not violate any applicable law or these Terms..." - "Ownership of content... (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output." Notice, since hidden CoT is not shown to us, we cannot be responsible for it, nor own it, which tells us it is not part of Output or Content, and thus exempt from training opt-out? 2. If we contrast to other similar services, e.g. OpenAI's own enterprise terms definition of Output, or Anthropic, the user receiving it is not mentioned. It is not clear whether hidden CoT is included in the opt-out in these or not either, but at least the earlier reasoning chain breaks in that "user receiving" is not explicitly mentioned, though similar statements about users owning and being responsible for Output apply. - OpenAI Services Agreement Definitions Section: "“Output” means output from the Services based on the Input." - Anthropic Consumer Terms: "Our Services may generate responses (we call these “Outputs”)" I am not a legal expert, and did not consult one, so it would be great if someone who knows better (or OpenAI) could confirm. I contacted the OpenAI data policy offer email provided 2 days back, and did not receive a clarification. Note that I am only pointing out an ambiguity in the guarantee, and not saying OpenAI definitely trains on hidden CoT. A clarification would be super useful in any case!
더 보기