Tried to build this a while ago and faced a few problems:
- video is incredibly dense and searching through it / deriving secondary signals from it is incredibly hard + expensive
- the consumer won’t want to pay that price when they can just type the answers themselves (free!)
- integrating applications with this seems obvious until they realize they’re giving away all the data to some chum start up when they could get it themselves
- i do believe video -> user representation is the solution to personalization, it is seamless and obvious. Technically though it is incredibly challenging