가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

ahhhhfs | AI & Open Source
@ahhhhfs
가입 October 2021
10 팔로잉 중    40.8K 팬
A lot of PDFs already have a text layer — running OCR again is just wasted work. Firecrawl’s open-source pdf-inspector first checks the PDF type locally (in tens of milliseconds), then decides which files can be extracted directly and which pages actually need OCR. Readable ones become Markdown; only the missing-text parts get OCR. Their data: about 54% of PDFs don’t need OCR at all. Perfect as a pre-check step for AI knowledge bases and document Q&A. 👉
더 보기