تخطي إلى المحتوى الرئيسي
Microsoft
separator
https://catalogartifact.azureedge.net/publicartifacts/upstage-marketplace.document-parse-c556a484-de0f-4acb-b2f1-105f257431b9/image3_symbolgradient.png

Document Parse

بواسطة Upstage

Parse PDFs, images, and scans into LLM-ready HTML & Markdown

Upstage Document Parse is an enterprise-ready REST API that transforms complex, multi-layout documents into structured markup so LLMs, search indexes, and downstream apps can "see" every element on the page.

  • Lightning-fast throughput. Internal tests show 0.6 s per page and 100-page PDFs processed in < 60 s - 10x faster than AWS Textract and Unstructured, and 4x faster than LlamaParse.
  • Best-in-class fidelity. On the DP-Bench benchmark it scores TEDS 93.48 / TEDS-S 94.16, beating Google & Microsoft by 5 pts+ in table and layout reconstruction.
  • Advanced element support. Beyond text and tables it now converts charts and other embedded images to structured HTML, expanding the data LLMs can reason over.
  • Scales to books. The synchronous endpoint accepts documents up to 100 pages, while the asynchronous job API parses up to 1000 pages in one call.

Typical use cases include RAG pipelines, contract ingestion, research-paper summarization, and any workflow that needs pixel-perfect markup for retrieval, QA, or summarization.

العربية (ليبيا)
أيقونة إلغاء الاشتراك في اختيارات خصوصيتك خيارات خصوصيتك
خصوصية صحة المستهلك خريطة الموقع اتصل بنا الخصوصية وملفات تعريف الارتباط شروط الاستخدام حول إعلاناتنا إدارة ملفات تعريف الارتباط