Delivering Document Conversion as a Cloud Service with High Throughput and Responsiveness

06/01/2022
by   Christoph Auer, et al.
17

Document understanding is a key business process in the data-driven economy since documents are central to knowledge discovery and business insights. Converting documents into a machine-processable format is a particular challenge here due to their huge variability in formats and complex structure. Accordingly, many algorithms and machine-learning methods emerged to solve particular tasks such as Optical Character Recognition (OCR), layout analysis, table-structure recovery, figure understanding, etc. We observe the adoption of such methods in document understanding solutions offered by all major cloud providers. Yet, publications outlining how such services are designed and optimized to scale in the cloud are scarce. In this paper, we focus on the case of document conversion to illustrate the particular challenges of scaling a complex data processing pipeline with a strong reliance on machine-learning methods on cloud infrastructure. Our key objective is to achieve high scalability and responsiveness for different workload profiles in a well-defined resource budget. We outline the requirements, design, and implementation choices of our document conversion service and reflect on the challenges we faced. Evidence for the scaling behavior and resource efficiency is provided for two alternative workload distribution strategies and deployment configurations. Our best-performing method achieves sustained throughput of over one million PDF pages per hour on 3072 CPU cores across 192 nodes.

READ FULL TEXT

page 1

page 8

research
05/15/2018

Corpus Conversion Service: A machine learning platform to ingest documents at scale [Poster abstract]

Over the past few decades, the amount of scientific articles and technic...
research
05/24/2018

Corpus Conversion Service: A Machine Learning Platform to Ingest Documents at Scale

Over the past few decades, the amount of scientific articles and technic...
research
07/04/2022

BusiNet – a Light and Fast Text Detection Network for Business Documents

For digitizing or indexing physical documents, Optical Character Recogni...
research
11/05/2020

Infer XPath

We propose reformulation of discovery of data structure within a web pag...
research
05/24/2023

ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents

Transforming documents into machine-processable representations is a cha...
research
03/25/2022

Whole Slide Image to DICOM Conversion as Event-Driven Cloud Infrastructure

The Digital Imaging and Communication in Medicine (DICOM) specification ...
research
11/19/2020

WAE: Workload Automation Engine for CDN-specialized Container Orchestration

Content Delivery Network (CDN) has been emerged as a compelling technolo...

Please sign up or login with your details

Forgot password? Click here to reset