Has anyone deployed Docling reliably on Databricks?

0
0
Asked By MapleQuill47 On

My team has started using Docling for document parsing in Databricks. The parsing quality is impressive, but we occasionally run into GIL-related runtime failures and instability, especially when processing larger workloads. Has anyone found a reliable way to run Docling with Databricks agent or AI workloads? We'd appreciate advice on batching, scaling, deployment patterns, or alternative services that work better in this environment.

3 Answers

Answered By CopperMango6 On

For a more production-friendly setup, consider running Docling as a separate serving process instead of embedding it directly in every Databricks task. That gives you a scalable endpoint, isolates parser failures, and can be more economical than repeatedly starting the full runtime.

Answered By SunnyKite82 On

We ran into similar problems when processing too much in a single job. Splitting the workload into smaller batches and spreading them across separate workers made the GIL failures less frequent. It still needs some monitoring, but the setup became usable.

Answered By VelvetOrbit31 On

We compared Docling with Kreuzberg, also known as xberg, on a cloud GPU platform. Kreuzberg was noticeably faster and more reliable in our tests, with better task performance. Docling is capable, but it may be worth benchmarking alternatives against your actual document mix before committing to it.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.