the.ai

Multimodal / Applications

verified

Document Understanding

An invoice, a form, a scientific paper: the words matter and so does where they sit. A total in the bottom right of a table means something different from the same number in a paragraph, and a model reading the text in reading order has thrown away the layout that made it interpretable.

Viz primitive · budget-splittabular-content = 20

tabular-content holds 33% of the budget; rest holds the remaining 67%.

Document content whose meaning depends on two-dimensional position, against content that survives reading order, in equal units. Drag the tabular share up to watch linearisation lose the document — a paragraph loses nothing, a grid loses its columns.

20

Reviewed by opendroid · 2026-08-18

  • arXiv:1912.13318 — LayoutLM: Pre-training of Text and Layout for Document Image Understanding
  • arXiv:2110.08518 — MarkupLM: Pre-training of Text and Markup Language for Visually-rich Document Understanding