AI Knowledge Hub

Can AI process unstructured bordereaux?

Quick answer

Sometimes. AI can help process irregular workbooks, PDFs, scans and narrative elements, but these are not equally unstructured. The workflow may need OCR, layout reconstruction, table detection and semantic mapping before validation. Illegible, missing or genuinely ambiguous information cannot be made reliable merely by applying AI.

What to remember

Key takeaways

  • Define what unstructured means for the actual source population.
  • OCR extracts text; it does not establish insurance meaning.
  • Retain a trace from every output value to its source.
  • Refer or reject content that cannot be supported reliably.

AI can help process some bordereaux described as unstructured, but the label covers very different problems.

A spreadsheet with moved columns is still tabular. A PDF with a visible table is semi-structured. A scanned image first needs character and layout extraction. A narrative description contains language rather than a fixed set of columns. Each source needs different methods and has different reliability limits.

The practical question is therefore which information is present, how clearly it can be recovered and what evidence is required before the result can be used.

Unstructured can mean several different things

Many supposedly unstructured bordereaux are irregular rather than truly unstructured. An Excel workbook may contain merged headings, notes above the table, several header rows, embedded subtotals or inconsistent sheet names. The underlying records still have rows and fields once the useful region is found.

PDF bordereaux are often semi-structured. They may show columns visually, but the stored text does not necessarily preserve table relationships. Tables can continue across pages, headers can repeat and values can be split by page boundaries.

Scans and images contain pixels rather than machine-readable characters. Narrative loss descriptions or free-text notes are closer to genuinely unstructured content because their meaning is expressed in language.

Before selecting a method, teams should profile the actual source population: file types, languages, image quality, table patterns, handwriting, embedded objects and required target fields. A claim to handle unstructured data is meaningless without that boundary.

Processing may require several interpretation stages

A scanned document may first pass through optical character recognition. OCR proposes characters and words; layout analysis then identifies their positions. Table reconstruction tries to recover rows, columns and relationships. These are separate quality stages.

AI can help recognise headers, distinguish transaction rows from totals and associate unfamiliar labels with target concepts. Language models may extract supported facts from narrative fields, provided the target definition and evidence requirements are clear.

Each stage can introduce error. A poorly recognised digit can change an amount. A page-break problem can attach a value to the wrong row. A plausible mapping can still confuse gross and net premium.

For that reason, the service should retain intermediate evidence: the original page or cell, extracted text, reconstructed table, proposed mapping and transformation. Confidence should be tied to the relevant stage rather than presented as one unexplained score for the whole document.

Validation must reconnect output to source

The transformed output should be validated against both the target requirement and the source. Mandatory-field and code checks show whether the target is populated correctly, while counts and totals help reveal omitted or duplicated rows.

Provenance lets a reviewer move from a target value back to the page region or workbook cell that supports it. Review can then focus on low-quality extraction, uncertain mappings and material values rather than rereading every page.

Validation should also test the boundaries between stages. A date may be valid in form but attached to the wrong claim. A set of numbers may add up while representing a repeated subtotal rather than transaction records.

Representative testing needs examples of poor scans, multi-page tables, changed layouts and unsupported cases—not only clean documents. Extraction accuracy, mapping accuracy and accepted-output quality should be measured separately so that the source of failure is visible.

Some inputs should be referred or rejected

AI cannot recover information that is absent. It should not guess an unreadable policy number, invent a missing currency or select an insurance meaning when the source provides insufficient context.

The operating design should define minimum quality and supported-source conditions. A case may be referred for targeted correction, returned to the sender for resubmission or rejected as unusable. Material uncertainty should not be concealed by a complete-looking output.

Security and privacy also limit processing choices. Bordereaux may contain personal or confidential data, so approved environments, access controls, retention and supplier arrangements remain relevant regardless of format.

AI expands the range of content that can be interpreted efficiently. It does not make every document processable, and a reliable service is distinguished as much by how it fails safely as by what it extracts.

Example

A scanned claims bordereau contains a table across several pages, repeated headers and a narrative loss description. OCR extracts the characters, and layout processing reconstructs the table while excluding repeated page headers.

AI maps supported columns and proposes structured content from the narrative. Each result retains a link to its page region. Control totals expose one missing row at a page boundary, and two low-quality amount cells are referred to a claims bordereaux reviewer.

The reviewer corrects the cells from the source image. Nothing is inferred for an unreadable policy reference; the sender is asked to confirm it before the record is accepted.

FAQs

What's next?

Bordereaux Myth Buster

Bordereaux Myth Buster

Many organisations delay AI because of misconceptions about risk. Test your thinking with five quick Myth Buster questions.

Our latest insurance insights