DocbyteFacebookPixel
Join the Docbyte Vault 26.3 Executive Preview — Thu 29 October, 12:00–18:00, Ghent — Request an invitation

Adding a LLM to your mailroom solution: Should the IDP platform move over?

[tta_listen_btn]

docbyte-idp-llm-mailroom-no-panels

Table of Content

Docbyte worked with a large insurer to add a large language model (LLM) as a contextual reasoning layer on top of an existing document workflow. The workflow itself stayed in place.

The insurer’s intelligent document processing (IDP) workflow already handled incoming customer and broker correspondence, but using more traditional technologies: Cloud-based machine-learning models combined with keyword rules. Nothing fancy, but it worked. It also needed constant adjustments, because document formats, products and terminology keep changing and document types with high similarity were hard to tell apart reliably. The machine- learning models lacked the required ability to reason through nuance and context that rigid rules simply couldn’t capture.

 

The problem: a workflow that kept needing attention

The workflow processes roughly 50.000 documents a working day. Around 97% of the incoming material is email from customers, insurers and brokers, often with several attachments, mostly digitally generated, with some scans and handwritten content mixed in. For every one of those documents, the process needs two answers: which existing claim it belongs to, and what type of request it is.

So, this is clearly an information classification job: Identifying the right document type, case type, and claim. Classification is harder than data extraction. Plenty of IDP platforms only provide extraction, because they think: “Hey, we just need to extract the claim number, and some keywords!”

The wrong type but the correct case is a quick reassignment. No case match at all sends the document to human review, because there is no
safe automated fallback for it. Type classification is important, but it is the secondary metric.

The insurer piloted the approach in its car insurance business unit: the highest-volume unit it has, and the one with the most complex, hardest-to-distinguish document set. This was deliberately chosen as a demanding test case. If the reasoning layer couldn’t clear the bar on the hardest problem the business had, it wasn’t worth building out further.

The pilot covered 16 business-defined document classes: quotation requests, proposals and contracts, policy modifications, registration
and cancellation requests, deregistration notices, information requests, death notifications, bankruptcy notices. These are not always visually
distinct templates. A notice of deregistration, where a licence plate is revoked, may also imply cancellation of the related insurance. It represents a specific business event rather than standard cancellation language. That is exactly the kind of ambiguity keyword and sample-trained machine learning models struggle to keep up with.

 

The baseline: a system that was slipping

Before testing anything new, the team measured where things actually stood. The existing model uses OCR as input, and had started around 60% accurate on this task and had degraded to roughly 40% as formats and terminology drifted. Retraining was not an option either. The document types had never been relabelled consistently, so there was no reliable source of truth to train against.

That decline mattered more than the raw number. A model that loses accuracy over time shifts more work onto manual review, and that shift
compounds the longer the model runs unchanged.

 

What was built: an LLM layered onto the existing platform

The project kept the controlled document workflow in place and added an LLM as a reasoning layer on top of it. The pipeline:

  1. Capture the email and its attachments, including scanned material.
  2. Convert the input into a PDF representation.
  3. Send it to the LLM with the allowed classifications, business definitions and decision rules.
  4. Return the case number, a structured classification and the model’s generated explanation.
  5. Validate the result and export it to the case-management environment, where the real downstream action happens.

 

The model is not given an open-ended task. To reduce the risk and operational impact of unsupported answers, every request starts a fresh, restricted session, and the classification must come from a closed set of permitted values. Deterministic services still control validation, routing, storage and workflow execution. That containment is what makes it possible to trust the model’s output without having to trust the
model itself.
Because claim assignment is the metric that matters, the pipeline treats it accordingly: the extracted case number goes through deterministic
validation and invalid results go to a person rather than being forced into a plausible-looking match. Type misclassification is handled more leniently, since it’s the cheaper error to fix. The model’s generated explanation is also shown in the interface. It is also not proof that the answer is correct but it enables the people handling the case to see why it landed there instead of ping-ponging the document around. This is more or less the same way as it should be done with human case workers, where one person would classify a document in a different way than his/her colleague. Additionally, this helps implementation teams spot systematic errors and refine document definitions.

 

The errors taught us something useful

The most visible confusion was between a notice of deregistration and a general cancellation. Rather than running a retraining cycle, this is fixed by sharpening the classification definition: deregistration as a specific subtype of cancellation, tied explicitly to licence-plate removal. Correcting a known error pattern requires editing a definition rather than training or retraining a model.

A separate test checked whether the LLM added value on the documents that mattered most: the ones the existing ML model had already got wrong. The LLM correctly classified 7% of the documents in this subset, selected precisely because the existing model had misclassified them. This isn’t a like-for-like overall benchmark, since the existing model’s accuracy on that same subset is 0% by construction, but it shows the reasoning layer recovered a meaningful share of the existing workflow’s failures. This means that the success rate changed from 0% to 73% for these exceptions.

 

The results so far

The numbers below measure document-type classification, the secondary metric. Case assignment itself isn’t scored the same way: the extracted case number either passes deterministic validation or it doesn’t, and anything uncertain goes to human review rather than showing up as a percentage.

Approach Accuracy
Existing ML model (degraded from an initial ~60%) 40%
LLM (Gemini 2.5 Flash) – full test sample 80%

The 80% figure was reached without additional model training or fine-tuning, on an initial test set checked manually and reviewed with the insurer’s business experts. The advantage came from the model’s ability to weigh the relationship between phrases across an email and its attachments, and to apply a business definition written in natural language rather than one baked into a proprietary, retrainable classifier.

 

What this means for the business

Day to day, the workflow now behaves differently in three concrete ways.

  • Multilingual processing: the solution reaches the same results in every language, without needing a document set for every language to train the model.
  • Rule refinement gets easier: a known confusion, like the deregistration/cancellation overlap, gets fixed by sharpening a definition, not by retraining a model or waiting for a release cycle.
  • Review gets more transparent: the model’s generated explanation shows up in the interface, so the people handling a case can see why it landed there instead of passing it around to find out.

 

A tip: get security and the DPO involved before the business case is finished

A strong business case is not enough on its own. If the business wants an LLM in the workflow but security or the data protection officer
hasn’t signed off on where the data goes, how it’s retained and who can access it, the project doesn’t roll out, no matter how good the accuracy numbers look. Bring those stakeholders in while the architecture is still being decided, not after a pilot has already proven the concept. That order is what most often stalls a working pilot before it reaches production: data flow, retention and processor terms get reviewed only once the concept already works, instead of while the architecture is still open.

Design and initial development took roughly three to four months. Most of that time went into redeveloping the pipeline itself, along with
discussions with the business about document types and what distinguishes one from another. The LLM does not remove the need for that domain knowledge. It makes it easier to express, test and refine, which is exactly the work now underway to move from 80% toward 90%.

IDP still provides capture, structure and control, and the LLM adds contextual interpretation on top of it. Business rules, validation and
people are what turn that combination into a process that can be trusted. Instead of focusing on classification accuracy in isolation, the key metric is how often the document reaches the right case on the first try.

 

Review how your document workflow handles the cases that don’t fit

We can help map where an LLM reasoning layer would help most in your document workflow: which classifications are ambiguous, what a safe fallback to human review looks like, and how case validation stays deterministic while the model handles interpretation.

 

FAQs

Does adding an LLM mean removing the existing IDP system?

No. The LLM sits on top of the existing workflow as a reasoning layer. Deterministic validation, routing and case-management still control what happens downstream.

Why measure case assignment instead of overall classification accuracy?

Because the two errors have different costs. A document in the right case but wrong type is a quick reassignment. A document that matches no case goes to manual review with no automated fallback.

Why start with the hardest document set instead of the easiest?

To find out whether the approach holds up under real conditions early, rather than proving a smaller point on an easier business unit and leaving the harder question open.

Picture of Frederik Rosseel
Frederik Rosseel

Hi, I’m Frederik, CEO of Docbyte. Having pioneered solutions in digital archiving and qualified trust services for years, I distill that invaluable experience into writing. My goal is to help businesses achieve robust data security and seamless regulatory compliance through crystal-clear insights

Contact Us


At Docbyte, we take your privacy seriously. We’ll only use your personal information to manage your account and provide the products and services you’ve requested from us.

Are you interested in contributing to our blog?
Recent Blogs