Process Automation · Email processing

Email inbox automation: orders, invoices and delivery notes captured by AI

Orders arrive as plain text in an email, as a PDF attachment, sometimes as a photo of a filled-in form. In the end, someone types them into the ERP. An AI pipeline reads these documents, checks the values against your master data and hands them to the system that needs them. Anything it cannot read with confidence goes to a person.

Julia Rose 10 min read
Office desk with an email inbox open on screen showing PDF attachments, next to paper invoices and delivery notes
Image: AI-generated

Key takeaways

  • Automation pays off most for document types with high volume and varying formats: orders, invoices, delivery notes, shipment and scheduling data.
  • Templates and rules work reliably as long as the format stays the same. Once customers write in their own words, every variant needs its own rule.
  • A solution that holds up reports a confidence level for every single value and hands uncertain cases to a person instead of guessing.
  • In our projects manual effort fell by 80 to 90%: interventions at Radaway by 90%, extraction effort at Fr. Meyer’s Sohn by 80%.

In almost every company that receives orders or invoices by email, the same routine runs: someone opens the message, reads the order, looks up the right product numbers and types everything into the ERP. At ten emails a day nobody notices, at two hundred someone spends half the morning on it, and as soon as that person is on holiday the inbox starts piling up.

Language models now read documents reliably, so large parts of this step can be automated without asking customers to type their orders into a portal first. Which documents make it worthwhile, how the capture works, what a solution has to handle and what a first project looks like is covered below.

Which documents are worth automating

Not every document is a case for automation. It gets interesting where many documents of the same kind arrive every day and each still looks a little different: customer orders, supplier invoices, delivery notes and order confirmations, and in logistics shipment, routing and scheduling data.

Two things decide it: how many documents arrive, and how different they look. Twenty identical orders a week are typed faster than a project takes. Two hundred orders a day in forty layouts are a different matter entirely.

The third point usually gets overlooked: what an error costs. A wrong quantity tends to surface only once the goods are at the customer. When a typo triggers a return shipment or throws production planning off course, automation often pays off far earlier than the pure time saving would suggest.

Why templates eventually cost more than they save

Classic text recognition works with templates: the order number sits top right, the quantity in the third column. As long as every document looks the same, that is fast, cheap and easy to trace.

The trouble starts with the exceptions. One customer switches layout, another writes the order in two sentences in the middle of a long email thread, a third sends a photo, and suddenly every variant needs its own rule, the maintenance grows with every customer. At the logistics company Fr. Meyer’s Sohn, which receives thousands of emails a day in German and English, rule-based extraction stopped keeping up at exactly this point.

Language models, by contrast, capture the content regardless of where a value sits and how it is phrased. That does not put rules out of work: in our projects they usually take over the checking, meaning format validation, recalculating totals and controlling mandatory fields, because a rule solves that more cheaply and more precisely than a model. The comparison on the AI agent development page shows when each approach fits.

How does automated capture from the inbox work?

Between the incoming email and the order in the ERP there are five steps, each of which can be measured and improved on its own. They are what shows later whether the system can really run without supervision.

Recognise the document type

First the system classifies every message: new order, change, cancellation, invoice, question. As trivial as the step looks, it prevents the most expensive mistake, because if a cancellation goes through as a new order, goods leave the warehouse that nobody ordered.

Read text and attachments together

Many orders sit in the attachment, as a PDF, Excel file or scan, and then the context matters: the email often carries the change to the attachment, such as a different delivery date. At Radaway, orders sent as attachments used to fall outside automated processing entirely; today they are covered.

Convert to a fixed format

The language model returns the data in a strictly defined schema, for example customer number, product, quantity and delivery date, because only then can the result be checked by machine. With free text, all you can do is read and hope.

Check against master data

Since customers rarely use the name from your catalogue, the system looks up the product that matches their wording, even when the words differ. A second step checks whether the product it found fits the rest of the order. Plausibility checks come on top: does the customer exist, is the price right, is the quantity realistic?

Hand over or hand back

Confident results go straight into the ERP. When the system is unsure about a value, it puts the case in front of a colleague for review. The whole benefit hangs on that distinction: a system that passes every result on unchecked saves time and builds in errors that only surface once the goods are on their way. One that queries every case saves nothing at all.

What it delivers in practice

In practice: Radaway, manufacturing

The bathroom equipment maker Radaway receives orders by email, in whatever wording and format the customer chooses. A first LLM system was already in use, yet many cases still had to be reworked by hand. After three weeks in which we improved prompts, product matching, attachments and message type detection, manual interventions fell by 90% and product matching accuracy sits above 95%.

Read the Radaway case study

At Fr. Meyer’s Sohn the subject was shipment, routing and planning data from thousands of emails a day. Because the data was not to leave the company, the pipeline runs on its own servers there and delivers the values in structured form in real time; manual extraction effort fell by 80%. The details are in the Fr. Meyer’s Sohn case study, with more applications on our pages AI for manufacturing and AI for logistics.

Neither figure is at a hundred percent, and that is deliberate. A remainder stays with people because it belongs there: exceptions, unreadable documents, orders with contradictory information.

Checklist: what should you look for when choosing a solution?

Whether you buy a solution or have one built, these points decide whether it works in daily use.

CriterionWhat matters
Attachments and scansDoes the solution read PDF, Excel and scanned documents in the same pass as the email text?
Multiple languagesAre documents captured in every language your customers and suppliers use?
Master data matchingAre customers, products and prices checked against your data, even when the names differ?
Confidence per fieldDoes the system know how certain it is about each individual value?
Review of unclear casesIs there a workflow where a person quickly checks and confirms uncertain cases?
Where processing runsDoes processing run on-premises, in the cloud or hybrid, in line with your security policies?
IntegrationDoes the solution write to your ERP through existing interfaces, whether SAP, another ERP or an in-house system?
Room to growCan further document types be added later without starting over?

Most offerings fail on the fourth row, because a system that only returns a result without saying how sure it is cannot be built into any approval process.

Getting the data into the ERP

The captured data goes through the interfaces your system already has, whether SAP, another ERP or an in-house system. Not everything has to run straight through at the start. Many teams begin with the system only making suggestions that an employee confirms with one click. That costs a few seconds per document and shows week by week where the capture is solid and where it is not, and as confidence grows, a larger share goes straight through. How the integration works technically is covered in our article Integrating AI with SAP, Microsoft 365 and Jira.

Where to start

With a single document type, and specifically the one with the highest volume or the most expensive errors, not the one with the most interesting use case. In a process analysis at a fixed price from €3k we map today’s workflow and propose an architecture with scope, timeline and cost; first working components arrive in six to eight weeks, built with your real documents.

If a system is already running but needs too much intervention in daily use, the first question is which of the five steps it stumbles on. At Radaway no rebuild was needed for that reason, three weeks of focused work were enough. Which language models suit tasks like these and where they can run is covered on the LLM development page.

Which documents does your team still type in by hand?

In the process analysis we map your workflows, rank the possible use cases by value and propose an architecture. Fixed price from €3k.

Request a process analysis

Frequently asked questions

Yes. A language model reads the email text and the attachments, picks out customer number, product, quantity and delivery date and hands them to the ERP in structured form. At the bathroom equipment maker Radaway this cut manual interventions by 90%, with product matching accuracy above 95%.

Yes. Scanned and photographed documents are read by text recognition or directly by a model that understands images. The quality of the original affects the result, and a good system flags illegible passages as uncertain and passes them to a person instead of guessing.

Nothing. Unlike templates and fixed rules, a language model captures the content regardless of layout, even when the order is written freely in the middle of an email thread. No new rule per customer is needed.

Through several checks: does the value fit the fixed data format, do the customer and product exist in the master data, are quantity and price plausible? Each field gets a confidence level from this. If a check fails, the case goes to a person for approval.

First working components arrive in six to eight weeks, built with real documents from your own inbox. If a system is already running it goes faster: at Radaway an existing LLM system was production-ready after three weeks of focused work.

That depends on the deployment. Processing can run on-premises on your own servers, as in the data extraction for Fr. Meyer’s Sohn, in the cloud or hybrid. With cloud providers we always make sure that the use of your data for training is switched off.

Julia Rose

About the author

Julia Rose

Marketing Lead, theBlue.ai

Julia has been part of theBlue.ai since 2019 and has accompanied the development of AI applications in the enterprise environment since the company’s early days. In her role as Marketing Lead, she works closely with the engineering and consulting teams and makes complex technical topics understandable and accessible for decision-makers.

In her articles, she writes about practical experience from enterprise AI projects, as well as the challenges and opportunities of using AI in companies.