Teaching AI to read logistics paperwork

Date

October 6, 2026

October 6, 2026

Author

Amari

Type

Engineering


Every shipment moves on paperwork: commercial invoices, packing lists, bills of lading, certificates of origin. It usually arrives as an email with a few PDFs, a product spreadsheet and maybe a phone photo of a stamped certificate. At each step, from booking to customs to delivery and billing, someone reads that pile and keys the same values into another system. At Amari, we help clear about $75 billion of goods into the US each year, roughly 2% of US imports. For every one of those shipments, an AI model reads the pile and drafts the customs entry, and a licensed broker corrects it and submits it.

That workflow gives us our evaluations: thousands of real shipments, each with the documents the model saw and the values a licensed broker filed. Customs entries make a demanding test of logistics paperwork in general: they pull values from every kind of document in the pile, and each one ends with values a professional checked and filed. Here is what a month of experiments on them showed, and what we expect to carry over to other logistics documents:

  • Presenting the documents. Text that keeps each table row together is enough. Fancier layout and page images added little or nothing.

  • Choosing a model. Of the dozen we tested, the newest weren't the best.

  • Context size. Test your longest requests: a window that fits isn't one that reads well.

  • Evaluating. Rerun your baseline whenever instructions change, and use enough shipments to tell models apart.

1. You don't need fancy document layout

VAREX (arXiv:2603.15118) tests extraction on about 1,770 synthetic single-page US government forms. It found that presentation matters: for models scoring above 90%, text that keeps the page layout beat plain text by 3 to 8 points (up to 18 for small models), and page images did about as well as good text.

The paper's plain text is scrambled, though. It comes from a PDF-reading tool that walks tables column by column, pulling each row apart. The tool we use already keeps each printed row on one line, which puts our default text close to the paper's layout-preserving text. So the question for us was whether going further pays: full layout, the PDFs themselves, or page images.

We ran those comparisons on 300 shipments, with ten models for the text versions, and scored two things:

  • Field accuracy: the share of the values the broker filed that a model reproduced exactly. It is steady enough to show effects under a point: the same model, run twice, moved by only 0.3 points (section 4).

  • Pass rate: the share of shipments a broker could have filed without changing anything. It is closer to what a broker experiences, but noisier.

Change in points against our default text, range across models:

What we changed

Field accuracy

Pass rate

Extra tokens

Layout-preserving text instead of our default text

−0.3 to +0.3

−1.7 to +1.7

+2–3%

Adding the PDFs themselves on top of the text

0 to +0.5

−0.7 to +2.7

+25–71%

Page images instead of text (DeepSeek V4.1 Flash, GLM 5.3 Flash)

both −1.7

−0.7 and −3.0

+12–58%

No field-accuracy effect exceeded 2 points, and every pass-rate change is within noise: each model's confidence interval includes zero.

This is consistent with the paper. The layout gain it measured is the gain from keeping rows together, and our text already does that. We didn't test the paper's scrambled text, so we can't put our own number on that gain. Past that point, the rest of the layout adds nothing, and the image results sit roughly within the paper's range.

Fancier presentation also buys little for a reason common to logistics paperwork: the documents repeat each other. The invoice, packing list and bill of lading restate the same goods, quantities and parties. We traced each of the 14,687 values filed on these shipments back to the documents:

  • 76% also appear in the email or a spreadsheet;

  • 7.5% appear only in a PDF;

  • 17% appear nowhere word for word, because the broker derived or reformatted them.

A quantity that's hard to read on a scanned packing list is usually sitting in a spreadsheet cell too. Presentation can only matter for the 7.5%, and those are mostly short shipment-level fields (vessel name, voyage number, bill of lading) that read the same in any layout. The image result fits: images cost the most on the 21 shipments with no spreadsheet, where the PDFs carried everything.

The image test is tilted against images. We render pages at about 90 DPI where VAREX used 200, and only two open-weight models, DeepSeek V4.1 Flash and GLM 5.3 Flash, ran it. Our live system renders pages the same way, so the result holds for us, not for images in general.

So where do the errors come from? Only about a third of the mistakes made by the Amari model, the model we run in production, trace back to how the text is presented. Most come from shipment-level fields: they make up a third of the filed values but contain over half the errors. The worst of these (export date, voyage, arrival date, consignee, port) aren't misreadings so much as conventions the model hasn't learned, like which of several dates, ports or parties to report — and conventions can be taught.

Keep each printed row intact, skip the fancy layout, and spend the effort on conventions. Images are a fallback for scans, and an expensive one.

2. You don't need the latest or biggest model

Next we compared the Amari model with cheaper and newer alternatives on the same shipments and instructions. Here we scored pass rate, because that is what a broker experiences. The baseline is the Amari model, run again with today's instructions (section 4 explains why that matters).

  • DeepSeek V4.1 Flash, Gemini 3.5 Flash and Kimi K3 showed no measurable difference from the Amari model. The best, the open-weight DeepSeek V4.1 Flash, costs about a twentieth as much per shipment.

  • Newer models trailed. Claude Opus 5.5 was about 4 points behind. That gap is small, but the Amari model won 74% of the shipments where the two disagreed, which is significant. The newer Gemini models trailed by more.

Most of the newer models' gap sat in one field: country of origin. Our instructions ask for it but say nothing about what to do when the documents don't state it, and often they don't. On the shipments where Gemini 3.6 Flash left it blank, only a quarter had an origin label anywhere. China still showed up as the port of loading, the export country or the shipper's address. The older Amari model infers China from that, and the broker filed China on 143 of those 150 shipments. Newer models leave the field blank.

A blank is a defensible answer, but the broker still has to fill it in. We applied one rule to the answers the models had already given: if origin is blank, use the export country and flag it for the broker. Opus 5.5 and Gemini 3.8 Flash both came to within a point of the Amari model. The rule fails on transshipment: goods made in one country and shipped from another get the wrong origin, and origin decides which duties apply, such as US Section 301 tariffs. That's why it flags rather than decides.

Newer and bigger didn't mean better here. A model at a twentieth of the cost matched our current one, and the newer models' deficit came down to one field our instructions never covered. Before paying for an upgrade, or ruling a model out, check which fields it misses.

3. Context windows: measure your longest requests, not your typical ones

Models are increasingly sold on context window size, now up to a million tokens. Logistics paperwork puts that claim to the test: one shipment can bring a hundred-page PDF or a spreadsheet with thousands of rows. Fitting it in, though, is not the same as reading it well. Long-context benchmarks (RULER, Chroma) find accuracy falling well before a model's advertised limit.

On our own shipments, length didn't hurt within the range we tested: the longest requests passed as often as the shortest for every model we tried, and only two Gemini Flash models slipped as documents grew.

The harder case is the rare giant shipment beyond that range. These carry little for their size: the very largest typically end up as one to three lines on the filing, so the task is finding a needle in a haystack. A 128,000-token model may run out of room on them, and a million-token model may fit them yet still miss the one value that matters.

There are several ways to handle these:

  • A larger-window model as a fallback, used only for the few requests that need it.

  • Selecting the most relevant input: pass along only the pages, sheets or rows likely to hold the filed values, picked by a quick first pass or a simple search, and leave out the rest.

  • Splitting the input across several requests, for example a giant PDF or spreadsheet in its own request or in pieces, then merging the answers into one draft. Each request stays in the range where models read accurately, at the cost of a merge step.

Choose a window, and a way to handle the giants, by testing on your longest requests. Your typical request will tell you little, and which approach works best depends on your documents.

4. You do need careful evaluations

Our evaluations move with the instructions. Our filing instructions keep evolving, for example which port to report when goods clear customs inland, how to list shipping documents and how to state quantities. So whenever they change, we refresh the baseline by rerunning the Amari model on today's instructions. On the latest refresh, it scored 6.7 ± 3.3 points (95% confidence interval) below its own answers from a few months earlier, with the difference falling on exactly the fields whose instructions had changed. So section 2 compares every model with the Amari model rerun on today's instructions, not with its old answers.

Telling models apart takes thousands of shipments. Our first quick check used 18 documents, and it put Gemini 3.5 Flash in a tie for last. On 300 real shipments, that model was level with the Amari model. With 18 documents, a single document moves the pass rate by more than 5 points.

Even 300 shipments only go so far. Models don't give identical answers every time: when we repeated six 300-shipment runs a week later, changing nothing, pass rate moved by up to 3 points. So on 300 shipments a gap needs to be about 5 points before we can trust it, which is why Opus 5.5's 4-point gap in section 2 needed a second check. Pinning a gap down to within ±1 point would take around 3,700 shipments.

Before trusting a comparison, run your baseline again and measure your noise. Evaluations older than your current instructions, or a sample too small to separate models, will rank them wrong.

What this means

Customs entries are one product of logistics paperwork. Bookings, delivery orders, freight invoices and proof-of-delivery records are others, drawn from the same documents. That paperwork differs from a benchmark in the ways ours did: values repeat across documents, some are conventions rather than text on the page, and the correct answers come from people working alongside an earlier model. Those differences, more than the model or the input format, decided our results. Teams automating any of it are likely to see the same, and should test on their own records before trusting a leaderboard or a spec sheet.

A person still reviews every shipment: in our case, a licensed broker who also files it. What changes is how much reading they have to redo, and what a good first draft costs.


Contact

Let's get started

Our partners

Office

575 Market ST, San Francisco, CA 94105

Mailing

28 Geary St, STE 650, San Francisco, CA 94108

Quick Contact

©2026. All right reserved

Amari

Contact

Let's get started

Our partners

Office

575 Market ST, San Francisco, CA 94105

Mailing

28 Geary St, STE 650, San Francisco, CA 94108

Quick Contact

©2026. All right reserved

Amari