Multi-Task Sentence Transformer for Receipt Processing

Natural Language Processing
Deep Learning
One frozen BGE-M3 encoder, two small heads. Each receipt line gets a sentence class and token-level name and price tags, trained and tested on the public CORD-v2 corpus.
Published

January 24, 2026

The Decision

Freeze the BGE-M3 backbone and train only the two task heads.

Takeaways

  • 97.8% sentence classification accuracy on 414 held-out receipt lines.
  • 0.898 macro-F1 on entity tagging, with labels from the public CORD-v2 receipt corpus.
  • Training both tasks to one stopping point means choosing which one to compromise.

A receipt is a list of short lines. For each line I want two answers. What kind of line is it, and which words are the product name and which are the price? This project answers both with one transformer encoder and two small heads on top of it. The encoder runs once per line and both heads read its output.

Source code: github.com/olivia-jackson-lambert/sentence-project

System Overview

Six boxes stacked top to bottom and joined by arrows: Receipt Image, OCR with Tesseract, Segmentation into lines, the BGE-M3 encoder (24 layers, frozen, 1024-dimensional states), the two task heads, and Structured Data holding a line class plus name and price spans. Brackets at the side group the boxes into OCR, Preprocessing and Multi-Task Model.

The receipt processing pipeline, from photograph to structured line data.

The OCR stage converts the photo to grayscale, denoises it, binarises it with Otsu thresholding and passes it to Tesseract. The text is then split into lines on newlines and periods.

import pytesseract
import cv2

def preprocess_receipt_image(image_path):
    """Preprocess receipt image for better OCR accuracy"""
    img = cv2.imread(image_path)
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    denoised = cv2.fastNlMeansDenoising(gray, h=10)
    _, thresh = cv2.threshold(denoised, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
    return thresh

def extract_text_from_receipt(image_path):
    """Extract text from receipt image using OCR"""
    return pytesseract.image_to_string(preprocess_receipt_image(image_path), config='--psm 6')

The results below do not use this stage. They run on CORD’s own transcribed lines, so OCR error is not measured. I come back to this under Limitations.

The Model

Backbone

The encoder is BGE-M3 from the Beijing Academy of Artificial Intelligence. It is an XLM-RoBERTa-large model with 24 layers, 1024 hidden dimensions, a 250,000 token vocabulary and a 2.3 GB checkpoint. That is big for a receipt parser. It gives strong general sentence representations, and its multilingual coverage matters here because the CORD receipts are Indonesian.

The “M3” in the name means multi-lingual, multi-granularity and multi-functionality. Multi-functionality refers to dense, sparse and ColBERT-style retrieval. The model has no built-in classification or NER. I added the multi-task structure on top.

Two Heads

Flow diagram. Input tokens go into the frozen BGE-M3 backbone (XLM-RoBERTa, 24 layers, 16 heads, hidden size 1024). The CLS token state of size 1024 feeds a classification head (Linear 1024 to 128, ReLU with dropout 0.1, Linear 128 to 4) that outputs four sentence classes. All token states feed an NER head (Linear 1024 to 5) that outputs one of five BIO tags per token. Only the heads are marked trainable.

The multi-task architecture. Both heads read the same frozen encoder output.

The classification head reads the CLS token state and passes it through a 128-unit hidden layer to four classes. The entity head is a single linear layer applied to every token state.

self.cls_head = nn.Sequential(
    nn.Linear(1024, 128), nn.ReLU(), nn.Dropout(0.1), nn.Linear(128, 4)
)
self.ner_head = nn.Linear(1024, 5)
Labels
Sentence class (4) ITEM, ITEM_OPTION, SUBTOTAL, TOTAL
Entity tags (5) O, B-NAME, I-NAME, B-PRICE, I-PRICE

The entity tags use the BIO scheme. B marks the first word of a name or price, I marks the words that continue it and O marks everything else. That lets the model pick out multi-word names such as Sambal Gabus Asin Cb.

The Data

Training data comes from CORD-v2, the Consolidated Receipt Dataset. It has 1,000 photographed Indonesian receipts with line-level annotations. CORD groups the lines of a receipt into rows with a group_id, and I use each row as one training sentence:

menu.cnt    "1"
menu.nm     "REAL GANACHE"    ->   "1 REAL GANACHE 16,500"
menu.price  "16,500"

This grouping is what makes the problem genuinely multi-task. The sentence class says what kind of row it is. The entity tags say which words inside it are the name and which are the price. Here is one training sentence:

[ITEM]  1/O  x/O  Nasi/B-NAME  Campur/I-NAME  Bali/I-NAME  75,000/B-PRICE

CORD tags the whole value field of a subtotal or total row as price, so a label word such as Subtotal also carries a price tag.

Split Sentences ITEM ITEM_OPTION SUBTOTAL TOTAL
Train 3,430 1,921 183 531 795
Validation 386 202 19 65 100
Test 414 229 22 63 100

Word tags are aligned to BGE-M3 subword tokens with the tokenizer’s word_ids() mapping. Only the first subword of each word carries the tag. Continuation subwords and padding are set to -100, so they add nothing to the loss.

Training

The backbone is frozen. Its token states are computed once, cached, and reused every epoch, so only the heads train. That is 137K trainable parameters against 568M frozen, or 0.024 percent of the model, which is what makes the run practical on a laptop. The loss is the sum of the two cross-entropy losses. I trained with AdamW at a learning rate of 0.001 and a batch size of 64 for 30 epochs.

The original notebook’s MultiTaskModel did not freeze anything. It never set requires_grad = False and passed model.parameters() to the optimiser. The training script behind these results, train_heads.py, adds the freeze.

Two line charts of cross-entropy loss by epoch, stacked on a shared epoch axis. Top, sentence classification: validation loss reaches its minimum of about 0.068 at epoch 4 and then climbs to about 0.11 by epoch 30, while training loss falls to near zero. Bottom, entity tagging: training and validation loss both fall steadily from about 0.7 to about 0.25 and are still falling at epoch 30.

Training and validation loss for each head over 30 epochs.

The two tasks want different amounts of training. Classification validation loss bottoms out at epoch 4 and then climbs while training loss goes to zero. The model has already solved that task and is now overfitting it. Entity tagging validation loss is still falling at epoch 30.

So one stopping point means compromising one task. Stopping early protects classification and leaves entity tagging undertrained. Running to 30 epochs does the reverse. Separate learning rates, or a loss weight that decays the classification term, would serve both.

Two line charts of score by epoch, stacked on a shared epoch axis. Top, sentence classification: validation accuracy and weighted F1 reach about 0.98 by epoch 2 and stay flat, while training accuracy rises to 1.0. Bottom, entity tagging: validation token accuracy and macro-F1 climb from about 0.79 to about 0.93 and are still rising at epoch 30.

Accuracy and F1 for each head over 30 epochs.

Results

Scores on the 414 held-out test lines:

Task Metric Score
Classification Accuracy 0.978
Classification F1 (weighted) 0.979
Entity tagging Token accuracy 0.902
Entity tagging F1 (macro) 0.898

Macro-F1 is the entity number I quote. It weights all five tags equally, so the rarer tags count as much as the common ones. In the test labels B-PRICE is the most common tag at about 30 percent of words and O is about 16 percent.

Example Predictions

The predictions below come from the trained heads running on held-out test lines.

Eight words of a receipt line, wrapped onto two rows of four, each in a coloured box with its predicted tag below: 1 is O; 35.000, Rp and 35.000 are each B-PRICE; NASI is B-NAME; GOR, SFD and TLR are I-NAME. The second 35.000 has a dashed outline because the reference tags it I-PRICE. The line reads: Sentence class ITEM (correct), tags matching the reference 7 of 8. A bar chart below puts nearly all class probability on ITEM.

One item line from the test set, with the entity tag and class probabilities the model predicted.

The name span is right and the sentence class is right. The one miss is the price. The reference treats Rp 35.000 as one span, but the model starts a new span at 35.000 instead of continuing it.

Top, a photograph of an Indonesian restaurant receipt from the CORD-v2 test set. Below it, a table of 13 lines with the predicted class, the line text, the name span and the price span. Eleven ITEM rows, one SUBTOTAL and one TOTAL, all classified correctly. Four rows are shaded and outlined with dashes where the tags differ from the reference: Nasi Tambah, Ikan Bawal Besa, the subtotal row and the total row.

A full test receipt run through the model one line at a time.

All 13 lines get the right sentence class. The tags match the reference exactly on 9 of them. The four misses are different kinds of error:

  • 1 Nasi Tambah 7.000 gives the name as Nasi and the price as Tambah 7.000. The model has pushed part of a product name into the price.
  • 1 Ikan Bawal Besa 28.500 extracts the right name, but tags Bawal as the start of a second name.
  • The subtotal row tags part of the food and drink breakdown as price. CORD leaves those words outside any span.
  • The total row pulls the item count 14 into the price span.

Limitations

The data is not what the system was designed for. CORD receipts are Indonesian, and the labels are CORD’s hierarchy mapped onto four classes and five tags. English grocery receipts would need their own labelled sample.

The encoder never adapts. With the backbone frozen, errors like Nasi Tambah cannot be fixed by the encoder, only by the heads. Unfreezing the top layers needs real GPU time instead of a cached forward pass.

Entity scoring is per token. A name with one wrong I-NAME tag still earns most of its credit. Strict span matching would count the whole entity as missed, and would likely give a lower number.

One stopping point serves two tasks badly. Classification overfits from epoch 4 while entity tagging is still improving at 30. This run picks 30.

OCR is outside the measured loop. Real Tesseract output would add errors to both tasks, and I have not measured how many. The segmentation step also splits on periods, and CORD prices use a period as the thousands separator (7.000). As written it would cut prices in two, so it needs fixing before the OCR path is used.

Next Steps

Unfreeze the top encoder layers. This is the most likely gain for entity tagging and a direct test of the frozen-backbone choice.

Weight the two losses. Decay the classification term once it saturates so training keeps serving entity tagging.

Score entities at span level. Report strict and relaxed span matches alongside token macro-F1.

Close the OCR loop. Fix the period split, run Tesseract over the CORD photographs and re-measure both tasks on its output.

Also covered
Machine LearningMulti-Task LearningPyTorchTransformersSentence EmbeddingsNamed Entity RecognitionTransfer LearningOCRPython