Multi-Task Sentence Transformer for Receipt Processing
The Decision
Freeze the BGE-M3 backbone and train only the two task heads.
Takeaways
- 97.8% sentence classification accuracy on 414 held-out receipt lines.
- 0.898 macro-F1 on entity tagging, with labels from the public CORD-v2 receipt corpus.
- Training both tasks to one stopping point means choosing which one to compromise.
A receipt is a list of short lines. For each line I want two answers. What kind of line is it, and which words are the product name and which are the price? This project answers both with one transformer encoder and two small heads on top of it. The encoder runs once per line and both heads read its output.
Source code: github.com/olivia-jackson-lambert/sentence-project
System Overview
The OCR stage converts the photo to grayscale, denoises it, binarises it with Otsu thresholding and passes it to Tesseract. The text is then split into lines on newlines and periods.
import pytesseract
import cv2
def preprocess_receipt_image(image_path):
"""Preprocess receipt image for better OCR accuracy"""
img = cv2.imread(image_path)
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
denoised = cv2.fastNlMeansDenoising(gray, h=10)
_, thresh = cv2.threshold(denoised, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
return thresh
def extract_text_from_receipt(image_path):
"""Extract text from receipt image using OCR"""
return pytesseract.image_to_string(preprocess_receipt_image(image_path), config='--psm 6')The results below do not use this stage. They run on CORD’s own transcribed lines, so OCR error is not measured. I come back to this under Limitations.
The Model
Backbone
The encoder is BGE-M3 from the Beijing Academy of Artificial Intelligence. It is an XLM-RoBERTa-large model with 24 layers, 1024 hidden dimensions, a 250,000 token vocabulary and a 2.3 GB checkpoint. That is big for a receipt parser. It gives strong general sentence representations, and its multilingual coverage matters here because the CORD receipts are Indonesian.
The “M3” in the name means multi-lingual, multi-granularity and multi-functionality. Multi-functionality refers to dense, sparse and ColBERT-style retrieval. The model has no built-in classification or NER. I added the multi-task structure on top.
Two Heads
The classification head reads the CLS token state and passes it through a 128-unit hidden layer to four classes. The entity head is a single linear layer applied to every token state.
self.cls_head = nn.Sequential(
nn.Linear(1024, 128), nn.ReLU(), nn.Dropout(0.1), nn.Linear(128, 4)
)
self.ner_head = nn.Linear(1024, 5)| Labels | |
|---|---|
| Sentence class (4) | ITEM, ITEM_OPTION, SUBTOTAL, TOTAL |
| Entity tags (5) | O, B-NAME, I-NAME, B-PRICE, I-PRICE |
The entity tags use the BIO scheme. B marks the first word of a name or price, I marks the words that continue it and O marks everything else. That lets the model pick out multi-word names such as Sambal Gabus Asin Cb.
The Data
Training data comes from CORD-v2, the Consolidated Receipt Dataset. It has 1,000 photographed Indonesian receipts with line-level annotations. CORD groups the lines of a receipt into rows with a group_id, and I use each row as one training sentence:
menu.cnt "1"
menu.nm "REAL GANACHE" -> "1 REAL GANACHE 16,500"
menu.price "16,500"
This grouping is what makes the problem genuinely multi-task. The sentence class says what kind of row it is. The entity tags say which words inside it are the name and which are the price. Here is one training sentence:
[ITEM] 1/O x/O Nasi/B-NAME Campur/I-NAME Bali/I-NAME 75,000/B-PRICE
CORD tags the whole value field of a subtotal or total row as price, so a label word such as Subtotal also carries a price tag.
| Split | Sentences | ITEM | ITEM_OPTION | SUBTOTAL | TOTAL |
|---|---|---|---|---|---|
| Train | 3,430 | 1,921 | 183 | 531 | 795 |
| Validation | 386 | 202 | 19 | 65 | 100 |
| Test | 414 | 229 | 22 | 63 | 100 |
Word tags are aligned to BGE-M3 subword tokens with the tokenizer’s word_ids() mapping. Only the first subword of each word carries the tag. Continuation subwords and padding are set to -100, so they add nothing to the loss.
Training
The backbone is frozen. Its token states are computed once, cached, and reused every epoch, so only the heads train. That is 137K trainable parameters against 568M frozen, or 0.024 percent of the model, which is what makes the run practical on a laptop. The loss is the sum of the two cross-entropy losses. I trained with AdamW at a learning rate of 0.001 and a batch size of 64 for 30 epochs.
The original notebook’s MultiTaskModel did not freeze anything. It never set requires_grad = False and passed model.parameters() to the optimiser. The training script behind these results, train_heads.py, adds the freeze.
The two tasks want different amounts of training. Classification validation loss bottoms out at epoch 4 and then climbs while training loss goes to zero. The model has already solved that task and is now overfitting it. Entity tagging validation loss is still falling at epoch 30.
So one stopping point means compromising one task. Stopping early protects classification and leaves entity tagging undertrained. Running to 30 epochs does the reverse. Separate learning rates, or a loss weight that decays the classification term, would serve both.
Results
Scores on the 414 held-out test lines:
| Task | Metric | Score |
|---|---|---|
| Classification | Accuracy | 0.978 |
| Classification | F1 (weighted) | 0.979 |
| Entity tagging | Token accuracy | 0.902 |
| Entity tagging | F1 (macro) | 0.898 |
Macro-F1 is the entity number I quote. It weights all five tags equally, so the rarer tags count as much as the common ones. In the test labels B-PRICE is the most common tag at about 30 percent of words and O is about 16 percent.
Example Predictions
The predictions below come from the trained heads running on held-out test lines.
The name span is right and the sentence class is right. The one miss is the price. The reference treats Rp 35.000 as one span, but the model starts a new span at 35.000 instead of continuing it.
All 13 lines get the right sentence class. The tags match the reference exactly on 9 of them. The four misses are different kinds of error:
1 Nasi Tambah 7.000gives the name asNasiand the price asTambah 7.000. The model has pushed part of a product name into the price.1 Ikan Bawal Besa 28.500extracts the right name, but tagsBawalas the start of a second name.- The subtotal row tags part of the food and drink breakdown as price. CORD leaves those words outside any span.
- The total row pulls the item count
14into the price span.
Limitations
The data is not what the system was designed for. CORD receipts are Indonesian, and the labels are CORD’s hierarchy mapped onto four classes and five tags. English grocery receipts would need their own labelled sample.
The encoder never adapts. With the backbone frozen, errors like Nasi Tambah cannot be fixed by the encoder, only by the heads. Unfreezing the top layers needs real GPU time instead of a cached forward pass.
Entity scoring is per token. A name with one wrong I-NAME tag still earns most of its credit. Strict span matching would count the whole entity as missed, and would likely give a lower number.
One stopping point serves two tasks badly. Classification overfits from epoch 4 while entity tagging is still improving at 30. This run picks 30.
OCR is outside the measured loop. Real Tesseract output would add errors to both tasks, and I have not measured how many. The segmentation step also splits on periods, and CORD prices use a period as the thousands separator (7.000). As written it would cut prices in two, so it needs fixing before the OCR path is used.
Next Steps
Unfreeze the top encoder layers. This is the most likely gain for entity tagging and a direct test of the frozen-backbone choice.
Weight the two losses. Decay the classification term once it saturates so training keeps serving entity tagging.
Score entities at span level. Report strict and relaxed span matches alongside token macro-F1.
Close the OCR loop. Fix the period split, run Tesseract over the CORD photographs and re-measure both tasks on its output.





