Global Invoice OCR Corpus
Invoices and receipts from 38 countries with 24 extracted fields per page, line-item tables and full bounding-box coordinates.
- Volume
- 1.2M pages
- Formats
- JSONL + PNG

Products
Each catalogue dataset ships with a datasheet, a licence covering commercial model training, a held-out evaluation split and a provenance record for every source.
Catalogue
Sample splits of 500 records are available for any dataset on request, at no cost.
Invoices and receipts from 38 countries with 24 extracted fields per page, line-item tables and full bounding-box coordinates.
Store-shelf photography with SKU-level boxes, facing counts, price-tag regions and out-of-stock gap annotations.
Consented contact-centre audio in 21 languages with verbatim transcripts, diarisation and accent metadata.
Human-authored instruction and response pairs across reasoning, coding, extraction and refusal categories, with rubric scores.
Synchronised camera and LiDAR sequences with 3D cuboids, tracking IDs, lane geometry and weather condition tags.
De-identified clinical intake and claims forms with structured field extraction, reviewed by medical coders.
Custom builds
We recruit consented contributors and capture net-new data to your protocol — imagery, speech, handwriting or behavioural.
Programmatic generation and augmentation to fill class imbalance, with human review on every synthetic sample kept.
Provenance tracking, consent records, PII scrubbing and a datasheet you can hand to your legal team.
Send us a spec, a schema, or a handful of raw files. We return a pilot batch of 1,000 records with a quality report within five working days — free of charge.
info@betalen.in