
Global invoice extraction corpus
1.2M double-keyed pages across 19 countries with normalised tax fields, delivered as JSONL plus page-level confidence.
99.6% field accuracy

Portfolio
Anonymised where clients require it, but the scope, volume and measured quality are exactly as shipped.
Case studies

1.2M double-keyed pages across 19 countries with normalised tax fields, delivered as JSONL plus page-level confidence.
99.6% field accuracy

480,000 annotated shelf frames with SKU polygons, facing counts and out-of-stock flags for a grocery vision model.
0.92 mean IoU

Authored 26,000 grounded instruction/response pairs and a rubric-scored eval suite for a deployed customer-support agent.
31% fewer escalations

3,100 hours of transcribed and diarised contact-centre audio across 11 languages with speaker and sentiment tags.
4.1% word error rate

310 hours of fused camera and LiDAR sequences focused on night, rain and fog edge cases for a robotics perception stack.
0.94 audit IoU

40,000 weekly pairwise comparisons from a calibrated rater panel with continuous agreement monitoring.
0.81 rater agreement
Partner network
We work with vetted data companies in China, Russia, the USA and the Middle East so language coverage, timezone coverage and surge capacity are never a bottleneck.
Shenzhen, China
Vision and OCR corpora for east-Asian scripts.
Hangzhou, China
Retail shelf and traffic-scene collection.
Moscow, Russia
Cyrillic NLP corpora and speech transcription.
St. Petersburg, Russia
Multi-accent audio capture and diarisation.
Austin, USA
Document extraction and fintech ground truth.
Seattle, USA
LiDAR cuboids and autonomy sequence review.
Dubai, UAE
Arabic annotation and bilingual QA pods.
Riyadh, Saudi Arabia
Regulated-sector entry and compliance review.
Send us a spec, a schema, or a handful of raw files. We return a pilot batch of 1,000 records with a quality report within five working days — free of charge.
info@betalen.in