Layered datasets rendered as structured data blocks

Portfolio

Selected work from our delivery floor.

Anonymised where clients require it, but the scope, volume and measured quality are exactly as shipped.

Projects delivered
260+
Records processed
180M+
Languages covered
42
Delivery centres
6

Case studies

Projects across every data modality

Global invoice extraction corpus
Document AI

Global invoice extraction corpus

1.2M double-keyed pages across 19 countries with normalised tax fields, delivered as JSONL plus page-level confidence.

99.6% field accuracy

Retail shelf perception set
Computer vision

Retail shelf perception set

480,000 annotated shelf frames with SKU polygons, facing counts and out-of-stock flags for a grocery vision model.

0.92 mean IoU

Support agent evaluation harness
Agents

Support agent evaluation harness

Authored 26,000 grounded instruction/response pairs and a rubric-scored eval suite for a deployed customer-support agent.

31% fewer escalations

Multilingual call transcription
Speech

Multilingual call transcription

3,100 hours of transcribed and diarised contact-centre audio across 11 languages with speaker and sentiment tags.

4.1% word error rate

Adverse-weather LiDAR sequences
Autonomy

Adverse-weather LiDAR sequences

310 hours of fused camera and LiDAR sequences focused on night, rain and fog edge cases for a robotics perception stack.

0.94 audit IoU

Preference ranking programme
RLHF

Preference ranking programme

40,000 weekly pairwise comparisons from a calibrated rater panel with continuous agreement monitoring.

0.81 rater agreement

Partner network

A delivery network across four regions

We work with vetted data companies in China, Russia, the USA and the Middle East so language coverage, timezone coverage and surge capacity are never a bottleneck.

Shenzhen, China

Silk Data Works

Vision and OCR corpora for east-Asian scripts.

Hangzhou, China

Hanyu Vision Labs

Retail shelf and traffic-scene collection.

Moscow, Russia

Volga Datalytics

Cyrillic NLP corpora and speech transcription.

St. Petersburg, Russia

Neva Speech Corp

Multi-accent audio capture and diarisation.

Austin, USA

Meridian Data Group

Document extraction and fintech ground truth.

Seattle, USA

Northline Labeling

LiDAR cuboids and autonomy sequence review.

Dubai, UAE

Gulf Cognitive

Arabic annotation and bilingual QA pods.

Riyadh, Saudi Arabia

Levant Data Partners

Regulated-sector entry and compliance review.

Need labelled data by next sprint?

Send us a spec, a schema, or a handful of raw files. We return a pilot batch of 1,000 records with a quality report within five working days — free of charge.

info@betalen.in