Language and text data
Language and text data, worked by people who speak it
Transcription, translation review, tagging and speech data — the work where a native speaker is the bottleneck and no model can fill the gap.
The work
What we take on
If your task isn't listed, ask anyway — the answer is either yes, or a straight no with a reason.
- Transcription and timestamping
- Translation and translation review
- Speech data collection and validation
- Text annotation, tagging and classification
- Transliteration and romanisation checks
- Prompt and response review for language models
How it's delivered
- Formats
- JSON, JSONL, CSV, SRT/VTT, or your own schema
- Unit of work
- Per file, per minute of audio, or per seat-hour
- Quality signal
- Sample pass rate against your gold set, reported per batch
- Turnaround
- Agreed per batch at scoping, then tracked against it daily
Honest scoping
When this is a good fit, and when it isn't
Taking work we'd do badly costs us more than turning it down. This is the list we'd give you on the call anyway.
Good fit
- You have volume in Indian languages and can't hire for it
- Your existing vendor's output needs a human review pass
- You need consistent people rather than a rotating crowd
Not us
- You need a language nobody on our floor speaks — we will tell you rather than guess
- You need certified legal or medical translation