Multilingual resume parsing: what breaks, and how to test a parser
Published
Multilingual resume parsing means reading a CV written in any language and returning the same structured fields you would get from an English one: name, contact details, jobs with dates, education, skills. The language count on a vendor's page tells you very little about whether that works for your candidates. What decides it is five specific things that go wrong on non-English CVs, and you can check all five in an afternoon with ten real files. This guide lists them, shows what three established vendors say about language coverage, and ends with a test that works on any parser, including ours.
We build ResumeJSON, a resume parsing API, so read this as written by an interested party. Every vendor figure below was read from the vendor's own page as of October 2026.
What the vendors publish
Language coverage is the one number every parser advertises, and the three below publish it differently.
| Vendor | What its page says about languages | Where it says it |
|---|---|---|
| Textkernel | Resume parsing in 29 languages, job parsing in 9 | Parser product page |
| Affinda | Resume parsing "across 50+ languages and formats" | Resume parser page, FAQ |
| RChilli | Parsing in "nearly 27+ languages", with English, French, German, Polish, Spanish, Turkish, Portuguese, Italian, Dutch, Japanese, Russian, Hebrew, Romanian and Thai named | Multilingual parser page for Oracle Recruiting |
Read the table with care, because the three numbers are not comparable. Textkernel gives separate counts for resumes and for job postings. Affinda's sentence folds languages and file formats into a single count. RChilli's figure appears on a page about one Oracle integration and says "nearly 27+", which is a floor with an approximation in front of it. None of the three says how well each language performs, and none can: accuracy on Thai is a different number from accuracy on German.
So a list of languages answers "was this language considered?" and nothing more. The questions that matter are below.
The five things that break on a non-English CV
1. Dates written in words
An English CV says "Aug 2021 to present". An Indonesian one says "Agustus 2021 sampai sekarang". A German one says "08/2021 bis heute". A Spanish one says "agosto de 2021 al presente". A parser that handles the month names but misses the phrase for "currently" returns an end date of null and a job that looks finished. That error is quiet: the field is simply wrong in the direction of "left the company".
What to check: the is_current flag or its equivalent on a CV whose current job is described in the candidate's own words.
2. Section headings you did not plan for
Parsers find the work history by recognising its heading. "Pengalaman Kerja", "Berufserfahrung", "Expérience professionnelle" and "職歴" all mean the same thing. A rule-based parser needs each one on a list. A parser that reads meaning handles headings it has never seen, but still has to cope with CVs that have no headings at all, which are common in some countries.
What to check: a CV where the headings are decorative or missing.
3. Names, titles and the order of both
Family name first is normal in Chinese, Japanese, Korean and Hungarian CVs, and a parser that assumes "given name, family name" splits them wrongly. Honorifics and academic titles (Dr., Ir., S.Kom., Dipl.-Ing.) sit inside the name line and are easy to absorb into the name.
What to check: the name field on three CVs with different conventions, compared with how the candidate writes their name on their own profile.
4. Scripts, right-to-left text and bad text layers
Latin text with accents is the easy case. Arabic and Hebrew run right to left, and the text layer of some PDFs stores such lines in reverse order. CJK documents sometimes use fonts whose embedded text comes out as meaningless characters when copied, even though the page looks fine on screen. If the parser starts from the extracted text, it starts from garbage.
What to check: copy the text out of the PDF yourself. If what you paste is unreadable, the parser needs to read the page as an image, which is the OCR route.
5. Mixed-language CVs
A CV in Spanish with English job titles, or a Dutch CV with an English skills list, is the normal case for anyone who works internationally. Parsers that detect one language per document and then apply that language's rules can mishandle the other half.
What to check: a CV that switches language in the middle, and whether the fields from both halves arrive.
How to test any parser in an afternoon
You do not need a benchmark. You need ten files and a spreadsheet.
- Collect ten real CVs in the languages your candidates use, including one scan, one two-column layout and one mixed-language file. Real files only: a translated English CV is easier than the real thing.
- Write the expected answer for four fields by hand: the candidate's name, the start date of their most recent job, whether that job is current, and the highest degree. Four fields keep the checking short.
- Run every file through the parser and put the four returned values beside your four expected ones.
- Mark each cell right or wrong and count wrong cells per language. Ten files and four fields give you forty cells, which is enough to see a pattern and too few to quote a percentage, so do not.
- Look at the wrong ones first. A wrong name is a parsing problem. A missing job is usually a heading or layout problem. A wrong
is_currentis the date-phrase problem from the first section.
Most vendors offer a demo or a free tier for exactly this. Textkernel's page lists 500 free credits for its API, and our own free resume parser needs no signup. Use whichever suits you and compare.
Reading a CV in one language and storing it in another
There are two separate jobs hiding under "multilingual", and mixing them up causes most confusion.
- Reading a CV in its own language and extracting the fields.
- Writing the extracted text in a language your team reads.
If your recruiters read English, a French CV that arrives as French fields forces someone to translate before they can search. ResumeJSON handles both in one request. By default the text values come back in the CV's own language. Add output_language (a name like English or a code like en) and the job titles, highlights, skills, degrees and locations are written in that language, while names of people, companies and schools, emails, phone numbers and URLs stay exactly as the CV has them. Dates are normalised to YYYY-MM in every case, which removes the "Agustus 2021" problem from your code entirely. A field the CV does not state comes back as null instead of a guess.
There is no fixed list of supported output languages to check, so the way to know whether yours works is the test above. Our article on translating a resume to English covers the output side in more detail, and extract information from a resume lists every field you get back.
Scanned and photographed CVs are read as images, so a paper CV in Arabic or Thai goes through the same request as a PDF. That removes failure number four for text layers, though it does not make a blurred photo readable.
Which option fits
| Your situation | A reasonable choice |
|---|---|
| Your candidates are in a handful of European languages and you want a vendor with a published, per-language engine | A specialist vendor such as Textkernel, whose page names the 29 languages up front |
| You need the widest advertised count and a managed recruiting platform | Affinda lists 50+, which is worth verifying on your own files |
| You run an Oracle Recruiting stack | RChilli publishes a page for exactly that integration |
| You are a developer who wants JSON from any language, with the output language under your control and per-parse pricing | ResumeJSON, after you have run the ten-file test |
When the manual way is enough
If you receive a few foreign-language CVs a month, a translator and a spreadsheet are cheaper than any integration. Paste the CV into a translation tool, read it, and type the four fields. Parsing earns its place when the volume makes typing the bottleneck, or when search and filters need every candidate in the same fields regardless of the language they wrote in. For a big backlog, bulk resume parsing covers running it in a loop, and resume to Excel turns the output into one sheet.
Try it on your hardest CV
Take the file you expect to fail: the scan, the right-to-left one, the mixed-language one. Drop it into the free resume parser, type English in the Output language box if you want it translated, and read the four fields against your expected answers. If they are right, ten more files will tell you whether it holds. If they are wrong, you learned that for free, which is the point of testing before you integrate.