Europass XML: what the CV format holds, and how to read it in code
Published
Europass XML is the machine-readable form of a Europass CV: an XML document, defined by a schema the European Commission publishes, that holds the candidate's name, contact details, work history, education and skills as tagged elements. A person who builds a CV in the Europass online editor can receive it as that XML, or as a PDF with the XML attached inside it. If your job board or ATS receives Europass CVs, the attached XML is the best data you will ever get from an upload, because nobody has to guess where the job title ends and the employer begins.
This guide is for the developer who has just been handed one of these files. It covers what the schema is built on, how to get the XML out of a Europass PDF, which elements hold the fields you actually want, the traps in reading them, and what to do with the much larger pile of CVs that carry no XML at all. We build ResumeJSON, a resume parsing API that returns JSON, so read the parts about it as written by an interested party. The schema facts below come from the Commission's own Europass2 CV XML Schema Documentation, version 3.0.0, as read in October 2026.
What the Europass CV schema is built on
The schema documentation states its purpose plainly: it describes "the Europass Curriculum Vitae (ECV) Data Standard Specification, which is the format used for the exchanges of CVs in Europass e-portfolio". It is published by DG Employment, Social Affairs and Inclusion, and it is not a format invented from scratch.
The Europass CV model is based on the HR-Open Standards and EURES specifications.
That inheritance shapes everything you will read in a file. The schema is layered across four namespaces:
| Layer | Namespace | What it contributes |
|---|---|---|
| Europass | http://www.europass.eu/1.0 | The main entry point and Europass-only elements |
| EURES | http://www.europass_eures.eu/1.0 | The EU job-mobility extensions |
| HR-Open | http://www.hr-xml.org/3 | Most of the vocabulary: names, employment, education |
| OAGIS | http://www.openapplications.org/oagis/9 | Low-level datatypes and attributes |
The documentation names Candidate.xsd in the org_europass/1.0/Developer/Nouns folder as the file to validate against; everything else is imported from there. If you have met an HR-XML resume before, most of the element names will look familiar, because they are the same ones.
How to get the XML out of a Europass PDF
The schema documentation says individuals using the Europass editors "have the option to receive the document in Europass XML format or PDF format with the XML attached". The second case is the one that trips people up, because the PDF looks like an ordinary CV. The XML is a file attachment embedded in the PDF, which is a standard PDF feature that most viewers hide in a side panel.
Check for it before you do anything else:
- List the attachments. With Poppler installed,
pdfdetach -list cv.pdfprints every embedded file. A Europass PDF with data in it shows an XML file there. - Save the attachment.
pdfdetach -saveall cv.pdfwrites it to the current directory. - Or do it in code. In Python,
pypdfexposes a reader's embedded files throughreader.attachments, a mapping from file name to the file's bytes. In Node,pdf-liband similar libraries can walk the document's embedded-files tree.
from pypdf import PdfReader
reader = PdfReader("cv.pdf")
for name, contents in reader.attachments.items():
if name.lower().endswith(".xml"):
xml_bytes = contents[0]
break
else:
xml_bytes = None # no Europass data: treat it as an ordinary PDFAlways branch on the result. A CV that says "Europass" in its header may have been retyped in a word processor, printed to PDF, or exported without the data. No attachment means you are holding an ordinary PDF, and it needs an ordinary parser.
Which elements hold the fields you want
The root of a Europass CV is Candidate. Under it, the documentation describes three main branches: CandidateSupplier (who supplied the record), CandidatePerson (who the candidate is) and CandidateProfile (what they have done). For a job board or ATS, almost everything lives in the last two.
| You want | Look under | Notes |
|---|---|---|
| Name | CandidatePerson/PersonName | GivenName, FamilyName, plus MiddleName and FormerFamilyName where given |
| Email, phone, address | CandidatePerson/Communication | Repeated, one element per channel |
| Jobs | CandidateProfile/EmploymentHistory/EmployerHistory | One per employer, with OrganizationName, EmploymentPeriod and nested PositionHistory |
| Education | CandidateProfile/EducationHistory/EducationOrganizationAttendance | OrganizationName, EducationDegree, and an attendance period |
| Skills and languages | CandidateProfile/PersonQualifications | Competencies with an optional proficiency level |
| Certificates, licences | CandidateProfile/Certifications, CandidateProfile/Licenses | Separate branches, both optional |
The profile branch goes well beyond a typical CV. The documentation lists sections for military history, patents, publications, speaking history, projects, digital skills, hobbies and more. Decide up front which of these your product stores, because a schema that covers everybody gives you dozens of optional branches, and mapping the ones you never display is wasted code.
Reading a job in Python
Namespaces are the first thing that breaks a naive parser. Element names are qualified, so a search for EmployerHistory with no namespace finds nothing. A tolerant approach matches on the local name instead:
import xml.etree.ElementTree as ET
def local(tag):
return tag.rsplit("}", 1)[-1]
def find_all(node, name):
return [el for el in node.iter() if local(el.tag) == name]
root = ET.fromstring(xml_bytes)
for employer in find_all(root, "EmployerHistory"):
org = next((el.text for el in employer.iter() if local(el.tag) == "OrganizationName"), None)
titles = [el.text for el in employer.iter() if local(el.tag) == "PositionTitle"]
print(org, titles)Matching on local names is less strict than binding each namespace properly, but it survives the case where one file uses a slightly different prefix or version from the next. Validate against the official XSD in a test, and parse tolerantly in production.
The traps in a Europass file
The format is well specified, and it still has edges worth knowing before your first import:
- Dates come at different precisions. The documentation's business rule BR-COM-05 allows
YYYY-MM-DD,YYYY-MM,YYYYor a full timestamp. A start date of2019is valid, so your date column has to accept a year with no month. Store the precision alongside the value, or you will invent a January. - Optional means absent, often. The documentation marks large parts of the tree as optional. Every lookup in your code is a "maybe missing, maybe repeated" case.
- Language is per profile. Rule BR-COM-02 says multiple profiles are allowed but must be in different languages, and BR-COM-01 makes English the default where none is given. A candidate can hand you the same history twice, in two languages. Pick one, or store both deliberately.
- Codes need look-up tables. Industries, for example, are coded against NACE in the schema's code lists. The XML gives you a code; turning it into a label your users can read is your job.
- Old files exist. Europass has changed editors over the years, and the 2020 schema documentation describes the format for the new platform release. Expect older files with a different structure, and test against a sample from each source you receive.
When a CV has no Europass XML
Here is the honest proportion: most CVs you receive will not be Europass files, and many Europass-looking PDFs will not carry the data. For a job board in the EU you might see Europass PDFs with XML, Europass PDFs without it, Word documents, scans and photos, all in the same inbox.
You have three options for everything without an attachment:
| Option | Good for | What you own |
|---|---|---|
| Ask candidates for the XML | A portal for a public employment service | The upload form and the rejection message |
| An open-source parser | Clean, text-based PDFs in one language | The extraction, the field rules, the upkeep |
| A parsing API | Mixed formats and languages | One HTTP call, and mapping its JSON to your tables |
The open source options and the Python libraries are covered in their own articles. For the API route, ResumeJSON takes a PDF, DOCX, plain text, or a JPEG, PNG or WebP image of the CV, and returns one fixed JSON shape: basics, work, education, skills, certifications and languages, with a published field reference. It reads the CV's visible content. It does not read a PDF's embedded attachments, so when the Europass XML is there, use it and skip the call.
The pipeline that works is a fallback, in this order: look for the attachment, use the XML when it exists, and send everything else to a parser. Both branches should write the same record in your database, so map Europass elements and parser fields onto one internal shape rather than storing two. If you need multilingual CVs handled, our notes on multilingual resume parsing cover what breaks.
Mapping both sources to one record
A small mapping table keeps the two branches honest:
| Your field | From Europass XML | From a parser's JSON |
|---|---|---|
| First name | PersonName/GivenName | the parser's name field |
| Employer | EmployerHistory/OrganizationName | work[].company or equivalent |
| Job title | PositionHistory/PositionTitle | work[].title or equivalent |
| Start date | EmploymentPeriod start, at its precision | the parser's start date |
| School | EducationOrganizationAttendance/OrganizationName | education[].institution or equivalent |
Write one test per row with a real file from each source. When the two disagree for the same candidate, the XML wins, because the candidate typed it into labelled boxes.
When you only need the XML
Skip a parser entirely when:
- Every CV comes from one channel that produces Europass files, such as a portal that asks applicants to upload their Europass export.
- You can reject what does not carry XML, and tell the candidate how to export it, without losing applicants you care about.
- You only need the fields the schema guarantees and can live with the optional ones being absent.
In that case a schema validator and the twenty lines of Python above are the whole integration. Everywhere else, the attachment check is the first step of a pipeline, and a parser handles the rest. You can try ours on one of your own CVs with the free parser, no signup needed.