Resume NER: datasets, models, and what NER cannot give you
Published
Resume NER is named entity recognition applied to CVs: a model reads the text of a resume and tags spans as a name, an email, a company, a job title, a degree or a skill. You can train one on a public dataset, or download one already fine-tuned from Hugging Face and run it in five lines of Python. What you get back is a flat list of labelled spans. Turning that list into a resume record, with each job title attached to the right employer and dates, is a second problem that NER does not solve. This article covers both halves: the datasets and models that exist, working code, and what to do about the gap.
We build ResumeJSON, a resume parsing API, so read this as written by an interested party. Every dataset size, license and score below was read from the dataset or model card on Hugging Face as of September 2026.
What resume NER actually does
A token classification model looks at each token of the input and assigns it a tag. Most resume datasets use BIO tagging:
B-ORGmarks the first token of an organisation;I-ORGmarks each following token of the same organisation;Omarks a token that belongs to no entity.
So "Senior Engineer at Acme Corp" might come back as B-TITLE I-TITLE O B-ORG I-ORG. Group the tags and you have two entities: a title and an organisation.
That is the whole output. The model has no notion of a "work entry". It found a title and it found a company. Whether they belong together, and which of the four date ranges on the page belongs to them, is left to you.
Public resume NER datasets
| Dataset | Language | Size | Labels | License |
|---|---|---|---|---|
yashpwr/resume-ner-training-data | English | 22,855 examples, BIO-tagged JSONL | PERSON, ORG, EMAIL, PHONE, LOCATION, EDUCATION, SKILLS, EXPERIENCE | MIT |
Mehyaar/Annotated_NER_PDF_Resumes | English | 5,029 CVs, JSON with character offsets | IT skills only | MIT |
PassbyGrocer/resume-ner | Chinese | 3,821 train, 463 validation, 477 test | NAME, ORG, TITLE, EDU, LOC, PRO, CONT, RACE | Apache-2.0 |
A few things the cards tell you that matter before you pick one:
- The label sets do not agree. One dataset calls a job title
EXPERIENCE, another calls itTITLE, and the skills dataset tags nothing but skills. Merging two datasets means writing a label map first. - The skills dataset is noisy by its own example. Its sample record tags "Building", "Knowledge" and "performance" as skills. That is a fair picture of what CV skill annotation looks like at scale, and a warning about what a model trained on it will return.
- The third dataset is Chinese. Its labels are useful as a reference schema, but it will not help a model read English CVs.
Pretrained resume NER models
You do not have to train anything to try this. The model card for yashpwr/resume-ner-bert-v2 describes a BERT token classifier with 25 entity types, and states:
This model achieves 90.87% F1 score and is trained on a comprehensive dataset of 22,542 resume samples from multiple sources.
That score is the author's own measurement on their own held-out data. Treat it as a claim about that data. Your CVs come from a different mix of countries, templates and industries, and the only F1 that matters is the one you measure on fifty of them.
Running it takes the standard transformers pipeline:
from transformers import pipeline
ner = pipeline(
"token-classification",
model="yashpwr/resume-ner-bert-v2",
aggregation_strategy="simple",
)
text = open("cv.txt").read()
for ent in ner(text):
print(ent["entity_group"], "|", ent["word"], "|", round(ent["score"], 2))aggregation_strategy="simple" merges the B and I tags into whole spans, so you get "Acme Corp" once instead of two tokens.
The 512-token wall
BERT models read at most 512 tokens at a time. A two-page CV is often longer than that. Unless you handle it, the tail of the resume is truncated or the call fails, and the tail is usually education and certifications. Split the text into overlapping windows (recent transformers versions accept a stride argument on this pipeline for that), run each, and de-duplicate spans that appear in the overlap.
Training your own
If you have labelled CVs of your own, fine-tuning is the standard recipe:
- Convert your data to tokens plus BIO tags, the same shape as the first dataset above.
- Load a base encoder (
bert-base-casedor similar) withAutoModelForTokenClassification, passing your label list. - Align labels to sub-word tokens. A word like "PostgreSQL" splits into several pieces; label the first and mask the rest with
-100. - Train with the
TrainerAPI and evaluate withseqeval, which scores whole entities rather than tokens. - Measure on CVs the model never saw, from sources the training set did not include.
Step 5 is where most resume NER projects find out their real accuracy. A model trained on one recruiting agency's CVs learns that agency's templates.
spaCy is the other common route. Its EntityRecognizer trains from the same kind of data, and the Mehyaar dataset ships character offsets that convert to spaCy's DocBin format with a short script. For skills specifically, a dictionary matcher is often a better start than a model; our article on extracting skills from a resume walks through that choice.
What NER cannot give you
Here is the output of a good resume NER model on a short work history, grouped:
| Span | Label |
|---|---|
| Acme Corp | ORG |
| Senior Engineer | TITLE |
| 2021 | DATE |
| Present | DATE |
| Beta Ltd | ORG |
| Engineer | TITLE |
| 2018 | DATE |
| 2021 | DATE |
Every span is right. You still do not have a work history. To build one you need:
- Grouping. Which title belongs to which company. Layout decides this, and layout varies: title above company, company above title, both on one line, two titles under one employer after a promotion.
- Date pairing. Which dates are a start and which an end, and that "Present" means the job is current.
- Normalisation. "Jan '21", "01/2021" and "2021" should become comparable values.
- Sections. A company named in the education section is a university. A company named in a project description may be a client.
Teams handle this with rules layered on the NER output: find section headings, then walk entities in reading order and open a new entry at each organisation. It works on the templates you wrote the rules for. Our resume parsing meaning piece compares this approach with the others in more detail.
Resume NER against a parsing API
| Your own NER model | A resume parsing API | |
|---|---|---|
| Output | Flat list of labelled spans | Nested record: contact details, work entries, education, skills |
| Grouping and dates | You write it | Done for you |
| PDF, DOCX and scan reading | You add it | Send the file |
| Data leaves your servers | No | Yes, to the API |
| Cost | GPU or CPU time, plus your engineering | Per call |
| Control over labels | Complete | The provider's schema |
Build it yourself when the CVs cannot leave your infrastructure, when you need a label the market does not offer (a clearance level, a licence number format specific to your country), or when the NER model is itself the research project.
Use an API when what you actually need is the record. ResumeJSON takes a PDF, DOCX, plain text or scanned CV and returns typed JSON with basics, work, education, skills, certifications and languages, where each work entry already carries its company, title, start and end dates and highlights. The quickstart shows the request, and the field reference lists every field.
When plain NER is enough
Not every job needs the full record. NER on its own is a good fit when:
- you only want contact details, to deduplicate applicants by email and phone;
- you are redacting names and contact details before a review, where finding every span matters and grouping does not (see blind resume screening);
- you are tagging skills for search, and a flat list of skills is exactly the output you want;
- you are counting how often employers or schools appear across a large archive.
In each of those, the flat list is the answer. The grouping problem only appears when a person or a program needs to read one candidate's history in order.
Try it on your own CV
The quickest way to see the difference between spans and a record is to look at both for the same resume. Run a model from the table above on your CV, then drop the same file into the free parser, no signup needed, and compare what you would have to write to get from the first output to the second.