ResumeJSON

Resume classification: sort CVs by category, with code

Resume classification: sorting CVs into categories you can use

Resume classification means assigning each CV to one or more categories from a fixed list: a job family such as "Data Science" or "Sales", a seniority band, or the team that should see it. There are three practical ways to do it. You can write rules over parsed fields (job titles, skills, years of experience), train a text classifier on labelled CVs, or ask a language model to pick from your list. For most products the right start is rules over clean fields, with a trained model added once the categories outgrow what rules can keep up with.

This article is for the developer adding resume classification to a job board, an applicant tracking system or an internal hiring tool, and for anyone building a classifier as a data project. It covers choosing labels, the three approaches with code, how to check the result, and when sorting by hand is still the better answer. We build ResumeJSON, a resume parsing API, so read the parsing parts as written by an interested party. The classification logic is yours to own whichever parser you use.

What resume classification is for

Classification answers "what kind of candidate is this?" before anyone asks a more specific question. In a real product it usually does one of these jobs:

Classification is a different job from matching. It puts a CV in a bucket; it says nothing about whether the person fits one particular role. For that, see candidate matching and automated resume screening.

Decide the labels before the method

The label list matters more than the algorithm. Two decisions come first.

Is it one label or several? A data engineer who also managed a team fits two families. If your product routes CVs, one label per CV is simpler. If it tags a pool for search, allow several and store them as a list.

Which axis are you classifying on? Job family, seniority and industry are three separate questions. Mixing them into one list ("Senior Sales", "Junior Sales", "Healthcare Sales") multiplies the categories and starves each one of examples. Classify each axis on its own.

AxisExample labelsStrongest signal in a CV
Job familySoftware engineering, Data, Sales, Finance, DesignRecent job titles, then skills
SeniorityEntry, Mid, Senior, LeadYears of experience, title words such as "Senior" or "Head of"
IndustryHealthcare, Retail, FintechEmployer names and descriptions in work history
LanguageEnglish, German, SpanishLanguages section, then the language the CV is written in

Always keep an "Other" label, and treat it as a signal. A growing "Other" bucket tells you the list is missing a category. If your families should line up with an established list of occupations or skills, skills taxonomy compares the public options.

Approach 1: rules over parsed fields

Rules are the simplest route to a classifier you can explain. They work well because the strongest signal, the job title, is short and repetitive: "Backend Developer", "Software Engineer II" and "Senior Python Engineer" all point at the same family.

The steps:

  1. Parse every CV into the same fields. You need each job's title and dates, the skills list and a total of years worked. Raw text makes rules fragile, because a title and an employer name look the same to a regular expression.
  2. Normalise titles. Lowercase, strip seniority words and numbering ("II", "Sr."), collapse punctuation.
  3. Map title patterns to families, checking the most recent job first.
  4. Fall back on skills when no title matches.
  5. Derive seniority separately, from years of experience and title words.
  6. Return "Other" with the reason when nothing matches.

Here it is in Python, reading a parsed CV as JSON. The field names are ResumeJSON's (work[].title, skills, total_years_experience), but any parser that returns per-job titles fits the same shape.

import re

FAMILIES = {
    "software_engineering": [r"\b(software|backend|frontend|full ?stack|web|mobile)\b.*\b(engineer|developer)\b", r"\bdevops\b", r"\bsre\b"],
    "data": [r"\bdata (scientist|engineer|analyst)\b", r"\bmachine learning\b", r"\bml engineer\b"],
    "sales": [r"\b(account executive|sales|business development)\b"],
    "design": [r"\b(ux|ui|product|graphic) designer\b"],
}
SKILL_HINTS = {
    "data": {"pandas", "sql", "spark", "tableau", "scikit-learn"},
    "software_engineering": {"java", "go", "react", "kubernetes", "node.js"},
}

def normalise(title):
    title = title.lower()
    title = re.sub(r"\b(senior|sr\.?|junior|jr\.?|lead|principal|ii|iii)\b", " ", title)
    return re.sub(r"[^a-z0-9 .]+", " ", title).strip()

def classify_family(resume):
    for job in resume["work"]:  # most recent first, as the parser returns it
        title = normalise(job.get("title") or "")
        for family, patterns in FAMILIES.items():
            if any(re.search(p, title) for p in patterns):
                return family, f"title: {job['title']}"
    skills = {s.lower() for s in resume["skills"]}
    best = max(SKILL_HINTS, key=lambda f: len(SKILL_HINTS[f] & skills))
    if SKILL_HINTS[best] & skills:
        return best, f"skills: {sorted(SKILL_HINTS[best] & skills)}"
    return "other", "no title or skill matched"

def classify_seniority(resume):
    years = resume.get("total_years_experience") or 0
    latest = (resume["work"][0].get("title") or "").lower() if resume["work"] else ""
    if re.search(r"\b(head of|director|vp|principal|lead)\b", latest):
        return "lead"
    if years >= 6 or "senior" in latest:
        return "senior"
    return "mid" if years >= 2 else "entry"

Every answer carries its reason, so a recruiter who disagrees can see why and you can fix the rule. That reason is the main thing rules give you that a model does not. The cost is upkeep: each new family and each odd title format is another pattern, and somewhere past twenty or thirty families the list becomes hard to maintain.

Approach 2: a trained text classifier

When you have labelled examples, a classic text classifier is cheap to train and runs in milliseconds. TF-IDF features with a linear model is the standard baseline, and it is hard to beat by much on short, repetitive text like job titles and skills.

The useful trick is to train on parsed fields instead of the whole document. Joining titles, skills and highlights gives the model the evidence and leaves out names, addresses, phone numbers and layout noise, which carry no information about the job family and can carry information you should not be classifying on.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.metrics import classification_report

def to_text(resume):
    titles = " ".join(j.get("title") or "" for j in resume["work"])
    highlights = " ".join(h for j in resume["work"] for h in j["highlights"])
    return f"{titles} {titles} {' '.join(resume['skills'])} {highlights}"  # titles weighted twice

X = [to_text(r) for r in resumes]  # parsed CVs
y = labels                         # one family per CV

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
model = make_pipeline(TfidfVectorizer(ngram_range=(1, 2), min_df=2), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train)
print(classification_report(y_test, model.predict(X_test)))

The hard part is the labels. Public datasets exist, mostly scraped from resume sites, and they are useful for learning the method. As of October 2026, the ResumeAtlas paper on arXiv describes a curated set of 13,389 resumes across 43 categories and notes that earlier datasets typically ran from 5 to 25 categories. Its abstract reports:

our best model achieving a top-1 accuracy of 92% and a top-5 accuracy of 97.5%

Treat figures like that as what is possible on that dataset's categories. Your own label list and your own applicants will score differently, so the only number that counts is the one you measure on a sample of your CVs, labelled by the people who will use the result.

Approach 3: a language model or embeddings

Two ways to classify with no training at all:

Both are quick to build. The trade-offs show up later: each CV costs a model call, the answer can change when the provider updates the model, and the explanation is whatever the model says about itself. A reasonable split is to let rules handle the confident cases and send only the "Other" bucket to a model, which keeps the model's share small and visible.

The three approaches compared

Rules over parsed fieldsTrained classifierLanguage model or embeddings
Needs labelled CVsNoYes, hundreds per categoryNo
Explains each answerYes, the matched rulePartly, through feature weightsOnly in its own words
Adding a categoryNew patternsRelabel and retrainEdit the list or a description
Running cost per CVNegligibleNegligibleOne model call
Stable over timeYesYes, until retrainedCan shift with the provider's model
Best forA short, stable list of familiesMany categories with labelled historyLong-tail titles, or the "Other" bucket

Checking that the classifier works

Whichever approach you pick, test it the same way:

  1. Hold out a labelled sample that nobody looked at while writing rules or training. Two hundred CVs labelled by recruiters beats a public dataset whose categories are not yours.
  2. Read precision and recall per category, not one overall accuracy figure. A classifier can look excellent overall while sending every designer to "Other".
  3. Look at the confusion matrix. Data analysts and data engineers swapping places is a label problem; designers landing in sales is a feature problem.
  4. Review the "Other" bucket every few weeks. It is where new categories show up first.
  5. Re-run the test after any change to rules, labels or model, and keep the scores beside the change.

Fairness and what to leave out

Classify on job evidence: titles, skills, work history and qualifications. Remove names, photos, addresses, dates of birth and anything else that can stand in for a protected characteristic before the text reaches a trained model, because a model will learn from any pattern in the data, including ones you would never write as a rule. Parsing first makes this easy, since you choose which fields go in. Blind resume screening covers the redaction step in detail.

Keep classification as a routing aid. A category decides which queue a CV lands in; a person still decides what happens to the candidate.

When the manual way is enough

Sorting by hand is the right call more often than the tutorials suggest:

If you are sorting a backlog by hand, putting the CVs into one sheet first makes it much faster. Resume to Excel shows how to turn a folder of CVs into rows you can sort and filter.

Where ResumeJSON fits

ResumeJSON does the parsing step and nothing else. It does not classify, rank or route anyone, because the label list and the rules belong in your product, where you can read and test them. It returns every CV, PDF or DOCX, in one typed schema in about two seconds: per-job titles and dates, skills, education, certifications, languages and a computed total_years_experience, which are the fields every approach above reads.

As of October 2026 the API is sold through RapidAPI. The Basic plan is $0 for 100 parses a month with a hard cap, Pro has no monthly fee and bills $0.05 a parse, Ultra is $29 a month for 1,000 parses and Mega is $99 a month for 5,000. The docs carry the current table and the full field reference, and the free parser shows the fields for a real CV before you write any code. If you have a backlog to classify, bulk resume parsing covers running it.

Resume classification works best in this order: fix the label list, parse every CV into the same fields, start with rules you can explain, and bring in a trained model or a language model only where the rules run out.

All articles