ResumeJSON

Extract information from a resume: every field, and how

To extract information from a resume, decide the fields first, then pick a method per field. Contact details (email, phone, links) come out reliably with patterns. Work history, education and skills need the document's structure, and that is where hand-written rules start to cost more than they save. This guide covers the whole record: the target shape, working Python for the pattern route, the CVs that break it, and how to get the same record from one API call.

We build ResumeJSON, a resume parsing API, so read the API section as written by an interested party. Everything before it needs nothing from us.

The information a resume actually holds

Most projects start with "get the name and email" and end up needing the whole record a month later. Pick the shape once:

GroupFieldsHow hard to extract
Contactfull name, email, phone, location, linksEmail and links are easy; name and phone are harder than they look
Work historycompany, title, start and end date, current or not, highlightsHard: every CV lays it out differently
Educationinstitution, degree, field of study, datesMedium: short entries, many formats
Skillsa flat list of skill namesMedium: scattered across the whole document
Certificationsname, issuer, dateMedium: often hidden under Education
Languageslanguage, proficiencyEasy when a section exists

Every field must be allowed to be empty. A CV with no phone number should give you null, never the first string of digits the code happened to find. A wrong value looks exactly like a right one in your database; an empty one tells you to look.

Step 1: get the text out

Every method starts with text. For a text-based PDF, pypdf is enough:

from pypdf import PdfReader

def pdf_text(path: str) -> str:
    return "\n".join(page.extract_text() or "" for page in PdfReader(path).pages)

For DOCX, python-docx reads the paragraphs. Two warnings before you build on that text. A two-column PDF often comes out with the columns interleaved line by line, so a job title can end up next to a skill from the sidebar. And a scanned CV has no text layer at all, so extract_text() returns an empty string that looks like an empty CV. OCR resume parser covers the scanned case.

Step 2: contact details with patterns

Email and links are the one part of a resume where patterns genuinely work, because their format is fixed by something other than the candidate.

import re

EMAIL = re.compile(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+")
URL = re.compile(r"(?:https?://|www\.)\S+|(?:linkedin\.com|github\.com)/\S+", re.I)
PHONE = re.compile(r"(?:\+\d{1,3}[\s.-]?)?(?:\(?\d{2,4}\)?[\s.-]?){2,4}\d{2,4}")

def contact(text: str) -> dict:
    email = EMAIL.search(text)
    phones = [p.strip() for p in PHONE.findall(text) if len(re.sub(r"\D", "", p)) >= 8]
    return {
        "email": email.group(0) if email else None,
        "phone": phones[0] if phones else None,
        "links": sorted(set(u.rstrip(".,;)") for u in URL.findall(text))),
    }

The phone pattern is the weak one. It also matches date ranges such as 2019 - 2023 2024, which is why the sketch demands at least eight digits and still gets it wrong now and then. If you only serve one country, a library such as phonenumbers with a default region is more reliable.

The name is the hardest contact field. The usual rule is "the first line that is two or three capitalised words", and it fails on CVs that open with a headline, a logo, "Curriculum Vitae", or a name written in capitals. Treat a name found this way as a guess and show it to a person before it goes anywhere important.

Step 3: split the document into sections

Everything else depends on knowing where each section starts. Find the headings, then read until the next one:

HEADINGS = {
    "work": r"(work|professional)?\s*experience|employment( history)?|career history",
    "education": r"education|academic background|qualifications",
    "skills": r"(technical |core )?skills|competencies|technologies",
    "languages": r"languages",
}

def sections(text: str) -> dict:
    found, current = {}, None
    for line in text.splitlines():
        clean = line.strip().rstrip(":").lower()
        name = next((k for k, p in HEADINGS.items() if re.fullmatch(p, clean)), None)
        if name:
            current = name
            found[name] = []
        elif current:
            found[current].append(line)
    return {k: "\n".join(v) for k, v in found.items()}

From there, each section gets its own extractor. Skills are usually a comma or bullet list, so splitting on ,, | and bullet characters gets you most of the way; extract skills from resume goes further, including skills mentioned only inside job descriptions. Education has its own walkthrough in extract education from a resume.

Step 4: work history, the hard part

A work entry is a company, a title, a date range and some bullets, in any order. Date ranges are the most reliable anchor, so most extractors split the section on them:

DATE_RANGE = re.compile(
    r"((?:jan|feb|mar|apr|may|jun|jul|aug|sep|oct|nov|dec)[a-z]*\.?\s+)?(\d{4})\s*[-to]+\s*"
    r"((?:jan|feb|mar|apr|may|jun|jul|aug|sep|oct|nov|dec)[a-z]*\.?\s+)?(\d{4}|present|current|now)",
    re.I,
)

def work_entries(section: str) -> list[dict]:
    entries, lines = [], section.splitlines()
    for i, line in enumerate(lines):
        m = DATE_RANGE.search(line)
        if m:
            header = " ".join(l.strip() for l in lines[max(0, i - 1): i + 1])
            entries.append({
                "header": DATE_RANGE.sub("", header).strip(" |,"),
                "start": m.group(2),
                "end": m.group(4),
                "is_current": m.group(4).lower() in ("present", "current", "now"),
            })
    return entries

Notice what this gives back: a header string with the company and the title still glued together. Separating them is where rules run out, because "Acme Ltd, Senior Engineer", "Senior Engineer at Acme" and a title on one line with the company on the next are all common, and a company name can look like a title ("Engineering Partners").

Where the pattern route breaks

These are the CVs that turn a weekend script into a maintenance job:

The dangerous failure is silence. None of these raise an error. They return a record with fewer entries, or the wrong entries, and it looks exactly like a thinner CV.

When the manual way is enough

Patterns are a fine answer when the input is under your control:

Extract information from a resume with an API

Once CVs come from strangers, in every layout and language, the section finder and the date rules become the project. A parsing API moves that work behind one call and returns the full record already split into fields.

Here is the same job against ResumeJSON, which returns typed JSON in about two seconds:

import requests

resp = requests.post(
    "https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse",
    headers={
        "x-rapidapi-key": "YOUR_KEY",
        "x-rapidapi-host": "resumejson-resume-cv-parser-api.p.rapidapi.com",
        "content-type": "application/json",
    },
    json={"text": text},          # or send the file: base64 or multipart
    timeout=30,
)
resume = resp.json()["resume"]
print(resume["basics"]["full_name"], resume["basics"]["email"], resume["basics"]["phone"])
for job in resume["work"]:
    print(job["company"], "|", job["title"], job["start_date"], "to", job["end_date"] or "now")

The response carries every group from the table at the top:

KeyWhat is in it
basicsfull_name, email, phone, location, headline, links
workcompany, title, start_date, end_date, is_current, location, highlights
educationinstitution, degree, field_of_study, start_date, end_date
skillsa flat list of names
certificationsname, issuer, date
languageslanguage, proficiency
total_years_experiencecomputed from work, never read off the page

Company and title arrive as separate fields. Dates come back as YYYY-MM, or YYYY when only a year is written, and a field the CV does not state is null rather than a guess. PDFs (two-column ones included), DOCX files and scanned or photographed CVs all go to the same endpoint. If you want the values in one language whatever the CV was written in, add ?output_language=English.

The honest gaps. The API needs a network connection, so it will not run offline. Skills come back as the CV writes them, so mapping them onto a fixed list is still your job; skills taxonomy covers that step.

The two routes, side by side

Patterns and section rulesResumeJSON
SetupYour code, your heading and date listsOne HTTP call
Contact detailsGood for email and links, weak for name and phoneAll fields returned
Company and titleGlued together unless you add rulesSeparate fields
Two-column and scanned CVsNeed extra toolingSame endpoint
Unknown headingsSilent empty sectionRead from the content
Cost at 100 CVs/monthYour time$0 on the Basic plan, a hard cap
Cost at 1,000 CVs/monthYour time$29/month, or $0.05 a parse pay-per-use
Runs offlineYesNo

Prices are the plans listed on our docs page.

What to do with the extracted record

With the information as data, the next steps are the ones you were extracting it for. Resume to Excel turns a folder of CVs into one spreadsheet. Candidate matching holds the record against a job, and automated resume screening builds rules on top of it. For a backlog of thousands, bulk resume parsing covers concurrency and retries.

How to choose

You can try the whole record on your own CV with the free parser, no signup needed.

All articles