Extract information from a resume: every field, and how
Published
To extract information from a resume, decide the fields first, then pick a method per field. Contact details (email, phone, links) come out reliably with patterns. Work history, education and skills need the document's structure, and that is where hand-written rules start to cost more than they save. This guide covers the whole record: the target shape, working Python for the pattern route, the CVs that break it, and how to get the same record from one API call.
We build ResumeJSON, a resume parsing API, so read the API section as written by an interested party. Everything before it needs nothing from us.
The information a resume actually holds
Most projects start with "get the name and email" and end up needing the whole record a month later. Pick the shape once:
| Group | Fields | How hard to extract |
|---|---|---|
| Contact | full name, email, phone, location, links | Email and links are easy; name and phone are harder than they look |
| Work history | company, title, start and end date, current or not, highlights | Hard: every CV lays it out differently |
| Education | institution, degree, field of study, dates | Medium: short entries, many formats |
| Skills | a flat list of skill names | Medium: scattered across the whole document |
| Certifications | name, issuer, date | Medium: often hidden under Education |
| Languages | language, proficiency | Easy when a section exists |
Every field must be allowed to be empty. A CV with no phone number should give you null, never the first string of digits the code happened to find. A wrong value looks exactly like a right one in your database; an empty one tells you to look.
Step 1: get the text out
Every method starts with text. For a text-based PDF, pypdf is enough:
from pypdf import PdfReader
def pdf_text(path: str) -> str:
return "\n".join(page.extract_text() or "" for page in PdfReader(path).pages)For DOCX, python-docx reads the paragraphs. Two warnings before you build on that text. A two-column PDF often comes out with the columns interleaved line by line, so a job title can end up next to a skill from the sidebar. And a scanned CV has no text layer at all, so extract_text() returns an empty string that looks like an empty CV. OCR resume parser covers the scanned case.
Step 2: contact details with patterns
Email and links are the one part of a resume where patterns genuinely work, because their format is fixed by something other than the candidate.
import re
EMAIL = re.compile(r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+")
URL = re.compile(r"(?:https?://|www\.)\S+|(?:linkedin\.com|github\.com)/\S+", re.I)
PHONE = re.compile(r"(?:\+\d{1,3}[\s.-]?)?(?:\(?\d{2,4}\)?[\s.-]?){2,4}\d{2,4}")
def contact(text: str) -> dict:
email = EMAIL.search(text)
phones = [p.strip() for p in PHONE.findall(text) if len(re.sub(r"\D", "", p)) >= 8]
return {
"email": email.group(0) if email else None,
"phone": phones[0] if phones else None,
"links": sorted(set(u.rstrip(".,;)") for u in URL.findall(text))),
}The phone pattern is the weak one. It also matches date ranges such as 2019 - 2023 2024, which is why the sketch demands at least eight digits and still gets it wrong now and then. If you only serve one country, a library such as phonenumbers with a default region is more reliable.
The name is the hardest contact field. The usual rule is "the first line that is two or three capitalised words", and it fails on CVs that open with a headline, a logo, "Curriculum Vitae", or a name written in capitals. Treat a name found this way as a guess and show it to a person before it goes anywhere important.
Step 3: split the document into sections
Everything else depends on knowing where each section starts. Find the headings, then read until the next one:
HEADINGS = {
"work": r"(work|professional)?\s*experience|employment( history)?|career history",
"education": r"education|academic background|qualifications",
"skills": r"(technical |core )?skills|competencies|technologies",
"languages": r"languages",
}
def sections(text: str) -> dict:
found, current = {}, None
for line in text.splitlines():
clean = line.strip().rstrip(":").lower()
name = next((k for k, p in HEADINGS.items() if re.fullmatch(p, clean)), None)
if name:
current = name
found[name] = []
elif current:
found[current].append(line)
return {k: "\n".join(v) for k, v in found.items()}From there, each section gets its own extractor. Skills are usually a comma or bullet list, so splitting on ,, | and bullet characters gets you most of the way; extract skills from resume goes further, including skills mentioned only inside job descriptions. Education has its own walkthrough in extract education from a resume.
Step 4: work history, the hard part
A work entry is a company, a title, a date range and some bullets, in any order. Date ranges are the most reliable anchor, so most extractors split the section on them:
DATE_RANGE = re.compile(
r"((?:jan|feb|mar|apr|may|jun|jul|aug|sep|oct|nov|dec)[a-z]*\.?\s+)?(\d{4})\s*[-to]+\s*"
r"((?:jan|feb|mar|apr|may|jun|jul|aug|sep|oct|nov|dec)[a-z]*\.?\s+)?(\d{4}|present|current|now)",
re.I,
)
def work_entries(section: str) -> list[dict]:
entries, lines = [], section.splitlines()
for i, line in enumerate(lines):
m = DATE_RANGE.search(line)
if m:
header = " ".join(l.strip() for l in lines[max(0, i - 1): i + 1])
entries.append({
"header": DATE_RANGE.sub("", header).strip(" |,"),
"start": m.group(2),
"end": m.group(4),
"is_current": m.group(4).lower() in ("present", "current", "now"),
})
return entriesNotice what this gives back: a header string with the company and the title still glued together. Separating them is where rules run out, because "Acme Ltd, Senior Engineer", "Senior Engineer at Acme" and a title on one line with the company on the next are all common, and a company name can look like a title ("Engineering Partners").
Where the pattern route breaks
These are the CVs that turn a weekend script into a maintenance job:
- Two-column layouts. The text arrives interleaved, so sections bleed into each other before your heading finder ever sees them.
- Headings nobody put on your list. "Where I've worked", "Background", or a heading in another language gives you an empty section, which reads as "no experience".
- Dates in every format.
03/2019,2019.03,Mar '19,Since 2021, and a Spanish or German month name each need a rule. - Several roles at one company. Two titles under one employer look like one entry with two date ranges.
- Certifications under Education, and projects under Experience, which inflate both.
- Scans and photos, which have no text at all until something reads the pixels.
The dangerous failure is silence. None of these raise an error. They return a record with fewer entries, or the wrong entries, and it looks exactly like a thinner CV.
When the manual way is enough
Patterns are a fine answer when the input is under your control:
- One template. Your own application form exported to PDF always has the same headings in the same order.
- One or two fields. If all you need is the email and the LinkedIn link, the Step 2 code is the whole project.
- A person checks every record. Pre-filling a form a recruiter reviews tolerates misses that an automated screen would not.
Extract information from a resume with an API
Once CVs come from strangers, in every layout and language, the section finder and the date rules become the project. A parsing API moves that work behind one call and returns the full record already split into fields.
Here is the same job against ResumeJSON, which returns typed JSON in about two seconds:
import requests
resp = requests.post(
"https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse",
headers={
"x-rapidapi-key": "YOUR_KEY",
"x-rapidapi-host": "resumejson-resume-cv-parser-api.p.rapidapi.com",
"content-type": "application/json",
},
json={"text": text}, # or send the file: base64 or multipart
timeout=30,
)
resume = resp.json()["resume"]
print(resume["basics"]["full_name"], resume["basics"]["email"], resume["basics"]["phone"])
for job in resume["work"]:
print(job["company"], "|", job["title"], job["start_date"], "to", job["end_date"] or "now")The response carries every group from the table at the top:
| Key | What is in it |
|---|---|
basics | full_name, email, phone, location, headline, links |
work | company, title, start_date, end_date, is_current, location, highlights |
education | institution, degree, field_of_study, start_date, end_date |
skills | a flat list of names |
certifications | name, issuer, date |
languages | language, proficiency |
total_years_experience | computed from work, never read off the page |
Company and title arrive as separate fields. Dates come back as YYYY-MM, or YYYY when only a year is written, and a field the CV does not state is null rather than a guess. PDFs (two-column ones included), DOCX files and scanned or photographed CVs all go to the same endpoint. If you want the values in one language whatever the CV was written in, add ?output_language=English.
The honest gaps. The API needs a network connection, so it will not run offline. Skills come back as the CV writes them, so mapping them onto a fixed list is still your job; skills taxonomy covers that step.
The two routes, side by side
| Patterns and section rules | ResumeJSON | |
|---|---|---|
| Setup | Your code, your heading and date lists | One HTTP call |
| Contact details | Good for email and links, weak for name and phone | All fields returned |
| Company and title | Glued together unless you add rules | Separate fields |
| Two-column and scanned CVs | Need extra tooling | Same endpoint |
| Unknown headings | Silent empty section | Read from the content |
| Cost at 100 CVs/month | Your time | $0 on the Basic plan, a hard cap |
| Cost at 1,000 CVs/month | Your time | $29/month, or $0.05 a parse pay-per-use |
| Runs offline | Yes | No |
Prices are the plans listed on our docs page.
What to do with the extracted record
With the information as data, the next steps are the ones you were extracting it for. Resume to Excel turns a folder of CVs into one spreadsheet. Candidate matching holds the record against a job, and automated resume screening builds rules on top of it. For a backlog of thousands, bulk resume parsing covers concurrency and retries.
How to choose
- One template, or only email and links → the pattern code above is enough.
- Many CV layouts, and you need the whole record → a parsing API, then your own rules on the fields it returns.
- Offline or on-premise only → the pattern route, plus an open source resume parser for the harder sections.
You can try the whole record on your own CV with the free parser, no signup needed.