Convert resume to JSON: three ways, with the code and the output
Published
To convert a resume to JSON, you extract the text from the file, then decide which part of that text is the name, which is a job, which is a degree, and write each into a named field. The first step is a solved problem. The second is the whole difficulty. This article walks through three routes: typing it by hand, a short script with regular expressions, and a parsing API that returns the fields for you. It shows the JSON each one produces and says when each is the right call.
We build ResumeJSON, a resume parsing API, so the third route is ours and you should read that section as written by an interested party. The first two need nothing from us, and for some jobs they are the better choice.
What "resume as JSON" has to mean
A resume converted to JSON is only useful if the shape is predictable. Before choosing a route, fix the target. A workable minimum looks like this:
{
"basics": {
"full_name": "Sarah Okonkwo",
"email": "sarah@example.com",
"phone": "+44 7700 900123",
"location": "Leeds, UK",
"headline": "Staff Engineer"
},
"work": [
{
"company": "Example Ltd",
"title": "Staff Engineer",
"start_date": "2021-04",
"end_date": null,
"is_current": true,
"highlights": ["Led the move to event-driven billing"]
}
],
"education": [
{ "institution": "University of Leeds", "degree": "BSc", "field_of_study": "Computer Science", "start_date": "2010", "end_date": "2013" }
],
"skills": ["TypeScript", "PostgreSQL", "Kubernetes"],
"total_years_experience": 12.5
}Three decisions matter more than the field names:
- Dates in one format.
2021-04,2021ornull. A CV writes the same month as "April 2021", "04/21" and "Apr '21", and a database that stores all three cannot sort. - Missing means
null. If the CV does not state a phone number, the field isnull, not an empty string and not a guess. - Arrays for repeating things. Jobs, schools and skills are lists. A flat object with
job1_titleandjob2_titlebreaks on the first person with seven roles.
If you want the community standard rather than your own shape, JSON Resume schema: a complete example covers it section by section, including how to map parsed output into it.
Route 1: by hand, when it is enough
For one CV, or ten, open the file beside an editor and type. It is slower than it sounds only after about twenty files.
- Open the document and copy the text out. For a PDF, select all and paste into a plain text editor. A two-column layout may paste out of order, so read the result once.
- Write the object. Start from the template above and fill it in. Keep the dates in
YYYY-MM. - Validate the JSON. Paste it into any JSON validator or run
python -m json.tool file.json. A trailing comma is the usual culprit.
By hand is enough when the number of resumes is small and the job is a one-off, such as filling a sample file for a test or migrating a single profile. It stops being enough when somebody else will keep sending you CVs.
Route 2: a script with text extraction and regexes
A short Python script can get the easy fields. Extract the text, then match the patterns that look the same on every resume.
import re
import json
import sys
import subprocess
def pdf_text(path):
# pdftotext ships with poppler; "-" writes to stdout
return subprocess.run(["pdftotext", "-layout", path, "-"],
capture_output=True, text=True, check=True).stdout
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
PHONE = re.compile(r"\+?\d[\d\s().-]{7,}\d")
text = pdf_text(sys.argv[1])
first_line = next((l.strip() for l in text.splitlines() if l.strip()), None)
out = {
"basics": {
"full_name": first_line,
"email": (EMAIL.search(text) or [None])[0],
"phone": (PHONE.search(text) or [None])[0],
}
}
print(json.dumps(out, indent=2))This works on a good share of clean, single-column PDFs for the contact block, and it is free and runs anywhere. Be honest with yourself about where it ends:
| Field | Regex approach | What goes wrong |
|---|---|---|
| Reliable | Several addresses on one CV; you take the first | |
| Phone | Fair | A date range like "2019-2023" can match the pattern |
| Name | Guess: the first line | A headline, a logo text or a "CURRICULUM VITAE" banner comes first |
| Jobs and dates | Poor | Every CV orders company, title and dates differently |
| Education | Poor | "BSc", "B.Sc.", "Bachelor of Science" and a bare university name |
| Skills | Poor | A paragraph in one CV, a table in another, a tag cloud in a third |
The regex route is a good prototype and a poor product. The moment you need the work history, you are writing a parser, and the rules multiply with every layout you meet. Scanned CVs make it worse, because there is no text to extract until you run OCR, which adds a second tool to maintain. If your input is mostly photographs and scans, OCR resume parser covers that case on its own.
For the libraries you can assemble yourself in Python, Python resume parser compares what exists and what each one returns.
Route 3: a parsing API
A parsing API takes the file and returns the object. You send one resume per request and get typed JSON back on the same request. With ResumeJSON, you post the file as multipart, as raw bytes, or as a JSON body when you already have the text:
curl -X POST 'https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse' \
-H 'x-rapidapi-key: YOUR_KEY' \
-H 'x-rapidapi-host: resumejson-resume-cv-parser-api.p.rapidapi.com' \
-F 'file=@cv.pdf'The response carries a resume object with basics, a work array, an education array, skills, certifications, languages and total_years_experience, plus a meta object. The full list of fields is in the field reference. A few properties are worth knowing before you build on it, all as of October 2026 and read from our own documentation:
- PDF, DOCX and plain text are accepted, and so are images. The file type is detected from the bytes, so an upload labelled as a generic binary still works. A PDF with no text layer, or a JPEG, PNG or WebP of the page, is read as images instead, and
meta.readtells you which path was taken. - Dates are normalised. Every date is
YYYY-MM,YYYYornull, so sorting bystart_dateis a string comparison. - A field the CV does not state is
null. It is never an empty string. total_years_experienceis computed from the work history, not read off the document. Overlapping roles are counted once.- The output language can be set. By default the values come back in the CV's own language; add
output_language=Frenchto the query string and the job titles, skills and similar text values are written in French, while names, emails, phone numbers, URLs and dates are left alone. - Limits. Uploads go up to 20 MB, one resume per request, and the response is synchronous. There is no batch endpoint.
The documentation quotes about 2.2 seconds for a typical parse. A scan takes longer, because it goes through the image path.
The same job in Python
import requests
URL = "https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse"
HEADERS = {
"x-rapidapi-key": "YOUR_KEY",
"x-rapidapi-host": "resumejson-resume-cv-parser-api.p.rapidapi.com",
}
with open("cv.pdf", "rb") as f:
resp = requests.post(URL, headers=HEADERS, files={"file": f}, timeout=60)
resp.raise_for_status()
resume = resp.json()["resume"]
print(resume["basics"]["full_name"], resume["total_years_experience"])Check the status before you read the body. A file that is not a resume or cannot be read comes back as a 422 with an error.code, and a transient upstream problem as a 503 with a Retry-After header. Retry the 503 after the delay; do not retry the 4xx, because the same file will fail the same way. Bulk resume parsing covers the retry and concurrency side if you are converting a backlog.
What each route costs
| Route | Money | Your time | Output you can trust |
|---|---|---|---|
| By hand | Free | Minutes per CV | As good as your care |
| Regexes | Free | Hours to start, then ongoing | Contact block only |
| Parsing API | See the plans below | Minutes to integrate | Every field, one fixed shape |
As of October 2026, the plans published on the docs page are:
| Plan | Monthly price | Parses included | Past the quota |
|---|---|---|---|
| Basic | $0 | 100 | Hard cap, never bills |
| Pro | $0 | None, pay per use | $0.05 a parse |
| Ultra | $29 | 1,000 | $0.045 a parse |
| Mega | $99 | 5,000 | $0.018 a parse |
Every call the service receives is metered, whatever the answer, so a refused file still counts. Drop duplicates and non-CV attachments before you run a batch.
Which route to pick
- Fewer than about twenty CVs, once: do it by hand. Nothing beats it for effort.
- A prototype that only needs name, email and phone: the regex script is enough, and you can ship it today.
- Work history, education, skills, scans, or a steady stream of uploads: use a parser, ours or anyone's. The point is that the shape is fixed and the dates are comparable, so the code that reads the JSON stays simple.
- You need the exact JSON Resume standard: parse first, then map one object to another, as shown in the JSON Resume article.
Try one file first
Before you write any code, see what a real CV of yours turns into. The free resume parser takes one file with no signup and shows the JSON the API would return, so you can check that the fields you planned for are the ones the document actually fills. If the JSON is going into a bigger system, How to build an ATS and Candidate database pick up from there.