ResumeJSON

Convert resume to JSON: three ways, with the code and the output

To convert a resume to JSON, you extract the text from the file, then decide which part of that text is the name, which is a job, which is a degree, and write each into a named field. The first step is a solved problem. The second is the whole difficulty. This article walks through three routes: typing it by hand, a short script with regular expressions, and a parsing API that returns the fields for you. It shows the JSON each one produces and says when each is the right call.

We build ResumeJSON, a resume parsing API, so the third route is ours and you should read that section as written by an interested party. The first two need nothing from us, and for some jobs they are the better choice.


What "resume as JSON" has to mean

A resume converted to JSON is only useful if the shape is predictable. Before choosing a route, fix the target. A workable minimum looks like this:

{
  "basics": {
    "full_name": "Sarah Okonkwo",
    "email": "sarah@example.com",
    "phone": "+44 7700 900123",
    "location": "Leeds, UK",
    "headline": "Staff Engineer"
  },
  "work": [
    {
      "company": "Example Ltd",
      "title": "Staff Engineer",
      "start_date": "2021-04",
      "end_date": null,
      "is_current": true,
      "highlights": ["Led the move to event-driven billing"]
    }
  ],
  "education": [
    { "institution": "University of Leeds", "degree": "BSc", "field_of_study": "Computer Science", "start_date": "2010", "end_date": "2013" }
  ],
  "skills": ["TypeScript", "PostgreSQL", "Kubernetes"],
  "total_years_experience": 12.5
}

Three decisions matter more than the field names:

If you want the community standard rather than your own shape, JSON Resume schema: a complete example covers it section by section, including how to map parsed output into it.

Route 1: by hand, when it is enough

For one CV, or ten, open the file beside an editor and type. It is slower than it sounds only after about twenty files.

  1. Open the document and copy the text out. For a PDF, select all and paste into a plain text editor. A two-column layout may paste out of order, so read the result once.
  2. Write the object. Start from the template above and fill it in. Keep the dates in YYYY-MM.
  3. Validate the JSON. Paste it into any JSON validator or run python -m json.tool file.json. A trailing comma is the usual culprit.

By hand is enough when the number of resumes is small and the job is a one-off, such as filling a sample file for a test or migrating a single profile. It stops being enough when somebody else will keep sending you CVs.

Route 2: a script with text extraction and regexes

A short Python script can get the easy fields. Extract the text, then match the patterns that look the same on every resume.

import re
import json
import sys
import subprocess

def pdf_text(path):
    # pdftotext ships with poppler; "-" writes to stdout
    return subprocess.run(["pdftotext", "-layout", path, "-"],
                          capture_output=True, text=True, check=True).stdout

EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
PHONE = re.compile(r"\+?\d[\d\s().-]{7,}\d")

text = pdf_text(sys.argv[1])
first_line = next((l.strip() for l in text.splitlines() if l.strip()), None)

out = {
    "basics": {
        "full_name": first_line,
        "email": (EMAIL.search(text) or [None])[0],
        "phone": (PHONE.search(text) or [None])[0],
    }
}
print(json.dumps(out, indent=2))

This works on a good share of clean, single-column PDFs for the contact block, and it is free and runs anywhere. Be honest with yourself about where it ends:

FieldRegex approachWhat goes wrong
EmailReliableSeveral addresses on one CV; you take the first
PhoneFairA date range like "2019-2023" can match the pattern
NameGuess: the first lineA headline, a logo text or a "CURRICULUM VITAE" banner comes first
Jobs and datesPoorEvery CV orders company, title and dates differently
EducationPoor"BSc", "B.Sc.", "Bachelor of Science" and a bare university name
SkillsPoorA paragraph in one CV, a table in another, a tag cloud in a third

The regex route is a good prototype and a poor product. The moment you need the work history, you are writing a parser, and the rules multiply with every layout you meet. Scanned CVs make it worse, because there is no text to extract until you run OCR, which adds a second tool to maintain. If your input is mostly photographs and scans, OCR resume parser covers that case on its own.

For the libraries you can assemble yourself in Python, Python resume parser compares what exists and what each one returns.

Route 3: a parsing API

A parsing API takes the file and returns the object. You send one resume per request and get typed JSON back on the same request. With ResumeJSON, you post the file as multipart, as raw bytes, or as a JSON body when you already have the text:

curl -X POST 'https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse' \
  -H 'x-rapidapi-key: YOUR_KEY' \
  -H 'x-rapidapi-host: resumejson-resume-cv-parser-api.p.rapidapi.com' \
  -F 'file=@cv.pdf'

The response carries a resume object with basics, a work array, an education array, skills, certifications, languages and total_years_experience, plus a meta object. The full list of fields is in the field reference. A few properties are worth knowing before you build on it, all as of October 2026 and read from our own documentation:

The documentation quotes about 2.2 seconds for a typical parse. A scan takes longer, because it goes through the image path.

The same job in Python

import requests

URL = "https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse"
HEADERS = {
    "x-rapidapi-key": "YOUR_KEY",
    "x-rapidapi-host": "resumejson-resume-cv-parser-api.p.rapidapi.com",
}

with open("cv.pdf", "rb") as f:
    resp = requests.post(URL, headers=HEADERS, files={"file": f}, timeout=60)

resp.raise_for_status()
resume = resp.json()["resume"]
print(resume["basics"]["full_name"], resume["total_years_experience"])

Check the status before you read the body. A file that is not a resume or cannot be read comes back as a 422 with an error.code, and a transient upstream problem as a 503 with a Retry-After header. Retry the 503 after the delay; do not retry the 4xx, because the same file will fail the same way. Bulk resume parsing covers the retry and concurrency side if you are converting a backlog.

What each route costs

RouteMoneyYour timeOutput you can trust
By handFreeMinutes per CVAs good as your care
RegexesFreeHours to start, then ongoingContact block only
Parsing APISee the plans belowMinutes to integrateEvery field, one fixed shape

As of October 2026, the plans published on the docs page are:

PlanMonthly priceParses includedPast the quota
Basic$0100Hard cap, never bills
Pro$0None, pay per use$0.05 a parse
Ultra$291,000$0.045 a parse
Mega$995,000$0.018 a parse

Every call the service receives is metered, whatever the answer, so a refused file still counts. Drop duplicates and non-CV attachments before you run a batch.

Which route to pick

Try one file first

Before you write any code, see what a real CV of yours turns into. The free resume parser takes one file with no signup and shows the JSON the API would return, so you can check that the fields you planned for are the ones the document actually fills. If the JSON is going into a bigger system, How to build an ATS and Candidate database pick up from there.

All articles