ResumeJSON

Candidate database: how to build one from the CVs you already have

A candidate database is one record per person, with the same fields on every record, so you can search past applicants instead of starting each role from zero. Most teams already hold the raw material: a shared drive of CVs, an inbox of applications, the export of an old job board. Building the database means choosing a small set of fields, picking somewhere to store them, filling the records from those CVs, and stopping the same person from appearing three times. This article walks through each step, for a spreadsheet and for a real database, and says where the manual way is enough.

It is written for two readers: the recruiter who wants to stop re-reading the same folder, and the developer asked to turn that folder into something searchable. We build ResumeJSON, a parser that turns a CV into structured fields, so read the filling step as written by an interested party. Every other step works with whatever tools you already use.


What a candidate database is for

Before choosing a tool, write down the two or three questions the database has to answer. They decide the fields. Typical ones:

A database that tries to hold everything on the CV is slow to fill and hard to search. Keep the fields you will actually filter on, and keep the original file for everything else.

Step 1: choose the fields

Split the fields into two kinds: what the CV tells you, and what your team adds.

FieldComes fromWhy keep it
Full nameCVShown on every result
EmailCVThe main key for spotting duplicates
PhoneCVThe second key, and a way to reach them
LocationCVFiltering by city or region
Current titleCV, latest roleFiltering by role
Years of experienceCounted from the work historyFiltering by seniority
SkillsCV, skills sectionThe search most recruiters run first
LanguagesCVFiltering for language requirements
Profile linksCVLinkedIn or GitHub, for a quick second look
Source fileYour folderEvery record leads back to the document
StatusYour teamNew, contacted, interviewing, placed, not now
Last contactedYour teamWho is due a follow up
Consent untilYour teamWhen the record must be reviewed or deleted

The last three columns matter as much as the first ten. A database with no status turns into a list of names nobody remembers calling.

Step 2: choose where it lives

Three shapes cover almost every team. Pick by size and by how many people search it.

OptionGood up toSearchDuplicatesWho it suits
Excel or Google SheetsA few hundred rowsFilter and SEARCHBy eye, or conditional formattingOne recruiter, one desk
Airtable or a similar table toolA few thousand recordsFiltered views, linked tablesFormula field on emailA small team sharing one list
Postgres or another SQL databaseHundreds of thousandsFull text search, indexesA unique constraintA developer building it into a product

If you already pay for an ATS, check its candidate search first. Most applicant tracking systems keep every applicant, and the database you want may already exist under a different name. Build your own when the ATS cannot search the way you need, when you are importing an archive from somewhere else, or when the database is part of a product you sell, such as a job board.

Step 3: fill it from the CVs

This is the slow part. A CV is a document laid out for a human, so every record has to be read off it.

By hand

For a few dozen CVs, typing is fine. Open the CV beside the sheet and fill one row. Three habits keep the result searchable:

  1. Write dates as YYYY-MM so they sort correctly.
  2. Keep skills in one cell, separated by ; , so a single row stays a single person.
  3. Count experience from the dates of each role, rather than trusting "10+ years" in a headline.

With a resume parser

Past a few dozen CVs, the reading is the bottleneck, and the errors start: a digit missed in a phone number, a skill left out because it sat inside a job description. A resume parser reads each file and returns the fields as structured data, which you then write into the database.

ResumeJSON returns, for each CV, a basics block with full_name, email, phone, location, headline and links; a work array where each role has company, title, start_date, end_date and is_current; plus education, skills, certifications, languages and total_years_experience, which is computed from the work history rather than copied from the CV. That maps onto the table in Step 1 almost column for column. The full list is in the field reference.

For a spreadsheet, Resume to Excel has a complete script that turns a folder of CVs into one sheet, with a second tab listing every file that failed. For a SQL database, the same loop writes rows instead:

create table candidates (
  id              bigserial primary key,
  full_name       text,
  email           text,
  phone           text,
  location        text,
  current_title   text,
  years_exp       numeric,
  skills          text[],
  languages       text[],
  links           text[],
  source_file     text not null,
  status          text not null default 'new',
  last_contacted  date,
  consent_until   date,
  raw             jsonb not null
);

create unique index candidates_email_key on candidates (lower(email)) where email is not null;
create index candidates_skills_idx on candidates using gin (skills);

Keep the whole parse in the raw column. The columns are what you filter on today. The JSON is what lets you add a column next year without opening every file again.

Scans and odd files

A real archive is never all clean PDFs. ResumeJSON detects the type from the file's own bytes, so a PDF with no text layer, or a JPEG, PNG or WebP photo of a page, is read as an image. DOCX and plain text work too, and each upload can be up to 20 MB. A file that is not a CV at all, such as a cover letter, comes back as an error with the code not_a_resume, so it lands in your failure list instead of becoming an empty record. If most of your archive is scans, OCR resume parser covers what changes.

Step 4: keep duplicates out

The same person applies twice, sends an updated CV, or turns up in two imported archives. Without a rule, the database fills with near copies, and the newest notes end up on the wrong one.

A workable order of checks:

  1. Same email, ignoring case. The strongest signal. The unique index above enforces it.
  2. Same phone, after stripping spaces, dashes and brackets. Compare the last nine or ten digits, because some CVs carry a country code and some do not.
  3. Same name and same current company. Weaker. Flag these for a person to merge rather than merging them automatically.

When a match is found, update the record instead of adding a new one: keep your team's status and notes, replace the CV fields with the newer parse, and keep both source files. In SQL that is an insert ... on conflict (lower(email)) do update that leaves status, last_contacted and consent_until alone.

Step 5: make it searchable

Search is the reason the database exists, so test it with the questions from the start.

Skills are written many ways ("JS", "JavaScript", "Javascript ES6"). Search gets much better once you map them to one name each. Skills taxonomy covers the lists you can map to, and Candidate matching goes further, ranking records against a job description.

Step 6: keep it current and lawful

A candidate database holds personal data about real people, and it goes stale quickly.

ResumeJSON parses the document and keeps no copy of it, so the record in your database is the only one the parse produces. The data handling page has the detail.

When the manual way is enough

You do not need a parser, a script or a SQL table if:

In those cases a well kept spreadsheet with the columns from Step 1 is the right tool. The point to automate is when filling the records takes longer than searching them saves.

What it costs to fill

As of October 2026, our published plans on RapidAPI are:

PlanMonthly priceParses includedPast the quota
Basic$0100Hard cap, never bills
Pro$0None, pay per use$0.05 a parse
Ultra$291,000$0.045 a parse
Mega$995,000$0.018 a parse

An archive of up to 100 CVs fits the free Basic plan. A one-off import of 2,000 CVs costs $100 on pay per use, and a steady flow of new applications usually fits one of the monthly plans. Every call is metered whatever the answer, so remove obvious duplicates and non-CV attachments from the folder first. For archives in the thousands, Bulk resume parsing covers running the backlog in parallel without restarting from zero after a crash.

Start with one CV

Before you design the table, see what one of your own CVs turns into. The free resume parser takes a single file with no signup and shows the JSON a script would receive, so you can check that the fields you planned in Step 1 are the ones your documents actually fill. If the database is the first piece of a larger system, How to build an ATS picks up from here.

All articles