Candidate database: how to build one from the CVs you already have
Published
A candidate database is one record per person, with the same fields on every record, so you can search past applicants instead of starting each role from zero. Most teams already hold the raw material: a shared drive of CVs, an inbox of applications, the export of an old job board. Building the database means choosing a small set of fields, picking somewhere to store them, filling the records from those CVs, and stopping the same person from appearing three times. This article walks through each step, for a spreadsheet and for a real database, and says where the manual way is enough.
It is written for two readers: the recruiter who wants to stop re-reading the same folder, and the developer asked to turn that folder into something searchable. We build ResumeJSON, a parser that turns a CV into structured fields, so read the filling step as written by an interested party. Every other step works with whatever tools you already use.
What a candidate database is for
Before choosing a tool, write down the two or three questions the database has to answer. They decide the fields. Typical ones:
- "Who have we already seen for this kind of role?" Needs title, skills and seniority.
- "Who is in this city, or open to remote?" Needs location.
- "When did we last talk to them, and what happened?" Needs a status and a date that you write, since no CV contains them.
A database that tries to hold everything on the CV is slow to fill and hard to search. Keep the fields you will actually filter on, and keep the original file for everything else.
Step 1: choose the fields
Split the fields into two kinds: what the CV tells you, and what your team adds.
| Field | Comes from | Why keep it |
|---|---|---|
| Full name | CV | Shown on every result |
| CV | The main key for spotting duplicates | |
| Phone | CV | The second key, and a way to reach them |
| Location | CV | Filtering by city or region |
| Current title | CV, latest role | Filtering by role |
| Years of experience | Counted from the work history | Filtering by seniority |
| Skills | CV, skills section | The search most recruiters run first |
| Languages | CV | Filtering for language requirements |
| Profile links | CV | LinkedIn or GitHub, for a quick second look |
| Source file | Your folder | Every record leads back to the document |
| Status | Your team | New, contacted, interviewing, placed, not now |
| Last contacted | Your team | Who is due a follow up |
| Consent until | Your team | When the record must be reviewed or deleted |
The last three columns matter as much as the first ten. A database with no status turns into a list of names nobody remembers calling.
Step 2: choose where it lives
Three shapes cover almost every team. Pick by size and by how many people search it.
| Option | Good up to | Search | Duplicates | Who it suits |
|---|---|---|---|---|
| Excel or Google Sheets | A few hundred rows | Filter and SEARCH | By eye, or conditional formatting | One recruiter, one desk |
| Airtable or a similar table tool | A few thousand records | Filtered views, linked tables | Formula field on email | A small team sharing one list |
| Postgres or another SQL database | Hundreds of thousands | Full text search, indexes | A unique constraint | A developer building it into a product |
If you already pay for an ATS, check its candidate search first. Most applicant tracking systems keep every applicant, and the database you want may already exist under a different name. Build your own when the ATS cannot search the way you need, when you are importing an archive from somewhere else, or when the database is part of a product you sell, such as a job board.
Step 3: fill it from the CVs
This is the slow part. A CV is a document laid out for a human, so every record has to be read off it.
By hand
For a few dozen CVs, typing is fine. Open the CV beside the sheet and fill one row. Three habits keep the result searchable:
- Write dates as
YYYY-MMso they sort correctly. - Keep skills in one cell, separated by
;, so a single row stays a single person. - Count experience from the dates of each role, rather than trusting "10+ years" in a headline.
With a resume parser
Past a few dozen CVs, the reading is the bottleneck, and the errors start: a digit missed in a phone number, a skill left out because it sat inside a job description. A resume parser reads each file and returns the fields as structured data, which you then write into the database.
ResumeJSON returns, for each CV, a basics block with full_name, email, phone, location, headline and links; a work array where each role has company, title, start_date, end_date and is_current; plus education, skills, certifications, languages and total_years_experience, which is computed from the work history rather than copied from the CV. That maps onto the table in Step 1 almost column for column. The full list is in the field reference.
For a spreadsheet, Resume to Excel has a complete script that turns a folder of CVs into one sheet, with a second tab listing every file that failed. For a SQL database, the same loop writes rows instead:
create table candidates (
id bigserial primary key,
full_name text,
email text,
phone text,
location text,
current_title text,
years_exp numeric,
skills text[],
languages text[],
links text[],
source_file text not null,
status text not null default 'new',
last_contacted date,
consent_until date,
raw jsonb not null
);
create unique index candidates_email_key on candidates (lower(email)) where email is not null;
create index candidates_skills_idx on candidates using gin (skills);Keep the whole parse in the raw column. The columns are what you filter on today. The JSON is what lets you add a column next year without opening every file again.
Scans and odd files
A real archive is never all clean PDFs. ResumeJSON detects the type from the file's own bytes, so a PDF with no text layer, or a JPEG, PNG or WebP photo of a page, is read as an image. DOCX and plain text work too, and each upload can be up to 20 MB. A file that is not a CV at all, such as a cover letter, comes back as an error with the code not_a_resume, so it lands in your failure list instead of becoming an empty record. If most of your archive is scans, OCR resume parser covers what changes.
Step 4: keep duplicates out
The same person applies twice, sends an updated CV, or turns up in two imported archives. Without a rule, the database fills with near copies, and the newest notes end up on the wrong one.
A workable order of checks:
- Same email, ignoring case. The strongest signal. The unique index above enforces it.
- Same phone, after stripping spaces, dashes and brackets. Compare the last nine or ten digits, because some CVs carry a country code and some do not.
- Same name and same current company. Weaker. Flag these for a person to merge rather than merging them automatically.
When a match is found, update the record instead of adding a new one: keep your team's status and notes, replace the CV fields with the newer parse, and keep both source files. In SQL that is an insert ... on conflict (lower(email)) do update that leaves status, last_contacted and consent_until alone.
Step 5: make it searchable
Search is the reason the database exists, so test it with the questions from the start.
- In a sheet: a filter on title and location, and
=FILTER(A:M, ISNUMBER(SEARCH("python", H:H)))for a skill. - In Airtable: one saved view per common search, such as "Backend, Berlin, contacted more than 6 months ago".
- In Postgres:
where skills && array['Python','Django']uses the GIN index, and atsvectorover title and skills handles free text.
Skills are written many ways ("JS", "JavaScript", "Javascript ES6"). Search gets much better once you map them to one name each. Skills taxonomy covers the lists you can map to, and Candidate matching goes further, ranking records against a job description.
Step 6: keep it current and lawful
A candidate database holds personal data about real people, and it goes stale quickly.
- Record why you hold each record and until when. That is the
consent_untilcolumn. Under GDPR and similar laws, keeping CVs indefinitely "in case" is hard to justify. - Run a monthly review. Records past their date get a re-consent email or get deleted, together with the source file.
- Limit who can export it. A spreadsheet of candidates travels much further than a folder does.
ResumeJSON parses the document and keeps no copy of it, so the record in your database is the only one the parse produces. The data handling page has the detail.
When the manual way is enough
You do not need a parser, a script or a SQL table if:
- you hire for a handful of roles a year and see fewer than about fifty CVs in total;
- one person does all the searching and remembers most candidates anyway;
- your ATS already keeps every applicant and its search answers your questions.
In those cases a well kept spreadsheet with the columns from Step 1 is the right tool. The point to automate is when filling the records takes longer than searching them saves.
What it costs to fill
As of October 2026, our published plans on RapidAPI are:
| Plan | Monthly price | Parses included | Past the quota |
|---|---|---|---|
| Basic | $0 | 100 | Hard cap, never bills |
| Pro | $0 | None, pay per use | $0.05 a parse |
| Ultra | $29 | 1,000 | $0.045 a parse |
| Mega | $99 | 5,000 | $0.018 a parse |
An archive of up to 100 CVs fits the free Basic plan. A one-off import of 2,000 CVs costs $100 on pay per use, and a steady flow of new applications usually fits one of the monthly plans. Every call is metered whatever the answer, so remove obvious duplicates and non-CV attachments from the folder first. For archives in the thousands, Bulk resume parsing covers running the backlog in parallel without restarting from zero after a crash.
Start with one CV
Before you design the table, see what one of your own CVs turns into. The free resume parser takes a single file with no signup and shows the JSON a script would receive, so you can check that the fields you planned in Step 1 are the ones your documents actually fill. If the database is the first piece of a larger system, How to build an ATS picks up from here.