Candidate matching: how to match candidates to jobs
Published
Candidate matching: ranking a pool, not scoring a pair
Candidate matching is the job of taking one role and finding the best people for it out of everyone you hold, or taking one person and finding the roles that fit them. It works in three passes: filter the pool on hard requirements, rank what is left on evidence from each CV, and show the reason beside every result. The hard part is rarely the ranking formula. It is getting every CV into the same fields so the first two passes have something to read.
This article is for the developer building candidate matching into a job board, an applicant tracking system or a recruiting tool. Comparing one CV against one posting is covered in compare resume and job description, and deciding what to do with one applicant is covered in automated resume screening. Here the question is the pool: thousands of candidates, hundreds of open roles, and a list that has to come back fast and make sense to a recruiter.
We build ResumeJSON, a parsing API that turns a CV into typed JSON. Read the parsing parts as written by an interested party. The matching itself is yours to own, whichever parser you use.
What candidate matching has to answer
A matching feature gets asked two questions, and a good design answers both from one data model.
| Direction | Who asks | The question | Typical surface |
|---|---|---|---|
| Role to candidates | Recruiter, hiring manager | "Who in our database fits this role?" | A ranked shortlist on the job page |
| Candidate to roles | Job seeker, or a recruiter holding one CV | "Which open roles fit this person?" | "Jobs for you", an alert email |
Both directions use the same comparison. They differ only in which side is fixed and which side is the pool. If you build the comparison once, as a function over two records, you get both directions for free.
Three properties separate a matcher people trust from one they ignore:
- It never hides a good candidate for a formatting reason. A CV with unreadable dates is unclear, never disqualified.
- It explains itself. Every result carries the requirements it met and the ones it missed.
- It is stable. The same pool and the same role give the same order tomorrow.
Step 1: get every candidate into the same fields
Matching reads fields, so every CV has to become the same record first. The fields that matching actually uses:
| Field | What matching does with it |
|---|---|
skills[] | Coverage of must-have and nice-to-have skills |
work[] with title, start_date, end_date, is_current | Relevant titles, recency, experience floors |
total_years_experience | A fast pre-filter on seniority |
education[] with degree, field_of_study | Degree requirements |
certifications[] with name | Licences and other hard requirements |
languages[] with language, proficiency | Language requirements |
basics.location | Location filters and relocation |
Parse once, at upload, and store the record beside the file. Never parse at match time. A matcher that reads a thousand PDFs per search is slow, and it gives a different answer whenever the parser changes.
You can build the parse yourself; the do-it-yourself routes are walked through for Python, Node.js and ChatGPT. If you already hold a backlog of files, bulk resume parsing covers running them through once. Or call a parsing API and store what comes back. ResumeJSON returns the fields above for PDF, DOCX, plain text and images of a page, with dates as YYYY-MM or YYYY and missing values as null rather than a guess.
Normalise skills before you store them
The single biggest cause of missed matches is spelling. "Postgres", "PostgreSQL" and "postgres db" are one skill. Keep a synonyms table in your own code and map every skill to a canonical id at write time:
const CANONICAL: Record<string, string> = {
postgres: 'postgresql',
'postgres db': 'postgresql',
js: 'javascript',
'node': 'nodejs',
'node.js': 'nodejs',
};
function canonicalSkill(raw: string): string {
const key = raw.trim().toLowerCase();
return CANONICAL[key] ?? key;
}Grow the table from real misses, which you find by reading the "missing" column of matches a recruiter says were good. If you want a published list to map to rather than your own, skills taxonomy compares the options.
Step 2: store requirements as data, not prose
The job side needs the same treatment. A posting written as a paragraph cannot be matched reliably, so ask the employer for requirements as structure when they create the role:
type Requirement =
| { kind: 'skill'; id: string; hard: boolean }
| { kind: 'years'; min: number; hard: boolean }
| { kind: 'certification'; name: string; hard: true }
| { kind: 'language'; language: string; hard: boolean }
| { kind: 'location'; city: string; remoteOk: boolean; hard: boolean };Each line is tagged required or preferred. That one flag is what lets the next step split the work into a cheap filter and a careful rank.
Step 3: filter the pool on hard requirements
Filtering is a database query, and it should stay one. Put the fields you filter on into columns or a search index, and let the database drop everyone who plainly cannot do the job before any scoring code runs.
SELECT c.id
FROM candidates c
WHERE (c.total_years_experience IS NULL OR c.total_years_experience >= 4)
AND EXISTS (
SELECT 1 FROM candidate_certifications cc
WHERE cc.candidate_id = c.id AND cc.name_canonical = 'registered nurse'
);Read the first line of the WHERE clause again. A NULL passes the filter. A candidate whose dates could not be read has not proved they lack the experience, so they go through to ranking, where they will show as "unclear" and sort below the candidates who clearly meet it. Filter out a NULL and you have built a matcher that punishes unusual CV layouts, which is exactly the fault people already complain about in ATS parsing.
Keep the hard filter short. Every requirement you move from "preferred" to "hard" removes people a recruiter might have wanted to see.
Step 4: rank what is left, with reasons
Ranking runs on the filtered set, so it can afford to be careful. Compute a mark per requirement, then sort on the marks:
type Mark = 'met' | 'missing' | 'unclear';
function marksFor(c: Candidate, reqs: Requirement[]) {
const skills = new Set(c.skills.map(canonicalSkill));
return reqs.map(req => {
switch (req.kind) {
case 'skill':
return { req, mark: (skills.has(req.id) ? 'met' : 'missing') as Mark };
case 'years': {
const y = c.total_years_experience;
if (y === null) return { req, mark: 'unclear' as Mark };
return { req, mark: (y >= req.min ? 'met' : 'missing') as Mark };
}
case 'certification':
return {
req,
mark: (c.certifications.some(x => x.name?.toLowerCase().includes(req.name)) ? 'met' : 'missing') as Mark,
};
case 'language':
return {
req,
mark: (c.languages.some(l => l.language.toLowerCase() === req.language) ? 'met' : 'missing') as Mark,
};
case 'location': {
const loc = c.basics.location;
if (req.remoteOk) return { req, mark: 'met' as Mark };
if (loc === null) return { req, mark: 'unclear' as Mark };
return { req, mark: (loc.toLowerCase().includes(req.city) ? 'met' : 'unclear') as Mark };
}
}
});
}Then sort with a key a recruiter could say out loud:
- Hard requirements met, most first.
- Preferred requirements met, most first.
- Fewest "unclear" marks.
- Most recent relevant role, newest first.
function rankKey(marks: ReturnType<typeof marksFor>, lastRelevantEnd: string | null) {
const count = (hard: boolean, m: Mark) =>
marks.filter(x => x.req.hard === hard && x.mark === m).length;
return [
-count(true, 'met'),
-count(false, 'met'),
marks.filter(x => x.mark === 'unclear').length,
lastRelevantEnd === null ? 1 : -Date.parse(lastRelevantEnd),
];
}A tuple sort beats a weighted percentage. With weights, a recruiter who asks "why is this person fourth?" gets an answer in decimals. With a tuple, the answer is "they meet every must-have and two of the four nice-to-haves". Store the marks with the result, and render them as a checklist beside each name.
Recency is the one tie-breaker worth adding early. Two people who both list Kubernetes are not equal when one used it last month and the other in 2016, and work[].end_date with is_current gives you that for free.
Step 5: run it the other way round
"Jobs for you" is the same function with the sides swapped. Fix the candidate, treat open roles as the pool, filter roles whose hard requirements the candidate cannot meet, and rank the rest with the same key.
Two differences are worth handling:
- Show the gap, not only the fit. A job seeker gains more from "you meet five of six; this role wants a Terraform certification" than from a match badge. The marks already carry this.
- Recompute on change, not on read. When a candidate uploads a new CV or a role is edited, rematch that one record against the other side and store the result. Alert emails then read stored matches rather than running a search per recipient.
Semantic matching: where embeddings help, and where they do not
Embeddings (turning text into vectors and comparing their distance) are the popular answer to "match on meaning, not on keywords". They are useful in one narrow place and risky in the rest.
| Use | Embeddings help? | Why |
|---|---|---|
| Finding skill synonyms for your table | Yes | Suggests "PostgreSQL" is near "Postgres" for a person to approve |
| Similar job titles ("backend engineer", "server-side developer") | Often | Titles vary more than skills, and a suggestion list is easy to review |
| The final ranking | Rarely | A distance score cannot say which requirement was met, and it shifts when the model changes |
| Hard filters | No | A licence is present or absent; closeness is the wrong question |
The sound pattern is to use similarity to suggest mappings that a person approves into your synonyms table, then keep the match itself on the approved, canonical fields. You get the recall of fuzzy matching with the explanation of exact matching.
The same reasoning applies to asking a language model for a match score. It is fast to build and hard to defend: the number moves between runs, and nobody can tell a rejected candidate what drove it. Use a model to read a CV into fields if you like, because extraction can be checked. Keep the decision in code.
Checking that the matcher is any good
A matcher that nobody measures drifts into one that everybody ignores. Three cheap checks:
- Ask recruiters for the answer first. For ten past roles, take the people who were actually interviewed and see where your ranking put them. If they sit on page three, look at the "missing" marks: the cause is usually a synonym or a title the table does not know.
- Watch the unclear rate. If a large share of the pool shows "unclear" on experience, the problem is the parse, not the ranking. Look at a few of those CVs directly.
- Freeze a test pool. Keep a fixed set of parsed CVs and roles in your test suite, and assert the order. A change to the synonyms table that reshuffles every role should fail a test, not surprise a customer.
When a simple search is enough
Not every product needs a matcher.
- Under a few hundred candidates, a recruiter reading a filtered list is faster than any ranking you will build, and a skills filter plus a location filter covers most of it.
- If employers will not enter structured requirements, you cannot rank on them. Start with a good filter on parsed fields and add requirements to the job form later.
- For one role with a handful of applicants, the manual method in compare resume and job description is the whole job.
What every version needs, even the simple one, is the CV as fields. A keyword search over raw PDF text gives you the matching quality you were trying to escape.
Where ResumeJSON fits
ResumeJSON does step 1 and nothing else. It does not match, rank or score, because that logic belongs in your product, where you can read it and test it. It gives you every CV in one typed schema, in about two seconds, ready to store and query.
As of September 2026 the API is sold through RapidAPI. The Basic plan is $0 for 100 parses a month with a hard cap, Pro has no monthly fee and bills $0.05 a parse, and Ultra is $29 a month for 1,000 parses. The docs carry the current table and the full field reference, and the free parser lets you try a real CV before you write any code. If you are building the rest of the system too, how to build an ATS covers the data model around it.
Candidate matching comes down to three passes over clean fields: filter on what is required, rank on the evidence, and show the reason for every result.