ResumeJSON

Candidate matching: how to match candidates to jobs

Candidate matching: ranking a pool, not scoring a pair

Candidate matching is the job of taking one role and finding the best people for it out of everyone you hold, or taking one person and finding the roles that fit them. It works in three passes: filter the pool on hard requirements, rank what is left on evidence from each CV, and show the reason beside every result. The hard part is rarely the ranking formula. It is getting every CV into the same fields so the first two passes have something to read.

This article is for the developer building candidate matching into a job board, an applicant tracking system or a recruiting tool. Comparing one CV against one posting is covered in compare resume and job description, and deciding what to do with one applicant is covered in automated resume screening. Here the question is the pool: thousands of candidates, hundreds of open roles, and a list that has to come back fast and make sense to a recruiter.

We build ResumeJSON, a parsing API that turns a CV into typed JSON. Read the parsing parts as written by an interested party. The matching itself is yours to own, whichever parser you use.


What candidate matching has to answer

A matching feature gets asked two questions, and a good design answers both from one data model.

DirectionWho asksThe questionTypical surface
Role to candidatesRecruiter, hiring manager"Who in our database fits this role?"A ranked shortlist on the job page
Candidate to rolesJob seeker, or a recruiter holding one CV"Which open roles fit this person?""Jobs for you", an alert email

Both directions use the same comparison. They differ only in which side is fixed and which side is the pool. If you build the comparison once, as a function over two records, you get both directions for free.

Three properties separate a matcher people trust from one they ignore:

Step 1: get every candidate into the same fields

Matching reads fields, so every CV has to become the same record first. The fields that matching actually uses:

FieldWhat matching does with it
skills[]Coverage of must-have and nice-to-have skills
work[] with title, start_date, end_date, is_currentRelevant titles, recency, experience floors
total_years_experienceA fast pre-filter on seniority
education[] with degree, field_of_studyDegree requirements
certifications[] with nameLicences and other hard requirements
languages[] with language, proficiencyLanguage requirements
basics.locationLocation filters and relocation

Parse once, at upload, and store the record beside the file. Never parse at match time. A matcher that reads a thousand PDFs per search is slow, and it gives a different answer whenever the parser changes.

You can build the parse yourself; the do-it-yourself routes are walked through for Python, Node.js and ChatGPT. If you already hold a backlog of files, bulk resume parsing covers running them through once. Or call a parsing API and store what comes back. ResumeJSON returns the fields above for PDF, DOCX, plain text and images of a page, with dates as YYYY-MM or YYYY and missing values as null rather than a guess.

Normalise skills before you store them

The single biggest cause of missed matches is spelling. "Postgres", "PostgreSQL" and "postgres db" are one skill. Keep a synonyms table in your own code and map every skill to a canonical id at write time:

const CANONICAL: Record<string, string> = {
  postgres: 'postgresql',
  'postgres db': 'postgresql',
  js: 'javascript',
  'node': 'nodejs',
  'node.js': 'nodejs',
};

function canonicalSkill(raw: string): string {
  const key = raw.trim().toLowerCase();
  return CANONICAL[key] ?? key;
}

Grow the table from real misses, which you find by reading the "missing" column of matches a recruiter says were good. If you want a published list to map to rather than your own, skills taxonomy compares the options.

Step 2: store requirements as data, not prose

The job side needs the same treatment. A posting written as a paragraph cannot be matched reliably, so ask the employer for requirements as structure when they create the role:

type Requirement =
  | { kind: 'skill'; id: string; hard: boolean }
  | { kind: 'years'; min: number; hard: boolean }
  | { kind: 'certification'; name: string; hard: true }
  | { kind: 'language'; language: string; hard: boolean }
  | { kind: 'location'; city: string; remoteOk: boolean; hard: boolean };

Each line is tagged required or preferred. That one flag is what lets the next step split the work into a cheap filter and a careful rank.

Step 3: filter the pool on hard requirements

Filtering is a database query, and it should stay one. Put the fields you filter on into columns or a search index, and let the database drop everyone who plainly cannot do the job before any scoring code runs.

SELECT c.id
FROM candidates c
WHERE (c.total_years_experience IS NULL OR c.total_years_experience >= 4)
  AND EXISTS (
    SELECT 1 FROM candidate_certifications cc
    WHERE cc.candidate_id = c.id AND cc.name_canonical = 'registered nurse'
  );

Read the first line of the WHERE clause again. A NULL passes the filter. A candidate whose dates could not be read has not proved they lack the experience, so they go through to ranking, where they will show as "unclear" and sort below the candidates who clearly meet it. Filter out a NULL and you have built a matcher that punishes unusual CV layouts, which is exactly the fault people already complain about in ATS parsing.

Keep the hard filter short. Every requirement you move from "preferred" to "hard" removes people a recruiter might have wanted to see.

Step 4: rank what is left, with reasons

Ranking runs on the filtered set, so it can afford to be careful. Compute a mark per requirement, then sort on the marks:

type Mark = 'met' | 'missing' | 'unclear';

function marksFor(c: Candidate, reqs: Requirement[]) {
  const skills = new Set(c.skills.map(canonicalSkill));
  return reqs.map(req => {
    switch (req.kind) {
      case 'skill':
        return { req, mark: (skills.has(req.id) ? 'met' : 'missing') as Mark };
      case 'years': {
        const y = c.total_years_experience;
        if (y === null) return { req, mark: 'unclear' as Mark };
        return { req, mark: (y >= req.min ? 'met' : 'missing') as Mark };
      }
      case 'certification':
        return {
          req,
          mark: (c.certifications.some(x => x.name?.toLowerCase().includes(req.name)) ? 'met' : 'missing') as Mark,
        };
      case 'language':
        return {
          req,
          mark: (c.languages.some(l => l.language.toLowerCase() === req.language) ? 'met' : 'missing') as Mark,
        };
      case 'location': {
        const loc = c.basics.location;
        if (req.remoteOk) return { req, mark: 'met' as Mark };
        if (loc === null) return { req, mark: 'unclear' as Mark };
        return { req, mark: (loc.toLowerCase().includes(req.city) ? 'met' : 'unclear') as Mark };
      }
    }
  });
}

Then sort with a key a recruiter could say out loud:

  1. Hard requirements met, most first.
  2. Preferred requirements met, most first.
  3. Fewest "unclear" marks.
  4. Most recent relevant role, newest first.
function rankKey(marks: ReturnType<typeof marksFor>, lastRelevantEnd: string | null) {
  const count = (hard: boolean, m: Mark) =>
    marks.filter(x => x.req.hard === hard && x.mark === m).length;
  return [
    -count(true, 'met'),
    -count(false, 'met'),
    marks.filter(x => x.mark === 'unclear').length,
    lastRelevantEnd === null ? 1 : -Date.parse(lastRelevantEnd),
  ];
}

A tuple sort beats a weighted percentage. With weights, a recruiter who asks "why is this person fourth?" gets an answer in decimals. With a tuple, the answer is "they meet every must-have and two of the four nice-to-haves". Store the marks with the result, and render them as a checklist beside each name.

Recency is the one tie-breaker worth adding early. Two people who both list Kubernetes are not equal when one used it last month and the other in 2016, and work[].end_date with is_current gives you that for free.

Step 5: run it the other way round

"Jobs for you" is the same function with the sides swapped. Fix the candidate, treat open roles as the pool, filter roles whose hard requirements the candidate cannot meet, and rank the rest with the same key.

Two differences are worth handling:

Semantic matching: where embeddings help, and where they do not

Embeddings (turning text into vectors and comparing their distance) are the popular answer to "match on meaning, not on keywords". They are useful in one narrow place and risky in the rest.

UseEmbeddings help?Why
Finding skill synonyms for your tableYesSuggests "PostgreSQL" is near "Postgres" for a person to approve
Similar job titles ("backend engineer", "server-side developer")OftenTitles vary more than skills, and a suggestion list is easy to review
The final rankingRarelyA distance score cannot say which requirement was met, and it shifts when the model changes
Hard filtersNoA licence is present or absent; closeness is the wrong question

The sound pattern is to use similarity to suggest mappings that a person approves into your synonyms table, then keep the match itself on the approved, canonical fields. You get the recall of fuzzy matching with the explanation of exact matching.

The same reasoning applies to asking a language model for a match score. It is fast to build and hard to defend: the number moves between runs, and nobody can tell a rejected candidate what drove it. Use a model to read a CV into fields if you like, because extraction can be checked. Keep the decision in code.

Checking that the matcher is any good

A matcher that nobody measures drifts into one that everybody ignores. Three cheap checks:

When a simple search is enough

Not every product needs a matcher.

What every version needs, even the simple one, is the CV as fields. A keyword search over raw PDF text gives you the matching quality you were trying to escape.

Where ResumeJSON fits

ResumeJSON does step 1 and nothing else. It does not match, rank or score, because that logic belongs in your product, where you can read it and test it. It gives you every CV in one typed schema, in about two seconds, ready to store and query.

As of September 2026 the API is sold through RapidAPI. The Basic plan is $0 for 100 parses a month with a hard cap, Pro has no monthly fee and bills $0.05 a parse, and Ultra is $29 a month for 1,000 parses. The docs carry the current table and the full field reference, and the free parser lets you try a real CV before you write any code. If you are building the rest of the system too, how to build an ATS covers the data model around it.

Candidate matching comes down to three passes over clean fields: filter on what is required, rank on the evidence, and show the reason for every result.

All articles