ResumeJSON

Greenhouse API: which one to use, and how to get resume data out

Greenhouse API: a developer's map, and the resume data inside it

The Greenhouse API is really several APIs, and the one you want depends on what you are building. For reading and writing a company's own recruiting data (jobs, candidates, applications, attachments) you use the Harvest API, which is now on version 3 and authenticates with OAuth 2.0. For a custom careers page you use the Job Board API, whose read endpoints are public. Sourcing tools that push candidates in use the Ingestion API, and anything that should react to events uses webhooks. If your goal is resume data, Harvest gives you the files and a few profile fields. Turning the file itself into structured work history, skills and education is a separate step you add.

This article is for the developer writing that integration. Everything about Greenhouse below was read from Greenhouse's own developer documentation at developers.greenhouse.io, harvestdocs.greenhouse.io and docs.greenhouse.io, and from its support site, as of October 2026. We build ResumeJSON, a resume parsing API, so read the parsing sections as written by an interested party.

The Greenhouse APIs at a glance

Greenhouse's developer portal lists these for Recruiting, plus a GraphQL API, a REST API and webhooks for its Onboarding product:

APIWhat it is forWho calls it
Harvest v3Read and write jobs, candidates, applications, attachments, interviewsThe customer's own integration, or a partner app
Job Board APIBuild a careers page, list jobs, submit applicationsYour website or job board
Ingestion APISubmit prospects and candidates from a sourcing toolSourcing partners
Recruiting webhooksBe told when something happens, such as a candidate being hiredAny integration that reacts to events
Assessment APIConnect a testing platform to the hiring flowAssessment vendors
Audit log APISee who accessed or edited whatSecurity and compliance tooling

Most searches for "Greenhouse API" end up at one of the first two. The rest of this guide covers those, then the resume question that brings many developers here.

Harvest v1 and v2 are gone

If you are reading an older tutorial, check its version. The developer portal states it plainly:

Harvest v1/v2 APIs were deprecated on August 31, 2026

and adds that "Harvest v1/v2 have stopped accepting requests." Code that sends Basic auth to a /v1/ Harvest URL is code that needs migrating. Greenhouse publishes separate READ and WRITE endpoint migration guides on the Harvest v3 docs site. One change worth knowing early: v3 flattens resources that v1 nested under a candidate. Work history that lived at /v1/candidates/{id}/employments is now its own list, /v3/candidate_employments, filtered by candidate.

Authenticating to Harvest v3

Harvest v3 uses OAuth 2.0, with two flows depending on who owns the integration:

For a custom integration the steps from Greenhouse's authentication guide are:

  1. In Greenhouse, open API Credentials, click Create new API credentials and choose Harvest V3 (OAuth).
  2. Configure the scopes the integration needs. Scopes are per resource and action, for example harvest:attachments:list or harvest:candidate_employments:create.
  3. Exchange the Client ID and Client Secret for a token with a server-to-server POST to https://auth.greenhouse.io/token, using HTTP Basic auth and grant_type=client_credentials.
  4. Call the API with Authorization: Bearer <token>, and request a new token when you get a 401.
curl https://auth.greenhouse.io/token \
  --user "$GH_CLIENT_ID:$GH_CLIENT_SECRET" \
  --header 'Content-Type: application/x-www-form-urlencoded' \
  --data 'grant_type=client_credentials'

Who the token acts as matters. Creating the credential also creates an integration service user, and omitting the sub parameter makes requests as that user. Passing sub with a user id limits the request to what that user can see. Greenhouse notes that "all list endpoints require authorization by a Site Admin", and that private data such as offers and salary needs an extra advanced permission. Without it, list endpoints for those resources come back empty rather than failing, which is easy to mistake for a clean result.

Paging and rate limits in Harvest v3

Two mechanics catch nearly every first integration.

Pagination is cursor based. Put your filters and per_page on the first request only, then follow the rel="next" URL in the response's Link header until there is none. The cursor must be the only query parameter on later requests; adding a filter alongside it returns a 422. Page size defaults to 100 and goes up to 500, and results come back by id in descending order.

Rate limits use a fixed 30-second window. Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset. Go over and you get a 429 with a Retry-After header in seconds. Requests for a new access token have their own 60-second window. Greenhouse's own advice is to slow down as X-RateLimit-Remaining falls, rather than waiting for the error.

SituationWhat Harvest v3 doesWhat your code should do
More results than one pageSends a Link header with rel="next"Follow it exactly, add nothing to it
Cursor plus another parameterReturns 422Put filters on the first request only
Too many requests in 30 secondsReturns 429 with Retry-AfterWait that many seconds, then retry
Token expiredReturns 401Fetch a new token and retry once
Missing permission on private dataReturns an empty listCheck the user behind sub before trusting it

The Job Board API: careers pages and applications

The Job Board API lives at https://boards-api.greenhouse.io/v1/boards/{board_token}/. Greenhouse's documentation says "Job Board data is publicly available, so authentication is not required for any GET endpoints." That covers jobs, a single job with its application questions, offices, departments and the education lists used on forms.

Submitting an application is different. POST /v1/boards/{board_token}/jobs/{id} requires HTTP Basic auth with your Job Board API key, and Greenhouse is direct about where that request must come from:

Any form posts should be proxied by your own servers.

So the browser posts to your server, and your server posts to Greenhouse. A resume goes up either as a file in a multipart/form-data request, or as resume_content with a resume_content_filename, or as plain resume_text. The questions array on each job tells you which fields that job's form expects, because application forms are job specific.

If you are building a job board that feeds Greenhouse, this is also the moment to fill the form for the applicant. Parse the CV when they upload it, show them the fields to correct, then submit. Job application autofill walks through that pattern.

Getting resume data out of Greenhouse

Greenhouse Recruiting parses resumes itself when a candidate is added. Its support article describes the feature as scanning an imported resume and auto-filling "appropriate fields with information it detects". The same article lists why a parse fails: the file is over 2.5MB, the resume contains fake or poorly disguised data, or the formatting confuses it. On a failure the resume is only attached, and someone fills in the profile by hand.

For an integration, the resume reaches you as an attachment. In Harvest v3:

  1. List the attachments. GET /v3/attachments accepts candidate_ids, application_ids and a type filter. The type enum includes resume, cover_letter and take_home_test for candidate documents, alongside offer and agreement types.
  2. Download promptly. Each attachment has a url, and the documentation says it is "a time-limited download link that expires after seven days". It may also redirect to a fresh short-lived file URL, so your HTTP client must follow redirects. Store the file, never the link.
  3. Read what the profile already holds. Work history is GET /v3/candidate_employments and education is GET /v3/candidate_educations, each filtered by candidate.

That last step is where many projects stall. A profile filled by Greenhouse's parse, by a recruiter, or by an application form gives you what that route captured. If your code needs every job with its dates, a skills list, languages or certifications in one consistent shape, you parse the file yourself.

When Greenhouse's own data is enough

Stay with what Greenhouse already holds when all of these are true:

That describes plenty of internal reporting and sync jobs. Adding a parser to repeat work the profile already shows gains you nothing.

Parsing Greenhouse resumes into structured JSON

The pattern that works is the same whatever parser you choose: pull the file from Harvest, parse it once, keep the full record on your side, and write back only what recruiters should see.

// 1. Find resume attachments for a set of candidates (Harvest v3)
const list = await fetch(
  'https://harvest.greenhouse.io/v3/attachments?type=resume&candidate_ids=' + ids.join(','),
  { headers: { Authorization: `Bearer ${ghToken}` } }
)
const attachments = await list.json()

for (const a of attachments) {
  // 2. Download now; the URL expires, and it may redirect
  const file = await fetch(a.url, { redirect: 'follow' })
  const form = new FormData()
  form.set('file', await file.blob(), a.filename)

  // 3. Parse the file into typed fields
  const res = await fetch('https://resumejson-resume-cv-parser-api.p.rapidapi.com/v1/parse', {
    method: 'POST',
    headers: {
      'x-rapidapi-key': process.env.RAPIDAPI_KEY,
      'x-rapidapi-host': 'resumejson-resume-cv-parser-api.p.rapidapi.com'
    },
    body: form
  })
  const { resume } = await res.json()

  // 4. Store the full record keyed on Greenhouse ids
  await save({ candidateId: a.candidate_id, applicationId: a.application_id, resume })
}

Write-back is optional and has rules of its own. POST /v3/candidate_employments needs candidate_id, company_name, title and start_date, with free-text company and title. POST /v3/candidate_educations is stricter: school, degree and discipline must be ids of existing custom field options, so you look those up first, or create them. Many teams keep skills and certifications in their own database and push only tags or a few employment rows back.

Two operational notes from the Harvest docs apply here. Use per_page up to 500 on the list call to spend fewer requests, and watch X-RateLimit-Remaining, because a backlog sync is exactly the job that hits the 30-second window. Bulk resume parsing covers running a large backlog through a parser without losing files.

Greenhouse data and a parsed CV, side by side

What your code needsFrom Greenhouse's APIFrom a parsed CV (ResumeJSON)
The original fileGET /v3/attachments, link valid seven daysYou send it
Contact detailsOn the candidate recordbasics.full_name, email, phone, links, location
Work historycandidate_employments rows, as entered or parsedwork[] with company, title, highlights[], is_current, dates
Educationcandidate_educations rows, tied to custom field optionseducation[] with institution, degree, field_of_study
SkillsNot a documented Harvest resourceskills[]
Languages and certificationsNot documented Harvest resourceslanguages[], certifications[]
Total experienceCompute it from the employment rowstotal_years_experience
Works before the candidate existsNoYes

Field names on the Greenhouse side are from the Harvest v3 reference as of October 2026. ResumeJSON reads PDF, DOCX, plain text and JPEG, PNG or WebP images of a page, and returns null for a value the CV does not state rather than guessing one. It is sold through RapidAPI: the Basic plan is $0 for 100 parses a month with a hard cap, Pro is pay per use at $0.05 a parse, and the monthly plans start at $29 a month for 1,000 parses; every plan is on the pricing table. You can check how a real CV comes out in the free parser before writing any code.

When you need a parser of your own

These are the cases where the Greenhouse API alone does not reach:

Summary

If you are designing where that record lives, how to build an ATS covers the data model, and JSON Resume schema covers a common format for storing it.

All articles