mirror of
https://github.com/MadsLorentzen/ai-job-search.git
synced 2026-09-17 08:36:25 +00:00
The skill queried /api/v1/jobs/search, whose `description` is the search index's truncated preview — and the CLI dropped it entirely, so a result carried only title/company/location/date/url. Reading a posting therefore meant a `detail` call per hit, which is exactly what job-scraper's Step 2 prescribes: "fetch full detail with that portal's `detail` command". freehire exposes a search endpoint for programmatic consumers, /api/v1/agent/jobs/search: same query, ranking, facets and pagination, but asked to (`include_description=true`) it replaces the preview with the posting's full description read from the database, rendered as `description_format=markdown|text|html`. Reproduce the difference: curl -s "https://freehire.me/api/v1/jobs/search?q=golang&limit=1" \ | jq -r '.data[0].description | length' # preview, capped curl -s "https://freehire.me/api/v1/agent/jobs/search?q=golang&limit=1\ &include_description=true&description_format=markdown" \ | jq -r '.data[0].description | length' # full text So `search` now calls that endpoint, always asking for full descriptions, and each JSON result carries `description` verbatim — no client-side HTML stripping, since the API already rendered it. Markdown is the default because it preserves the headings and requirement lists /rank reasons over; `--description-format text|html` selects the others. The flag is validated client-side: the API answers an unrecognized format with raw HTML rather than an error, so a typo would silently change the output instead of failing. `table` and `plain` stay description-free — a full posting body would swamp a scannable list — and `detail` is untouched, for looking one posting up by slug (including a closed one, absent from search). One behaviour change beyond the endpoint: a 404 from the search path used to be folded into an empty result set. On the agent endpoint a 404 means the instance predates it — a self-hosted freehire behind FREEHIRE_API_URL — so it is now reported as an error naming the path, instead of a plausible "no results" that hides the misconfiguration. Tests cover the requested URL and params, verbatim (unstripped) markdown, the null-when-absent case, the 404-is-an-error contract, and the flag validation. All network-free.
146 lines
7.0 KiB
Markdown
146 lines
7.0 KiB
Markdown
# freehire.me API reference
|
|
|
|
The endpoints, parameters, and response shapes this skill depends on. This is the
|
|
file to update if the freehire API changes. Base URL defaults to
|
|
`https://freehire.me` and is overridable via the `FREEHIRE_API_URL` env var.
|
|
|
|
## Authentication
|
|
|
|
None for reads. `GET /api/v1/jobs/*` and `/companies/*` are public; only per-user
|
|
tracking mutations (`apply`/`save`/`me`) require a bearer API key, and this skill
|
|
does not use them.
|
|
|
|
Verified against the live API:
|
|
|
|
| Endpoint | Status |
|
|
|----------|--------|
|
|
| `GET /api/v1/agent/jobs/search` | 200 |
|
|
| `GET /api/v1/jobs/search` | 200 (the web variant; not used by this skill) |
|
|
| `GET /api/v1/jobs/facets` | 200 |
|
|
| `GET /api/v1/jobs/{slug}` | 200 |
|
|
| `GET /api/v1/auth/me` | 401 (auth required — not used here) |
|
|
|
|
## Envelope
|
|
|
|
Every response is `{ "data": ..., "meta": {...}, "error": "..." }`. Lists put the
|
|
array in `data` and pagination in `meta` (`{ total, limit, offset }`); a single
|
|
item puts the object in `data`. Errors are `{ "error": "<message>" }` with a 4xx/5xx
|
|
status (e.g. 404 → `{ "error": "not found" }`).
|
|
|
|
## `GET /api/v1/agent/jobs/search`
|
|
|
|
The endpoint the skill's `search` command uses. Full-text + facet search over open
|
|
jobs, returning `data: [job, …]` with `meta.total` = the estimated match count.
|
|
|
|
It runs the **same query** as the web-facing `/api/v1/jobs/search` — same `q`, same
|
|
facets, same ranking, same pagination guard (`offset + limit ≤ 10000`) — and differs
|
|
in one respect: asked to, it replaces the search index's truncated `description`
|
|
preview with the posting's **full** description read from the database. That is what
|
|
lets a search of N roles stay one request instead of N + 1.
|
|
|
|
Two extra parameters control it:
|
|
|
|
| Param | Maps to CLI flag | Notes |
|
|
|-------|------------------|-------|
|
|
| `include_description` | (always `true`) | Without it the endpoint serves the index preview, same as the web search. |
|
|
| `description_format` | `--description-format` | `markdown` (the skill's default), `text`, or `html`. **An unrecognized value is not an error** — the API falls back to `html`, so the CLI validates the flag itself. |
|
|
|
|
Hydration is best-effort per hit: a result whose row has vanished from the database
|
|
(the index lagging a just-removed job) keeps the preview rather than being dropped,
|
|
so `description` is a full text in practice but never guaranteed to be.
|
|
|
|
A `404` from this path means the instance predates the endpoint (a self-hosted
|
|
freehire behind `FREEHIRE_API_URL`), not a missing job; the CLI reports it as an
|
|
error naming the path rather than as an empty result set.
|
|
|
|
## `GET /api/v1/jobs/search`
|
|
|
|
The web variant of the same search — identical query surface, but `description` is
|
|
always the index's truncated preview. The skill does not call it; it is listed here
|
|
because the shared parameters below are documented against both.
|
|
|
|
Query parameters used by the skill:
|
|
|
|
| Param | Maps to CLI flag | Notes |
|
|
|-------|------------------|-------|
|
|
| `q` | `--query` / `-q` | Keyword full-text query. |
|
|
| `limit` | `--limit` / `-n` | Page size. Default 25 in the CLI. |
|
|
| `offset` | (derived) | `offset = (page - 1) * limit`. |
|
|
| `semantic_ratio` | (fixed `0`) | Keyword search; the semantic index is opt-in. |
|
|
| `posted_within_days` | `--jobage` | Restrict to postings from the last N days. |
|
|
| `regions` | `--region` | Repeatable; OR within the facet. Values like `global`, `eu`, `us`, `apac`, `latam`, `cis`. |
|
|
| `countries` | `--country` | Repeatable; ISO-3166 alpha-2 (lowercased). |
|
|
| `cities` | `--city` | Repeatable; display-name city. |
|
|
| `seniority` | `--seniority` | Repeatable; `junior`, `middle`, `senior`, `staff`, … |
|
|
| `category` | `--category` | Repeatable; `backend`, `frontend`, `fullstack`, `devops`, `ml_ai`, … |
|
|
| `skills` | `--skill` | Repeatable; canonical skill names. |
|
|
| `company_slug` | `--company` | Single company. |
|
|
| `work_mode` | `--remote` | `remote` \| `hybrid` \| `onsite`. |
|
|
| any facet param | `--facet key=value` | Escape hatch for the long tail (e.g. `salary_min`, `visa_sponsorship`, `employment_type`, `english_level`). |
|
|
|
|
Repeated params (`?seniority=senior&seniority=staff`) are ORed within a facet;
|
|
different facets are ANDed (geography ORs into one location group). Deep paging is
|
|
bounded server-side (`offset + limit ≤ 10000`).
|
|
|
|
### Job object (the fields the skill reads)
|
|
|
|
```jsonc
|
|
{
|
|
"public_slug": "golang-zensar-2bxu6dxm", // -> result.id, and detail's <slug>
|
|
"source": "oracle",
|
|
"external_id": "…",
|
|
"url": "https://…", // the real posting URL (ATS host)
|
|
"title": "GOLANG",
|
|
"company": "Zensar",
|
|
"company_slug": "zensar",
|
|
"location": "India", // free-text ATS location
|
|
"description": "- …", // agent search: full text in the requested
|
|
// format; elsewhere HTML, stripped client-side
|
|
"skills": ["go", "kubernetes", …], // dictionary facet (top-level)
|
|
"work_mode": "remote", // may be absent
|
|
"regions": ["apac"], // dictionary/hybrid facet
|
|
"countries": ["in"],
|
|
"cities": [],
|
|
"collections": [],
|
|
"posted_at": "2026-07-06T00:00:00Z", // -> result.date (nullable)
|
|
"created_at": "2026-07-06T15:25:…Z",
|
|
"enrichment": { // nested, typed; {} when unenriched
|
|
"seniority": "senior",
|
|
"category": "backend",
|
|
"employment_type": "full_time",
|
|
"salary_min": 90000, "salary_max": 120000, "salary_currency": "EUR"
|
|
}
|
|
}
|
|
```
|
|
|
|
The internal numeric id is deliberately never exposed; `public_slug` is the stable
|
|
identifier.
|
|
|
|
## `GET /api/v1/jobs/{slug}`
|
|
|
|
A single job by its `public_slug`. Returns the same job object in `data`. A closed
|
|
posting is still served here (with a non-null `closed_at`); a missing slug is 404
|
|
`{ "error": "not found" }`. The skill's `detail` command maps a 404 to a
|
|
`NOT_FOUND` error on stderr.
|
|
|
|
## `GET /api/v1/jobs/facets`
|
|
|
|
The market's facet-value distributions under an optional filter — each facet's live
|
|
values with counts. `data.facets` is `{ <facet>: { <value>: <count> } }`. This skill
|
|
does not call it programmatically, but it is the vocabulary source the SKILL.md
|
|
points users to (`?q=<role>` scopes the counts). Example:
|
|
`GET /api/v1/jobs/facets?q=react`.
|
|
|
|
## Parsing notes
|
|
|
|
- The response is JSON, so there is no HTML card parsing (unlike the scraping
|
|
portals). The only markup handling left client-side is `detail`'s: `/jobs/{slug}`
|
|
serves HTML, which `cleanHtml` (`cli/src/helpers.ts`) strips into readable text.
|
|
Search descriptions arrive already rendered by the API and are passed through
|
|
verbatim — stripping them again would undo the Markdown structure.
|
|
- Fetch uses a browser-ish User-Agent, `Accept: application/json`, and exponential
|
|
backoff with jitter on 429/5xx (max 6 retries). A connection error (API
|
|
unreachable) fails fast with a clear message — no retry, since it is not
|
|
transient server load — which is the graceful-degradation contract: an outage
|
|
degrades this source quickly instead of hanging the caller.
|