mirror of
https://github.com/MadsLorentzen/ai-job-search.git
synced 2026-09-17 00:26:26 +00:00
* fix(scrape): persist each posting's publication date in seen_jobs.json (#390) Step 2's contract guarantees a `date` on every portal CLI's search output and CI enforces it in test_scrape_contract.py; Step 3 uses that date to scope a run to the last 14 days. Step 4's storage schema then dropped it, so a posting's age was unrecoverable the moment the run ended - `first_seen` records when the scraper saw an entry, not when the employer posted it. /rank reads the stored entry rather than the run, so it had no age signal to weigh. A freehire-search posting dated 2024-05-13 was scraped 27 months later and ranked Strong Fit at position 1 of 133. The scoring note observed the listing "may be long stale" in prose nothing reads, and an /apply run drafted a tailored CV and cover letter against it. The schema gains `posted_date` (null when the portal returned no date, never inferred or backfilled), documented alongside `deadline` with the same never-backfill rule. Three new cases, each verified to fail on the unfixed spec. Closes #390 * fix(scrape): correct the 14-day scoping cross-reference, restore EOF newline Review follow-up on #391. The 14-day scoping is Step 1b's list item 3, not Step 3 - Step 3 is Quick Fit Assessment and never touches dates. The "3." list item had been promoted to a step number. Corrected in the new SKILL.md paragraph (both occurrences), the CHANGELOG entry, and the test class docstring; a wrong pointer in a file agents execute as instructions actively misleads. Also restores the trailing newline on tests/test_scrape_contract.py (the nit left for a future touch in #344) and adds the (#390) ref to the CHANGELOG entry to match its siblings.
This commit is contained in:
@@ -147,6 +147,7 @@ For each new job, do a rapid fit check (NOT the full evaluation from `04-job-eva
|
||||
"company": "...",
|
||||
"url": "...",
|
||||
"first_seen": "YYYY-MM-DD",
|
||||
"posted_date": "YYYY-MM-DD" | null,
|
||||
"deadline": "YYYY-MM-DD" | null,
|
||||
"fit": "high/medium/low",
|
||||
"status": "new/skipped/ranked/expired",
|
||||
@@ -165,6 +166,8 @@ The `source` field records which mechanism produced the entry: `cli` for Step 1b
|
||||
|
||||
`deadline` is a base field rather than a `/rank` extension: Step 2's detail fetch already extracts the application deadline, so it is written when the job is first seen and refreshed by `/rank` Step 4 when a scoring agent returns a different value. `null` means the posting states no deadline; a missing key means the entry predates this field - **never infer a deadline** from either, and never backfill by guessing.
|
||||
|
||||
`posted_date` is the posting's own publication date, taken from the `date` field Step 2's contract already guarantees on every portal CLI's search output. Step 1b uses that date to scope the run to the last 14 days and then drops it, so nothing downstream can distinguish a posting published yesterday from one published two years ago - `first_seen` is when this scraper first saw the entry, not when the employer posted it. Persisting it makes Step 1b's window auditable after the run and gives `/rank` a freshness signal to weigh, instead of rediscovering the date and recording it in prose that nothing reads. That gap landed for real: a freehire-search posting dated 2024-05-13 was scraped and ranked Strong Fit at position 1 of 133, its own scoring note observing the listing "may be long stale" with nothing able to act on it. `null` means the portal returned no date for that result (the CLIs emit `date: null` when a listing omits it); a missing key means the entry predates this field - **never infer a posting date** from either, and never backfill by guessing.
|
||||
|
||||
2. Only present jobs NOT already in the seen list or tracker.
|
||||
|
||||
### Step 4.5: Generate Referral Contact Links (High & Medium Fit Only)
|
||||
|
||||
Reference in New Issue
Block a user