mirror of
https://github.com/MadsLorentzen/ai-job-search.git
synced 2026-09-17 00:26:26 +00:00
fix(web-research): stop treating a WebFetch 403 as a dead posting (#277)
* fix(web-research): stop treating a WebFetch 403 as a dead posting WebFetch sends a bot user agent, and many bank and corporate sites answer with HTTP 403 while serving the same page to a browser normally. Every command treated that as "page unavailable" and degraded silently rather than failing loudly: - /rank marked live postings `expired` - /apply fell back to search snippets, or to vague cover-letter prose - /scrape stored listing-page `#fragment` URLs, which fetch fine and return unrelated jobs, so every later /rank and /apply run on that entry failed Adds 09-web-research.md as the single reference: the trust boundary, a curl browser-header retry with a tag-stripping extractor, a four-step escalation order, the login-wall case, why the employer's own careers posting beats an aggregator listing (the requisition ID and the grade survive there), and the rule that a search-result snippet is a lead rather than a source. Wires it into /apply, /rank, /interview, /outcome, /notion-sync, the job-scraper skill, and writing-style rule 5. Bumps 03-writing-style.md to 1.2.0; 09-web-research.md starts at 1.0.0. Aggregator examples are given generically (LinkedIn, Indeed, national job boards) so the guidance holds in any market. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web-research): gate the browser-header retry on robots.txt Addresses review feedback on #277. WebFetch identifies itself as Claude-User and honors robots.txt, so a 403 has two very different causes and they must not be treated the same: a WAF default on a site whose published policy allows access, or a site that has actually declined. Retrying with browser headers in the second case circumvents the very opt-out mechanism site owners are told they can rely on, and the core framework cannot hold a looser standard than it asks of community forks. The escalation now runs tools/robots_check.py before the retry. A disallow for "*" or for "Claude-User" skips the retry entirely and goes to step 3 (find the employer's own posting). The rule is stated plainly in 09-web-research.md so later edits do not erode it: the retry exists to get past bot-filtering firewalls on sites whose robots.txt permits access; it is never used to override a site that has said no. Two findings from testing the gate against live sites, both pinned by tests/test_robots_check.py (15 offline cases): - The WAF usually blocks robots.txt too. privatebank.barclays.com returns 403 on the policy file to Claude-User and 200 to a browser, so a naive gate would block the retry on exactly the sites the retry is for. The checker reads the policy as a browser when the honest request is refused, then obeys it strictly - a policy you are prevented from reading cannot be honored, and robots.txt is not the protected resource. - urllib.robotparser cannot be used. It ends a record at a blank line and matches rules in file order, so Barclays' real file (blank lines between "User-agent: *" and its rules, "Allow: /" before "Disallow: /cs/") reads as everything-allowed. That fails open, in the one direction that matters. The checker implements RFC 9309 longest-match instead, with ties resolved to Disallow rather than Allow. Verified live: barclays /careers/ allowed and /cs/ blocked, ubs.com allowed, jobup.ch /api/ blocked while /en/jobs/ stays allowed. 09-web-research.md 1.0.0 to 1.1.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: kgb <kevingblackman@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
kgb
parent
9aea6e7a44
commit
fcefb8150f
@@ -21,6 +21,36 @@ per-file diff commands.
|
||||
during #275 (vetoes reported in console output but `language_gate: null` on every persisted
|
||||
entry). Mirrors the existing `gaps`/`strengths` pinning pattern. No behavior change.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **A `WebFetch` 403 is no longer treated as a dead posting** - `WebFetch` sends a bot user
|
||||
agent, and many bank and corporate sites answer it with HTTP 403 while serving the same
|
||||
page to a browser normally. Every command read that as "page unavailable" and degraded
|
||||
silently instead of failing loudly: `/rank` marked live postings `expired`, `/apply` fell
|
||||
back to search-result snippets or to vague cover-letter prose, and `/scrape` stored
|
||||
listing-page `#fragment` URLs that fetch fine but return unrelated jobs, breaking every
|
||||
later run on that entry. New `09-web-research.md` (`framework_version` 1.0.0) is the
|
||||
single reference: the trust boundary, a curl browser-header retry with a tag-stripping
|
||||
extractor, a four-step escalation order, the login-wall case, why the employer's own
|
||||
careers posting beats an aggregator listing (the requisition ID and the grade survive
|
||||
there), and the rule that a search snippet is a lead rather than a source. Wired into
|
||||
`/apply`, `/rank`, `/interview`, `/outcome`, `/notion-sync`, the job-scraper skill, and
|
||||
writing-style rule 5 (`03-writing-style.md` 1.1.0 to 1.2.0).
|
||||
|
||||
**The retry is gated on `robots.txt`.** `WebFetch` identifies itself as `Claude-User`
|
||||
and honors `robots.txt`, so a 403 means either a WAF default on a site whose published
|
||||
policy allows access, or a site that has actually declined. New `tools/robots_check.py`
|
||||
tells them apart and the escalation runs it before retrying: a disallow for `*` or
|
||||
`Claude-User` skips the retry entirely and goes straight to finding the employer's own
|
||||
posting. The rule is stated in the file so later edits do not erode it - *the retry
|
||||
exists to get past bot-filtering firewalls on sites whose robots.txt permits access; it
|
||||
is never used to override a site that has said no.* Two findings are pinned by
|
||||
`tests/test_robots_check.py` (15 offline cases): the WAF usually blocks `robots.txt`
|
||||
itself, so the policy is read as a browser when the honest request is refused and then
|
||||
obeyed strictly; and `urllib.robotparser` cannot be used, because it ends a record at a
|
||||
blank line and matches in file order, which reads a real-world policy as
|
||||
"everything allowed".
|
||||
|
||||
## [1.3.0] - 2026-08-03
|
||||
|
||||
### Added
|
||||
|
||||
Reference in New Issue
Block a user