fix(web-research): stop treating a WebFetch 403 as a dead posting (#277)

* fix(web-research): stop treating a WebFetch 403 as a dead posting

WebFetch sends a bot user agent, and many bank and corporate sites answer
with HTTP 403 while serving the same page to a browser normally. Every
command treated that as "page unavailable" and degraded silently rather
than failing loudly:

- /rank marked live postings `expired`
- /apply fell back to search snippets, or to vague cover-letter prose
- /scrape stored listing-page `#fragment` URLs, which fetch fine and
  return unrelated jobs, so every later /rank and /apply run on that
  entry failed

Adds 09-web-research.md as the single reference: the trust boundary, a
curl browser-header retry with a tag-stripping extractor, a four-step
escalation order, the login-wall case, why the employer's own careers
posting beats an aggregator listing (the requisition ID and the grade
survive there), and the rule that a search-result snippet is a lead
rather than a source.

Wires it into /apply, /rank, /interview, /outcome, /notion-sync, the
job-scraper skill, and writing-style rule 5. Bumps 03-writing-style.md
to 1.2.0; 09-web-research.md starts at 1.0.0.

Aggregator examples are given generically (LinkedIn, Indeed, national
job boards) so the guidance holds in any market.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web-research): gate the browser-header retry on robots.txt

Addresses review feedback on #277.

WebFetch identifies itself as Claude-User and honors robots.txt, so a 403 has
two very different causes and they must not be treated the same: a WAF default
on a site whose published policy allows access, or a site that has actually
declined. Retrying with browser headers in the second case circumvents the very
opt-out mechanism site owners are told they can rely on, and the core framework
cannot hold a looser standard than it asks of community forks.

The escalation now runs tools/robots_check.py before the retry. A disallow for
"*" or for "Claude-User" skips the retry entirely and goes to step 3 (find the
employer's own posting). The rule is stated plainly in 09-web-research.md so
later edits do not erode it: the retry exists to get past bot-filtering
firewalls on sites whose robots.txt permits access; it is never used to
override a site that has said no.

Two findings from testing the gate against live sites, both pinned by
tests/test_robots_check.py (15 offline cases):

- The WAF usually blocks robots.txt too. privatebank.barclays.com returns 403
  on the policy file to Claude-User and 200 to a browser, so a naive gate would
  block the retry on exactly the sites the retry is for. The checker reads the
  policy as a browser when the honest request is refused, then obeys it
  strictly - a policy you are prevented from reading cannot be honored, and
  robots.txt is not the protected resource.
- urllib.robotparser cannot be used. It ends a record at a blank line and
  matches rules in file order, so Barclays' real file (blank lines between
  "User-agent: *" and its rules, "Allow: /" before "Disallow: /cs/") reads as
  everything-allowed. That fails open, in the one direction that matters. The
  checker implements RFC 9309 longest-match instead, with ties resolved to
  Disallow rather than Allow.

Verified live: barclays /careers/ allowed and /cs/ blocked, ubs.com allowed,
jobup.ch /api/ blocked while /en/jobs/ stays allowed. 09-web-research.md
1.0.0 to 1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: kgb <kevingblackman@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kblackma
2026-08-04 20:36:55 +02:00
committed by GitHub
co-authored by Claude Opus 5 kgb
parent 9aea6e7a44
commit fcefb8150f
12 changed files with 384 additions and 11 deletions
+30
View File
@@ -21,6 +21,36 @@ per-file diff commands.
during #275 (vetoes reported in console output but `language_gate: null` on every persisted
entry). Mirrors the existing `gaps`/`strengths` pinning pattern. No behavior change.
### Fixed
- **A `WebFetch` 403 is no longer treated as a dead posting** - `WebFetch` sends a bot user
agent, and many bank and corporate sites answer it with HTTP 403 while serving the same
page to a browser normally. Every command read that as "page unavailable" and degraded
silently instead of failing loudly: `/rank` marked live postings `expired`, `/apply` fell
back to search-result snippets or to vague cover-letter prose, and `/scrape` stored
listing-page `#fragment` URLs that fetch fine but return unrelated jobs, breaking every
later run on that entry. New `09-web-research.md` (`framework_version` 1.0.0) is the
single reference: the trust boundary, a curl browser-header retry with a tag-stripping
extractor, a four-step escalation order, the login-wall case, why the employer's own
careers posting beats an aggregator listing (the requisition ID and the grade survive
there), and the rule that a search snippet is a lead rather than a source. Wired into
`/apply`, `/rank`, `/interview`, `/outcome`, `/notion-sync`, the job-scraper skill, and
writing-style rule 5 (`03-writing-style.md` 1.1.0 to 1.2.0).
**The retry is gated on `robots.txt`.** `WebFetch` identifies itself as `Claude-User`
and honors `robots.txt`, so a 403 means either a WAF default on a site whose published
policy allows access, or a site that has actually declined. New `tools/robots_check.py`
tells them apart and the escalation runs it before retrying: a disallow for `*` or
`Claude-User` skips the retry entirely and goes straight to finding the employer's own
posting. The rule is stated in the file so later edits do not erode it - *the retry
exists to get past bot-filtering firewalls on sites whose robots.txt permits access; it
is never used to override a site that has said no.* Two findings are pinned by
`tests/test_robots_check.py` (15 offline cases): the WAF usually blocks `robots.txt`
itself, so the policy is read as a browser when the honest request is refused and then
obeyed strictly; and `urllib.robotparser` cannot be used, because it ends a record at a
blank line and matches in file order, which reads a real-world policy as
"everything allowed".
## [1.3.0] - 2026-08-03
### Added