Files
ai-job-search/.claude/commands/notion-sync.md
T
fcefb8150f fix(web-research): stop treating a WebFetch 403 as a dead posting (#277)
* fix(web-research): stop treating a WebFetch 403 as a dead posting

WebFetch sends a bot user agent, and many bank and corporate sites answer
with HTTP 403 while serving the same page to a browser normally. Every
command treated that as "page unavailable" and degraded silently rather
than failing loudly:

- /rank marked live postings `expired`
- /apply fell back to search snippets, or to vague cover-letter prose
- /scrape stored listing-page `#fragment` URLs, which fetch fine and
  return unrelated jobs, so every later /rank and /apply run on that
  entry failed

Adds 09-web-research.md as the single reference: the trust boundary, a
curl browser-header retry with a tag-stripping extractor, a four-step
escalation order, the login-wall case, why the employer's own careers
posting beats an aggregator listing (the requisition ID and the grade
survive there), and the rule that a search-result snippet is a lead
rather than a source.

Wires it into /apply, /rank, /interview, /outcome, /notion-sync, the
job-scraper skill, and writing-style rule 5. Bumps 03-writing-style.md
to 1.2.0; 09-web-research.md starts at 1.0.0.

Aggregator examples are given generically (LinkedIn, Indeed, national
job boards) so the guidance holds in any market.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web-research): gate the browser-header retry on robots.txt

Addresses review feedback on #277.

WebFetch identifies itself as Claude-User and honors robots.txt, so a 403 has
two very different causes and they must not be treated the same: a WAF default
on a site whose published policy allows access, or a site that has actually
declined. Retrying with browser headers in the second case circumvents the very
opt-out mechanism site owners are told they can rely on, and the core framework
cannot hold a looser standard than it asks of community forks.

The escalation now runs tools/robots_check.py before the retry. A disallow for
"*" or for "Claude-User" skips the retry entirely and goes to step 3 (find the
employer's own posting). The rule is stated plainly in 09-web-research.md so
later edits do not erode it: the retry exists to get past bot-filtering
firewalls on sites whose robots.txt permits access; it is never used to
override a site that has said no.

Two findings from testing the gate against live sites, both pinned by
tests/test_robots_check.py (15 offline cases):

- The WAF usually blocks robots.txt too. privatebank.barclays.com returns 403
  on the policy file to Claude-User and 200 to a browser, so a naive gate would
  block the retry on exactly the sites the retry is for. The checker reads the
  policy as a browser when the honest request is refused, then obeys it
  strictly - a policy you are prevented from reading cannot be honored, and
  robots.txt is not the protected resource.
- urllib.robotparser cannot be used. It ends a record at a blank line and
  matches rules in file order, so Barclays' real file (blank lines between
  "User-agent: *" and its rules, "Allow: /" before "Disallow: /cs/") reads as
  everything-allowed. That fails open, in the one direction that matters. The
  checker implements RFC 9309 longest-match instead, with ties resolved to
  Disallow rather than Allow.

Verified live: barclays /careers/ allowed and /cs/ blocked, ubs.com allowed,
jobup.ch /api/ blocked while /en/jobs/ stays allowed. 09-web-research.md
1.0.0 to 1.1.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: kgb <kevingblackman@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 20:36:55 +02:00

12 KiB

/notion-sync - Push Ranked Jobs and Applications to a Notion Database

You are publishing a read-only view of the job search into the user's Notion workspace: one database row per job, with a detailed page per shortlisted match. The repo files stay the system of record - job_scraper/seen_jobs.json owns scraped/ranked jobs and job_search_tracker.csv owns applications. Notion is a disposable presentation layer on top of them; nothing ever syncs back.

This command requires the Notion MCP server (OAuth). It reads state, upserts pages, and stops - it never ranks, applies, or edits repo files. Notion is the in-tree reference binding; the sync contract itself is tool-agnostic (see "Adapting to Another Tool" at the end - only the two sections marked (Notion binding) are tool-specific).

Lane: /html-report vs /notion-sync

Both present the same tracker data; they own different moments. /html-report is the deep-review lane: a self-contained offline dashboard with charts and a filterable table, regenerated at your desk. /notion-sync is the glanceable lane: the current state of the pipeline, reachable anywhere Notion runs (desktop, web, phone). They compose rather than compete - after /outcome records a result, re-run either or both to refresh the views.

Follow these steps in order.


Step 0: Parse Input

$ARGUMENTS may contain:

  • Nothing → sync ranked jobs with score ≥ 60 (Good Fit and above) plus every tracked application
  • --min-score <N> → override the score threshold
  • --all → sync every ranked job regardless of score
  • --rebuild → re-fetch and rewrite page bodies too (see Step 5 - normally bodies are write-once)

Step 1: Preflight the Connection (Notion binding)

The command is silently optional: when the destination is not reachable, the outcome is one clear message and a clean exit - nothing else in the framework notices this command exists.

  1. Check that Notion MCP tools are available in this session (tool names starting with mcp__notion__ or similar). Determine this from the session's own tool list only - never by running shell commands like claude mcp list, which would interrupt the user with a permission prompt before the graceful exit. If the tools are not available, stop and tell the user how to connect:

    Notion MCP isn't connected. Run claude mcp add --transport http notion https://mcp.notion.com/mcp, then start a new session (servers added mid-session are only picked up on restart), run /mcp there to complete the OAuth login, and re-run /notion-sync.

  2. Verify the connection with one cheap call (e.g. a workspace search). An auth error → tell the user to re-authenticate via /mcp and stop. Never retry in a loop.
  3. The Notion MCP server is interactively authenticated, so "connected but not authenticable right now" (expired OAuth, headless/CI context where the login flow cannot run) gets the same graceful exit as "not configured": state the reason in one line and stop. This includes the configured-but-unauthenticated state where the server exposes only its auth handshake and no data tools - never initiate the OAuth flow from this command and never ask whether to authenticate now; the one line points at /mcp and the command ends there. Authenticating is the user's move, made outside this command.

Step 2: Build the Sync Set (local data only - no external calls yet)

Validate the cheap, local precondition before creating anything external. A run with nothing to sync must exit with zero side effects - no database created, no state file written.

  1. Read job_scraper/seen_jobs.json and job_search_tracker.csv (either may be missing).
  2. Select seen_jobs.json entries with status ranked whose rank_score meets the threshold from Step 0. --all lifts the threshold entirely.
  3. Every tracker row joins the sync set (an applied-to job always syncs, ranked or not), matched to seen_jobs.json entries case-insensitively on company + role where possible. Tracker rows with no seen_jobs.json entry sync too - build their Key as <company>_<role> lowercased with underscores.
  4. Status precedence: the tracker wins. A job that is ranked in seen_jobs.json but interview in the tracker syncs as interview. Jobs only in seen_jobs.json keep their stored status.
  5. If the sync set is empty (no ranked entries meet the threshold and there are no tracker rows), say "Nothing to sync - run /scrape and /rank first" (or, when jobs exist but all score below the threshold, say so and suggest --min-score/--all) and stop.
  6. State the counts before touching the destination: how many rows will be created or checked, and the threshold in effect.

Step 3: Load Sync State and Locate the Database (Notion binding)

  1. Read job_scraper/notion_sync.json. Structure:

    { "database_id": "...", "database_url": "...", "last_sync": "YYYY-MM-DD" }
    
  2. If it exists, verify the database id still resolves in Notion. If the database was deleted, treat this as a first run.

  3. First run: search the workspace for a database named "Job Search Pipeline". If none exists, ask the user where to create it (top-level page or an existing page they name), then create it with exactly these properties:

    Property Type Values / notes
    Name title <Role> — <Company>
    Company rich text
    Score number 0-100 from rank_score
    Verdict select Strong Fit / Good Fit / Moderate Fit / Weak Fit / Poor Fit
    Status select ranked / applied / interview / offer / hired / rejected / no response / withdrawn / expired
    Fit select high / medium / low (scraper quick-fit)
    Deadline date omit when unknown
    First seen date
    Ranked date rank_date from seen_jobs.json; omit when not ranked
    Applied on date tracker date column; omit when not in the tracker
    Channel select tracker channel column (e.g. portal / email / referral); options grow as values appear
    CV file rich text tracker cv_file column - the filename only, never document content
    Cover letter rich text tracker cover_letter_file column - the filename only, never document content
    URL url posting URL
    Key rich text the job's key in seen_jobs.json - dedup anchor, never edited by hand

    The tracker-sourced properties (Applied on, Channel, CV file, Cover letter) stay empty for jobs that have no tracker row - they fill in once /outcome records the application. Only filenames ever sync; document contents stay local.

  4. Existing database with missing properties: if the located database predates a schema addition (a property from the table above does not exist), add the missing properties to the database before upserting. Never remove or retype existing properties.

  5. Write job_scraper/notion_sync.json with the database id and URL. This file is personal state and is gitignored - never commit it.


Step 4: Upsert Database Rows

For each job in the sync set:

  1. Query the database for a page whose Key equals the job's key.
  2. No match → create the page with all properties from the Step 3 table, then write its body (Step 5).
  3. Match → update properties only: Status, Score, Verdict, Deadline, Ranked, Applied on, Channel, CV file, Cover letter. Properties are the always-current surface (bodies are write-once), so tracker updates recorded by /outcome reach the destination exclusively through them. Do not touch the page body - the user may have added their own notes there, and clobbering them breaks trust in the whole view. (--rebuild is the sole exception.)
  4. Never delete or archive pages, even for jobs that turned expired - set Status to expired instead. Rows the user added to the database by hand (no Key value) are invisible to this command.

Batch politely: if the MCP server rate-limits, back off and continue; report any page that failed rather than retrying indefinitely.


Step 5: Write the Detail Page (new pages only)

The page body is what makes a row worth clicking. Build it only from stored data and actually fetched content:

  1. Fit summary - a short section from seen_jobs.json fields: score, verdict, quick-fit level, first-seen and ranked dates. If the job is in the tracker, add the application timeline (date applied, channel, current status, dated notes from the notes column) and name the submitted documents from cv_file/cover_letter_file (filenames only - the documents themselves never sync).
  2. The posting - WebFetch the job URL and write a readable digest: what the role is, key requirements, practical details (location, deadline, salary if stated). Retry a 403 with browser headers per .claude/skills/job-application-assistant/09-web-research.md first. If the fetch still fails or redirects to a listing page, write "Posting no longer available (checked YYYY-MM-DD)" - never reconstruct a posting from memory.
  3. Links - the posting URL; if documents/applications/<company>_<role>/ exists locally, name it as the local archive path (plain text - the destination cannot link into the filesystem).

Keep the page under ~40 blocks; this is a briefing, not a mirror of the posting.


Step 6: Report

## Pipeline Sync - YYYY-MM-DD

Database: <database_url>
Synced <N> jobs (threshold: score ≥ <T>): <C> created, <U> updated, <S> unchanged, <F> failed.

| | Title | Company | Status | Score |
|---|-------|---------|--------|-------|
| ✚ | ... | ... | ranked | 78 |
| ↻ | ... | ... | interview | 71 |

List failures with their error and the suggestion to re-run - the upsert is idempotent, so a re-run only touches what failed. Update last_sync in the sync-state file.

Remind the user once (first run only): the repo files remain the source of truth - edits made in the destination never flow back, and /outcome is still how application results get recorded.


Important Rules

  1. One-way, always. Destination content never flows back into seen_jobs.json, the tracker, or any repo file. This command reads repo state and writes the destination - both repo files are read-only to it, and the gitignored sync-state file is its only local write.
  2. Idempotent upsert on Key. Re-running creates nothing twice; matching is on the stored key, never on fuzzy title comparison.
  3. Page bodies are write-once. Property updates keep rows current; bodies belong to the user after creation. Only --rebuild may rewrite them, and it says so before doing it.
  4. Never fabricate. A dead posting URL gets an explicit "unavailable" note, not a reconstruction. Every page claim traces to stored state or fetched content.
  5. Job data only. The candidate profile, behavioral notes, and evaluation framework never sync - this is a pipeline view, not a profile export.
  6. Documents never leave the machine. CVs and cover letters sync as filenames only - never upload, attach, or embed the documents themselves, nor HTML/text renditions of their content, into the destination. The local repo and documents/applications/ archive are the only home for application documents; the row's page names them so the user knows what to open locally.

Adapting to Another Tool (forks)

The sync contract is tool-agnostic; only the two sections marked (Notion binding) are tool-specific. A fork targeting a different destination (Airtable, Google Sheets, Linear, ...) keeps Steps 0, 2, 4, 5, and 6 and every Important Rule unchanged - build the same sync set, upsert on the same Key, keep bodies write-once and documents local - and swaps only:

  • Step 1 (connection preflight) for the target tool's MCP server or access check, keeping the silently-optional bar: not configured or not authenticable both end in one message and a clean exit
  • Step 3 (locate/create the database) for the equivalent container in the target tool, using the same property table and a renamed sync-state file

Like the portal skills, tool bindings beyond this Notion reference live in forks, where their maintainers can test them against a live workspace.