Extends the canonical Subfolder-naming rule by citation to all six archive derivation sites (apply, gmail-sync, interview, notion-sync, outcome, assistant SKILL.md), adds a fail-closed guard for an empty derived name, and pins every site with mutation-verified tests. framework_version 1.3.3 -> 1.3.4. jakob1379 independently specified the same fix in his fork's issue #22 before this PR's rework. Co-authored-by: ayobamiseun <66267222+ayobamiseun@users.noreply.github.com>
14 KiB
/notion-sync - Push Ranked Jobs and Applications to a Notion Database
You are publishing a read-only view of the job search into the user's Notion workspace: one database row per job, with a detailed page per shortlisted match. The repo files stay the system of record - job_scraper/seen_jobs.json owns scraped/ranked jobs and job_search_tracker.csv owns applications. Notion is a disposable presentation layer on top of them; nothing ever syncs back.
This command requires the Notion MCP server (OAuth). It reads state, upserts pages, and stops - it never ranks, applies, or edits repo files. Notion is the in-tree reference binding; the sync contract itself is tool-agnostic (see "Adapting to Another Tool" at the end - only the two sections marked (Notion binding) are tool-specific).
Lane: /html-report vs /notion-sync
Both present the same tracker data; they own different moments. /html-report is the deep-review lane: a self-contained offline dashboard with charts and a filterable table, regenerated at your desk. /notion-sync is the glanceable lane: the current state of the pipeline, reachable anywhere Notion runs (desktop, web, phone). They compose rather than compete - after /outcome records a result, re-run either or both to refresh the views.
Follow these steps in order.
Step 0: Parse Input
$ARGUMENTS may contain:
- Nothing → sync ranked jobs with score ≥ 60 (Good Fit and above) plus every tracked application
--min-score <N>→ override the score threshold--all→ sync every ranked job regardless of score--rebuild→ re-fetch and rewrite page bodies too (see Step 5 - normally bodies are write-once)
Step 1: Preflight the Connection (Notion binding)
The command is silently optional: when the destination is not reachable, the outcome is one clear message and a clean exit - nothing else in the framework notices this command exists.
- Check that Notion MCP tools are available in this session (tool names starting with
mcp__notion__or similar). Determine this from the session's own tool list only - never by running shell commands likeclaude mcp list, which would interrupt the user with a permission prompt before the graceful exit. If the tools are not available, stop and tell the user how to connect:Notion MCP isn't connected. Run
claude mcp add --transport http notion https://mcp.notion.com/mcp, then start a new session (servers added mid-session are only picked up on restart), run/mcpthere to complete the OAuth login, and re-run/notion-sync. - Verify the connection with one cheap call (e.g. a workspace search). An auth error → tell the user to re-authenticate via
/mcpand stop. Never retry in a loop. - The Notion MCP server is interactively authenticated, so "connected but not authenticable right now" (expired OAuth, headless/CI context where the login flow cannot run) gets the same graceful exit as "not configured": state the reason in one line and stop. This includes the configured-but-unauthenticated state where the server exposes only its auth handshake and no data tools - never initiate the OAuth flow from this command and never ask whether to authenticate now; the one line points at
/mcpand the command ends there. Authenticating is the user's move, made outside this command.
Step 2: Build the Sync Set (local data only - no external calls yet)
Validate the cheap, local precondition before creating anything external. A run with nothing to sync must exit with zero side effects - no database created, no state file written.
- Read
job_scraper/seen_jobs.jsonandjob_search_tracker.csv(either may be missing). - Select
seen_jobs.jsonentries with statusrankedwhoserank_scoremeets the threshold from Step 0.--alllifts the threshold entirely. - Every tracker row joins the sync set (an applied-to job always syncs, ranked or not), matched to
seen_jobs.jsonentries case-insensitively on company + role where possible. Tracker rows with noseen_jobs.jsonentry sync too - build their Key as<company>_<role>lowercased with underscores. - Status precedence: the tracker wins. A job that is
rankedinseen_jobs.jsonbutinterviewin the tracker syncs asinterview. Jobs only inseen_jobs.jsonkeep their stored status. Deadline precedence: the tracker wins too - the tracker'sdeadline(written by/applyfrom the posting the application was actually built on) overrides theseen_jobs.jsonvalue; jobs only inseen_jobs.jsonkeep the scraper's stored deadline. Omit the property when neither states one, and never reconcile the two by picking the earlier or later date - both were read from the posting at different times, and the safe-lookingmin()substitutes a date the user never applied against. - If the sync set is empty (no ranked entries meet the threshold and there are no tracker rows), say "Nothing to sync - run
/scrapeand/rankfirst" (or, when jobs exist but all score below the threshold, say so and suggest--min-score/--all) and stop. - State the counts before touching the destination: how many rows will be created or checked, and the threshold in effect.
Step 3: Load Sync State and Locate the Database (Notion binding)
-
Read
job_scraper/notion_sync.json. Structure:{ "database_id": "...", "database_url": "...", "last_sync": "YYYY-MM-DD" } -
If it exists, verify the database id still resolves in Notion. If the database was deleted, treat this as a first run.
-
First run: search the workspace for a database named "Job Search Pipeline". If none exists, ask the user where to create it (top-level page or an existing page they name), then create it with exactly these properties:
Property Type Values / notes Name title <Role> — <Company>Company rich text Score number 0-100 from rank_scoreVerdict select Strong Fit / Good Fit / Moderate Fit / Weak Fit / Poor Fit Status select ranked/drafted/applied/interview/offer/hired/rejected/no_response/offer_declined/withdrawn/expired— canonical tracker spellings per Tracker status vocabulary in/outcome; Notion options grow to match as values appearFit select high / medium / low (scraper quick-fit) Deadline date tracker deadlinecolumn, falling back toseen_jobs.json'sdeadlinewhen the row has none; omit when neither states oneFirst seen date Ranked date rank_datefromseen_jobs.json; omit when not rankedApplied on date tracker datecolumn; omit when not in the tracker, and omit when the status isdraftedChannel select tracker channelcolumn (e.g. portal / email / referral); options grow as values appearCV file rich text tracker cv_filecolumn - the filename only, never document contentCover letter rich text tracker cover_letter_filecolumn - the filename only, never document contentURL url posting URL Key rich text the job's key in seen_jobs.json- dedup anchor, never edited by handThe tracker-sourced properties (Applied on, Channel, CV file, Cover letter) stay empty for jobs that have no tracker row. CV file and Cover letter fill in once
/applyrecords the draft; Applied on stays empty until/outcomerecords the submission. Only filenames ever sync; document contents stay local. -
Existing database with missing properties: if the located database predates a schema addition (a property from the table above does not exist), add the missing properties to the database before upserting. Never remove or retype existing properties.
-
Write
job_scraper/notion_sync.jsonwith the database id and URL. This file is personal state and is gitignored - never commit it.
Step 4: Upsert Database Rows
For each job in the sync set:
- Query the database for a page whose
Keyequals the job's key. - No match → create the page with all properties from the Step 3 table, then write its body (Step 5).
- Match → update properties only: Status, Score, Verdict, Deadline, Ranked, Applied on, Channel, CV file, Cover letter. Properties are the always-current surface (bodies are write-once), so tracker updates recorded by
/outcomereach the destination exclusively through them. Do not touch the page body - the user may have added their own notes there, and clobbering them breaks trust in the whole view. (--rebuildis the sole exception.) - Never delete or archive pages, even for jobs that turned
expired- set Status toexpiredinstead. Rows the user added to the database by hand (noKeyvalue) are invisible to this command.
Normalise the Status value before writing. The tracker may hold legacy space spellings (no response, offer declined) from before the canonical forms were locked. Map them to no_response / offer_declined per the Tracker status vocabulary in /outcome before setting Status on create or update - never push a space form to Notion, which would auto-create a separate select option per unique string. Pre-existing space-form options in an existing database simply go unused; Notion never auto-removes select options.
Batch politely: if the MCP server rate-limits, back off and continue; report any page that failed rather than retrying indefinitely.
Step 5: Write the Detail Page (new pages only)
The page body is what makes a row worth clicking. Build it only from stored data and actually fetched content:
- Fit summary - a short section from
seen_jobs.jsonfields: score, verdict, quick-fit level, first-seen and ranked dates. If the job is in the tracker, add the application timeline (date applied, channel, current status, dated notes from thenotescolumn) and name the submitted documents fromcv_file/cover_letter_file(filenames only - the documents themselves never sync). When the status isdrafted, write "drafted YYYY-MM-DD, not yet submitted" instead of a date applied, and call the files drafts rather than submitted documents (page bodies are write-once - Step 4.3). - The posting - WebFetch the job URL and write a readable digest: what the role is, key requirements, practical details (location, deadline, salary if stated). Retry a 403 with browser headers per
.claude/skills/job-application-assistant/09-web-research.mdfirst. If the fetch still fails or redirects to a listing page, write "Posting no longer available (checked YYYY-MM-DD)" - never reconstruct a posting from memory. - Links - the posting URL; derive
<company>_<role>by the Subfolder naming rule indocuments/README.md, and if that archive exists locally, name its path (plain text - the destination cannot link into the filesystem).
Keep the page under ~40 blocks; this is a briefing, not a mirror of the posting.
Step 6: Report
## Pipeline Sync - YYYY-MM-DD
Database: <database_url>
Synced <N> jobs (threshold: score ≥ <T>): <C> created, <U> updated, <S> unchanged, <F> failed.
| | Title | Company | Status | Score |
|---|-------|---------|--------|-------|
| ✚ | ... | ... | ranked | 78 |
| ↻ | ... | ... | interview | 71 |
List failures with their error and the suggestion to re-run - the upsert is idempotent, so a re-run only touches what failed. Update last_sync in the sync-state file.
Remind the user once (first run only): the repo files remain the source of truth - edits made in the destination never flow back, and /outcome is still how application results get recorded.
Important Rules
- One-way, always. Destination content never flows back into
seen_jobs.json, the tracker, or any repo file. This command reads repo state and writes the destination - both repo files are read-only to it, and the gitignored sync-state file is its only local write. - Idempotent upsert on
Key. Re-running creates nothing twice; matching is on the stored key, never on fuzzy title comparison. - Page bodies are write-once. Property updates keep rows current; bodies belong to the user after creation. Only
--rebuildmay rewrite them, and it says so before doing it. - Never fabricate. A dead posting URL gets an explicit "unavailable" note, not a reconstruction. Every page claim traces to stored state or fetched content.
- Job data only. The candidate profile, behavioral notes, and evaluation framework never sync - this is a pipeline view, not a profile export.
- Documents never leave the machine. CVs and cover letters sync as filenames only - never upload, attach, or embed the documents themselves, nor HTML/text renditions of their content, into the destination. The local repo and
documents/applications/archive are the only home for application documents; the row's page names them so the user knows what to open locally.
Adapting to Another Tool (forks)
The sync contract is tool-agnostic; only the two sections marked (Notion binding) are tool-specific. A fork targeting a different destination (Airtable, Google Sheets, Linear, ...) keeps Steps 0, 2, 4, 5, and 6 and every Important Rule unchanged - build the same sync set, upsert on the same Key, keep bodies write-once and documents local - and swaps only:
- Step 1 (connection preflight) for the target tool's MCP server or access check, keeping the silently-optional bar: not configured or not authenticable both end in one message and a clean exit
- Step 3 (locate/create the database) for the equivalent container in the target tool, using the same property table and a renamed sync-state file
Like the portal skills, tool bindings beyond this Notion reference live in forks, where their maintainers can test them against a live workspace.