Files
ai-job-search/.claude/skills/job-scraper/SKILL.md
T
Mads LorentzenandClaude Fable 5 a5ffcc39ff chore: untrack tracker CSV, scope scraper Bash permission, fix portal SKILL.md paths (#71)
- Untrack job_search_tracker.csv: it was both tracked and listed in
  .gitignore (same inconsistency class as the settings.local.json fix
  in #27). Users' personal rows risked merge conflicts on every pull;
  commands already create the file with the standard header when it
  is missing.
- Scope job-scraper's allowed-tools Bash entry (from #52) to
  'bun --version' and the portal-CLI invocation pattern, adopting the
  tighter form proposed in #65.
- Fix all five portal SKILL.mds documenting 'bun run skills/...'
  paths that do not resolve from the repo root ('.agents/skills/...'
  is correct) - now load-bearing since #52 wired /scrape to read
  these docs for CLI invocations. Surfaced in #66.
- Teach tools/lint_skills.py to glob-expand allowed-tools bun run
  targets so scoped wildcard permissions lint correctly.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 17:14:39 +02:00

6.4 KiB

name, description, allowed-tools
name description allowed-tools
job-scraper Scrapes Danish job sites for new positions matching your profile. Deduplicates across runs. Triggers on: job scrape, find jobs, search jobs, new jobs, job search, scrape jobs, /scrape Read, Write, Edit, Glob, Grep, Bash(bun --version), Bash(bun run .agents/skills/*/cli/src/cli.ts *), WebFetch, WebSearch, Agent, AskUserQuestion

Job Scraper


How It Works

This skill searches multiple Danish job sites using targeted queries based on your profile, deduplicates against previously seen jobs and the application tracker, and presents new matches with a quick fit assessment.

Invocation

The user triggers this skill by saying things like:

  • "Find new jobs"
  • "Scrape for jobs"
  • "Any new positions?"
  • "/scrape"

Optional arguments:

  • A focus area, e.g. "/scrape data science" or "/scrape geophysics"
  • "broad" to run all search categories, e.g. "/scrape broad"

Execution Steps

Step 0: Load State

  1. Read job_scraper/seen_jobs.json (create if missing - start with {"seen": {}})
  2. Read job_search_tracker.csv to extract already-applied companies+roles
  3. Read search-queries.md (this directory) for the search strategy

Read search-queries.md (this directory) for the search strategy. By default, run the top 3 priority query categories. If the user said "broad", run all categories. If the user specified a focus area (e.g. "data science"), prioritize queries from that category.

Use the installed CLI tools as the primary search mechanism. Fall back to WebSearch only for portals that do not have a CLI skill, or if bun is unavailable on the system.

1a. Check bun availability

bun --version

If this fails (bun not installed), skip to 1c (WebSearch fallback) for all portals and note the fallback in the Step 5 output.

1b. Run CLI tools (primary — run these in parallel where possible)

Discover all installed portal CLI skills by reading every SKILL.md found under .agents/skills/*/SKILL.md. Each file documents that portal's exact CLI flags and usage examples. Use each portal's own documented interface — do not guess flags. This approach automatically includes any new portals added via /add-portal without requiring changes to this file.

For each installed portal skill:

  1. Read its SKILL.md to find the correct bun run … invocation and supported flags.
  2. Translate the query terms from search-queries.md into that portal's flag format (e.g. --key, --search-string, --query, filter codes — whatever the portal's SKILL.md specifies).
  3. Scope to the last 14 days using the portal's supported recency flag (--jobage, --since <YYYY-MM-DD>, --order PublicationDate, etc. — as documented per portal).
  4. Cap results to ~20 per call using the portal's limit flag.
  5. Use --format json for machine-readable output.

Run all portal CLI calls in parallel where possible using the Agent tool. Collect all results arrays into a single pool for Step 2.

If a CLI tool exits with a non-zero code, log the error message and continue — do not abort the whole search.

1c. WebSearch fallback

Use WebSearch for:

  • Portals listed in search-queries.md that do not have a corresponding directory under .agents/skills/
  • Any portal whose CLI fails at runtime
  • When bun is unavailable (Step 1a failed)

Use the site-specific query strings from search-queries.md directly as WebSearch queries for these portals.

Step 2: Fetch & Parse

For each promising result from Step 1:

  • Use WebFetch to retrieve the job posting page
  • Extract: job title, company, location, posting date (or "recent"), URL, key requirements (brief), application deadline (if listed)
  • Skip if the URL or company+title combo already exists in seen_jobs.json
  • Skip if the company+role already appears in job_search_tracker.csv

Step 3: Quick Fit Assessment

For each new job, do a rapid fit check (NOT the full evaluation from 04-job-evaluation.md - just a quick signal):

  • High match: Role directly involves your core skills
  • Medium match: Role is adjacent to your experience
  • Low match: Role requires significant skills you lack

Step 4: Deduplicate & Store

  1. Add ALL fetched jobs (new and skipped) to seen_jobs.json with structure:
{
  "seen": {
    "<url_or_company_title_key>": {
      "title": "...",
      "company": "...",
      "url": "...",
      "first_seen": "YYYY-MM-DD",
      "fit": "high/medium/low",
      "status": "new/skipped/evaluated/ranked/expired"
    }
  }
}
  1. Only present jobs NOT already in the seen list or tracker.

Step 5: Present Results

Present new jobs in a table sorted by fit (high first):

## New Job Matches - YYYY-MM-DD

Found X new positions (Y high, Z medium, W low match).

| # | Fit | Title | Company | Location | Deadline | URL |
|---|-----|-------|---------|----------|----------|-----|
| 1 | High | ... | ... | ... | ... | [Link](...) |

### High-Match Highlights
For each high-match job, add 2-3 bullet points:
- Why it matches your profile
- Key requirements to check
- Any red flags

After presenting, ask:

"Want me to evaluate any of these in detail? Just give me the number(s)."

If the user picks a number, invoke the job-application-assistant skill workflow (fit evaluation first, then CV + cover letter if approved).

If the run found many new jobs (roughly 8+), also suggest /rank - it batch-scores all new postings against the full fit framework and returns a ranked shortlist, which beats eyeballing a long table. (/rank sets the ranked and expired status values in seen_jobs.json; treat both as already-seen for dedup purposes.)

Step 6: Update Tracker (Optional)

If the user decides to apply to any job, add a row to job_search_tracker.csv.


Important Rules

  1. Never fabricate job postings. Only present jobs found via actual WebSearch/WebFetch results.
  2. Respect deduplication. Always check seen_jobs.json AND job_search_tracker.csv before presenting.
  3. Focus on configured geographic area. Skip jobs that require relocation or are clearly outside commute range.
  4. Only open positions. Skip postings with expired deadlines or those marked as closed.
  5. Be efficient with WebFetch. Don't fetch every search result - use titles and snippets to pre-filter before fetching.
  6. Parallel searches. Use the Agent tool or parallel WebSearch calls to speed up the search phase.