Files
ai-job-search/.claude/skills/job-scraper/SKILL.md
T
PRATHAM KUMARandrajpratham1 1ac677dd7d This PR fixes two genuine gaps in the framework — a documented command with no backing file, and CLI tools that were installed but never called. (#52)
* feat: add /upskill command file to wire the upskill skill into Claude Code

The upskill skill and its full SKILL.md workflow already existed in
.claude/skills/upskill/SKILL.md, but there was no corresponding command
file in .claude/commands/. Without it, running /upskill in Claude Code
had zero structured behaviour — Claude would improvise with no defined
steps, mode detection, or output format.

This commit adds .claude/commands/upskill.md as the thin orchestration
layer that was missing:

- Step 0: Parses \ to determine aggregate mode (no args,
  analyses all jobs in job_search_tracker.csv) vs. targeted mode
  (a URL is passed, analyses that single posting). Unrecognised input
  triggers a clarifying prompt rather than silently misbehaving.

- Step 1: In aggregate mode, reads the tracker and exits early with a
  helpful message if it is empty, so the user is never dropped into a
  broken analysis with no data.

- Step 2: Delegates all analysis work to the existing upskill SKILL.md
  (hard skill diff, LLM synthesis, heatmap, web-searched resources,
  study order, report save). No analysis logic is duplicated here.

- Step 3: Presents a concise post-run summary — critical/high gaps,
  total estimated study time, and next-step suggestions (/scrape,
  /apply, review the saved report).

Design principle: the command is intentionally a thin driver. All
substantive logic lives in SKILL.md so it remains in one place and
is easy to update independently of the command shell.

* fix: wire CLI tools into /scrape as primary search mechanism

The repo ships five Bun CLI search tools under .agents/skills/:
  - jobindex-search   (Jobindex.dk — largest Danish board)
  - jobbank-search    (Akademikernes Jobbank — academic/professional)
  - jobdanmark-search (Jobdanmark.dk — broad coverage)
  - jobnet-search     (Jobnet.dk — government portal)
  - linkedin-search   (LinkedIn public jobs-guest API — country-agnostic)

Before this fix, none of them were ever called during /scrape. The
job-scraper SKILL.md told Claude to run WebSearch for everything,
meaning the CLIs were installed and documented but sat in dead-code
limbo with no callers.

Changes to .claude/skills/job-scraper/SKILL.md:

1. Added Bash to allowed-tools so the bun CLI commands are permitted
   by Claude Code's tool-permission system. Without this, any attempt
   to shell out would be blocked regardless of the instruction text.

2. Replaced the single WebSearch-only Step 1 with a three-part search
   strategy:

   Step 1a — bun availability check
   Runs \un --version\ first. If bun is not installed the skill
   gracefully degrades to WebSearch for all portals (Step 1c) and
   notes the fallback in the results output, rather than crashing.

   Step 1b — CLI tools as primary mechanism
   For each query term extracted from search-queries.md, runs all five
   CLIs with \--jobage 14 --limit 20 --format json\. Flags are
   consistent with each tool's documented contract so output is
   predictable. Each CLI call is independent: a non-zero exit or empty
   result on one portal does not abort searches on the others. Results
   are collected and merged before deduplication.

   Step 1c — WebSearch fallback
   Used for portals without a CLI skill (karriere.dk, jobfinder.dk,
   company career pages via site: searches) and as the universal
   fallback when bun is unavailable. This preserves backwards
   compatibility for users who have not installed bun yet.

The net effect: /scrape now actually uses the CLI infrastructure the
repo was built around. WebSearch remains available for portals outside
the shipped skill set and for users on environments without bun.

* fix: rework /scrape CLI wiring + drop /upskill command file

Two changes addressing maintainer feedback on PR #52.

--- /scrape: use portal SKILL.md as source of truth ---

The previous approach hardcoded per-portal bun invocations directly
in job-scraper/SKILL.md. This broke in practice:
  - jobbank requires --key (not --query); --query is not a valid flag
  - jobnet uses --search-string and region/occupation filters; passing
    --query silently returns the full unfiltered job firehose
  - --sort date and uniform --jobage 14 are not supported by all portals

The fix removes all hardcoded per-portal flag examples. Instead, Step
1b now instructs the agent to:
  1. Discover installed portal skills via .agents/skills/*/SKILL.md
  2. Read each portal's own SKILL.md for its documented CLI interface
  3. Translate search-queries.md terms into that portal's flag format
  4. Use each portal's supported recency and limit flags

This makes the scraper self-maintaining: new portals added via
/add-portal are automatically included without any changes to this
file, and the scraper can never drift from the CLIs again.

The bun availability check (Step 1a) and Bash in allowed-tools are
preserved - both are still needed. The WebSearch fallback (Step 1c)
is preserved and cleaned up to cover: portals without a CLI skill,
any portal whose CLI fails at runtime, and the bun-unavailable case.

--- /upskill: drop command file ---

.claude/commands/upskill.md is removed. /upskill is deliberately
skill-hosted: .claude/skills/upskill/SKILL.md is the backing file
and parses its own /upskill vs /upskill <URL> modes (same pattern
as /scrape, which also has no command file). The command file created
a second entry point that duplicated the skill's argument parsing,
violating the single-source-of-truth principle established in #44
and #49.

---------

Co-authored-by: rajpratham1 <your-email@example.com>
2026-07-08 17:08:12 +02:00

154 lines
6.3 KiB
Markdown

---
name: job-scraper
description: >
Scrapes Danish job sites for new positions matching your profile. Deduplicates across runs.
Triggers on: job scrape, find jobs, search jobs, new jobs, job search, scrape jobs, /scrape
allowed-tools: Read, Write, Edit, Glob, Grep, Bash, WebFetch, WebSearch, Agent, AskUserQuestion
---
# Job Scraper
---
## How It Works
This skill searches multiple Danish job sites using targeted queries based on your profile, deduplicates against previously seen jobs and the application tracker, and presents new matches with a quick fit assessment.
## Invocation
The user triggers this skill by saying things like:
- "Find new jobs"
- "Scrape for jobs"
- "Any new positions?"
- "/scrape"
Optional arguments:
- A focus area, e.g. "/scrape data science" or "/scrape geophysics"
- "broad" to run all search categories, e.g. "/scrape broad"
---
## Execution Steps
### Step 0: Load State
1. Read `job_scraper/seen_jobs.json` (create if missing - start with `{"seen": {}}`)
2. Read `job_search_tracker.csv` to extract already-applied companies+roles
3. Read `search-queries.md` (this directory) for the search strategy
### Step 1: Search
Read `search-queries.md` (this directory) for the search strategy. By default, run the top 3 priority query categories. If the user said "broad", run all categories. If the user specified a focus area (e.g. "data science"), prioritize queries from that category.
**Use the installed CLI tools as the primary search mechanism.** Fall back to `WebSearch` only for portals that do not have a CLI skill, or if `bun` is unavailable on the system.
#### 1a. Check bun availability
```bash
bun --version
```
If this fails (bun not installed), skip to **1c (WebSearch fallback)** for all portals and note the fallback in the Step 5 output.
#### 1b. Run CLI tools (primary — run these in parallel where possible)
Discover all installed portal CLI skills by reading every `SKILL.md` found under `.agents/skills/*/SKILL.md`. Each file documents that portal's exact CLI flags and usage examples. **Use each portal's own documented interface — do not guess flags.** This approach automatically includes any new portals added via `/add-portal` without requiring changes to this file.
For each installed portal skill:
1. Read its `SKILL.md` to find the correct `bun run …` invocation and supported flags.
2. Translate the query terms from `search-queries.md` into that portal's flag format (e.g. `--key`, `--search-string`, `--query`, filter codes — whatever the portal's SKILL.md specifies).
3. Scope to the last 14 days using the portal's supported recency flag (`--jobage`, `--since <YYYY-MM-DD>`, `--order PublicationDate`, etc. — as documented per portal).
4. Cap results to ~20 per call using the portal's limit flag.
5. Use `--format json` for machine-readable output.
Run all portal CLI calls in parallel where possible using the Agent tool. Collect all `results` arrays into a single pool for Step 2.
If a CLI tool exits with a non-zero code, log the error message and continue — do not abort the whole search.
#### 1c. WebSearch fallback
Use `WebSearch` for:
- Portals listed in `search-queries.md` that do **not** have a corresponding directory under `.agents/skills/`
- Any portal whose CLI fails at runtime
- When bun is unavailable (Step 1a failed)
Use the site-specific query strings from `search-queries.md` directly as WebSearch queries for these portals.
### Step 2: Fetch & Parse
For each promising result from Step 1:
- Use `WebFetch` to retrieve the job posting page
- Extract: **job title**, **company**, **location**, **posting date** (or "recent"), **URL**, **key requirements** (brief), **application deadline** (if listed)
- Skip if the URL or company+title combo already exists in `seen_jobs.json`
- Skip if the company+role already appears in `job_search_tracker.csv`
### Step 3: Quick Fit Assessment
For each new job, do a rapid fit check (NOT the full evaluation from `04-job-evaluation.md` - just a quick signal):
- **High match**: Role directly involves your core skills
- **Medium match**: Role is adjacent to your experience
- **Low match**: Role requires significant skills you lack
### Step 4: Deduplicate & Store
1. Add ALL fetched jobs (new and skipped) to `seen_jobs.json` with structure:
```json
{
"seen": {
"<url_or_company_title_key>": {
"title": "...",
"company": "...",
"url": "...",
"first_seen": "YYYY-MM-DD",
"fit": "high/medium/low",
"status": "new/skipped/evaluated/ranked/expired"
}
}
}
```
2. Only present jobs NOT already in the seen list or tracker.
### Step 5: Present Results
Present new jobs in a table sorted by fit (high first):
```
## New Job Matches - YYYY-MM-DD
Found X new positions (Y high, Z medium, W low match).
| # | Fit | Title | Company | Location | Deadline | URL |
|---|-----|-------|---------|----------|----------|-----|
| 1 | High | ... | ... | ... | ... | [Link](...) |
### High-Match Highlights
For each high-match job, add 2-3 bullet points:
- Why it matches your profile
- Key requirements to check
- Any red flags
```
After presenting, ask:
> "Want me to evaluate any of these in detail? Just give me the number(s)."
If the user picks a number, invoke the **job-application-assistant** skill workflow (fit evaluation first, then CV + cover letter if approved).
If the run found many new jobs (roughly 8+), also suggest `/rank` - it batch-scores all new postings against the full fit framework and returns a ranked shortlist, which beats eyeballing a long table. (`/rank` sets the `ranked` and `expired` status values in `seen_jobs.json`; treat both as already-seen for dedup purposes.)
### Step 6: Update Tracker (Optional)
If the user decides to apply to any job, add a row to `job_search_tracker.csv`.
---
## Important Rules
1. **Never fabricate job postings.** Only present jobs found via actual WebSearch/WebFetch results.
2. **Respect deduplication.** Always check seen_jobs.json AND job_search_tracker.csv before presenting.
3. **Focus on configured geographic area.** Skip jobs that require relocation or are clearly outside commute range.
4. **Only open positions.** Skip postings with expired deadlines or those marked as closed.
5. **Be efficient with WebFetch.** Don't fetch every search result - use titles and snippets to pre-filter before fetching.
6. **Parallel searches.** Use the Agent tool or parallel WebSearch calls to speed up the search phase.